> Markdown version of [/jobs/ext/1451944-aws-devops-sre-engineer-datadog-aiops](https://www.wearedevelopers.com/jobs/ext/1451944-aws-devops-sre-engineer-datadog-aiops). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AWS DevOps / SRE Engineer Datadog & AIOps - **Company:** VDart, Inc. - **Location:** Atlanta, GA, United States - **Experience:** Expert - **Salary:** $101,500.0 - $169,100.0 - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Agile Methodology, Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Confluence, JIRA, Automation of Tests, Bash Shell, Cloud Computing Security, Cloud Engineering, Code Review, Continuous Integration, Linux, DevOps, Github, Monitoring of Systems, Identity and Access Management, Issue Tracking Systems, Python (Programming Language), Key Management, Machine Learning, Octopus Deploy, Windows PowerShell, Role-Based Access Control, Reliability Engineering, Site Reliability Engineering Practices, Ansible, Prometheus, TypeScript, Datadog, Scripting, Autoscaling, Large Language Models, Grafana, Git, Cloudformation, Amazon Relational Database Service, Containerization, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Route53, BIG-IP Access Policy Manager (APM), Functional Programming, Cloudwatch, Terraform, Splunk, AWS EKS, Docker, Jenkins, Static Application Security Testing, Vulnerability Analysis, Golang, Programming Languages, Dynamic Application Security Testing - **Published:** July 26, 2026 - **Apply:** https://www.careerjet.com/jobad/us96396720ece594b6f43870caf28b439f ## About the Role * A hands-on engineer who balances delivery velocity with production reliability and operational discipline. * A proactive problem-solver who uses data, automation, and observability to prevent recurring issues. * A collaborative technical leader who can influence teams to adopt secure cloud, DevOps, SRE, and AI-assisted engineering practices. * Someone comfortable owning high-visibility production platforms and driving improvements from design through operations., * Minimum 5+ years of experience in DevOps, Cloud Engineering, Platform Engineering, or a related role supporting enterprise production systems. * Minimum 3+ years of hands-on experience designing, deploying, and operating solutions on AWS. * Strong experience building and supporting production-grade CI/CD pipelines; GitLab CI experience is preferred. * Demonstrated SRE experience with SLOs/SLIs, error budgets, incident response, on-call operations, reliability engineering, capacity planning, and blameless post-incident reviews. * Strong hands-on Datadog experience across metrics, logs, APM/tracing, dashboards, alerting, monitors, synthetics, integrations, and service-level reporting. * Experience with Terraform or CloudFormation and repeatable Infrastructure as Code practices. * Experience with Docker and Kubernetes; AWS EKS experience is strongly preferred. * Proficiency in at least one scripting or programming language such as Python, Bash, PowerShell, Go, or JavaScript/TypeScript. * Strong understanding of Linux, networking, IAM, secrets management, cloud security, and production troubleshooting. * Experience working in Agile environments and using tools such as Jira, Confluence, and Git. * Strong communication, collaboration, documentation, and problem-solving skills with a high degree of ownership. Preferred Qualifications: * Hands-on experience applying AIOps, machine learning, or generative AI to CI/CD, observability, incident management, automated testing, code review, or root-cause analysis. * Experience integrating LLM-based assistants or agents with developer platforms, repositories, ticketing systems, observability tools, or operational runbooks using secure enterprise controls. * Experience with progressive delivery and GitOps tools such as Argo CD, Flux, feature flags, canary deployments, and blue/green deployments. * Experience with OpenTelemetry, Prometheus, Grafana, Splunk, CloudWatch, or other observability platforms in addition to Datadog. * Knowledge of DORA metrics and experience improving deployment frequency, lead time for changes, change-failure rate, and mean time to recovery. * Experience supporting regulated, financial-services, or other highly controlled enterprise environments. * AWS, Kubernetes, Terraform, Datadog, or relevant DevOps/SRE certifications. * Experience leading technical initiatives or mentoring engineering teams. * Key Success Measures * Improved CI/CD speed, stability, reuse, and developer adoption. * Reduced deployment failures, alert noise, incident recurrence, and recovery time. * Clear service-health reporting through Datadog dashboards, SLOs, and actionable alerts. * Increased automation across provisioning, testing, release controls, and operational remediation. * Safe, measurable adoption of AIOps and AI capabilities without weakening security or governance. ## Description * We are seeking a hands-on AWS DevOps Engineer with strong Site Reliability Engineering capabilities and deep Datadog experience. This role will design and improve secure, scalable CI/CD pipelines; increase platform reliability through observability, automation, and SLO-driven practices; and introduce practical AIOps and generative AI capabilities that improve build quality, deployment safety, incident response, and engineering productivity. * Day-to-Day Job Duties * Design, build, and maintain resilient AWS environments using services such as EKS, EC2, S3, IAM, Lambda, RDS, CloudWatch, Route 53, ALB/NLB, and Secrets Manager. * Build, standardize, and optimize CI/CD pipelines using GitLab CI, GitHub Actions, Jenkins, or similar platforms, with automated testing, quality gates, approvals, rollback, and progressive-delivery controls. * Apply SRE practices by defining service-level indicators, service-level objectives, error budgets, availability targets, and operational-readiness criteria. * Implement and administer Datadog capabilities including infrastructure monitoring, APM, log management, Real User Monitoring, synthetics, dashboards, monitors, service maps, and incident workflows. * Create actionable observability and alerting strategies that reduce noise, improve mean time to detect and recover, and support rapid root-cause analysis. * Automate infrastructure provisioning and configuration using Terraform, CloudFormation, Ansible, or equivalent Infrastructure as Code tools. * Operate containerized workloads using Docker and Kubernetes/EKS, including autoscaling, health checks, resource optimization, and cluster reliability. * Integrate security and compliance controls into CI/CD, including secrets management, IAM least privilege, vulnerability scanning, SAST/DAST, dependency checks, artifact integrity, and audit evidence. * Use AIOps and generative AI to improve pipeline efficiency through intelligent failure analysis, configuration review, test generation, anomaly detection, change-risk scoring, and remediation recommendations. * Develop automation and operational tooling using Python, Bash, PowerShell, or similar scripting languages. * Lead production troubleshooting, incident response, post-incident reviews, problem management, and permanent corrective-action tracking. * Partner with application, platform, security, QA, and product teams to improve deployment frequency, change-failure rate, lead time, reliability, and recovery performance. * Maintain runbooks, architecture diagrams, operational procedures, and engineering standards in Confluence, Jira, or similar tools. * Provide technical guidance and mentor engineers on cloud reliability, observability, automation, DevOps, and SRE practices. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Industrializing your Data Science capabilities](https://www.wearedevelopers.com/videos/178-industrializing-your-data-science-capabilities) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 162: AI careers, MCP, AWS best practices & floppy sweaters](https://www.wearedevelopers.com/magazine/571-dev-digest-162-ai-careers-mcp-aws-best-practices-floppy-sweaters)