> Markdown version of [/jobs/ext/1459136-devops-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/1459136-devops-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # DevOps / Site Reliability Engineer (SRE) - **Company:** THE NEXTAMP LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $72,800.0 - $101,920.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Automation of Tests, Microsoft Azure, Backup Devices, Cloud Computing, Cloud Engineering, Computer Programming, Directed Acyclic Graph (Directed Graphs), Software Debugging, Linux, DevOps, Disaster Recovery, Github, Python (Programming Language), Key Management, Linux System Administration, Operational Databases, Performance Tuning, Reliability Engineering, Cloud Services, Prometheus, Software Engineering, Data Logging, Spring Cloud, Istio, System Availability, Delivery Pipeline, Grafana, Reliability of Systems, Cloudformation, Kubernetes, Performance Monitor, Hashicorp, Linkerd (Service Mesh), Terraform, Dynatrace, Docker - **Published:** July 27, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e5769a2f6742b9b0 ## About the Role * 5-8 years of experience in DevOps or Site Reliability Engineering. * Strong programming experience in Python. * Hands-on experience migrating Apache Airflow v2 to Airflow v3. * Experience working with the Astronomer platform. * Strong knowledge of Airflow DAG development, scheduling, monitoring, debugging, and optimization. * Experience building and managing CI/CD pipelines using GitHub Actions. * Experience designing and maintaining CI testing infrastructure. * Strong experience with Docker and Kubernetes. * Hands-on experience managing workloads on Amazon EKS (Elastic Kubernetes Service). * Strong knowledge of Infrastructure as Code using Terraform, CloudFormation, or similar tools. * Experience with AWS cloud services and cloud-native architectures. * Hands-on experience with the Grafana Stack, including: * Grafana * Grafana Alloy * Grafana Beyla * Deep understanding of eBPF for Linux observability, networking, and performance monitoring. * Experience implementing monitoring, logging, tracing, and alerting solutions. * Strong Linux administration, networking, and security fundamentals. * Strong troubleshooting skills with production incident management experience. Preferred Skills * Experience with Helm, ArgoCD, or FluxCD. * Experience with OpenTelemetry and distributed tracing. * Knowledge of Prometheus, Loki, Tempo, and modern observability platforms. * Experience with service mesh technologies such as Istio or Linkerd. * Familiarity with secrets management solutions such as HashiCorp Vault or AWS Secrets Manager. * Knowledge of SRE principles including SLIs, SLOs, Error Budgets, and reliability engineering. * Experience supporting highly available, distributed production systems. * Experience implementing cloud governance and infrastructure cost optimization. * Knowledge of disaster recovery, backup, and business continuity strategies. Nice to Have * Experience supporting 24×7 production environments. * Experience working with large-scale distributed platforms and cloud-native applications. * AWS Certified DevOps Engineer. * AWS Solutions Architect. * Certified Kubernetes Administrator (CKA). * Grafana or Kubernetes-related certifications. * Microsoft Azure DevOps Engineer Expert (if applicable). ## Description We are looking for an experienced DevOps / Site Reliability Engineer (SRE) with strong expertise in cloud infrastructure, automation, observability, and production operations. The ideal candidate will have hands-on experience with Python, Apache Airflow, Astronomer, GitHub Actions, Amazon EKS, and the Grafana observability stack. This role involves building scalable infrastructure, improving CI/CD pipelines, implementing modern monitoring solutions, and ensuring the reliability and performance of mission-critical applications., * Design, implement, and maintain scalable and secure cloud infrastructure on AWS. * Build, automate, and optimize CI/CD pipelines using GitHub Actions and other modern DevOps tools. * Lead migration of Apache Airflow v2 to Airflow v3 on the Astronomer platform. * Design, develop, optimize, and troubleshoot Airflow DAGs for production data workflows. * Deploy, manage, and maintain Kubernetes workloads on Amazon EKS. * Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools. * Implement and maintain observability solutions using the Grafana Stack, including Grafana Alloy and Grafana Beyla. * Utilize eBPF technologies for deep infrastructure and application observability. * Build and maintain automated testing infrastructure supporting CI/CD pipelines. * Monitor application and infrastructure health, respond to production incidents, and perform Root Cause Analysis (RCA). * Automate operational processes to improve deployment reliability and reduce manual effort. * Collaborate closely with software engineering teams to improve deployment processes, platform stability, and system reliability. * Implement logging, monitoring, alerting, and performance optimization best practices. * Ensure high availability, scalability, security, and operational excellence across production environments. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)