> Markdown version of [/jobs/ext/585726-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/585726-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Attain - **Location:** Greenville, SC, United States (Remote available) - **Experience:** Expert - **Salary:** $100,000.0 - $120,000.0 - **Contract:** Permanent contract - **Skills:** .NET Framework, Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Microsoft Azure, Bash Shell, C Sharp (Programming Language), Cloud Computing, Code Review, Continuous Integration, Software Debugging, DevOps, Disaster Recovery, Github, Identity and Access Management, Python (Programming Language), Uptime, PCI Data Security Standards, Reliability Engineering, Prometheus, Runbook, Software Engineering, Scripting, Autoscaling, Grafana, Amazon Virtual Private Cloud (VPC), Kubernetes, Information Technology, Cloudwatch, Terraform - **Published:** June 22, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=cb4fd37f5d5683cc ## About the Role Do you have experience in Tooling?, * Kubernetes, ArgoCD, Helm, Terraform, Python. Deep hands-on production experience. * Hands-on AWS. Operate and debug EKS, ECS, EC2, ECR, IAM/IRSA, VPC networking, ALB/NLB, CloudWatch, Secrets Manager, and KMS. * GitHub Actions and/or Azure DevOps. Build and operate CI/CD at scale. * Grafana and the observability stack. Hands-on with Grafana dashboards and alerting, and the metrics, logs, and traces stack (Prometheus/Mimir, Loki, Tempo, OpenTelemetry). * Strong scripting. Python and Bash, with the ability to grow into systems-level coding. * Production troubleshooting. Comfortable getting into a system under load, finding root cause, and fixing it. * Production ownership. Uptime and reliability accountability. * Incident response. You respond and help drive postmortems that yield real improvements. * Standards contribution. You contribute to engineering standards and best practices. * Compliance awareness. Experience in regulated or high-rigor environments or implementing audit and access controls in pipelines. * Mentorship.Through code review, examples, and pairing. * 5+ years in site reliability, platform, DevOps, or software engineering, with production ownership of systems or pipelines. Preferred Qualifications * Advanced GitOps. ArgoCD (or Flux), reusable Helm patterns, Argo Rollouts. * CI consolidation or migration. Moving between CI systems, such as Azure DevOps to GitHub Actions. * Self-hosted observability at scale. Running Grafana, Mimir, Loki, and Tempo in production. * Supply chain security. SBOMs, artifact signing (Sigstore/cosign), SLSA provenance. * Platform migrations. Contributing to modernization with minimal disruption. * .NET / C#. Enough to containerize and reason about application workloads. * Low-level Kubernetes. Cilium/eBPF, Karpenter, or self-hosted networking and autoscaling. * Resilience testing. Chaos/failure injection or disaster recovery drills. * AI-assisted tooling. Responsible use with output validation. * Certification. AWS Solutions Architect, AWS DevOps Engineer, or CKA/CKAD. * Degree in computer science or equivalent practical experience. ## Description We're looking for a Senior Site Reliability Engineer to help drive the reliability and operational excellence of how we build, ship, and run software. You'll work hands-on across AWS, Kubernetes (EKS), ArgoCD, Helm, Terraform, GitHub Actions, Azure DevOps, Grafana, and Python, building and operating the delivery systems that move our applications safely and reliably into production., * Build and operate the delivery platform. Work across AWS, EKS, ArgoCD, Helm, GitHub Actions, Azure DevOps, Terraform, and Python. * Fix the problems you own. Find root cause across the AWS and Kubernetes stack, fix it, and harden it so it stays fixed. * Respond to incidents. Help stabilize during outages, drive root-cause analysis, and ship corrective actions for your systems. * Standardize how we build and ship. Define reproducible container builds and GitOps paths on ArgoCD and Helm that replace manual deployment. * Help consolidate the CI estate. Standardize pipelines across GitHub Actions and Azure DevOps for your services - remove brittle steps and silent failures and improve visibility. * Support platform adoption. Build golden-path templates and tooling and help teams move services onto the platform. * Use progressive delivery. Canary and blue green deploys (Argo Rollouts) and automated rollback for the services you operate. * Build observability in. Wire golden-signal metrics, logs, and traces (Prometheus/Mimir, Loki, Tempo, OpenTelemetry) into your services, surfaced in Grafana with SLOs for your domain. * Operate production systems. Troubleshoot failed to deploy, respond to alerts, and improve behavior from real incidents. * Help meet SLOs and carry on call. Track reliability metrics for the services you operate and share the rotation. * Built across environments. Design dev, test, and prod for safe promotion, recovery from failed deployments, and zero-downtime upgrades. * Help set the standard. Build reference implementations for build, deploy, GitOps, promotion gates, and observability. * Uphold compliance with the pipeline. Support deployment traceability, approval trails, and segregation of duties for PCI DSS, SOC 2, SOX, and GLBA. * Cut toil and cost. Automate repetitive ops work and help tune EKS compute, CI runners, and observability cardinality. * Unblock across teams. Get hands-on with Cloud, Security, Application Engineering, Data, and Product to keep delivery moving. * Kill knowledge silos. Write docs, runbooks, and incident learnings, so engineers operate independently. ## Related Videos - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [Platform Engineering vs. DevOps Why not both?](https://www.wearedevelopers.com/videos/885-platform-engineering-vs-devops-why-not-both) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)