> Markdown version of [/jobs/ext/3018079-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3018079-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Activision Blizzard, Inc. - **Location:** Atlanta, GA, United States (Remote available) - **Experience:** Expert - **Salary:** $102,800.0 - $190,204.0 - **Contract:** Temporary contract - **Skills:** Amazon Web Services, Data Analysis, Android Software Development, Apple IOS, Big Data, Cloud Computing, Computer Programming, Continuous Delivery, Continuous Integration, Data Integration, Linux, Distributed Systems, Video Game Development, Github, Home Automation, Python (Programming Language), Unix Shell, Load Testing, Machine Learning, Microsoft Software, Octopus Deploy, Reliability Engineering, Prometheus, Azure Machine Learning, Software Engineering, Data Streaming, Xbox Linux, Scripting, Cloud Platform System, Grafana, Model Validation, Build Management, Kubernetes, Data Analytics, Apache Kafka, Machine Learning Operations, Terraform, Jenkins, Golang - **Published:** September 20, 2026 - **Apply:** https://www.careerbuilder.com/job-details/senior-site-reliability-engineer-data-analytics-atlanta-ga--c3d284df-e015-4785-ba3a-99858890b929 ## About the Role * Experience operating reliable, distributed systems in SRE, platform, or similar roles * Experience with data, analytics, ML, or large-scale distributed workloads * Strong knowledge of Linux, containers, Kubernetes, and cloud infrastructure * Experience building automation or internal tools (Python, Go, shell, etc.) * Experience with infrastructure-as-code (e.g., Terraform) * Experience with CI/CD or GitOps systems (e.g., Jenkins, GitHub Actions, ArgoCD) * Familiarity with observability (metrics, logs, traces, alerting, incident response) * Solid understanding of SRE concepts (SLIs, SLOs, error budgets, postmortems) * Experience using modern development and automation practices to improve reliability and efficiency * Experience building internal tooling, automation, or developer productivity systems * Strong communication skills with technical and cross-functional partners Bonus Points * Experience with data and ML systems (training pipelines, model serving, GPU workloads) * Experience with distributed systems and messaging (Kafka, Pub/Sub) * Experience working in Kubernetes-based environments * Familiarity with observability tools (Prometheus, Grafana) * Experience operating systems in cloud environments (GCP, AWS), Access Control, Amazon Web Services (AWS), Android, Apache Kafka, Automation, Budgeting, Business Support, Cloud Computing, Communication Skills, Continuous Deployment/Delivery, Continuous Integration, Cross-Functional, Data Analysis, Distributed Computing, Diversity, Entertainment and Media, GCP (Good Clinical Practices), Gaming, Geography, GitHub, Go Programming Language (Golang), Home Automation, Identify Issues, Incident Response, Jenkins, Linux Operating System, Load Testing, Machine Tool, Messaging Technology, Metrics, Microsoft Product Family, Model Validation, Multimedia, On Call, Operating Systems, PlayStation, Problem Solving Skills, Productivity Model, Python Programming/Scripting Language, Reliability Engineering, Software Development, Unix Shell Programming, Video Games, Xbox, iOS ## Description This Senior Site Reliability Engineer role is on our Data & Analytics team, partnering with data, analytics, ML, and platform engineering to improve the reliability, scalability, and performance of large-scale data platforms, analytics pipelines, ML training pipelines, and inference services. In addition to core SRE responsibilities, this role will build operational and automation tooling that reduces toil, speeds up issue resolution, and improves engineering velocity. This includes contributing to internal platform services such as shared tooling, data integrations, and access-control patterns used across Blizzard. The ideal candidate is a production-minded SRE or platform engineer who is comfortable operating critical systems, writing software, and building tools that improve engineering efficiency without compromising reliability. This role is open to candidates based in Irvine, CA or Albany, NY (hybrid or on-site), as well as fully remote candidates. Responsibilities * Participate in an on-call rotation and drive incidents to resolution * Lead blameless postmortems and identify systemic reliability improvements * Partner with data, ML, and platform teams to improve batch, streaming, training, and inference workloads * Support ML training pipelines and inference services, including GPU workloads * Help define how data and ML services run on Kubernetes * Design and build automation and operational tooling (e.g., workflows, diagnostic tooling, runbooks) to reduce on-call burden * Build and evolve centralized platform services, including shared tooling, data integrations, and access controls * Diagnose and resolve reliability, performance, and cost issues across distributed systems * Champion automation, documentation, and practices that reduce toil * Maintain infrastructure using Terraform and infrastructure-as-code principles * Improve CI/CD and GitOps workflows (Jenkins, GitHub Actions, ArgoCD) * Operate and improve containerized services on Kubernetes * Define and measure reliability using SLIs, SLOs, and error budgets * Run load tests, capacity modeling, and production validation * Build internal tools and paved paths that help teams operate safely and efficiently ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Most Popular Web Developer Jobs in Europe](https://www.wearedevelopers.com/magazine/163-7-most-popular-web-developer-jobs-in-europe) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss)