> Markdown version of [/jobs/ext/1394864-senior-ii-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1394864-senior-ii-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior II Site Reliability Engineer - **Company:** Akamai View all jobs - **Location:** Cambridge, MA, United States - **Experience:** Expert - **Salary:** $146,400.0 - $263,600.0 - **Contract:** Permanent contract - **Skills:** Adobe InDesign, Artificial Intelligence, Cloud Computing, Code Review, Distributed Systems, Python (Programming Language), Reliability Engineering, Site Reliability Engineering Practices, Akamai, Autoscaling, Containerization, AI Platforms, Kubernetes, Machine Learning Operations, Hardware Infrastructure, Serverless Computing - **Published:** July 23, 2026 - **Apply:** https://www.careerjet.com/job/us468c19e5e3c9d5d8640586b9438c914b/eaa ## About the Role * 8+ years of experience in SRE, infrastructure engineering, or platform engineering, working with large-scale distributed systems * Possess a proven track record of defining SLO/SLI frameworks, building observability platforms, and running incident management processes at scale * Have extensive Kubernetes and containerization experience at scale, including autoscaling, resource scheduling, and container orchestration for compute-intensive workloads * Have experience building automation and tooling in Python or Go, with familiarity in CI/CD pipelines, deployment safety, and infrastructure-as-code * Possess the ability to lead technical initiatives across teams, mentor other engineers, and drive complex reliability problems to resolution independently * Have experience with or exposure to AI/ML infrastructure, model serving, or GPU workloads ## Description Do you want to shape reliability practices for a new AI inference platform? Are you a senior technical leader who drives solutions across teams? Join the Akamai Inference Cloud Team! The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications. Partner with the best In this role, you'll lead reliability workstreams for Akamai's serverless inference platform, design SRE tooling and automation, and drive technical decisions. Opportunities exist to mentor other SREs, influence architecture decisions with product engineering teams, and shape SRE practices for AI inference workloads and GPU infrastructure at scale. As a Senior II Site Reliability Engineer, you will be responsible for: * Taking ownership of observability strategy for the serverless inference platform, designing telemetry, dashboards, and alerts, defining SLO/SLI frameworks, and driving improvements when targets are missed * Building production-grade automation and tooling that reduces operational toil, improves incident response, and sets patterns that other SREs adopt * Owning incident management integration for inference workloads, designing frameworks, leading incident response during on-call rotations, and driving systemic improvements from post-mortems * Defining and implementing deployment safety practices including progressive rollouts, canary analysis, and rollback automation, establishing standards for the team * Partnering with product engineering teams to influence architecture decisions, ensure operational readiness, and represent the SRE perspective in design reviews * Mentoring Senior and mid-level SREs through code reviews, design discussions, and hands-on problem-solving ## Related Videos - [What the Heck is Edge Computing Anyway?](https://www.wearedevelopers.com/videos/593-what-the-heck-is-edge-computing-anyway) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Build a CI/CD pipeline to automate code reviews and ensure code quality](https://www.wearedevelopers.com/videos/349-build-a-ci-cd-pipeline-to-automate-code-reviews-and-ensure-code-quality) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)