Senior II Site Reliability Engineer

Akamai View all jobs
Cambridge, MA, United States
about 2 months ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$146,400.0 - $263,600.0
Working hours
Regular working hours

Tech stack

Adobe InDesign Artificial Intelligence Cloud Computing Code Review Distributed Systems Python (Programming Language) Reliability Engineering Site Reliability Engineering Practices Akamai Autoscaling Containerization AI Platforms
+4 more
Kubernetes Machine Learning Operations Hardware Infrastructure Serverless Computing

Job description

Do you want to shape reliability practices for a new AI inference platform? Are you a senior technical leader who drives solutions across teams? Join the Akamai Inference Cloud Team! The Akamai Inference Cloud team is part of Akamai’s Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications. Partner with the best In this role, you’ll lead reliability workstreams for Akamai’s serverless inference platform, design SRE tooling and automation, and drive technical decisions. Opportunities exist to mentor other SREs, influence architecture decisions with product engineering teams, and shape SRE practices for AI inference workloads and GPU infrastructure at scale. As a Senior II Site Reliability Engineer, you will be responsible for:

  • Taking ownership of observability strategy for the serverless inference platform, designing telemetry, dashboards, and alerts, defining SLO/SLI frameworks, and driving improvements when targets are missed
  • Building production-grade automation and tooling that reduces operational toil, improves incident response, and sets patterns that other SREs adopt
  • Owning incident management integration for inference workloads, designing frameworks, leading incident response during on-call rotations, and driving systemic improvements from post-mortems
  • Defining and implementing deployment safety practices including progressive rollouts, canary analysis, and rollback automation, establishing standards for the team
  • Partnering with product engineering teams to influence architecture decisions, ensure operational readiness, and represent the SRE perspective in design reviews
  • Mentoring Senior and mid-level SREs through code reviews, design discussions, and hands-on problem-solving

Requirements

  • 8+ years of experience in SRE, infrastructure engineering, or platform engineering, working with large-scale distributed systems
  • Possess a proven track record of defining SLO/SLI frameworks, building observability platforms, and running incident management processes at scale
  • Have extensive Kubernetes and containerization experience at scale, including autoscaling, resource scheduling, and container orchestration for compute-intensive workloads
  • Have experience building automation and tooling in Python or Go, with familiarity in CI/CD pipelines, deployment safety, and infrastructure-as-code
  • Possess the ability to lead technical initiatives across teams, mentor other engineers, and drive complex reliability problems to resolution independently
  • Have experience with or exposure to AI/ML infrastructure, model serving, or GPU workloads

Benefits & conditions

Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $146,400 - $263,600/year; a candidate’s salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply., + $94,000-140,000 per year Build your best future with the Johnson Controls team! Who we are: Johnson Controls is global leader in smart, healthy, and sustainable buildings. Our mission is to reimagine t…

  • 7 days ago + *

Sr. Process Engineer Johnson Controls

  • Burlington, MA
  • $97,000-162,000 per year Build your best future with the Johnson Controls team! Who we are: Johnson Controls is global leader in smart, healthy, and sustainable buildings. Our mission is to reimagine t…

  • 18 days ago +

About the company

At Akamai, we make life better for billions of people, trillions of times a day. Whether you’re streaming live events, scrolling social media, watching your favorite series, or managing your savings, we’re the engine behind the scenes. We provide the world’s most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone. Our focus is simple: Cloud and Edge: Running apps closer to users for instant performance. Security: Neutralizing threats before they ever reach your data. Content Delivery: Scaling the world’s biggest moments without a glitch. AI: Enabling our customers to build, secure, and scale AI apps on the world’s most distributed cloud platform. At Akamai, we don’t just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we’re the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works. Benefits at Akamai: We support your health, well-being, finances, and life beyond work. FlexBase adapts to your job’s needs Akamai’s FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It’s not about telling employees where to work; it’s about supporting employees to do their best work. We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both. Connect with us on social and see what life at Akamai is like!, Johnson Controls

  • Burlington, MA
  • $76,100-114,100 per year Johnson Controls, a global leader in thermal management, mission-critical building systems, energy efficiency, and decarbonization, helps customers use energy more productively, re…

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:34 min

Leveraging Akamai edge workers for broad geographic scale

Austin Gil · LIVE

3:39 min

Addressing code review surrender and process exploitation

Laura Tacho Laura Tacho · World Congress 2026 Europe

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

56 sec

The hidden costs of delayed peer code reviews

Tim Gilboy Tim Gilboy

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

1:58 min

Application performance and its direct business impact

Jérôme Vieilledent · LIVE

Videos

See all

Related articles

See all