Site Reliability Engineer II
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
In this role, responsibilities will include automation, monitoring, incident response, and working collaboratively with skilled team members. Candidates should possess expertise in Linux systems, automation, and SRE practices. Daily activities involve coding, improving dashboards, enhancing alerts, and minimizing repetitive tasks. Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai’s serverless inference platform.
As an Site Reliability Engineer II, you will be responsible for:
-
Building and maintaining dashboards, alerts, and monitoring for inference workloads using Akamai’s existing observability platform
-
Writing automation and tooling in Python or Go to reduce operational toil and improve system reliability
-
Building and improving runbooks for inference-specific operational procedures, integrating into Akamai’s existing incident management processes
-
Contributing to SLO tracking and reporting, identifying trends and areas for improvement
-
Supporting CI/CD pipeline maintenance, deployment safety checks, and rollback procedures
-
Collaborating with product engineering teams to troubleshoot complex problems across the stack
-
Participating in on-call rotations, responding to production incidents, and conducting blameless post-mortems
Requirements
-
Have 2+ years of experience in Site Reliability Engineering and a Bachelor’s Degree or its equivalent experience
-
Demonstrate coding ability in at least one programming language (Python or Go) with experience writing automation
-
Have experience with Linux systems administration and the ability to troubleshoot complex infrastructure issues
-
Show familiarity with Kubernetes and containerization concepts
-
Have experience with monitoring and observability tools such as Prometheus, Grafana, or similar
-
Have exposure to CI/CD pipelines and infrastructure-as-code tools (Terraform, SaltStack, or equivalent)
-
Show a willingness to learn
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Find a Developer Job: 12 Best Job Sites For Developers
The Best Job Search Websites of 2025
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again