Site Reliability Engineer II

Akamai Technologies
Topeka, KS, United States
2 months ago
Apply on www.kansasworks.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Continuous Integration Linux Monitoring of Systems Python (Programming Language) Linux System Administration Reliability Engineering Site Reliability Engineering Practices Prometheus Akamai Saltstack Grafana
+6 more
Reliability of Systems Kubernetes Infrastructure Automation Frameworks Hardware Infrastructure Terraform Serverless Computing

Job description

In this role, responsibilities will include automation, monitoring, incident response, and working collaboratively with skilled team members. Candidates should possess expertise in Linux systems, automation, and SRE practices. Daily activities involve coding, improving dashboards, enhancing alerts, and minimizing repetitive tasks. Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai’s serverless inference platform.

As an Site Reliability Engineer II, you will be responsible for:

  • Building and maintaining dashboards, alerts, and monitoring for inference workloads using Akamai’s existing observability platform

  • Writing automation and tooling in Python or Go to reduce operational toil and improve system reliability

  • Building and improving runbooks for inference-specific operational procedures, integrating into Akamai’s existing incident management processes

  • Contributing to SLO tracking and reporting, identifying trends and areas for improvement

  • Supporting CI/CD pipeline maintenance, deployment safety checks, and rollback procedures

  • Collaborating with product engineering teams to troubleshoot complex problems across the stack

  • Participating in on-call rotations, responding to production incidents, and conducting blameless post-mortems

Requirements

  • Have 2+ years of experience in Site Reliability Engineering and a Bachelor’s Degree or its equivalent experience

  • Demonstrate coding ability in at least one programming language (Python or Go) with experience writing automation

  • Have experience with Linux systems administration and the ability to troubleshoot complex infrastructure issues

  • Show familiarity with Kubernetes and containerization concepts

  • Have experience with monitoring and observability tools such as Prometheus, Grafana, or similar

  • Have exposure to CI/CD pipelines and infrastructure-as-code tools (Terraform, SaltStack, or equivalent)

  • Show a willingness to learn

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.kansasworks.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:34 min

Leveraging Akamai edge workers for broad geographic scale

Austin Gil · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

2:01 min

Introduction to aki agency and operations context

Martin Beránek · LIVE

1:58 min

Application performance and its direct business impact

Jérôme Vieilledent · LIVE

Videos

See all

Related articles

See all