> Markdown version of [/jobs/ext/2112771-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2112771-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Akamai Technologies - **Location:** Providence, RI, United States - **Experience:** Expert - **Salary:** $121,400.0 - $218,600.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Border Gateway Protocol, Cloud Computing, Computer Engineering, Network Topologies, Systems Analysis, IPv4, IPv6, Python (Programming Language), Reliability Engineering, Prometheus, Software Systems, Systems Integration, Virtualization Technology, Private Cloud Environment, Network Switches, Network Routing, Scripting, Large Language Models, Grafana, Information Technology, Slack, Low Latency, Bare Metal, Hardware Acceleration, Api Management, Pagerduty - **Published:** August 19, 2026 - **Apply:** https://www.careerbuilder.com/job-details/senior-site-reliability-engineer-providence-ri--70ff4f54-4692-47df-a87b-6fbce51f7ca0 ## About the Role + Have 5 years of relevant experience and a Bachelor's degree in Computer Engineering, Computer Science or equivalent + Possess tooling and coding ability in languages like Python to construct scalable operational tools, API integrations, and automation frameworks. + Show hands-on experience with modern observability stacks and timeseries engines, like Prometheus, Grafana, OpenTelemetry, and Loki. + Possess a working understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks. + Have experience acting as a key designer for new service rollouts, including establishing operational readiness criteria, telemetry baselines, and alerting thresholds. + Demonstrate extensive experience building technical runbooks, leading complex incident response bridges, and driving comprehensive, blameless post-mortems. + Display a proven ability to take absolute ownership of ambiguous technical problems, coordinate cross-functional teams, and drive for production-grade solutions., Alliance/Partner Marketing, Application Programming Interface (API), Artificial Intelligence (AI), Automation, BGP, Cloud Computing, Computer Engineering, Computer Science, Content Delivery/Distribution, Cross-Functional, IPv4, IPv6, Incident Management, Incident Response, Machine Tool, Metrics, Network Routing, Network Switching, Network Topology, On Call, On Site Support, Performance Metrics, Private Cloud, Python Programming/Scripting Language, Reliability Engineering, Reporting Dashboards, Scripting (Scripting Languages), Slack, Systems Analysis, Telemetry, Time Management, Vehicle Fleets, Virtualization ## Description The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner with the best In this role, you'll play a part in pioneering the reliability an elite, high-density hardware and software infrastructure spanning the globe. You'll collaborate with product teams from the earliest stages of development to ensure the reliability, scalability, and performance of our systems. You'll define key performance indicators and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: + Developing and scaling robust programmatic tooling and infrastructure-as-code utilities in Python to eliminate operational toil and automate fleet-wide provisioning. + Integrating automated workflows across disconnected corporate ticketing systems to optimize time-to-mitigate metrics for hardware and network break-fix events. + Leveraging advanced AI utilities and LLM-assisted development paradigms where appropriate to accelerate technical execution, script authorship, and system analysis + Working on cutting-edge private cloud and compute technologies to improve the availability, latency, and overall systemic health of high-density hardware environments. + Designing and implementing telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI-based anomaly detection tailored for bare-metal and virtualized environments. + Participating in 24x7x365 on-call rotations, spearheading real-time incident management, and managing high-severity service disruption protocols via automated PagerDuty and Slack workflows. + Partnering directly with third-party infrastructure vendors and coordinating on-site field technicians to facilitate uptime activities. ## Related Videos - [IP Authentication: A Tale of Performance Pitfalls and Challenges in Prod](https://www.wearedevelopers.com/videos/1413-ip-authentication-a-tale-of-performance-pitfalls-and-challenges-in-prod) - [IoT: The road to sustainability](https://www.wearedevelopers.com/videos/557-iot-the-road-to-sustainability) - [Stack Overflow: Community and AI](https://www.wearedevelopers.com/videos/600-stack-overflow-community-and-ai) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Agentic employees in world's most downloaded FinTech app](https://www.wearedevelopers.com/videos/100123-agentic-employees-in-world-s-most-downloaded-fintech-app) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)