> Markdown version of [/jobs/ext/2872174-senior-site-reliability-engineer-ai-agents-automation](https://www.wearedevelopers.com/jobs/ext/2872174-senior-site-reliability-engineer-ai-agents-automation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer, AI Agents & Automation - **Company:** ServiceTitan, Inc. - **Location:** United States - **Experience:** Expert - **Salary:** $137,900.0 - $221,400.0 - **Contract:** Permanent contract - **Skills:** ASP.NET, Java (Programming Language), .NET Framework, Artificial Intelligence, Amazon Web Services, Computing Platforms, Microsoft Azure, Cloud Computing, Cloud Engineering, Computer Programming, Databases, Continuous Integration, Cursor (Graphical User Interface Elements), Distributed Systems, Elasticsearch, Github, IP Addressing, Subnetting, Python (Programming Language), Networking Basics, Reliability Engineering, Software Tools, Prometheus, Web Applications, Datadog, GitHub Copilot, Flask (Web Framework), Grafana, Fastapi, Gitlab-ci, Kubernetes, Teamcity, Human in the Loop - **Published:** September 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=f35b58795197a9c9 ## About the Role * Kubernetes (must-have): strong, hands-on understanding of Kubernetes as a system. * AI-native SRE practice (must-have): hands-on, personal experience using AI tools and agents (e.g., Claude Code, GitHub Copilot Workspace, Cursor, or custom agents built on Claude/MCP servers) to diagnose, automate, and resolve infrastructure and reliability work. + Agentic depth: You've built or operated agents that autonomously monitor infrastructure and take action (e.g., an agent that watches system load and scales, remediates, or escalates without a human in the loop) - not just general AI coding assistance. You can speak concretely to how these agents are actually built and operated (MCP protocol, what a harness is, how to differentiate agent-design strategies rather than just naming tools), to context management (e.g., progressive-disclosure strategies for surfacing the right information without dumping everything into context), and to securing agent actions (scoping authorization, guardrails, what's available in the AI infra ecosystem to enforce it). + SRE application: That agentic work is pointed at infrastructure and reliability problems specifically - you use AI to move at a materially faster pace in the SRE domain, not as a general-purpose coding aid. * SRE principles: practical experience with SLIs, SLOs, and error budgets - able to speak to how you've defined and monitored these on real systems, not just definitions. * Cloud engineering & networking: solid grounding in AWS or Azure, including networking fundamentals (subnetting, IP addressing). * Observability: deep experience with at least one modern observability stack (OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch) and the ability to translate that understanding across tools. * CI/CD: strong understanding of a CI/CD system - GitHub Actions preferred, but TeamCity, Azure DevOps, or GitLab CI experience is acceptable. * Programming: strong programming skills with the ability to build web applications - ideally with solid working knowledge of .NET and ASP.NET. We're also open to strong Python (Flask, FastAPI) or Java (Spring) backgrounds. The coding assessment will be tailored to whichever language/framework you're most comfortable in. * Experience with distributed systems and their common failure modes (retries, timeouts, cascading failures). * Strong production troubleshooting skills - comfortable diagnosing issues under pressure. * 8-10+ years of relevant hands-on experience. * Nice-to-have: database experience (not mandatory - databases are monitored by the same team, not owned individually)., You're someone who enjoys being directly accountable for the reliability of a business-critical, large-scale enterprise system. You're comfortable guiding and making decisions with limited information, and capable of operating within the trade-offs between solving for immediate needs versus bigger-scale solutions. You feel rewarded by developing an operability culture in a quickly growing and changing environment, and you're comfortable owning a wide and diverse set of problem areas. ## Description We're looking for a Senior Site Reliability Engineer to join our Site Reliability & Infrastructure Engineering team. We run entirely on the cloud, and this team owns the reliability and health of the applications running on top of it - designing the signals that tell us when something's wrong, and building the systems that keep ServiceTitan running better, faster, and cheaper as we scale. We make a huge impact on thousands of companies in the U.S. and abroad by enabling them to be more efficient and effective at running their business. Our Site Reliability and Infrastructure Engineering team centralizes the concerns of measurement and guidance so every engineer can improve availability and efficiency in their own area of the ServiceTitan cloud. We have a cultural foundation built on diversity, inclusion, and innovation, and we want you and your ideas to thrive at ServiceTitan. Come join us. What You'll Do * Participate in an on-call rotation, using runbooks and playbooks to diagnose and resolve production issues (e.g., adjusting Horizontal Pod Autoscaler rules in response to load). * Design, build, and maintain observability dashboards and alerting grounded in Service Level Indicators (SLIs) and Service Level Objectives (SLOs). * Operate and improve our Kubernetes-based compute platform, which runs the large majority of our infrastructure. * Work across cloud networking and infrastructure (Azure/AWS) to support reliable, scalable systems. * Investigate and resolve production incidents, including root-cause analysis and follow-up remediation work. * Build and operate AI agents and automation that take on manual, repetitive SRE work directly - not just tools that assist a human doing the work. * Partner with product engineering teams to review architecture and infrastructure decisions before they ship. * Write and maintain runbooks and documentation so on-call knowledge is shared across the team, not siloed with one person. * Help define non-functional requirements - scalability, availability, performance - for new systems as they're designed. * Collaborate across engineering teams to adopt best practices in reliability and observability. * Contribute to CI/CD pipelines and help teams ship changes safely and quickly. ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Intro to FastAPI](https://www.wearedevelopers.com/videos/462-intro-to-fastapi) - [Building Multi-Tenant ASP.NET Core Applications: Best Practices and Real-World Solutions](https://www.wearedevelopers.com/videos/1552-building-multi-tenant-asp-net-core-applications-best-practices-and-real-world-solutions) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Building and Deploying Multi-Agent Systems with ADK and Vertex AI](https://www.wearedevelopers.com/videos/1918-building-and-deploying-multi-agent-systems-with-adk-and-vertex-ai) ## Related Articles - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai)