> Markdown version of [/jobs/ext/138622-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/138622-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** BLITZY INC. - **Location:** Cambridge, MA, United States - **Experience:** Expert - **Salary:** $160,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Application Release Automation, Microsoft Azure, Bash Shell, DevOps, Distributed Systems, Fault Tolerance, Python (Programming Language), Open Source Technology, Reliability Engineering, Prometheus, Software Engineering, Datadog, Data Logging, Pulumi, Cloud Platform System, DevOps Tools - Open-source, Istio, Delivery Pipeline, Grafana, Mttr, Kubernetes, Infrastructure Automation Frameworks, Linkerd (Service Mesh), Terraform, Golang - **Published:** May 27, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=1d70cb5029446383 ## About the Role Do you have experience in Tooling?, * 5+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering roles. * Strong proficiency in at least one major cloud platform (AWS preferred); experience with Kubernetes and container orchestration at scale. * Hands-on experience with infrastructure-as-code tools (Terraform, Pulumi, or equivalent). * Proven track record designing and maintaining high-availability, distributed systems. * Deep expertise in observability tooling, incident management, and on-call practices. * Strong scripting and automation skills (Python, Go, Bash, or similar). * Excellent communication skills with the ability to collaborate across engineering teams and present technical findings to leadership. What Makes You Stand Out * Experience supporting AI/ML workloads or GPU-accelerated infrastructure. * Prior experience in a high-growth startup environment where you wore multiple hats. * Familiarity with eBPF, service mesh technologies (Istio, Linkerd), or advanced networking. * Contributions to open-source SRE/DevOps tooling or communities. * Experience building global, multi-region infrastructure with strict latency and availability requirements. ## Description As a Senior Site Reliability Engineer at Blitzy's Pune headquarters, you will be the backbone of our platform's reliability, scalability, and operational excellence. You'll work at the intersection of software engineering and infrastructure, ensuring our AI-powered development platform remains highly available and performant as we scale rapidly. This is a high-impact, hands-on role for an engineer who thrives in a fast-moving environment and takes deep ownership of the systems they build. What Success Looks Like * In 30 days: You have a deep understanding of Blitzy's infrastructure architecture, have identified key reliability risks, and are actively contributing to on-call rotations. * In 90 days: You have shipped meaningful improvements to observability, incident response workflows, and deployment pipelines that measurably reduce MTTR and increase system uptime. * In 6 months: You have driven at least one major reliability initiative from inception to production, established SLO/SLA frameworks for critical services, and are a trusted technical voice shaping our infrastructure roadmap. Areas of Ownership * Design, build, and operate scalable, fault-tolerant infrastructure across cloud environments (AWS, GCP, or Azure). * Define and enforce SLOs, SLAs, and error budgets; lead blameless postmortems and drive systemic improvements. * Build and maintain robust CI/CD pipelines, release automation, and deployment infrastructure. * Own observability: design and maintain logging, metrics, tracing, and alerting stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry). * Partner closely with software engineering teams to embed reliability practices into the development lifecycle. * Drive capacity planning, performance benchmarking, and cost optimization across our infrastructure. * Champion security best practices within the infrastructure and deployment layers. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)