Senior Site Reliability Engineer

BLITZY INC.
Cambridge, MA, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$160,000.0 - $180,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Application Release Automation Microsoft Azure Bash Shell DevOps Distributed Systems Fault Tolerance Python (Programming Language) Open Source Technology Reliability Engineering Prometheus
+15 more
Software Engineering Datadog Data Logging Pulumi Cloud Platform System DevOps Tools - Open-source Istio Delivery Pipeline Grafana Mttr Kubernetes Infrastructure Automation Frameworks Linkerd (Service Mesh) Terraform Golang

Job description

As a Senior Site Reliability Engineer at Blitzy’s Pune headquarters, you will be the backbone of our platform’s reliability, scalability, and operational excellence. You’ll work at the intersection of software engineering and infrastructure, ensuring our AI-powered development platform remains highly available and performant as we scale rapidly. This is a high-impact, hands-on role for an engineer who thrives in a fast-moving environment and takes deep ownership of the systems they build.

What Success Looks Like

  • In 30 days: You have a deep understanding of Blitzy’s infrastructure architecture, have identified key reliability risks, and are actively contributing to on-call rotations.
  • In 90 days: You have shipped meaningful improvements to observability, incident response workflows, and deployment pipelines that measurably reduce MTTR and increase system uptime.
  • In 6 months: You have driven at least one major reliability initiative from inception to production, established SLO/SLA frameworks for critical services, and are a trusted technical voice shaping our infrastructure roadmap.

Areas of Ownership

  • Design, build, and operate scalable, fault-tolerant infrastructure across cloud environments (AWS, GCP, or Azure).
  • Define and enforce SLOs, SLAs, and error budgets; lead blameless postmortems and drive systemic improvements.
  • Build and maintain robust CI/CD pipelines, release automation, and deployment infrastructure.
  • Own observability: design and maintain logging, metrics, tracing, and alerting stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry).
  • Partner closely with software engineering teams to embed reliability practices into the development lifecycle.
  • Drive capacity planning, performance benchmarking, and cost optimization across our infrastructure.
  • Champion security best practices within the infrastructure and deployment layers.

Requirements

Do you have experience in Tooling?, * 5+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering roles.

  • Strong proficiency in at least one major cloud platform (AWS preferred); experience with Kubernetes and container orchestration at scale.
  • Hands-on experience with infrastructure-as-code tools (Terraform, Pulumi, or equivalent).
  • Proven track record designing and maintaining high-availability, distributed systems.
  • Deep expertise in observability tooling, incident management, and on-call practices.
  • Strong scripting and automation skills (Python, Go, Bash, or similar).
  • Excellent communication skills with the ability to collaborate across engineering teams and present technical findings to leadership.

What Makes You Stand Out

  • Experience supporting AI/ML workloads or GPU-accelerated infrastructure.
  • Prior experience in a high-growth startup environment where you wore multiple hats.
  • Familiarity with eBPF, service mesh technologies (Istio, Linkerd), or advanced networking.
  • Contributions to open-source SRE/DevOps tooling or communities.
  • Experience building global, multi-region infrastructure with strict latency and availability requirements.

About the company

You won’t be maintaining legacy systems or fighting fires in a sprawling monolith. At Blitzy, you’re building reliability into a greenfield AI platform that is redefining how the world creates software. You’ll have direct influence over architectural decisions, work side-by-side with world-class engineers, and see the tangible impact of your work as we scale to serve Fortune 500 customers. As a founding member of the Pune SRE team, you’ll help shape the culture and technical standards of a team that will grow with the company., Led by two pioneering co-founders we are one of the fastest growing companies in the U.S., creating our own category of enterprise autonomous software development. We automate thousands of hours of software development for our customers, which includes strong representation within the Fortune 500.

How we work:

We move Blitzy Fast: Time is both our company’s and our clients’ most precious asset. We move quickly and decisively to innovate internally and deliver exceptional software externally.

Championship Mindset: We operate like a professional sports team. We win as a team by holding ourselves and each other to high standards, collaborating in-person, and remaining focused on the mission.

Passion for Invention: We’re pushing the frontier of what’s possible, requiring constant innovation and iteration.

We Work for the Customer: We focus on delivering outsized value to the customers we work with and expanding those relationships into deep, meaningful partnerships.

We believe in being ‘everyday athletes’-taking care of ourselves so we can bring our best minds to work. We promote great sleep, movement, and restorative activities for optimal mental performance. It makes for a happier and more productive team.

Blitzy is an equal opportunity employer committed to building a diverse and inclusive team. We believe different perspectives make us stronger.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · WWC Europe 2026

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

1:33 min

Case study on adopting Kubernetes and Golang effectively

Andrew Holway · LIVE

7:15 min

Installing Istio programmatically with bash scripts

Thomas Südbröcker · LIVE

Videos

See all

Related articles

See all