> Markdown version of [/jobs/ext/2727671-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2727671-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Braze - **Location:** Saint Paul, MN, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Computer Programming, Software Debugging, Linux, DevOps, Distributed Systems, PostgreSQL, Linux Kernel, Unix Shell, MongoDB, Ruby on Rails, Redis, Reliability Engineering, Ruby, Delivery Pipeline, Kubernetes, Apache Kafka, Terraform, Docker, Pagerduty - **Published:** September 5, 2026 - **Apply:** https://job-boards.greenhouse.io/braze/jobs/8114074 ## About the Role * 3+ years of experience as a Software, DevOps, or Site Reliability Engineer * You think about systems - interfaces, boundaries, edge cases, failure modes, behaviors, specific implementations * Have an urge to collaborate, document, and deliver quickly * Collaborating across the global remote teams, often working asynchronously * Document everything so you don't need to learn the same thing (or plan the same work) twice * Delivering fast to delight our customers - even internal ones Have an enthusiastic, go-for-it attitude. When you see something broken, you can't help but fix it Have a desire to solve everyday challenges facing software engineers and automate their toil away Have an excellent ability to manage multiple tasks and expectations at once Know your way around Linux and Unix Shell Have strong programming skills - Ruby and/or Go preferred Have experience with Docker, Kubernetes, Terraform, or similar IaC technologies Have experience with MongoDB, Redis, Kafka, Postgres, or similar data technologies #LI-Hybrid ## Description Site Reliability Engineers (SREs) are responsible for keeping all internal-facing services and platforms running smoothly. In a nutshell, SREs ensure site uptime. SREs blend sensible system administrators and software engineers who apply sound engineering principles, operational discipline, and mature automation to the environments and infrastructure services we provide. We specialize in systems-whether it be networking, the Linux kernel, or some more specific interest in scaling-algorithms or distributed systems. Our team helps to improve automation, infrastructure reliability, and empowers Braze's other engineering teams to leverage the infrastructure products and platforms we create easily. Braze operates at a massive scale with over 3.3 billion monthly active users across our customers, collecting hundreds of billions of data points each month, and sending billions of messages to end-users daily. We use a diverse technology stack rooted in Ruby on Rails, MongoDB, Redis, Kafka, Kubernetes, and more. As a Site Reliability Engineer at Braze, you will collaborate with your team and consumer engineering teams to continuously improve the infrastructure, automation, and tooling that build internal products from these technologies. Main responsibilities: * Partner with Braze's engineering teams on: * Architecting products to effectively utilize infrastructure platforms in a scalable, reliable manner * Debugging reliability and scalability issues across all stack layers, including the products built using our infrastructure platforms * Make monitoring and alerting alerts on symptoms and not on outages * Ensure that Braze meets our strict enterprise-grade SLAs with customers Develop Braze's internal platform infrastructure: * Create Infrastructure as code using Chef, Terraform, and Kubernetes * Develop deployment pipelines for applications in multiple languages using Docker, Kubernetes, etc * Provide centralized/common tooling, services, and automation frameworks that are critical for scaling operations, capacity management, reducing operational pain, and improving the day-to-day workflow of Braze's engineering teams Manage incidents: * Be on a PagerDuty rotation to respond to availability incidents and provide support for other engineers * Use your on-call shift to prevent incidents from ever happening * Retrospect everything that happens to turn lessons into system improvements/changes, automation, etc ## Related Videos - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Navigating the Corporate Jungle: Life as a Developer in a large Company](https://www.wearedevelopers.com/videos/621-navigating-the-corporate-jungle-life-as-a-developer-in-a-large-company) - [Coroutine explained yet again 60 years later](https://www.wearedevelopers.com/videos/690-coroutine-explained-yet-again-60-years-later) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)