> Markdown version of [/jobs/ext/2124814-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2124814-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Kalshi Inc. - **Location:** New York, NY, United States - **Experience:** Experienced - **Salary:** $100,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon Elastic Compute Cloud, Build Automation, Microsoft Azure, Software Debugging, Performance Tuning, Service-Oriented Architecture, Software Engineering, Datadog, Kubernetes, Low Latency, Terraform, Docker - **Published:** August 19, 2026 - **Apply:** https://job-boards.greenhouse.io/kalshi/jobs/6567367003 ## About the Role * You have at least 4 years of experience in software engineering. * You've designed, built, scaled and maintained production services, and know how to compose a service oriented architecture. * You write high quality, well tested code to meet the needs of your customers. * You're passionate about building an open financial system that brings the world together. * You possess strong technical skills for system design and coding. * Excellent written and verbal communication skills, and a bias toward open, transparent cultural practices. * Strong skills around observability, debugging and performance tuning. * Strong interpersonal skills working with engineers from junior to principal levels * Demonstrated critical thinking under pressure. * A willingness to dive into understanding, debugging, and improving any layer of the stack. * On-call availability to ensure swift resolution of issues., * Experience designing and building reliable systems capable of handling high throughput and low latency. * Experience with Datadog. * Experience with Rust, Go and Terraform. * Experience with AWS, GCP, or Azure. * Experience working in a highly regulated environment. * Experience writing company-facing blog posts and training materials. ## Description * Improve observability, reliability and availability by defining and measuring key metrics. * Build automation and improve systems to eliminate toil and operations work. * Collaborate with our core infrastructure team to performance tune and optimize our cloud deployments. (Think Docker, Terraform, Kubernetes, EC2, etc.) * Collaborate with product teams to reduce service disruptions and automate incident response. * Proactively find and analyze reliability problems across our business units and stack, then design and implement software to create step-function improvements. * Educate, mentor and hold accountable the engineering team to improve the reliability of our systems and make reliability a core value of the Kalshi engineering culture. * Write high quality, well tested code to meet the needs of your customers. * Debugging extremely difficult technical problems, and making systems and products both work better and are easier to deploy, own, operate and diagnose. * Review all feature designs within your product area and across the company for cross-cutting projects. * Be an owner of the security, safety, scale, operational integrity, and architectural clarity of these designs. * Build integrations with 3rd party vendors. * Participate in an on-call support rotation to provide timely troubleshooting and resolution of urgent issues. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)