Site Reliability Engineer

Kalshi Inc.
New York, NY, United States
18 days ago
Apply on job-boards.greenhouse.io
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$100,000.0 - $250,000.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Build Automation Microsoft Azure Software Debugging Performance Tuning Service-Oriented Architecture Software Engineering Datadog Kubernetes Low Latency Terraform
+1 more
Docker

Job description

  • Improve observability, reliability and availability by defining and measuring key metrics.
  • Build automation and improve systems to eliminate toil and operations work.
  • Collaborate with our core infrastructure team to performance tune and optimize our cloud deployments. (Think Docker, Terraform, Kubernetes, EC2, etc.)
  • Collaborate with product teams to reduce service disruptions and automate incident response.
  • Proactively find and analyze reliability problems across our business units and stack, then design and implement software to create step-function improvements.
  • Educate, mentor and hold accountable the engineering team to improve the reliability of our systems and make reliability a core value of the Kalshi engineering culture.
  • Write high quality, well tested code to meet the needs of your customers.
  • Debugging extremely difficult technical problems, and making systems and products both work better and are easier to deploy, own, operate and diagnose.
  • Review all feature designs within your product area and across the company for cross-cutting projects.
  • Be an owner of the security, safety, scale, operational integrity, and architectural clarity of these designs.
  • Build integrations with 3rd party vendors.
  • Participate in an on-call support rotation to provide timely troubleshooting and resolution of urgent issues.

Requirements

  • You have at least 4 years of experience in software engineering.
  • You’ve designed, built, scaled and maintained production services, and know how to compose a service oriented architecture.
  • You write high quality, well tested code to meet the needs of your customers.
  • You’re passionate about building an open financial system that brings the world together.
  • You possess strong technical skills for system design and coding.
  • Excellent written and verbal communication skills, and a bias toward open, transparent cultural practices.
  • Strong skills around observability, debugging and performance tuning.
  • Strong interpersonal skills working with engineers from junior to principal levels
  • Demonstrated critical thinking under pressure.
  • A willingness to dive into understanding, debugging, and improving any layer of the stack.
  • On-call availability to ensure swift resolution of issues., * Experience designing and building reliable systems capable of handling high throughput and low latency.
  • Experience with Datadog.
  • Experience with Rust, Go and Terraform.
  • Experience with AWS, GCP, or Azure.
  • Experience working in a highly regulated environment.
  • Experience writing company-facing blog posts and training materials.

Benefits & conditions

Salary Range: $100,000 to $250,000 annually plus equity and benefits.

About the company

We are building a next-generation financial ecosystem (think NYSE or CME from scratch). We are a small team, which means your responsibilities scale very rapidly, and your contributions are clear and visible, not marginal. There is still a lot of green field at Kalshi and a lot of it (including entire systems) can be yours.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on job-boards.greenhouse.io
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all