> Markdown version of [/jobs/ext/463121-site-reliability-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/463121-site-reliability-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability / Infrastructure Engineer - **Company:** Medal L.P. - **Location:** New York, NY, United States - **Salary:** $150,000.0 - $275,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), BigQuery, C Sharp (Programming Language), C++ (Programming Language), Cloud Computing, Databases, Continuous Integration, Relational Databases, Shard (Database Architecture), Elasticsearch, Github, Identity and Access Management, Key Management, PostgreSQL, MySQL, Network Segmentation, Performance Tuning, Query Optimization, RabbitMQ, Redis, Reliability Engineering, Management of Software Versions, Data Logging, Google Cloud, Delivery Pipeline, Amazon Virtual Private Cloud (VPC), Backend, Kotlin, Kubernetes, Terraform, Virtual Private Clouds, Data Pipelines - **Published:** June 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=7dd6b80058693cb4 ## About the Role Do you have experience in Virtual Private Clouds?, * You've worked at startups and are comfortable in an environment of rapid growth where scaling up is a priority * You have great judgment - you know the difference between a durable, sustainable fix vs. a patch that buys you a week * You have deep, hands-on experience scaling and sharding relational databases in production environments * You know GCP maybe a little too well: Kubernetes, VPC, IAM, Cloud Logging, and the managed services ecosystem * You are fluent in Terraform and have owned real infrastructure-as-code at scale * You've operated Elasticsearch in production and know how to keep a cluster healthy * You have strong incident response instincts: you can work a P0 calmly, communicate clearly under pressure, and run a postmortem that prevents recurrence. * You've worked with GitHub Actions in a production CI/CD environment. * You have excellent communication skills (this is crucial!) and can both flag issues clearly and rapidly during incidents, and lead / write actionable postmortems Our Stack Google Cloud Platform Terraform, Salt, GitHub Actions Java, Redis, RabbitMQ, ElasticSearch, BigQuery, Kubernetes for backend Electron+React C# and C++ for native windows recording & more Swift for iOS, Kotlin for Android ## Description Medal's infrastructure handles billions of clips, video ingestion pipelines, and social features at a massive scale most engineers never get to touch. We're looking for an SRE who cares deeply about reliability and scalability. The work centers on reliability, incident response, scaling, and making sure our infrastructure keeps up with our growth. You'll own the on-call rotation, drive postmortems, and work directly with engineering teams to meet their infra needs. The right person probably came through startups and scale-ups. You've been in the room when things broke at 2am, you've scaled databases under pressure, and you know the difference between a durable fix and a patch that buys you a week., * Own reliability across our GCP infrastructure: Kubernetes clusters, managed services, and data pipelines, driving measurable improvements to availability and latency * Lead incident response end-to-end: on-call rotations, runbooks, postmortems, and the follow-through that makes sure the same thing doesn't happen twice * Architect and execute database scaling strategies (sharding, replication, query optimization, and capacity planning) across MySQL and Postgres at meaningful scale * Partner with product engineering to translate feature requirements into infrastructure designs that hold up as we grow * Manage and evolve our Terraform-managed GCP environment and Kubernetes cluster configurations * Own our Elasticsearch cluster end-to-end: capacity planning, sharding strategy, index lifecycle management, version upgrades, and performance tuning at production scale * Build and maintain observability across the stack: metrics, dashboards, alerting, and tracing * Constantly improve CI/CD reliability and delivery pipelines across GitHub Actions * Harden IAM, secrets management, and network segmentation as part of normal infra hygiene ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Kotlin Multiplatform - True power of native code reuse](https://www.wearedevelopers.com/videos/4-kotlin-multiplatform-true-power-of-native-code-reuse) - [Navigating the Corporate Jungle: Life as a Developer in a large Company](https://www.wearedevelopers.com/videos/621-navigating-the-corporate-jungle-life-as-a-developer-in-a-large-company) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)