Site Reliability Engineer

Kong Inc.
United States
8 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Software as a Service Continuous Integration Data as a Services Disaster Recovery PostgreSQL Network Control Platform as a Service (PAAS) Redis Prometheus Datadog
+10 more
Istio Delivery Pipeline Grafana Backend Build Management Kubernetes Information Technology Deployment Automation Terraform Serverless Computing

Job description

Kong is building Project Volcano, an internal developer platform purpose-built for Kong’s engineering ecosystem. Volcano will provide teams with on-demand preview environments, edge deployments, managed PostgreSQL, auth, realtime, and storage APIs all deeply integrated with Kong products.

As the Senior SRE for Volcano, you will be the reliability voice for this platform. This role is a strategic initiative driven by the Office of the CTO (OCTO). You will partner directly with engineering leadership to define the platform’s reliability posture. This is a high-visibility, high-impact role with direct influence on Kong’s next generation developer platform.

What You’ll Do:

  • Own reliability for Volcano end-to-end: Define and drive SLOs, error budgets, and incident response practices for all Volcano services - edge deployments, managed Postgres, auth, realtime, storage, and the control plane.
  • Contribute to the platform’s infrastructure: Design and build the multi-region Kubernetes infrastructure, networking, and data plane that powers Volcano’s edge deployment pipeline and backend-as-a-service capabilities.
  • Build the GitOps and CI/CD backbone: Establish deployment automation, canary pipelines, and preview environment provisioning using ArgoCD, Helm, and Terraform/Terragrunt - setting patterns the broader team will follow.
  • Scale managed data services: Design, operate, and harden multi-tenant PostgreSQL clusters, Redis caching layers, and object storage - with a focus on data isolation, performance, and disaster recovery.
  • Drive observability from day one: Instrument every Volcano service with meaningful SLIs; build dashboards, alerts, and runbooks using Datadog, Prometheus, and Grafana before services go live, not after incidents.
  • Lead cross-functional reliability work: Collaborate with the OCTO team, product engineering, and security to bake reliability and compliance into Volcano’s architecture - not bolt it on later.
  • Evaluate and adopt emerging technologies: Given Volcano’s greenfield nature, evaluate and make architectural decisions on edge runtimes, serverless compute, vector databases, and AI-native infrastructure components.

Requirements

If you don’t think you meet all of the criteria below but are still interested in the job, please apply. Nobody checks every box - we’re looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others., * BS in Computer Science or equivalent; substantial experience at Staff or Principal IC level in SRE/Platform Engineering.

  • Proven track record building SRE or platform engineering practices for developer-facing platforms or PaaS/SaaS products - ideally at greenfield stage.
  • Kubernetes expertise: multi-tenant cluster design, networking (CNI, service mesh, ingress), autoscaling, and security hardening.

About the company

Kong Inc., the AI Connectivity Company, is building the connectivity layer of AI. Trusted by the Fortune 500® and AI-native startups alike, Kong’s unified API and AI platform enables organizations to secure, manage, accelerate, govern, and monetize the flow of intelligence across APIs and AI traffic - on any model, any cloud. For more information, visit www.konghq.com.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

2:32 min

Refactoring bulk frontend operations into scalable backend methods

Noam Honig · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · World Congress 2022

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

Videos

See all

Related articles

See all