Senior DevOps Engineer

Mission Pet Health
United States
3 months ago
Apply on indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$180,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Amazon S3 Application Services Automation of Tests Bash Shell Software as a Service Cloud Computing Static Program Analysis Computer Networks Linux DevOps
+25 more
Github Identity and Access Management Python (Programming Language) Key Management MongoDB OpenID Redis Reliability Engineering Prometheus Software Deployment Software Engineering Datadog Scripting Load Balancing Autoscaling Grafana Gitlab-ci Kubernetes Apache Kafka Terraform New Relic (SaaS) AWS EKS Docker Static Application Security Testing Dynamic Application Security Testing

Job description

We’re looking for a Senior DevOps Engineer to own our cloud infrastructure end-to-end - from operating a large multi-tenant Kubernetes environment to building CI/CD pipelines that teams actually trust. You’ll work across AWS, drive infrastructure-as-code standards, and lead our migration toward GitLab CI and a Grafana-based observability stack while keeping production environments stable.

What You’ll Do

  • Operate and scale a multi-tenant AWS EKS cluster where each client runs an isolated set of application services - owning tooling to onboard, scale, and observe hundreds of service instances reliably
  • Build and improve CI/CD pipelines in GitLab CI and GitHub Actions with automated testing, static analysis, and build-gated releases; maintain ArgoCD GitOps workflows for production deployments
  • Lead the migration from Datadog to a self-managed Grafana observability stack (Grafana, Loki, Mimir/Prometheus, Tempo) - dashboards, SLOs, alert routing, and on-call integration
  • Manage secrets, IAM, and security scanning pipelines using AWS KMS, Secrets Manager, external-secrets operator, and Auth0/Dex OIDC - enforcing least-privilege across all environments
  • Own and evolve the Redpanda (Kafka-compatible) streaming layer and its integrations with application workers
  • Drive cloud cost optimization through right-sizing, autoscaling, and shared infrastructure patterns on EKS
  • Document infrastructure with automated tooling (terraform-docs) and maintain standards that scale across teams
  • Automate operational toil - certificate renewal, clinic environment provisioning, deployment validation, runbook automation

Requirements

Do you have experience in WAF?, Required

  • 5+ years in DevOps or infrastructure engineering
  • 3+ years operating Kubernetes in production - AWS EKS preferred - including CSI drivers, cluster autoscaling, network policy (Calico), and pod identity
  • 3+ years hands-on with AWS core services (IAM, S3, KMS, Secrets Manager, STS, EKS, Load Balancer Controller, ECR)
  • Strong Terraform experience; GitOps experience with ArgoCD or Flux
  • Hands-on experience with GitLab CI and/or GitHub Actions
  • Scripting proficiency in Python and Bash
  • Experience with IAM design and security best practices (SAST/DAST, secret scanning, OIDC federation)
  • Familiarity with streaming or message-queue infrastructure (Redpanda, Kafka, or equivalent)

Nice to Have

  • Experience migrating from a SaaS observability tool (Datadog, New Relic) to a self-hosted Grafana stack
  • Grafana stack depth - Loki for logs, Mimir or Thanos for metrics, Tempo for traces, Alertmanager for routing
  • Experience with Redpanda specifically, or deep Kafka operations knowledge
  • Background in multi-tenant SaaS platforms or per-customer service isolation patterns
  • AWS certification
  • Familiarity with chaos engineering tooling (chaos-mesh or LitmusChaos)
  • Background in software engineering or scripting-heavy roles

Tech Stack

Current production: AWS (EKS, S3, KMS, Secrets Manager, STS, Load Balancer) · Terraform · GitHub Actions · ArgoCD · Kubernetes · Traefik · Coraza WAF · Redis HA · MongoDB · Auth0 · Dex · external-secrets · Datadog · Docker · Python · Bash · Linux

Where we’re going: GitLab CI · Redpanda · Grafana · Loki · Prometheus/Mimir · Tempo · Alertmanager

Platform components you’ll operate: ArgoCD · Traefik · Coraza WAF · Auth0 · Dex · Redis HA · MongoDB · API servers · client-facing portals · internal tooling

Benefits & conditions

  • Own infrastructure across a real multi-tenant platform serving production clinic environments
  • Lead the observability and streaming migrations - greenfield decisions with lasting impact
  • Collaborative engineering culture with high trust and low bureaucracy
  • Competitive salary, benefits, and flexible work arrangements

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · World Congress 2022

Videos

See all

Related articles

See all