Senior Site Reliability Engineer (REMOTE)

Discogs
United States
about 2 months ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$140,000.0
Working hours
Shift work
Job source

Tech stack

Amazon Web Services Amazon S3 Cloud Computing Cloud Engineering Databases Continuous Integration Relational Databases Software Design Patterns DevOps Elasticsearch Github Java Database Connectivity
+26 more
Python (Programming Language) Memcached MySQL Redis Reliability Engineering Software Engineering SQLAlchemy Datadog Scripting Infrastructure as Code (IaC) Amazon Virtual Private Cloud (VPC) Fastapi Amazon Relational Database Service Debezium Kubernetes Information Technology Apache Flink Sentry Hashicorp Apache Kafka Graphql Data Management Virtual Agents Cloud Optimization Restful APIs Terraform

Job description

The Discogs Platform team is focused on several objectives: building and supporting performant, cost-effective, reliable cloud infrastructure; data administration of complex CDC workflows; and developer experience tooling and mentorship, including our agentic AI discipline. As a Platform member, the Senior Site Reliability Engineer will contribute to the Platform team’s centralized infrastructure, including maintenance, monitoring, and automation of services ranging from databases to Kubernetes; lead incident response and postmortem efforts; and work closely with other engineering teams to understand their needs and drive improvements to both our technologies and processes. Location While we are a remote company we are only hiring for the following locations: OR, WA, CA, CO, TX, IL, What You’ll Accomplish Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.

  • Owning tasks and larger projects from planning to production rollout
  • Learning new technologies and building expertise with the goal of teaching and mentoring others; mentoring with the goal of force-multiplying through docs and tools
  • Maintaining organization cloud presence in AWS
  • Automating and deploying infrastructure configurations using Infrastructure as Code (IAC)
  • Mentoring engineering squads on Platform best practices for Kubernetes, MySQL, Kafka, and other software development lifecycle areas
  • Assisting engineering squads with capacity planning, on-call preparation, and production readiness
  • Writing documentation and runbooks that contribute to the engineering organization’s knowledge base
  • Implementing monitoring and alerting systems with Discogs observability tools
  • Working in a containerized, orchestrated environment
  • Participating in on-call rotation, responding to incidents, and troubleshooting data and other operations issues
  • Contributing to the reliability and design patterns of our Kafka CDC and event workflows
  • Contributing to agentic AI best practices and tooling, including skills, agents, and safety

Requirements

Required (or equivalent tools; listed is our current stack):

  • Infrastructure-as-code (Terraform)
  • CI/CD (GitHub Actions)
  • Kubernetes (EKS, Kustomize, Karpenter, administration, application manifests)
  • AWS and cloud development (VPC, EKS, RDS, S3)
  • FinOps and cloud cost optimization
  • Observability (Datadog, Sentry)
  • Agentic AI (Claude Code)
  • Scripting (Shell, Python)
  • Track record of collaboration and mentorship
  • Excellent written communication and documentation skills
  • Continuous learning
  • Ownership and proactive approach to solving large problems

Preferred:

  • Kafka: Cluster administration (Strimzi), Kafka Connect (Debezium, JDBC)
  • Flink
  • Relational database administration and performance (MySQL, Percona Server, AWS RDS)
  • Elasticsearch (ECK administration, scaling, performance)
  • Python (SQLAlchemy, FastAPI)
  • GraphQL (schema design, Apollo federation)
  • REST API
  • GitOps (ArgoCD)
  • Hashicorp Vault
  • Redis
  • Memcached

Education & Experience:

  • A Bachelor’s Degree in Computer Science or similar area of focus, or equivalent relevant work experience.
  • 5+ years experience in Ops, DevOps, Site Reliability, Platform or other systems roles.

Benefits & conditions

5.05.0 out of 5 stars United States Remote $140,000 a year - Full-time, Pulled from the full job description

  • Paid parental leave
  • Parental leave
  • Health insurance
  • 401(k) matching
  • Dental insurance
  • Adoption assistance
  • Work from home, Fixed Position Rate: $140,000 This role carries a single, non-negotiable Fixed Position Rate to ensure absolute equity and eliminate negotiation bias., What We Provide

  • Competitive compensation: salary, plus performance-related bonus program
  • 401(k) with employer match
  • 100% company-paid medical and dental insurance benefits for you and your dependents
  • 4 weeks paid vacation, increasing based on tenure
  • 18 weeks paid leave for birth moms
  • 8 weeks paid parental leave, including for adoption
  • Monthly wellness allowance
  • Annual professional and personal development allowance
  • Work from home office set-up and expense allowances
  • Flexible work location opportunities
  • Employer matching toward charitable contributions

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

1:22 min

Overview of the Sentry error and performance monitoring platform

Priscila Oliveira · World Congress 2023

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · World Congress 2022

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all