Site Reliability Engineer

Circle (nyse: Crcl)
San Francisco, CA, United States
20 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$152,500.0 - $205,000.0
Working hours
Regular working hours
Job source

Tech stack

JavaScript (Programming Language) Artificial Intelligence Cloud Computing Code Review Continuous Integration DevOps Distributed Systems Domain Name System (DNS) Identity and Access Management Python (Programming Language) Routing Reliability Engineering
+10 more
Software Engineering TypeScript Load Balancing Cloud Platform System Computer Network Technologies Backend Containerization Kubernetes Deployment Automation Terraform

Job description

As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical digital-assets, AI, and application workloads. You will bring an engineering mindset to production operations: writing and maintaining services and automation, developing reliable Kubernetes platforms, and using Terraform to make infrastructure repeatable, auditable, and easy to evolve.

You’ll work closely with platform, product, and application engineering teams to translate workload requirements into resilient technical designs across hybrid and public-cloud environments. This role is for an experienced SRE or infrastructure engineer who enjoys solving hard distributed-systems problems, taking ownership of production outcomes, and raising the reliability, performance, security, and cost-effectiveness of the systems our customers depend on.

What you’ll work on:

  • Design, build, and operate Kubernetes platforms that provide secure, highly available, and scalable foundations for critical production services across hybrid and public-cloud environments.
  • Build infrastructure as code with Terraform, creating reusable modules, safe delivery workflows, and well-governed infrastructure changes.
  • Develop backend services, internal tools, and operational automation in Go, Python, or JavaScript/TypeScript to eliminate manual work and improve the developer experience.
  • Partner with engineering and product teams to understand workload requirements and design pragmatic solutions for reliability, performance, capacity, security, and cost.
  • Improve the production lifecycle through reliable CI/CD, deployment automation, progressive delivery, and clear operational ownership.
  • Define and evolve observability practices across metrics, logs, traces, alerting, and dashboards so teams can detect issues early and troubleshoot effectively.
  • Own production reliability by participating in on-call, leading incident response, performing root-cause analysis, and driving blameless postmortems and durable corrective actions.
  • Establish and maintain reliability targets through meaningful SLIs, SLOs, error budgets, capacity planning, disaster-recovery testing, and resilience improvements.
  • Embed security and compliance into platform operations, partnering with Security to protect infrastructure, workloads, and data while meeting applicable regulatory requirements.
  • Apply AI-assisted and data-driven operational techniques to improve signal detection, reduce alert noise, accelerate root-cause analysis, and surface opportunities for automation.
  • Raise the bar for the team through thoughtful code reviews, documentation, knowledge sharing, and mentorship.
  • Mentor and support team growth, fostering collaboration and scalability.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a closely related software engineering role supporting production systems.
  • Deep, hands-on Kubernetes expertise: designing, operating, securing, and troubleshooting production clusters and containerized workloads at scale.
  • Strong Terraform experience, including authoring reusable modules, managing state and environments, and delivering infrastructure changes through reviewable, automated workflows.
  • Production software-development experience in Go, Python, or JavaScript/TypeScript, with the ability to build maintainable backend services, tooling, and automation-not only scripts.
  • Demonstrated success improving the reliability, performance, scalability, or cost efficiency of distributed systems in production.
  • Experience with cloud infrastructure and core networking concepts, including IAM, DNS, load balancing, routing, service networking, and secure connectivity.
  • Strong observability and troubleshooting skills using metrics, logs, traces, alerting, and incident data to diagnose complex systems.
  • Experience defining and operating against SLIs, SLOs, error budgets, incident-management processes, postmortems, and disaster-recovery practices.
  • Familiarity with CI/CD, GitOps or deployment automation, and safe rollout strategies such as canary or blue-green deployments.
  • A security-minded approach to infrastructure and a track record of partnering effectively with Security and engineering teams in regulated or high-availability environments.
  • Clear written and verbal communication, strong ownership, and the judgment to balance speed, risk, and operational excellence.
  • Experience applying AI-assisted tooling to engineering or operations workflows is a plus.

Circle is on a mission to create an inclusive financial future, with transparency at our core. We consider a wide variety of elements when crafting our compensation ranges and total compensation packages.

Starting pay is determined by various factors, including but not limited to: relevant experience, skill set, qualifications, and other business and organizational needs. Please note that compensation ranges may differ for candidates in other locations.

About the company

Circle (NYSE: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open, global economy through digital assets, payment applications, and programmable blockchain infrastructure. Circle’s platform includes the world’s largest regulated stablecoin network anchored by USDC, Circle Payments Network for global money movement, and Arc, an enterprise-grade blockchain designed to become the Economic OS for the internet. Enterprises, financial institutions, and developers use Circle to power trusted, internet-scale financial innovation. Learn more at circle.com.

What you’ll be part of:

Circle is committed to visibility and stability in everything we do. As we grow as an organization, we’re expanding into some of the world’s strongest jurisdictions. Speed and efficiency are motivators for our success and our employees live by our company values: High Integrity, Future Forward, Multistakeholder, Mindful, and Driven by Excellence. We have built a flexible work environment where new ideas are encouraged and everyone is a stakeholder.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos · LIVE

3:24 min

Evaluating remote software roles and compensation structures

Nacho Iacovino · World Congress 2021

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

4:41 min

Scale and diversity of software development teams

Bastian Heilemann Bastian Heilemann +1 · World Congress 2025

Videos

See all

Related articles

See all