Site Reliability Engineer

Digitive LLC
United States
8 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Systems Engineering Audit Trail Bash Shell Software as a Service Continuous Delivery DevOps Distributed Systems Federal Information Processing Standards (FIPS) Python (Programming Language) Reliability Engineering Ruby
+10 more
Runbook Software Vulnerability Management Scripting Core Api Backend Build Management Kubernetes Deployment Automation Terraform Golang

Job description

Looking for a Senior Site Reliability Engineer to help build and operate the platform capabilities that enable FedRAMP High and IL5 environments. This role is for an independent senior engineer who can own large features or bounded platform systems with minimal guidance. You will design and build centralized platform APIs, reusable CI components, continuous delivery capabilities, federal promotion workflows, and paved-road service onboarding patterns. You should be comfortable identifying the right technical approach when the problem is known but the solution is unclear, raising reliability and security standards, and acting as a trusted resource for engineers with less experience. What You’ll Do

  • Build and harden platform services required for FedRAMP High and IL5, including centralized platform APIs, reusable CI components, and continuous delivery capabilities.
  • Help move federal promotion workflows from manual operations toward automated, gated, auditable deployments.
  • Support controlled rollout paths for federal staging and production environments, including validation gates, rollback safety, and operational readiness.
  • Build paved-road platform patterns that help teams onboard services into federal environments safely and consistently.
  • Partner with other teams to unblock federal build-out and ensure platform capabilities are secure, reliable, operable, and easy to adopt.
  • Build infrastructure and automation that improves reliability, security, repeatability, and operator experience.

  • Participate in production support and incident response for platform services, including issues that primarily affect federal clusters.
  • Contribute documentation, runbooks, dashboards, and operational handoffs so federal systems can be supported sustainably.
  • Provide guidance and technical feedback to engineers with less experience.

Requirements

  • 8+ years of experience in SRE, infrastructure engineering, DevOps, platform engineering, or backend systems engineering.
  • Strong coding or scripting experience in Python, Go, Ruby, Bash, or similar languages.
  • Experience with AWS, Kubernetes, Terraform, CI/CD systems, deployment automation, or service orchestration.
  • Comfort operating production systems with observability, alerting, incident response, and post-incident follow-through.
  • Experience building secure systems with auditability, change control, and compliance requirements in mind.
  • Good judgment in distributed systems, deployment safety, reliability tradeoffs, and operational risk.
  • Strong written communication and a habit of leaving clear runbooks, design notes, and implementation plans behind.

Preferred

  • Experience with FedRAMP, DoD IL5, FIPS, GovCloud, vulnerability management, or regulated SaaS environments.
  • Experience with platform APIs, CI components, continuous delivery systems, artifact promotion, or progressive delivery.
  • Experience turning manual operational workflows into automated, supportable systems.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

50 sec

Why developer happiness matters in web frameworks

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

3:30 min

Falling in love with Ruby and creating Basecamp

David Heinemeier Hansson David Heinemeier Hansson +1 · Coffee With Developers

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

Videos

See all

Related articles

See all