Site Reliability Engineer

Onebrief, Inc.
Colorado Springs, CO, United States
1 day ago
Apply on www.builtincolorado.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$180,000.0 - $220,000.0
Working hours
Regular working hours

Tech stack

Proxmox Artificial Intelligence Amazon Web Services Bash Shell Continuous Integration DevOps Distributed Systems Github Hyper-V Python (Programming Language) Networking Basics Reliability Engineering
+17 more
Ansible Prometheus Toolchain Datadog Data Logging Scripting Istio Grafana Cloudformation Gitlab-ci Kubernetes Linkerd (Service Mesh) Nutanix Terraform Elk Stack Jenkins Vmware

Job description

We are hiring a Site Reliability Engineer to join our Infrastructure & Security team. You’ll work closely with fellow SREs, security, and customer success.

You will be the first line of support for our mission critical deployments, and responsible for ensuring best-in-class service quality and issue resolution. You will work in both on-premise DoD environments and AWS cloud environments. Your lessons from the field will shape how our team works, from policy to implementation.

In addition to working at the customer, you will contribute directly to solutions that increase stability, performance, and security of our deployments, and improve the overall experience of deploying and managing Onebrief on premise.

About You

You care deeply about reliability and treat it as a core feature of any application or platform, with a bias toward “reliability over novelty.” You think about infrastructure and operability as products to be automated, well-documented, and continuously improved, and you aim to leave systems easier to operate than you found them.

You are equally comfortable leading a post-incident review, or diving into a kubectl shell to triage a complex production issue. You don’t just fix problems; you translate constraints and failure modes into clear, automated guardrails and scalable, resilient architecture. For you, robust monitoring, actionable alerting, and insightful runbooks are core parts of the engineering process, not afterthoughts.

You mentor others, fostering a culture of blameless postmortems and proactive reliability. You collaborate naturally with application and platform teams, helping them move quickly but safely by building the tools, processes, and observability that make “fast recovery” a reality.

What You’ll Do

You’ll own the reliability, scalability, and security of the production application and/or platform. You will do this by:

  • Implementing a World-Class Observability Platform: Design, implement, and manage our monitoring, logging, and alerting stack (e.g., Prometheus, Loki, Alloy, and Grafana). You won’t just track metrics; you’ll create the actionable insights and automated alerting that allow teams to identify and resolve issues before they impact users.
  • Defining and Upholding Reliability: Define, measure, and own alerting that feeds into our Service Level Indicators (SLIs) and Service Level Objectives (SLOs), increasing trust internally and externally. You will be the organization’s expert on what it means for our systems to be reliable and how to measure it.
  • Leading Incident Response: Act as the incident responder and potentially incident commander during critical incidents who will lead blameless post-mortems / After Action Reviews (AARs) that identify true root causes and drive automated, long-term solutions to prevent recurrence.
  • Automating for Scale and Security: Partner with platform engineers to design, build, and manage secure, resilient Kubernetes clusters and cloud/on-prem environments using Infrastructure-as-Code (Terraform, Ansible). You will embed security and compliance controls (RMF, STIGs) directly into this automation.
  • Eliminating Toil and Scaling the Team: Proactively identify and eliminate operational toil by building automation. You will partner with other teams to share best practices for air-gapped environments and support their readiness for production.

Requirements

  • An active Top Secret clearance
  • 5+ years in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus.
  • Proven partner to DevOps/Platform and application teams; collaborates well across functions and shares context openly.
  • A deep understanding of incident response processes, with experience conducting thorough root cause analyses and driving continuous improvement.

Technical expertise

  • Infrastructure as Code: Terraform (or CloudFormation), Ansible.
  • Containers and orchestration: Kubernetes design, deployment, and operations.
  • CI/CD: experience building and maintaining pipelines (GitLab CI/CD, Jenkins, GitHub Actions).
  • Scripting: proficiency with at least one of Python, Go, or Bash.
  • Cloud: Familiarity with AWS or AWS GovCloud.
  • Observability: Grafana stack, ELK stack, or Datadog.
  • Networking fundamentals: core protocols and secure configurations.

Bonus points (nice to have)

  • Experience in DoD environments and compliance frameworks (RMF, STIGs, ICD 503).
  • GitOps practices and toolchains.
  • Security-minded design for sensitive environments.
  • Experience designing and implementing meaningful SLIs/SLOs (including error budgets) for complex, distributed systems.
  • Familiarity with on-prem virtualization(VMware, Proxmox, Nutanix, Hyper-V, etc).
  • Service mesh exposure (Istio, Linkerd).
  • Relevant certifications (e.g., AWS DevOps Engineer, CKA/CKAD).
  • Active Security+ or another DoD 8570.01-approved security credential, or the ability to obtain the valid credentials within 3 months of employment.

Benefits & conditions

Own reliability, scalability, security, and observability for production applications across on-premises DoD and AWS environments. Build monitoring and alerting, define SLIs and SLOs, lead incident response and blameless postmortems, and automate resilient Kubernetes infrastructure using Terraform and Ansible. Partner with platform, application, security, and customer success teams to eliminate operational toil, strengthen compliance controls, and improve deployment readiness in air-gapped environments. The summary above was generated by AI Consequential Work. Dedicated People. About Onebrief

Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination., In-Office Colorado Springs, CO, USA 150K-180K Annually Senior level 150K-180K Annually Senior level Software * Defense Own customer relationships at U.S. Space Command, expand Onebrief usage across workflows, support exercises, provide face-to-face and remote customer support, coordinate incidents with engineering, and translate user needs to product teams while navigating large government bureaucracy. Top Skills: Classified NetworksJwicsOnebrief

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

About the company

Military planning is complex by nature, requiring teams to coordinate information, people, and decisions across systems and locations. Onebrief brings planning, collaboration, simulation, and AI into one connected environment, helping teams test strategies, adapt to changing conditions, and make decisions with greater clarity when the stakes are real.

We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world.

Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.builtincolorado.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · World Congress 2024

Videos

See all

Related articles

See all