Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+20 more
Job description
Weâre hiring a Site Reliability Engineer to join our Infrastructure & Security team. Youâll work closely with product engineers, fellow SREs, security, and customer success.
This is an SRE role for someone whoâs comfortable in application code. Much of the reliability and performance work happens in the codebase (primarily TypeScript), so youâll fix problems at the source rather than working around them in the infrastructure. Youâll be a first line of support for our mission-critical deployments across on-prem DoD and AWS environments, and what you learn in the field will feed directly back into the product.
Youâll ship code that makes Onebrief more stable, faster, and easier to deploy and operate. The work sits at the seam between engineering and operations, and itâs weighted toward engineering., Youâll help make our production application reliable, scalable, and secure by improving the software itself, not just the systems it runs on. Day to day that looks like:
- Improving the application: Work directly in the codebase (primarily TypeScript) to fix reliability and performance problems at the source. Youâll partner with product engineers on design decisions, review code with reliability and security in mind, and treat âmake the app betterâ as a first-class part of the job rather than something you hand off.
- Building observability that developers actually use: Design and run our monitoring, logging, and alerting (Prometheus, Loki, Alloy, Grafana). The goal is alerts and dashboards tied to real application behavior, so teams catch issues before users do.
- Owning reliability targets: Define and measure SLIs and SLOs, wire up alerting that feeds them, and be the person who can say what âreliableâ means for our systems and prove it with data.
- Leading incident response: Act as incident responder, and incident commander when needed. Run blameless post-mortems (AARs) that find the actual root cause and turn it into a code or process fix so it doesnât happen again.
- Automating away toil: Spot the repetitive operational work and write software to kill it. Share what works with other teams, including those running in air-gapped environments, and help them get production-ready.
Requirements
You treat reliability as a feature, not an afterthought, and youâd rather fix a problem in the code than route around it. You understand the full software development lifecycle (design, review, testing, release) and you know where reliability fits into each step.
Youâre comfortable reading and writing application code, and youâre just as happy dropping into a kubectl shell to triage a production issue. You turn failure modes into guardrails, and you think monitoring, alerting, and clear runbooks are part of building software, not extra credit.
You mentor others and push a culture of blameless postmortems. You work naturally with product and platform teams, helping them move fast without breaking things by giving them the tools, tests, and observability that make quick recovery real., * An active Secret clearance
- 5+ years in software engineering, SRE, or a related role, with real time spent writing and shipping application code
- Strong TypeScript (or comparable modern language experience with willingness to work primarily in TypeScript)
- Solid grasp of the full SDLC: design, code review, testing, release, and how reliability fits into each stage
- Experience with incident response, root cause analysis, and turning findings into lasting fixes
- A collaborator who works well across product, platform, and DevOps teams and shares context openly
Technical expertise
- Application development in TypeScript (Node and/or a modern front-end framework)
- CI/CD: building and maintaining pipelines (GitHub Actions, GitLab CI/CD, Jenkins)
- Testing and quality practices as part of the delivery process
- Comfort with at least one of Python, Go, or Bash for tooling and automation
- Working knowledge of containers and Kubernetes (enough to debug and deploy, not necessarily to stand up clusters from scratch)
- Networking fundamentals and secure configuration basics
Bonus points (nice to have)
- Observability: Grafana stack, ELK, or Datadog
- Infrastructure as Code (Terraform, Ansible) and cloud experience (AWS or AWS GovCloud)
- Kubernetes cluster design and operations
- Designing meaningful SLIs/SLOs with error budgets for distributed systems
- GitOps practices and toolchains
- DoD environments and compliance frameworks (RMF, STIGs, ICD 503)
- Service mesh (Istio, Linkerd)
- On-prem virtualization (VMware, Proxmox, Nutanix, Hyper-V)
- Relevant certs (AWS DevOps Engineer, CKA/CKAD)
Benefits & conditions
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on-prem DoD and AWS environments. The summary above was generated by AI Consequential Work. Dedicated People. About Onebrief
About the company
Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination.
Today, many critical planning workflows still rely on fragmented systems, static documents, and disconnected tools that make collaboration and decision-making unnecessarily difficult. Onebrief brings modern software, AI, and real-time collaboration into those environments, helping teams operate with greater clarity, coordination, and adaptability in situations where decisions carry real-world consequences.
We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world.
Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on jobs.ashbyhq.comGood distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Dev Digest 120 - Apple and peers
Why Upskilling And Reskilling is Important For Developers
Dev Digest 121 - AI goes offline