Acquire-Site Reliability Engineer in Denver

Energy Jobline
Denver, CO, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$90,000.0 - $105,000.0
Working hours
Regular working hours

Tech stack

JavaScript (Programming Language) Artificial Intelligence Amazon Web Services Amazon Cloudfront Apple IOS Asana Audit Trail Command-Line Interface Software as a Service Continuous Integration Customer Data Management Cursor (Graphical User Interface Elements)
+26 more
DevOps Disaster Recovery Github Google Talk Identity and Access Management MongoDB Node.Js Operational Databases Role-Based Access Control Reliability Engineering Software Tools Amazon Simple Notification Service (SNS) TypeScript Datadog Data Logging Postman ReactJS Grafana Backend AWS ECS Integration Tests Playwright Sentry Google Play Cloudwatch Terraform

Job description

Acquire is hiring its first dedicated Site Reliability Engineer, a mid-level role with a clear path to Lead SRE as we grow. You will own production health, a trustworthy release pipeline, and the reliability surface of the codebase, while raising release-quality risk and acting as the customer-facing escalation point. You report to the Lead Engineer and work regularly with the CEO and CTO. You will not inherit a mature SRE team or a thick runbook library; you will help build them.

Just as important as the technical background: we want a product-focused thinker. AI tooling makes raw implementation cheaper, so the scarce skill is judgment about why the product matters to clinicians and what actually needs building. Treat reliability and ops as ways to keep a great product healthy, and step into feature work when the team needs it.

This is a broad role today by design. As the team grows it narrows toward Lead SRE: reliability strategy, incident response, and how ops, observability, and release engineering work at Acquire.

What You’ll Do

  • Production reliability: own day-to-day health across AWS and MongoDB Atlas. Triage and respond to alerts (CloudWatch, Sentry, Google Chat ops-alerts), run root-cause and incident comms, and turn retros into runbooks and alerting improvements. Close HIPAA-aware observability gaps: PHI-safe logging, auditability, access controls, incident evidence.
  • DevOps and release engineering: operate and improve our GitHub Actions deploy pipelines (backend, webapp, ), maintain Terraform infrastructure and review infra PRs for safety, improve CI signal quality, coordinate mobile releases via TestFlight and Google Play, and harden rollback, restore, and break-glass paths.
  • Reliability-focused code (TypeScript/Node): repair scripts, migrations and index management, observability instrumentation, tenant-scoped operational tooling, background job lifecycle, deploy tooling, and e2e (Playwright) and integration tests. This is your primary lane, but with agentic tooling we expect you to help on product features when it counts rather than treating “that’s not infra” as a boundary.
  • Release quality and QA: validate release candidates, walk core clinical workflows on web, iOS, and Android, run our test suites (Playwright, Jest, Postman), and drive release checklists and post-release verification.
  • Customer escalations: be the first internal contact for customer-reported issues. Reproduce, isolate, document, prioritize by clinical impact, and own the loop back to the customer.

Where We Need Help

HIPAA-aware operations, multi-tenant architecture (tenant isolation, org-safe migrations and diagnostics), disaster recovery and restore confidence, data-integrity operations, security and production-access hygiene (IAM, secrets, least privilege), incident-response maturity (severity levels, SLOs, alerting standards), and scale and cost visibility across AWS and MongoDB.

Requirements

  • 3-5 years that meaningfully includes SRE, DevOps, production, platform, or infrastructure engineering.
  • AWS proficiency is a hard requirement. Real, hands-on experience operating production workloads on AWS, and you can talk through it in specifics.
  • A product-focused thinker who asks why the product matters and what needs building, not only how the infrastructure runs, and who can contribute to feature work with agentic tooling when needed.
  • Self-sufficient and a fast ramper. Dropped into an unfamiliar codebase, cloud account, or toolchain, you have a real strategy to get productive on your own, and you should expect to demonstrate it live during interviews.
  • Curious about ABA, our product, and AI. Genuine interest in the clinical work Acquire supports is , and we want people who are eager to learn the domain and to explore how AI can move it forward.
  • Honest about your background, with recent, checkable professional references.
  • Comfortable with CI/CD, Terraform or comparable IaC, MongoDB or another production database, and observability tooling (CloudWatch, Datadog, Sentry, Grafana).
  • Fluent enough in Node/TypeScript for reliability work: repair scripts, migrations, instrumentation, job lifecycle, deploy tooling, and e2e/integration tests.
  • At home on the command line, in production logs, and in cloud consoles, with real incident experience you can talk through.
  • Careful about production access, customer data, and tenant boundaries, and clear on the difference between a helpful diagnostic and an accidental data leak.
  • Comfortable on a small team where some process exists, some needs creating, and everyone stays close to the product.

Using AI Tools

We use AI where it genuinely helps. You do not need to be an AI expert, and we are not looking for someone who treats AI as a substitute for judgment. We want someone comfortable with tools like Claude or Cursor to move faster (log triage, drafting repair scripts and tests, investigating support cases, verification checklists, working past the edge of their expertise) while still checking the result like an engineer. If a tool creates uncertainty, know when to slow down and verify.

Bonus Points

Healthcare or other HIPAA-regulated experience; multi-tenant SaaS (tenant isolation, support tooling, data repair); disaster recovery, SLOs, or incident command; familiarity with our stack (AWS ECS/CloudFront/IAM/CloudWatch/SNS, MongoDB Atlas, Terraform, GitHub Actions, Node, TypeScript, Playwright, Sentry, ClickUp); React/React /Express in production and monorepos (pnpm, Turborepo); release coordination; and prior work introducing SRE/DevOps/QA practices on a small team.

About the company

Acquire Learning is a learning management platform built for ABA (Applied Behavior Analysis) therapy. Clinicians and behavior technicians use it daily with clients on the autism spectrum, and the data it captures shapes real treatment decisions. We are a small, product-focused team in a HIPAA-regulated environment. When a clinician is mid-session with a child, the platform behaving predictably is the difference between productive therapy and a disrupted session. Your work affects the communities we serve., In the first months you learn the product, infrastructure, and highest-impact reliability surfaces, take over the alert-response loop, become a reliable hand on deploys and incidents, and start contributing to reliability code. Over the first year you take on reliability strategy, observability and post-incident hygiene, release engineering, and the operational and QA playbooks that keep Acquire trustworthy as it scales. The goal is Lead SRE at Acquire. Comp is revisited as you grow into that scope.

Compensation & Benefits

  • Base salary: $90,000-$105,000, revisited as the role grows into Lead SRE ownership.
  • Semi-annual performance reviews with raise potential.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.energyjobline.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:22 min

Overview of the Sentry error and performance monitoring platform

Priscila Oliveira · WWC 2023

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

3:21 min

Installing and configuring the Sentry JavaScript SDK for applications

Priscila Oliveira · WWC 2023

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

Videos

See all

Related articles

See all