Site Reliability Engineer

The Outpost
Latin America, United States
12 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$75,000.0 - $90,000.0
Working hours
Regular working hours
Languages
English

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Bash Shell Databases Data Infrastructure Monitoring of Systems Python (Programming Language) PostgreSQL Machine Learning Performance Tuning Query Optimization Prometheus
+14 more
Next.js Smart Devices TypeScript Zabbix Datadog Google Cloud ReactJS Grafana Indexer Backend Containerization Bare Metal Terraform Docker

Job description

Our platform combines AI-powered gate automation, computer vision, and operational software to help logistics operators run smarter, faster facilities. We’re a small, high-conviction team shipping real software that ends up in real yards, at real gates, moving real freight; if our system goes down, trucks stop moving and a customer’s yard stops running. What we build is mission-critical to the people who depend on it.

We’re scaling fast, with load expected to 10X over the next 18 months, and reliability is now core to whether customers trust us to run their gates. We need an SRE to own uptime and incident response as the system grows, and to help the team get proactive about issues instead of reactive.

Project Details:

  • Own reliability targets across our backend/API, worker services, applications and CV pipeline; MTD, MTM, MTR, and follow-through on root causes.
  • Level up our monitoring and alerting, and build out auto-remediation, so on-call load scales with automation, not headcount.
  • Partner with our agentic engineering work to build agents that triage alerts and handle routine remediation.
  • Harden and optimize our GCP infrastructure (Cloud Run, Cloud SQL, GCS) for cost and performance as load scales.
  • Own database scale and performance; connection pooling, query optimization and indexing, read replicas, and capacity planning, so Postgres doesn’t become the bottleneck as data volume grows.
  • Improve the reliability of our ML training and monitoring infrastructure, in partnership with the CV/ML team.
  • Run blameless postmortems and drive fixes for root causes, not just symptoms.
  • Participate in on-call rotation., TypeScript / Node.js · Next.js · React · Apollo Server · Express · PostgreSQL · GCP · Docker · Python (ML/data workloads). We’re pragmatic - the right tool matters more than the familiar one., This role is a full-time contract position. You’ll work closely with our core engineering team - embedded in our sprints, standups, and Slack channels - but employment is managed through the agency. We’ve built this model successfully with engineers in Latin America, and it’s been a great fit for both sides.

Requirements

  • 4+ years in an SRE, infrastructure, or backend engineering role with production on-call ownership.
  • Deep experience with a major cloud provider (GCP preferred); compute, managed databases, object storage, networking.
  • Experience building monitoring/alerting/observability stacks (Grafana, Prometheus, Zabbix, Datadog, or similar).
  • Strong scripting/automation skills (Python, Bash, or similar).
  • Comfortable with containerized workloads (Docker) and CI/CD pipelines.
  • Track record of reducing incident volume or improving reliability metrics - not just responding to incidents.
  • Strong communication skills, comfortable working with both technical and non-technical stakeholders, know when and how to escalate urgency, and build strong working relationships across teams.
  • Strong communication skills in English - you write clearly and engage well async.

Preferred Qualifications:

  • Experience with ML/data infrastructure - training pipelines, model monitoring, feature stores.
  • Experience building or integrating AI agents for operational automation (alert triage, auto-remediation).
  • Infrastructure-as-code experience (Terraform or similar).
  • PostgreSQL performance tuning at scale.
  • Background supporting physical/IoT systems (edge devices, cameras, on-site hardware).
  • Experience with bare-metal infrastructure in colocation environments, hardware monitoring, redundancy, and failover/high-availability configuration.

About the company

Outpost is building the backbone of freight. We’re reinventing how supply chain infrastructure works in America with carrier agnostic truck terminals. As a vertically integrated real estate, operations, and technology company, we acquire and operate mission-critical real estate across the country to serve the largest logistics providers in the world. Backed by $1B from Greenpoint Partners, we’re scaling and building the most valuable logistics network in the country.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.ashbyhq.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:38 min

Transitioning into backend engineering from web development

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto · World Congress 2024

Videos

See all

Related articles

See all