DevOps engineer

FABRION
Bodega Bay, CA, United States
2 days ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Audit Trail Microsoft Azure Cloud Computing Cloud Engineering Continuous Integration DevOps Github Identity and Access Management Key Management Uptime
+20 more
Open Source Technology Role-Based Access Control Prometheus Policy as Code Pulumi ReactJS Large Language Models Grafana Multi-Agent Systems Amazon Virtual Private Cloud (VPC) Backend Kubernetes Infrastructure Automation Frameworks Low Latency Deployment Automation Sentry Machine Learning Operations Terraform Software Version Control Docker

Job description

We’re building an AI-native, multi-tenant enterprise platform for complex domains in industrial verticals. In this architecture, DevOps isn’t just about shipping features - it’s about operationalizing intelligent agents, ensuring traceability across AI systems, and supporting mission-critical ML infrastructure at scale., * Build and maintain scalable cloud infrastructure across AWS/GCP/Azure with a focus on secure, tenant-isolated deployments

  • Own and evolve CI/CD systems (e.g. GitHub Actions, ArgoCD) with progressive rollout, testing, and rollback flows
  • Establish observability tooling across services, agents, and pipelines (OpenTelemetry, Prometheus, Grafana, Sentry)
  • Implement policy-as-code (OPA, Rego) for deployment safety, RBAC, audit logging, and approval workflows
  • Define and enforce SLAs, uptime targets (99.99%+), incident response, and remediation workflows
  • Secure infrastructure: IAM, VPC, encryption, key management, image scanning, secrets rotation
  • Automate deployments, infrastructure provisioning (Terraform, Helm), and environment replication, DevOps is the nervous system of the platform - every agent, every data fabric component, every pipeline flows through what you build. This is a rare opportunity to design that system early, the right way, and future-proof it for scale, compliance, and trust.

Requirements

Core Experience:

  • 4-10+ years in DevOps, platform engineering, or SRE in production-grade systems
  • Strong experience with Docker, Kubernetes (EKS/GKE), Terraform or Pulumi
  • Hands-on experience deploying and monitoring distributed cloud-native systems
  • Familiar with GitOps practices, CI/CD design, progressive delivery, and secure SDLC
  • Clear understanding of how to implement monitoring, alerting, and failure simulation in dynamic environments

Engineering Mindset:

  • Obsessed with reliability, latency, uptime, and repeatability
  • Security-aware and compliance-conscious
  • Proactive - you don’t wait for alerts to fix things
  • Comfortable collaborating with backend, AI, and data teams

Bonus: Agent-Native / ML Ops Capabilities

  • We’re building an agentic, AI-native platform from the ground up. Experience here isn’t required, but would be a strong differentiator:
  • Experience running LLM orchestration frameworks (e.g. LangChain, LangGraph, Dust, ReAct agents)
  • Building retrieval-augmented generation (RAG) pipelines - and deploying them safely and repeatably
  • Familiarity with vector DBs (Weaviate, Qdrant, Pinecone) and embedding pipelines
  • Monitoring and governing long-running or multi-agent chains
  • Auditability and replay systems for agent decision-making
  • Serving fine-tuned or open-source LLMs with model versioning and GPU scaling (e.g. vLLM, TGI)
  • Interest in auto-remediation using agents (e.g. observability + alert * insight * response via LLM)

About the company

Backed by 8VC, we’re building a world-class team to tackle one of the industry’s most critical infrastructure problems.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

1:22 min

Overview of the Sentry error and performance monitoring platform

Priscila Oliveira · World Congress 2023

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:45 min

Fusing developer experience and platform engineering for agentic SDLC

Julia Kordick Julia Kordick · World Congress 2026 Europe

3:20 min

Overview of infrastructure as code tools

Alexander Bubeck · World Congress 2023

3:21 min

Installing and configuring the Sentry JavaScript SDK for applications

Priscila Oliveira · World Congress 2023

Videos

See all

Related articles

See all