> Markdown version of [/jobs/ext/3136297-devops-engineer](https://www.wearedevelopers.com/jobs/ext/3136297-devops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # DevOps engineer - **Company:** FABRION - **Location:** Bodega Bay, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Audit Trail, Microsoft Azure, Cloud Computing, Cloud Engineering, Continuous Integration, DevOps, Github, Identity and Access Management, Key Management, Uptime, Open Source Technology, Role-Based Access Control, Prometheus, Policy as Code, Pulumi, ReactJS, Large Language Models, Grafana, Multi-Agent Systems, Amazon Virtual Private Cloud (VPC), Backend, Kubernetes, Infrastructure Automation Frameworks, Low Latency, Deployment Automation, Sentry, Machine Learning Operations, Terraform, Software Version Control, Docker - **Published:** September 29, 2026 - **Apply:** https://www.juju.com/job/16_5d34bc4913 ## About the Role Core Experience: * 4-10+ years in DevOps, platform engineering, or SRE in production-grade systems * Strong experience with Docker, Kubernetes (EKS/GKE), Terraform or Pulumi * Hands-on experience deploying and monitoring distributed cloud-native systems * Familiar with GitOps practices, CI/CD design, progressive delivery, and secure SDLC * Clear understanding of how to implement monitoring, alerting, and failure simulation in dynamic environments Engineering Mindset: * Obsessed with reliability, latency, uptime, and repeatability * Security-aware and compliance-conscious * Proactive - you don't wait for alerts to fix things * Comfortable collaborating with backend, AI, and data teams Bonus: Agent-Native / ML Ops Capabilities * We're building an agentic, AI-native platform from the ground up. Experience here isn't required, but would be a strong differentiator: * Experience running LLM orchestration frameworks (e.g. LangChain, LangGraph, Dust, ReAct agents) * Building retrieval-augmented generation (RAG) pipelines - and deploying them safely and repeatably * Familiarity with vector DBs (Weaviate, Qdrant, Pinecone) and embedding pipelines * Monitoring and governing long-running or multi-agent chains * Auditability and replay systems for agent decision-making * Serving fine-tuned or open-source LLMs with model versioning and GPU scaling (e.g. vLLM, TGI) * Interest in auto-remediation using agents (e.g. observability + alert * insight * response via LLM) ## Description We're building an AI-native, multi-tenant enterprise platform for complex domains in industrial verticals. In this architecture, DevOps isn't just about shipping features - it's about operationalizing intelligent agents, ensuring traceability across AI systems, and supporting mission-critical ML infrastructure at scale., * Build and maintain scalable cloud infrastructure across AWS/GCP/Azure with a focus on secure, tenant-isolated deployments * Own and evolve CI/CD systems (e.g. GitHub Actions, ArgoCD) with progressive rollout, testing, and rollback flows * Establish observability tooling across services, agents, and pipelines (OpenTelemetry, Prometheus, Grafana, Sentry) * Implement policy-as-code (OPA, Rego) for deployment safety, RBAC, audit logging, and approval workflows * Define and enforce SLAs, uptime targets (99.99%+), incident response, and remediation workflows * Secure infrastructure: IAM, VPC, encryption, key management, image scanning, secrets rotation * Automate deployments, infrastructure provisioning (Terraform, Helm), and environment replication, DevOps is the nervous system of the platform - every agent, every data fabric component, every pipeline flows through what you build. This is a rare opportunity to design that system early, the right way, and future-proof it for scale, compliance, and trust. ## Related Videos - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [From Doubt to Confidence: How Sentry Uses Verdaccio to Bulletproof SDK Releases](https://www.wearedevelopers.com/videos/739-from-doubt-to-confidence-how-sentry-uses-verdaccio-to-bulletproof-sdk-releases) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) ## Related Articles - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)