> Markdown version of [/jobs/ext/2580269-senior-devops-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2580269-senior-devops-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior DevOps / Infrastructure Engineer - **Company:** MONAD LABS, INC. - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $180,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Cloud Computing, Computer Programming, Software Debugging, Linux, DevOps, Domain Name System (DNS), Python (Programming Language), Ansible, Blockchain, Prometheus, Large Language Models, Grafana, Build Management, Kubernetes, Information Technology, Virtual Agents, Terraform - **Published:** August 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=73a59b26eff40652 ## About the Role * You have 5+ years in DevOps, SRE, or Infrastructure Engineering, operating production systems at scale. * You have strong Linux, systemd, networking, and shell fundamentals, and you're comfortable debugging live systems over SSH. * You have deep, hands-on infrastructure-as-code experience with Ansible and Terraform. * You have experience with observability stacks (Prometheus, Grafana, Loki, or equivalents). * You have hands-on fluency with AI-assisted engineering: you use coding agents and LLM tooling in your daily workflow and have judgment on where it helps and where it's risky. * You have experience designing automation with safe guardrails, and you bring calm, methodical incident response. * You have programming and scripting experience (e.g., Python, bash). * Experience with Kubernetes and GitOps (Flux or Argo) is a plus. * Experience building AI agent tooling, MCP servers, or agent orchestration frameworks is a plus. * Experience serving inference, either locally or as a service is a plus. * Previous experience with blockchain clients or node operations is a plus. * A Bachelor of Science in Computer Science, Engineering, or a related field is a plus. ## Description We're looking for a Senior DevOps / Infrastructure Engineer to operate the infrastructure behind Monad, and to push how much of that operation can be driven by AI. You'll keep our globally-distributed validator, full node, and archive fleet healthy across mainnet and testnet, own our infrastructure-as-code and observability, and build the agentic tooling and guardrails that let a small team safely operate a large fleet. As more of our engineering shifts toward AI, this role is central to designing the workflows, deterministic guardrails, and security perimeters within which autonomous agents operate our infrastructure; you'll also help stand up and operate the infrastructure behind our own growing model workloads., * Operate the Monad node fleet: health, sync, upgrades, and recovery across validators, full nodes, archive/historical, and indexer nodes on mainnet and testnet, including safe, staged rollouts and incident response. * Own our infrastructure-as-code: Ansible for fleet configuration, Terraform + Atlantis for cloud and DNS, and Kubernetes/Flux (GitOps) for platform services. * Build and operate observability and alerting (Prometheus, Grafana, Loki); create dashboards and alerts that catch problems before they page while minimizing false positives. * Automate the release pipeline: node upgrades, canary rollouts, snapshot/restore, and the guardrails that bound blast radius (e.g., protecting validators from automated changes). * Design and build agentic operations: develop AI agents, tooling (e.g., MCP servers), and runbooks-as-code that let agents safely investigate, diagnose, and execute routine operations, with deterministic guardrails and human oversight. * Codify operational knowledge into tools and automation that the whole team, and its agents, can reuse. * Harden nodes and services, manage secrets, and continuously drive down manual toil. ## Related Videos - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)