> Markdown version of [/jobs/ext/523210-senior-staff-cloudops-engineer](https://www.wearedevelopers.com/jobs/ext/523210-senior-staff-cloudops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior/Staff CloudOps Engineer - **Company:** Cloudzero Inc. - **Location:** Boston, MA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Software Debugging, Distributed Systems, Monitoring of Systems, Python (Programming Language), Prometheus, Datadog, Pulumi, Google Cloud, Deployment Automation, Terraform, Serverless Computing - **Published:** June 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=43808563d12dbfd2 ## About the Role Do you have experience in Technical documentation?, * 3 to 5+ years of experience building and operating distributed systems in AWS * Strong skills in Python and Infrastructure as Code using Pulumi or Terraform * Experience with frontier AI models such as Claude, Codex, or Gemini * Hands-on experience with monitoring tools such as Prometheus or Datadog * Proven ability to debug production issues under pressure * Values thoughtful, reliable system design over reactive hero efforts * Strong documentation habits to support long-term team clarity and system stability * Ability to clearly explain complex technical issues to non-technical stakeholders * Excited to take ownership of infrastructure and solve operational challenges at scale ## Description CloudZero is growing fast. Our customer base is expanding, the data challenges we're solving are getting more complex, and the platform is scaling to match. As a CloudOps Engineer you'll be a force multiplier for our engineering organization, owning the performance, reliability, and observability of CloudZero's infrastructure and empowering teams to ship features that help customers understand and optimize their cloud spend. This is real infrastructure work at real scale, not a ticket-closing role or a console-clicking job. CloudZero processes billions of events daily across AWS, Azure, and GCP. Our customers rely on real-time, accurate cost data to make business-critical decisions, and any instability in our system impacts their planning. Built entirely on a unique serverless architecture with no EC2s or containers, our platform demands infrastructure that scales gracefully, fails predictably, and recovers automatically. If you thrive on hard operational problems, care deeply about reliability and performance, and want to see your work matter to customers in direct and measurable ways, this role was built for you. What You'll Do Infrastructure as Code * Design and maintain Pulumi modules that provision reliable, cost-efficient cloud resources * Own infrastructure end to end with no clicking through consoles Observability * Instrument systems so that failures surface quickly and debugging happens with data, not guesswork * Build observability into everything so you know about problems before customers do Automation * Automate deployments, scaling, backups, and limit changes; if humans are doing it repeatedly, build a system to do it instead * Balance automation intelligently, building solutions to real problems rather than automating for its own sake Partner with Product Engineering * Help teams design resilient services, review architectures for operational complexity, and build deployment pipelines that enable safe and fast shipping * Optimize for cost and performance; CloudZero's business is helping others optimize cloud costs, and we should be exemplars of efficient cloud usage ourselves ## Related Videos - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Terraform for Developers](https://www.wearedevelopers.com/videos/3-terraform-for-developers) - [Seriously gaming your cloud expertise: from cloud tourist to cloud native](https://www.wearedevelopers.com/videos/373-seriously-gaming-your-cloud-expertise-from-cloud-tourist-to-cloud-native) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 162: AI careers, MCP, AWS best practices & floppy sweaters](https://www.wearedevelopers.com/magazine/571-dev-digest-162-ai-careers-mcp-aws-best-practices-floppy-sweaters) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)