Senior/Staff CloudOps Engineer

Cloudzero Inc.
San Francisco, CA, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Software Debugging Distributed Systems Monitoring of Systems Python (Programming Language) Prometheus Datadog Pulumi Google Cloud Deployment Automation
+2 more
Terraform Serverless Computing

Job description

CloudZero is growing fast. Our customer base is expanding, the data challenges we’re solving are getting more complex, and the platform is scaling to match. As a CloudOps Engineer you’ll be a force multiplier for our engineering organization, owning the performance, reliability, and observability of CloudZero’s infrastructure and empowering teams to ship features that help customers understand and optimize their cloud spend.

This is real infrastructure work at real scale, not a ticket-closing role or a console-clicking job. CloudZero processes billions of events daily across AWS, Azure, and GCP. Our customers rely on real-time, accurate cost data to make business-critical decisions, and any instability in our system impacts their planning. Built entirely on a unique serverless architecture with no EC2s or containers, our platform demands infrastructure that scales gracefully, fails predictably, and recovers automatically.

If you thrive on hard operational problems, care deeply about reliability and performance, and want to see your work matter to customers in direct and measurable ways, this role was built for you.

What You’ll Do

Infrastructure as Code

  • Design and maintain Pulumi modules that provision reliable, cost-efficient cloud resources
  • Own infrastructure end to end with no clicking through consoles

Observability

  • Instrument systems so that failures surface quickly and debugging happens with data, not guesswork
  • Build observability into everything so you know about problems before customers do

Automation

  • Automate deployments, scaling, backups, and limit changes; if humans are doing it repeatedly, build a system to do it instead
  • Balance automation intelligently, building solutions to real problems rather than automating for its own sake

Partner with Product Engineering

  • Help teams design resilient services, review architectures for operational complexity, and build deployment pipelines that enable safe and fast shipping
  • Optimize for cost and performance; CloudZero’s business is helping others optimize cloud costs, and we should be exemplars of efficient cloud usage ourselves

Requirements

Do you have experience in Technical documentation?, * 3 to 5+ years of experience building and operating distributed systems in AWS

  • Strong skills in Python and Infrastructure as Code using Pulumi or Terraform
  • Experience with frontier AI models such as Claude, Codex, or Gemini
  • Hands-on experience with monitoring tools such as Prometheus or Datadog
  • Proven ability to debug production issues under pressure
  • Values thoughtful, reliable system design over reactive hero efforts
  • Strong documentation habits to support long-term team clarity and system stability
  • Ability to clearly explain complex technical issues to non-technical stakeholders
  • Excited to take ownership of infrastructure and solve operational challenges at scale

About the company

Cloud cost management is one of the biggest challenges organizations face today. As cloud adoption continues to accelerate, so do the complexities and costs associated with it, and macroeconomic conditions only increase pressure to prove cloud efficiency.

CloudZero is a SaaS platform at the intersection of next-generation cloud cost management and FinOps. We ingest billing and usage data from all cloud, SaaS, and PaaS providers, organize it in real time according to our customers’ business structures, and empower organizations to make more informed business decisions.

Since our founding in 2016, our mission has been to make efficient innovation a reality for every cloud-driven organization. We believe every engineering decision is a buying decision, and we’re applying proven reliability engineering principles to financial efficiency.

We believe the best AI empowers users with clear insights and confident decisions, transforming complex cloud cost data into actionable intelligence that drives meaningful business outcomes.

To date, we’ve raised over $56 million from leading venture capital firms. We’re solving problems of massive scale, business importance, and complexity in a space that needs it more than ever.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:44 min

Career transition into cloud native and data management

Michael Cade · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

5:34 min

Managing token budgets and enterprise usage of coding agents

Chris Heilmann +2 · LIVE

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

3:20 min

Overview of infrastructure as code tools

Alexander Bubeck · WWC 2023

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all