> Markdown version of [/jobs/ext/1895466-software-engineer-production-engineering-cloud-on-prem](https://www.wearedevelopers.com/jobs/ext/1895466-software-engineer-production-engineering-cloud-on-prem). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Production Engineering (Cloud & On-Prem - **Company:** Coreweave, Inc. - **Location:** United States - **Experience:** Expert - **Salary:** $139,000.0 - $185,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Bash Shell, Software as a Service, Cloud Computing, Configuration Management, Databases, Continuous Integration, DevOps, Github, Python (Programming Language), PostgreSQL, MySQL, Ansible, Prometheus, Datadog, Google Cloud, Grafana, Multi-Cloud, Cloudformation, Containerization, Kubernetes, Vertica, Terraform, Pagerduty, Golang - **Published:** August 2, 2026 - **Apply:** https://www.dice.com/job-detail/a8590a5f-32cd-45bc-a1bb-28f40d8cfe17 ## About the Role * Extensive engineering experience designing, building, deploying, and operating critical production infrastructure and services across a large enterprise. * Expert in one or more major public clouds (AWS, Google Cloud Platform, or Azure), with real experience operating on-premises or hybrid environments, and in multi-cloud environments with multi-account strategies * Strong skills in infrastructure-as-code, automation, and configuration management (Terraform, CloudFormation, CDK, Ansible, or similar). * Hands-on expertise with Kubernetes and containerized workloads in production. * Proficient in a systems/scripting language (Go, Python, Bash, or similar) and comfortable writing tooling and automation, not just operating it. * Deep experience with CI/CD (GitHub Actions) and observability / monitoring tooling (Prometheus, Grafana, Datadog, etc.). * Proven ability to evolve designs to meet increasingly challenging scale, reliability, and performance requirements. * Comfortable owning on-call for services you build, and passionate about reducing on-call burden through architecture and automation rather than heroics. * Proven ability to provide technical leadership, work effectively with senior and principal engineers, influence technical direction, and contribute to the success of cross-functional stakeholders. Preferred * Experience building and operating SaaS products. * Experience building and operating reliability, incident, or developer-productivity platforms (service catalogs / Backstage, PagerDuty tooling, SLO frameworks). * Experience operating stateful systems and databases in production (PostgreSQL, MySQL, ClickHouse, etc.). * Cloud certifications (AWS Solutions Architect, Azure Architect, etc.). * Familiarity with security and access-management best practices in hybrid environments. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. * You love making production systems reliable, observable, and boring - and giving other engineers the tools to run their own services confidently. * You're curious about reducing on-call burden through better architecture and automation rather than more heroics. * You're comfortable across the DevOps stack - cloud and on-prem, Kubernetes, Terraform, CI/CD, and observability - and you like automating away toil. ## Description We are seeking a Senior Production Engineer with deep expertise across both cloud and on-premises environments to design, build, and operate the reliability platform at the core of CoreWeave's engineering organization. This is a hands-on DevOps/SRE-flavored role: you'll work across infrastructure-as-code, CI/CD, observability, and incident systems - automating away toil and giving service teams the visibility and guardrails they need to run their own code in production. You'll evolve our systems to meet increasingly challenging scale, reliability, and performance demands, and own large, ambiguous operational problems end to end - from error attribution and alert routing to release safety and capacity planning. You'll also provide technical leadership across the team, partnering with senior and principal engineers to influence direction and drive outcomes for cross-functional stakeholders. We care deeply about a sustainable on-call culture, so you'll participate in on-call rotations while championing the architecture and automation that reduce on-call burden over time. In this role, you will: * Design, build, deploy, and operate critical reliability and infrastructure services across AWS, Google Cloud Platform, Azure, and on-premises / hybrid environments. * Improve error attribution and alert routing, automatically routing errors and pages to the team that owns the affected service, so engineers are only on-call for what they own. * Own observability patterns, SLI/SLO frameworks, service catalog metadata, and dashboards that give teams visibility into their services from code change through production. * Build release-safety systems - canary deployments, smoke tests, staged rollouts, and reliable roll-back / roll-forward - so shipping to production is fast and safe. * Advance the incident and on-call program - incident tooling, on-call rotations, runbooks, and operational readiness reviews; measure incident volume by team to focus reliability investment. * Reduce on-call burden through better architecture, automation, and observability - and participate in on-call rotations yourself. * Provision and manage infrastructure with Terraform and drive manual, incident-time operations toward automated, repeatable infrastructure-as-code. * Break down and solve large, ambiguous operational problems, turning them into well-scoped, shippable engineering work. * Provide technical leadership and mentorship, work alongside senior and principal engineers, and influence the broader technical direction of the platform. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Building a Cloud Platform Where Everything is Just Another Kubernetes Resource](https://www.wearedevelopers.com/videos/100137-building-a-cloud-platform-where-everything-is-just-another-kubernetes-resource) ## Related Articles - [What Makes WeAreDevelopers World Congress Different From Every Other Tech Event?](https://www.wearedevelopers.com/magazine/701-what-makes-wearedevelopers-world-congress-different-from-every-other-tech-event) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How to Answer the Interview Question: “Why Do You Want to Be a Software Engineer?”](https://www.wearedevelopers.com/magazine/392-how-to-answer-the-interview-question-why-do-you-want-to-be-a-software-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)