> Markdown version of [/jobs/ext/1978451-senior-software-engineer-observability](https://www.wearedevelopers.com/jobs/ext/1978451-senior-software-engineer-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, Observability - **Company:** Coreweave, Inc. - **Location:** New York, NY, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Automation of Tests, Distributed Systems, Python (Programming Language), YAML, Data Logging, Graphics Processing Unit (GPU), Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Production Code, Microservices - **Published:** August 7, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/senior-software-engineer-observability-new-york-ny-usa-58833625 ## About the Role clusters of thousands of GPUs * Contribute to platform strategy and participate in on-call rotations to keep production systems robust Tasks * 5+ years in software or infrastructure engineering with large-scale distributed systems * Proficient in Go or Python with production-ready code * Hands-on Kubernetes, containerization, and microservices in production * Ability to design and deliver scalable, robust systems with automated testing and progressive release strategies * Skill in decomposing complex distributed problems into manageable work * Experience with Helm and YAML configurations; infrastructure-as-code practices * On-call experience for critical production systems * Bachelor's degree in Computer Science, Electrical Engineering, Mathematics, or related field Key requirements * Medical, dental, and vision insurance * 401(k) with generous employer match * Flexible PTO * Tuition Reimbursement * Employee Stock Purchase Program (ESPP) * Paid Parental Leave ## Description Experteer Overview As a Senior Software Engineer on CoreWeave's Observability team, you will design, build, and maintain core observability infrastructure for large-scale AI workloads. You will partner with internal engineering teams to embed best practices, tackle performance and reliability challenges across GPU-heavy clusters, and help shape platform strategy while participating in on-call rotations. This role blends system design, reliability, and scalable telemetry across metrics, logging, and tracing to give users deep visibility. You'll work in a fast-paced, hyper-growth environment that values curiosity, ownership, and collaboration. Compensation / Benefits * Design, build, and maintain observability infrastructure spanning metrics, logging, tracing, and telemetry pipelines * Develop highly reliable and scalable systems for production use * Collaborate with internal engineering teams to embed observability best practices * Address performance and reliability challenges across clusters of thousands of GPUs * Contribute to platform strategy and participate in on-call rotations to keep production systems robust Tasks * 5+ years in software or infrastructure engineering with large-scale distributed systems * Proficient in Go or Python with production-ready code * Hands-on Kubernetes, containerization, and microservices in production * Ability to design and deliver scalable, robust systems with automated testing and progressive release strategies * Skill in decomposing complex distributed problems into manageable work * Experience with Helm and YAML configurations; infrastructure-as-code practices * On-call experience for critical production systems * Bachelor's degree in Computer Science, Electrical Engineering, Mathematics, or related field Key requirements * Medical, dental, and vision insurance * 401(k) with generous employer match * Flexible PTO * Tuition Reimbursement * Employee Stock Purchase Program (ESPP) * Paid Parental Leave ## Related Videos - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [CI/CD with Github Actions](https://www.wearedevelopers.com/videos/856-ci-cd-with-github-actions) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Crypto-secure Data Management with In-Database Blockchain](https://www.wearedevelopers.com/videos/632-crypto-secure-data-management-with-in-database-blockchain) - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) - [Startup Presentation: Achieving True Developer Self-Service in Kubernetes](https://www.wearedevelopers.com/videos/1181-startup-presentation-achieving-true-developer-self-service-in-kubernetes) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Makes WeAreDevelopers World Congress Different From Every Other Tech Event?](https://www.wearedevelopers.com/magazine/701-what-makes-wearedevelopers-world-congress-different-from-every-other-tech-event) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)