Senior Software Engineer, Observability
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Experteer Overview As a Senior Software Engineer on CoreWeave’s Observability team, you will design, build, and maintain core observability infrastructure for large-scale AI workloads. You will partner with internal engineering teams to embed best practices, tackle performance and reliability challenges across GPU-heavy clusters, and help shape platform strategy while participating in on-call rotations. This role blends system design, reliability, and scalable telemetry across metrics, logging, and tracing to give users deep visibility. You’ll work in a fast-paced, hyper-growth environment that values curiosity, ownership, and collaboration. Compensation / Benefits * Design, build, and maintain observability infrastructure spanning metrics, logging, tracing, and telemetry pipelines * Develop highly reliable and scalable systems for production use * Collaborate with internal engineering teams to embed observability best practices * Address performance and reliability challenges across clusters of thousands of GPUs * Contribute to platform strategy and participate in on-call rotations to keep production systems robust Tasks * 5+ years in software or infrastructure engineering with large-scale distributed systems * Proficient in Go or Python with production-ready code * Hands-on Kubernetes, containerization, and microservices in production * Ability to design and deliver scalable, robust systems with automated testing and progressive release strategies * Skill in decomposing complex distributed problems into manageable work * Experience with Helm and YAML configurations; infrastructure-as-code practices * On-call experience for critical production systems * Bachelor’s degree in Computer Science, Electrical Engineering, Mathematics, or related field Key requirements * Medical, dental, and vision insurance * 401(k) with generous employer match * Flexible PTO * Tuition Reimbursement * Employee Stock Purchase Program (ESPP) * Paid Parental Leave
Requirements
clusters of thousands of GPUs * Contribute to platform strategy and participate in on-call rotations to keep production systems robust Tasks * 5+ years in software or infrastructure engineering with large-scale distributed systems * Proficient in Go or Python with production-ready code * Hands-on Kubernetes, containerization, and microservices in production * Ability to design and deliver scalable, robust systems with automated testing and progressive release strategies * Skill in decomposing complex distributed problems into manageable work * Experience with Helm and YAML configurations; infrastructure-as-code practices * On-call experience for critical production systems * Bachelor’s degree in Computer Science, Electrical Engineering, Mathematics, or related field Key requirements * Medical, dental, and vision insurance * 401(k) with generous employer match * Flexible PTO * Tuition Reimbursement * Employee Stock Purchase Program (ESPP) * Paid Parental Leave
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Why Attend a Developer Event in 2026?
Dev Digest 121 - AI goes offline
Highest Paying Tech Companies for Developers
What Makes WeAreDevelopers World Congress Different From Every Other Tech Event?