> Markdown version of [/jobs/ext/2706148-software-engineer-ai-infrastructure-performance-insights-observability](https://www.wearedevelopers.com/jobs/ext/2706148-software-engineer-ai-infrastructure-performance-insights-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - AI Infrastructure Performance Insights & Observability - **Company:** Coreweave, Inc. - **Location:** Sunnyvale, United States - **Experience:** Expert - **Salary:** $182,000.0 - $242,000.0 - **Contract:** Permanent contract - **Skills:** Apache HTTP Server, C++ (Programming Language), Nvidia CUDA, Databases, Continuous Integration, Data Centers, Information Engineering, Data Governance, Distributed Computing Environment, Distributed Systems, InfiniBand, Python (Programming Language), Meta-Data Management, Query Optimization, Remote Direct Memory Access, Prometheus, AI Infrastructure, Parquet, Graphics Processing Unit (GPU), Computer Networking Systems, Pytorch, Grafana, Apache Spark, Build Management, Data Lakes, Kubernetes, Avro, Machine Learning Operations, Hardware Infrastructure, Looker Analytics - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-software-engineer-ai-infrastructure-performance-insights-observability-coreweave-2-8997634 ## About the Role * 5+ years of experience building distributed systems, observability platforms, or performance engineering tooling, ideally for infrastructure or ML systems rather than general-purpose BI. * Strong coding in Python or Go (C++ a plus) and deep familiarity with networked systems, GPU infrastructure, and performance analysis. * Hands-on experience with Kubernetes at production scale, CI/CD, and observability stacks (Prometheus, Grafana, OpenTelemetry) used to monitor and diagnose infrastructure, not just report on it. * Working knowledge of time-series databases and fluency in PromQL or MetricsQL for building real-time alerting and anomaly detection, not only historical dashboards. * Familiarity with data lake architectures and modern table formats (Iceberg, Parquet, Avro) sufficient to support an insights platform, though this is not the primary skill this role is hiring for. * Comfortable working close to hardware and workload behavior: GPU utilization patterns, interconnect fabric health, distributed training/inference performance characteristics. * Strong communicator comfortable collaborating with cross-functional teams and external partners. Nice to have * Experience with time-series databases, LSM-based storage engines, or custom telemetry pipelines. * Experience running MLPerf submissions or similar large-scale audited benchmarks. * Contributions to OSS projects such as Apache Iceberg, Apache Spark, Trino, llm-d, vLLM, or PyTorch. * Direct experience benchmarking or monitoring large GPU fleets or multi-region clusters. * Experience with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies. * Familiarity with data cataloging, lineage tools, or data governance frameworks. ## Description We're looking for a Senior Engineer to be a driving force on CoreWeave's Benchmarking & Performance team, with a focus on building the performance insights and observability systems that make our AI infrastructure legible at every layer, from individual GPUs and NVLink/InfiniBand fabrics up through distributed training and inference workloads. You will own how we detect, diagnose, and surface performance signals across every data center in our global infrastructure, turning billions of raw telemetry events into real-time insight that engineers, product teams, and executives can act on with confidence. This is not a straightforward data engineering or BI role. The role will focus on building the observability and insight tooling that lets us answer, in near real time, whether a GPU fleet, a fabric, or a training run is performing the way it should, and why it isn't when it's not. If you're energized by building the systems that turn raw infrastructure telemetry into trusted, actionable performance intelligence, and you want that work to sit closer to the hardware and the workload than to a dashboard, this role was built for you. What you'll do * Performance Insights & Observability - Design and build the systems that continuously assess AI infrastructure health and performance: GPU utilization and efficiency, interconnect (NVLink, InfiniBand, RoCE) fabric behavior, distributed training and inference throughput, and hardware degradation signals. Build the detection and diagnosis logic that surfaces anomalies and regressions before they become incidents, not just dashboards that report on them after the fact. * Time-Series & Metrics Infrastructure - Own and extend our time-series database (TSDB) layer as the backbone of real-time observability. Write and optimize PromQL/MetricsQL queries that power alerting, anomaly detection, and trend analysis across thousands of GPUs and hundreds of benchmark runs. Bridge streaming metrics and batch-analytical workloads so engineers get sub-second answers during live incidents and analysts get complete historical context for root cause work. * Fabric & GPU Telemetry - Build and validate the pipelines and metrics that make network fabric and GPU-level behavior observable and comparable across racks, clusters, and hardware generations, including gray failure detection, congestion and error-rate signals, and health scoring that holds up under audit. * Data Lake Architecture (in support of insight work) - Design and build the performance data lake that underpins the above: table formats (Apache Iceberg, Parquet, Avro), hot/cold tiering, and schema evolution for latency distributions, throughput metrics, GPU utilization, cost-per-token, and hardware health signals. This is foundational infrastructure, not the end product. * Query Optimization & Performance - Profile and tune query engines against columnar and time-series stores so that the observability layer meets its own strict P99 latency and freshness SLAs. Benchmark the benchmarking infrastructure itself. * BI & Reporting (secondary) - Where needed, build self-service views (Grafana, Looker, or similar) for engineers, product managers, and executives, but as a downstream output of the insight and observability work above, not the primary deliverable. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [From event streaming to event sourcing 101](https://www.wearedevelopers.com/videos/91-from-event-streaming-to-event-sourcing-101) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)