Data Engineer, Fleet Monitoring & Analysis

Coreweave, Inc.
United States
27 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Compensation
$153,000.0 - $204,000.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Airflow Amazon Web Services Business Analytics Applications Data Analysis Apache HTTP Server Microsoft Azure Big Data Databases Data as a Services Information Engineering
+22 more
Data Governance Data Infrastructure Extract Transform Load (ETL) Data Security Data Stores Data Visualization Data Warehousing Database Queries Python (Programming Language) Meta-Data Management NoSQL Performance Tuning SQL Databases Data Streaming Data Processing Google Cloud Apache Spark Indexer Data Lakes Information Technology Data Pipelines Programming Languages

Job description

As a Senior Data Engineer you will own and evolve the data lake and analytics stack that powers observability and decision-making for CoreWeave’s global hardware fleet. You’ll maintain, monitor, and upgrade our data lake infrastructure (Apache Iceberg, Trino, Apache Airflow, Apache Spark, Apache Superset) and ETL pipelines, while delivering ad-hoc analysis and executive-ready reporting for team, director, and leadership stakeholders. You’ll also create visualizations, documentation, and integrations that make fleet monitoring data reliable, discoverable, and actionable across the organization. In this role, you will:

  • Design, develop, and maintain robust and scalable data pipelines to collect, process, and store data from various sources, including APIs, databases, and third-party services.
  • Maintain, monitor, and upgrade CoreWeave’s data lake infrastructure, including Apache Iceberg, the Trino query layer, Apache Airflow, Apache Spark, and Apache Superset.
  • Maintain, monitor, and upgrade ETL/ELT pipelines to ensure reliable, performant, and observable data flows across batch and (where applicable) streaming workloads.
  • Create and optimize data models and data products to support analytics and reporting, ensuring data accuracy, consistency, and performance.
  • Provide ad-hoc analysis and reporting for team, director, and executive-level stakeholders, translating business questions into data-driven insights and clear narratives.
  • Create visualizations and dashboards (e.g., in Apache Superset or similar tools) that surface key metrics, trends, and operational KPIs for a variety of internal audiences.
  • Develop and maintain documentation and runbooks for data pipelines, data lake infrastructure, data models, and usage patterns to support knowledge sharing and troubleshooting.
  • Implement data security and governance best practices to protect sensitive information and comply with data privacy regulations.
  • Collaborate with cross-functional teams to integrate data into applications and analytics platforms, helping to visualize performance metrics and identify opportunities for improvement.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field.
  • 4 - 7 years of experience as a Data Engineer or in a similar data-focused role in a fast-paced environment.
  • Strong SQL skills for data manipulation, modeling, and querying large datasets.
  • Proficiency in at least one programming language commonly used for data engineering such as Python, Java, or Scala.
  • Hands-on experience with data pipeline orchestration tools (e.g., Apache Airflow) and big data technologies (e.g., Apache Spark).
  • Experience designing, operating, and optimizing data lake and/or data warehouse solutions, with a solid understanding of data modeling and performance tuning.
  • Knowledge of cloud platforms (e.g., AWS, Google Cloud Platform, Azure) and related data services (e.g., object storage, managed databases, analytics services).
  • Familiarity with database systems (e.g., SQL and NoSQL) and data warehousing concepts, including partitioning, indexing, and schema design.
  • Experience building, maintaining, and monitoring ETL/ELT pipelines in production environments, including alerting and observability.
  • Experience creating and maintaining reporting and analytics solutions (dashboards, reports, and metrics) for technical and non-technical audiences.

Preferred:

  • Experience with modern data lakehouse technologies and table formats such as Apache Iceberg (or similar technologies like Delta Lake or Apache Hudi).
  • Experience with Trino or other distributed SQL query engines at scale.
  • Experience with Apache Superset or other BI/visualization tools for building self-service analytics.
  • Experience with data quality frameworks, data observability tooling, and/or metadata management.
  • Experience supporting executive-level reporting and KPI design in partnership with business and finance stakeholders.

Wondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren’t a 100% skill or experience match. Here are a few qualities we’ve found compatible with our team. If some of this describes you, we’d love to talk.

  • You love owning data infrastructure end-to-end-from ingestion and modeling to analytics and visualization.
  • You’re curious about modern data lake and lakehouse architectures and enjoy working with open-source data tooling.
  • You’re an expert at turning loosely defined business questions into concrete data products, metrics, and dashboards that drive decisions.

Benefits & conditions

The base salary range for this role is $153,000 to $204,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility)., In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings; for roles in other locations, benefits vary and are shared during the hiring process. These include:

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long-term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health
  • Family-Forming support provided by Carrot
  • Paid Parental Leave
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption

About the company

CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at

What You’ll Do:

The Fleet Monitoring and Analysis (FMA) team builds and operates the forward-deployed monitoring and observability layer for CoreWeave’s ever-expanding global hardware fleet; continually improving node and environmental visibility to support automated provisioning and high-reliability operations.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon ¡ World Congress 2026 Europe

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva ¡ JS Congress

2:36 min

Analyzing limitations with PostgreSQL bitmap heap scans

Dharin Shah Dharin Shah ¡ World Congress 2025

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon ¡ World Congress 2026 Europe

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes ¡ LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt ¡ LIVE

Videos

See all

Related articles

See all