Senior Software Engineer, ML Workflows - Weights & Biases

Weights & Biases
Bellevue, United States of America
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 220K

Job location

Bellevue, United States of America

Tech stack

Query Performance
C
Java
API
C Sharp (Programming Language)
C++
Cloud Computing
Programming Tools
Distributed Systems
Python
Query Optimization
Queueing Systems
Azure
Software Engineering
Systems Architecture
TypeScript
Rust
React
Technical Debt
Backend
Data Layers
Event Driven Architecture
Kubernetes
Kafka
GraphQL
Machine Learning Operations
Front End Software Development
Vertica
Api Design
Dropbox
Code Restructuring

Job description

Reporting to the Senior Engineering Manager for ML Workflows, this Senior Software Engineer will own the architecture and evolution of core systems in the Weights & Biases platform.

ML Workflows owns what customers use to move models and data through W&B. Artifacts is the versioned data layer that every object logged to W&B runs through. Registry sits on top of it as the organization-wide source of truth for what is production-ready, and it is used by teams at Pinterest, Dropbox, and Canva. Automations is the event and action layer the rest of the platform builds on. Launch is execution infrastructure that runs customer ML and agent workloads.

You will own systems end to end, from design through production reliability, working across the W&B backend, the Python SDK, and the frontend. You will partner closely with product managers, designers, and ML platform engineers, and you will talk directly to enterprise ML teams about their workflows. We hire senior engineers into the organization and match them to a team based on their strengths and interests, which we discuss with you during the interview process.

W&B is part of CoreWeave, so these products are being built onto CoreWeave infrastructure, and there is real room to shape how that happens.

What You'll Do

  • Own the architecture and evolution of a core platform area, designing systems that scale to billions of artifacts and high-volume event and job throughput.
  • Dive deep into system architecture to find optimization opportunities, solve complex bugs, and make the difficult calls that balance short-term delivery against long-term platform health.
  • Work across the stack as the problem requires, from Go backend services and GraphQL APIs to the Python SDK and the TypeScript frontend.
  • Lead migrations and re-architecture of systems carrying live customer traffic, keeping the experience seamless for the customers who depend on them.
  • Identify and refactor technical debt, improving maintainability and developer productivity.
  • Establish engineering patterns and best practices that hold up as the platform grows.
  • Collaborate with product and design to turn complex user requirements into clean technical implementations.
  • Work directly with customers to understand their ML workflow challenges, and bring that feedback into the product cycle.
  • Mentor engineers and raise the technical bar across the team.

Requirements

We expect all of these:

  • 6+ years of software engineering experience, including time leading significant technical initiatives or architectural changes
  • Strong proficiency in Go, or in another compiled language (C, C++, C#, Rust, Java) with the ability to ramp on Go quickly
  • Working proficiency in Python
  • Comfort diving into complex systems you did not write, to diagnose and fix difficult bugs across service boundaries
  • A track record of mentoring engineers and raising the technical bar around you
  • Strong communication, including the ability to explain complex technical decisions to different audiences

And real depth in at least one or two of these, which is also how we work out which team you join:

  • Distributed systems and infrastructure. Designing and operating services in production, containers and Kubernetes, cloud infrastructure, orchestration.
  • Frontend and full-stack. TypeScript and React, complex state management, frontend performance work, taking a feature all the way through to the UI.
  • Data at scale. Schema design, query performance, large-scale storage, and analytical stores such as ClickHouse.
  • APIs and developer surfaces. API design, GraphQL query optimization and data fetching strategies, SDKs and client libraries.
  • Event-driven systems. Message queues such as Kafka or PubSub, delivery semantics, idempotency, and retries.

We do not expect one person to have all of this. Depth in a couple of these areas and the range to work outside them is what we are after. Experience in ML infrastructure, MLOps, or developer tools is a plus.

Benefits & conditions

Pulled from the full job description

  • Tuition reimbursement
  • Paid parental leave
  • Employee stock purchase plan
  • Parental leave
  • Health insurance
  • 401(k) matching
  • Paid time off, The base pay and target total cash for this position range from $165,000 to $220,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility), In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings; for roles in other locations, benefits vary and are shared during the hiring process. These include:
  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long-term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health
  • Family-Forming support provided by Carrot
  • Paid Parental Leave
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption

About the company

CoreWeave, the AI Hyperscaler , acquired Weights & Biases to create the most powerful end-to-end platform to develop, deploy, and iterate AI faster. Since 2017, CoreWeave has operated a growing footprint of data centers covering every region of the US and across Europe, and was ranked as one of the TIME100 most influential companies of 2024. By bringing together CoreWeave's industry-leading cloud infrastructure with the best-in-class tools AI practitioners know and love from Weights & Biases, we're setting a new standard for how AI is built, trained, and scaled. The integration of our teams and technologies is accelerating our shared mission: to empower developers with the tools and infrastructure they need to push the boundaries of what AI can do. From experiment tracking and model optimization to high-performance training clusters, agent building, and inference at scale, we're combining forces to serve the full AI lifecycle - all in one seamless platform. Weights & Biases has long been trusted by over 1,500 organizations - including AstraZeneca, Canva, Cohere, OpenAI, Meta, Snowflake, Square,Toyota, and Wayve - to build better models, AI agents and applications. Now, as part of CoreWeave, that impact is amplified across a broader ecosystem of AI innovators, researchers, and enterprises. As we unite under one vision, we're looking for bold thinkers and agile builders who are excited to shape the future of AI alongside us. If you're passionate about solving complex problems at the intersection of software, hardware, and AI, there's never been a more exciting time to join our team.

Apply for this position