> Markdown version of [/jobs/ext/1481079-senior-software-engineer-ml-workflows-weights-biases](https://www.wearedevelopers.com/jobs/ext/1481079-senior-software-engineer-ml-workflows-weights-biases). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, ML Workflows - Weights & Biases - **Company:** Weights & Biases - **Location:** Bellevue, WA, United States - **Experience:** Expert - **Salary:** $165,000.0 - $220,000.0 - **Contract:** Permanent contract - **Skills:** Query Performance, C (Programming Language), Java (Programming Language), Application Programming Interfaces (APIs), C Sharp (Programming Language), C++ (Programming Language), Cloud Computing, Programming Tools, Distributed Systems, Python (Programming Language), Query Optimization, Queueing Systems, Azure Machine Learning, Software Engineering, Systems Architecture, TypeScript, Rust (Programming Language), ReactJS, Technical Debt, Backend, Data Layers, Event Driven Architecture, Kubernetes, Apache Kafka, Graphql, Machine Learning Operations, Front End Software Development, Vertica, Api Design, Dropbox, Code Restructuring - **Published:** July 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a203618ee18b5cf4 ## About the Role We expect all of these: * 6+ years of software engineering experience, including time leading significant technical initiatives or architectural changes * Strong proficiency in Go, or in another compiled language (C, C++, C#, Rust, Java) with the ability to ramp on Go quickly * Working proficiency in Python * Comfort diving into complex systems you did not write, to diagnose and fix difficult bugs across service boundaries * A track record of mentoring engineers and raising the technical bar around you * Strong communication, including the ability to explain complex technical decisions to different audiences And real depth in at least one or two of these, which is also how we work out which team you join: * Distributed systems and infrastructure. Designing and operating services in production, containers and Kubernetes, cloud infrastructure, orchestration. * Frontend and full-stack. TypeScript and React, complex state management, frontend performance work, taking a feature all the way through to the UI. * Data at scale. Schema design, query performance, large-scale storage, and analytical stores such as ClickHouse. * APIs and developer surfaces. API design, GraphQL query optimization and data fetching strategies, SDKs and client libraries. * Event-driven systems. Message queues such as Kafka or PubSub, delivery semantics, idempotency, and retries. We do not expect one person to have all of this. Depth in a couple of these areas and the range to work outside them is what we are after. Experience in ML infrastructure, MLOps, or developer tools is a plus. ## Description Reporting to the Senior Engineering Manager for ML Workflows, this Senior Software Engineer will own the architecture and evolution of core systems in the Weights & Biases platform. ML Workflows owns what customers use to move models and data through W&B. Artifacts is the versioned data layer that every object logged to W&B runs through. Registry sits on top of it as the organization-wide source of truth for what is production-ready, and it is used by teams at Pinterest, Dropbox, and Canva. Automations is the event and action layer the rest of the platform builds on. Launch is execution infrastructure that runs customer ML and agent workloads. You will own systems end to end, from design through production reliability, working across the W&B backend, the Python SDK, and the frontend. You will partner closely with product managers, designers, and ML platform engineers, and you will talk directly to enterprise ML teams about their workflows. We hire senior engineers into the organization and match them to a team based on their strengths and interests, which we discuss with you during the interview process. W&B is part of CoreWeave, so these products are being built onto CoreWeave infrastructure, and there is real room to shape how that happens. What You'll Do * Own the architecture and evolution of a core platform area, designing systems that scale to billions of artifacts and high-volume event and job throughput. * Dive deep into system architecture to find optimization opportunities, solve complex bugs, and make the difficult calls that balance short-term delivery against long-term platform health. * Work across the stack as the problem requires, from Go backend services and GraphQL APIs to the Python SDK and the TypeScript frontend. * Lead migrations and re-architecture of systems carrying live customer traffic, keeping the experience seamless for the customers who depend on them. * Identify and refactor technical debt, improving maintainability and developer productivity. * Establish engineering patterns and best practices that hold up as the platform grows. * Collaborate with product and design to turn complex user requirements into clean technical implementations. * Work directly with customers to understand their ML workflow challenges, and bring that feedback into the product cycle. * Mentor engineers and raise the technical bar across the team. ## Related Videos - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [GraphQL + Apollo + Next.js: A Lovely Trio](https://www.wearedevelopers.com/videos/311-graphql-apollo-next-js-a-lovely-trio) - [The state of MLOps - machine learning in production at enterprise scale](https://www.wearedevelopers.com/videos/369-the-state-of-mlops-machine-learning-in-production-at-enterprise-scale) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [Event based cache invalidation in GraphQL](https://www.wearedevelopers.com/videos/433-event-based-cache-invalidation-in-graphql) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)