Senior Software Engineer, ML Platform

NxT Level
San Francisco, CA, United States
21 days ago
Apply on www.wayup.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

A/B Testing Airflow Amazon Web Services Data Architecture Python (Programming Language) Machine Learning Azure Machine Learning Service Design Software Deployment Software Engineering Software Systems SQL Databases
+14 more
Strategies of Testing Real Time Systems Feature Engineering Apache Spark Model Validation Caching Parallel Computation Rate Limiting Pyspark Apache Kafka Machine Learning Operations Video Streaming Unsupervised Learning Databricks

Job description

Our client is hiring a Senior Software Engineer, ML Platform to lead the evolution of its machine learning platform. This person will design, build, and maintain the core systems that support model development, production deployment, batch inference, real-time inference, feature stores, observability, and underwriting infrastructure. You’ll work closely with Data Science and Platform Engineering to turn research workflows into reliable software systems. This is a strong fit for an engineer who enjoys building developer-friendly platforms, creating clean abstractions, and owning infrastructure that powers real business decisions. What You’ll Do

  • Own and evolve the company’s ML platform end-to-end
  • Turn data science notebooks into reusable, tested, production-ready software components
  • Build libraries, pipelines, templates, SDKs, and CLIs that help data scientists move faster
  • Create developer-friendly abstractions for feature definition, model training, evaluation, deployment, and monitoring
  • Build and scale low-latency real-time model serving infrastructure
  • Expand batch ML inference systems across scheduling, parallelism, cost controls, observability, failure handling, and rollback
  • Own and improve the feature store, including offline and online feature definitions
  • Design systems for high read/write throughput and consistent offline/online semantics
  • Instrument training and inference workflows for latency, throughput, accuracy, drift, data quality, and cost
  • Build alerting, dashboards, and observability systems for platform health
  • Support production underwriting systems across batch and real-time workflows
  • Partner with Data Science on model interfaces, SLAs, safety checks, and product integrations
  • Drive incident response, postmortems, and long-term reliability improvements

Requirements

  • 5+ years of software engineering experience
  • Experience building ML platform, MLOps, model training, model deployment, or feature pipeline systems
  • Strong Python experience
  • Strong software design, testing, and platform engineering fundamentals
  • Proficiency with SQL
  • Hands-on experience with Spark or PySpark
  • Strong understanding of ML fundamentals, including probability, statistics, supervised and unsupervised learning, feature engineering, validation strategies, model evaluation, drift, stability, and monitoring
  • Experience with modern data and ML infrastructure such as AWS, Databricks, MLflow, model registries, model serving, Airflow, or similar orchestration tools
  • Experience building real-time systems, including service design, caching, rate limiting, backpressure, and low-latency architecture
  • Experience building batch pipelines at scale
  • Practical knowledge of feature store concepts, including offline and online stores, backfills, point-in-time correctness, experiment tracking, and evaluation frameworks
  • Strong ownership mindset and proactive approach to platform reliability
  • Excellent communication and collaboration skills across engineering and data science teams Bonus Experience

  • Deep Databricks experience, including MLflow, workflows, lakehouse architecture, or model serving
  • Experience with feature stores such as Tecton, Feast, or similar platforms
  • Experience with streaming technologies such as Kafka or Kinesis
  • Experience in fintech, risk, lending, underwriting, or regulated financial systems
  • Familiarity with model safety checks, rejection flows, override flows, and auditability
  • Experience with A/B testing platforms, shadow deployments, canary releases, and automated rollback
  • Experience building low-latency inference systems

About the company

About Our Client Our client is building financial infrastructure that helps small businesses access the capital and products they need to grow. Their platform uses data, machine learning, and modern underwriting systems to power financial products at scale. As the company continues to expand, the infrastructure behind model experimentation, training, evaluation, inference, and retraining is becoming increasingly critical. This is an opportunity to join a high-impact infrastructure team and own the ML platform that enables data scientists to safely and quickly ship high-quality models into production., + Own a critical ML platform that powers underwriting and other ML-driven products

  • Build infrastructure that helps data scientists ship models safely and quickly
  • Work across real-time inference, batch inference, feature stores, model evaluation, and platform observability
  • Partner closely with Data Science and Platform Engineering on high-impact systems
  • Build developer-friendly tools that create leverage across the technical organization
  • Work on meaningful infrastructure tied directly to financial access for small businesses
  • Step into a senior role with end-to-end ownership over core ML platform systems

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.wayup.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

2:33 min

Maintaining prompt structures for prefix caching

Douglas Reiser Douglas Reiser · Europe 2026 Virtual

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all