> Markdown version of [/jobs/ext/1911885-software-engineer-mlops-machine-learning](https://www.wearedevelopers.com/jobs/ext/1911885-software-engineer-mlops-machine-learning). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, MLOps - Machine Learning - **Company:** Bodner, Aaron C Md - **Location:** San Francisco, CA, United States - **Salary:** $162,000.0 - $216,000.0 - **Contract:** Permanent contract - **Skills:** A/B Testing, Amazon Web Services, Batch Processing, Big Data, Cloud Computing, Information Engineering, Extract Transform Load (ETL), Data Systems, Distributed Computing Environment, Distributed Systems, Monitoring of Systems, Python (Programming Language), Machine Learning, Azure Machine Learning, Software Engineering, Cloud Platform System, Caching, Kubernetes, Production Code, Operational Systems, Machine Learning Operations - **Published:** August 4, 2026 - **Apply:** https://www.sanfranciscogigs.com/job.asp?id=3341047469&tx=FK5956FFI&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Advanced proficiency coding in production-grade Python at an L4 or L5 level * Experience working in an environment where production code directly impacts operations * Ability to build and maintain reliable software across modeling, infrastructure, and automation workflows Distributed Systems Expertise * Strong background in distributed computing, scalable ML infrastructure, and high-performance engineering * Experience building or maintaining systems that support data-intensive and ML workloads * Familiarity with big-data systems, batch processing, caching, and cloud infrastructure Machine Learning / MLOps * Experience implementing, deploying, and productionizing machine-learning algorithms * Hands-on experience with data engineering, distributed training, model monitoring, and experiment tracking * Experience with model retraining, redeployment, serving, and lifecycle management * Strong SQL knowledge and caching experience * Experience with model lifecycle platforms such as SageMaker is a plus and should be confirmed with Fabian as a must-have versus preferred qualification, * Experience implementing, deploying, monitoring, and maintaining machine-learning models in production. * Experience with Kubernetes and cloud infrastructure, preferably AWS. * Familiarity with ML and data technologies such as Kubeflow, Iceberg, Feast, or SageMaker. * Experience with batch prediction, model serving, distributed training, experiment tracking, caching, or feature stores. * Experience building scalable, self-serving infrastructure for machine-learning teams. * Experience integrating ML platforms with broader production or operational systems. * Previous experience in a technically rigorous environment such as a large-scale technology company, infrastructure organization, or high-growth engineering team. * Experience in logistics, transportation, freight, or supply chain is a plus but not required. ## Description As a Software Engineer on Baton's Machine Learning Pod, you will build and maintain the production infrastructure that supports the full machine-learning lifecycle. You will work across production software engineering, distributed systems, MLOps, and model development to help the team bring new models online and operate them reliably at scale. Baton's primary ML infrastructure is established, and the team is now building the next layer of MLOps capabilities on top of that foundation. You will help automate model monitoring, retraining, redeployment, experimentation, and drift detection as the number of production models continues to grow. This is a hands-on individual contributor role for an engineer who can work across both infrastructure and modeling. You will build on the patterns and templates the team has already established, improve integration between the ML platform and Baton's core transportation management platform, and make it easier for engineers to develop, ship, and maintain models end to end. Responsibilities * Build and Expand MLOps Infrastructure: + Build automated capabilities for model monitoring, retraining, redeployment, champion/challenger testing, A/B testing, and drift detection. + Improve experiment tracking and model lifecycle management as the number of production models increases. * Develop and Productionize Machine-Learning Models: + Bring new machine-learning models into production, including developing select models from initial concept through deployment. + Support models across development, deployment, monitoring, maintenance, and iteration. + Build scalable batch-prediction capabilities alongside real-time machine-learning workflows. * Create Self-Serving ML Infrastructure: + Build on existing infrastructure patterns and templates to create reliable and reusable ML workflows. + Make it easier for engineers to ship and maintain models end to end with less manual intervention. + Improve development velocity while maintaining production reliability and operational quality. * Strengthen Distributed ML Systems: + Design and maintain distributed systems that support data-intensive and machine-learning workloads. + Improve the scalability, performance, and reliability of production ML infrastructure. + Contribute to batch processing, caching, data movement, and cloud-native infrastructure. * Connect ML Systems with Baton's Core Platform: + Strengthen the integration between the ML platform and Baton's core transportation management platform. + Replace manual integration workflows with scalable and maintainable infrastructure. + Enable machine-learning capabilities to support transportation workflows and operational decision-making. * Collaborate Across the ML Lifecycle: + Partner with engineers and cross-functional stakeholders to identify opportunities for automation and model productionization. + Contribute across software engineering, ML development, infrastructure, and production operations based on the needs of the team. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Event based cache invalidation in GraphQL](https://www.wearedevelopers.com/videos/433-event-based-cache-invalidation-in-graphql) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)