Sr. Software Backend Engineer, Machine Learning Data Pipelines
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+25 more
Job description
In this role, you will focus on backend development in Python and Go for critical internal engineering tools and machine learning platform capabilities worldwide. You will design, develop, and maintain the backend systems that power training-data pipelines for AI agents, agentic workflows, and real-time anomaly detection. That includes the APIs, state management, scheduling, and observability needed for reliable, production-grade agent behavior. You will work closely with cross-functional teams to ensure seamless integration and deployment, troubleshoot and resolve complex technical issues, and stay up to date with industry trends and emerging technologies to drive innovation in Industry 4.0, equipment sensor data, vehicle data, and production ML/agent systems. What You’ll Do
- Design, develop, and maintain scalable and efficient backend systems using Python and Go, including internal web applications, engineering tools, real-time software systems, machine learning data pipelines, and agentic workflow backends
- Build and operate training-data pipelines and platform services that enable AI agents and agentic workflows-including support for stateful execution, failure recovery, retries, and long-running orchestration
- Enable production anomaly detection and support migrations and multi-region expansion for ML and agent-related services
- Develop system architectures, databases, and integration systems that meet application requirements and integrate cleanly with frontend, ML, and agent systems
- Write elegant, scalable APIs (REST and gRPC) for machine learning services, AI agents, tool/function calling, applications, systems, and databases
- Design and develop real-time software systems using technologies such as Apache Kafka and Apache Flink to feed both anomaly detection and agentic workflows
- Establish scheduling, metadata, and observability patterns so training, inference, and agent workflow data remain trustworthy, on time, auditable, and recoverable
- Create logging and monitoring systems to track system and agent health and optimize performance, scalability, reliability, and resilience
- Collaborate with cross-functional teams (frontend, DevOps, data engineering, and ML) to ship and operate backend systems that power interactive and agent-driven solutions through rapid iteration
Requirements
- Master’s Degree in Computer Science, Computer Engineering, or equivalent experience, with at least 5 years of relevant backend software development experience with Python and Go
- Expert-level proficiency designing and developing scalable, efficient, and secure APIs using Python and Go; strong knowledge of REST and gRPC; experience with frameworks such as Flask, FastAPI (Python) and Gin or Echo (Go)
- Experience building and operating data/ML pipelines and backend services that support AI agents or agentic workflows (stateful execution, scheduling, retries/failure recovery, metadata, and observability)
- Experience with reliable delivery of training and inference data for ML and agent systems
- Experience with SQL (e.g., MySQL, Microsoft SQL), NoSQL (e.g., MongoDB, Elasticsearch), and column-store databases (e.g., ClickHouse), with strong query optimization skills
- Experience with Docker and Kubernetes; ability to thrive in a fast-paced environment with multiple tight deadlines
- Strong attention to detail and diligence in unit and integration testing; excellent problem-solving skills
- Strong communication and collaboration skills with cross-functional teams and multiple stakeholders
Benefits & conditions
Along with competitive pay, as a full-time Tesla employee, you are eligible for the following benefits at day 1 of hire:
- Medical plans > plan options with $0 payroll deduction
- Family-building, fertility, adoption and surrogacy benefits
- Dental (including orthodontic coverage) and vision plans, both have options with a $0 paycheck contribution
- Company Paid (Health Savings Accounts) HSA Contribution when enrolled in the High-Deductible medical plan with HSA
- Healthcare and Dependent Care Flexible Spending Accounts (FSA)
- 401(k) with employer match, Employee Stock Purchase Plans, and other financial benefits
- Company paid Basic Life, AD&D
- Short-term and long-term disability insurance (90 day waiting period)
- Employee Assistance Program
- Sick and Vacation time (Flex time for salary positions, Accrued hours for Hourly positions), and Paid Holidays
- Back-up childcare and parenting support resources
- Voluntary benefits to include: critical illness, hospital indemnity, accident insurance, theft & legal services, and pet insurance
- Weight Loss and Tobacco Cessation Programs
- Tesla Babies program
- Commuter benefits
- Employee discounts and perks program
Expected Compensation $140,000 - $252,000/annual salary + cash and stock awards + benefits
About the company
The Manufacturing Quality Data Engineering team is a high-impact, high-priority, and high-visibility team laser-focused on addressing safety-critical issues and expanding critical services to Gigafactories worldwide. As part of Tesla’s Vehicle Engineering organization, you will have access to a vast array of data sources across design, manufacturing, and vehicle data, empowering you to design, develop, and deploy innovative data services, automation, and machine learning tools into production.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
How to Become an AI Engineer
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again