> Markdown version of [/jobs/ext/2057997-software-engineer-autotagging](https://www.wearedevelopers.com/jobs/ext/2057997-software-engineer-autotagging). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - AutoTagging - **Company:** Torc Robotics, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $160,800.0 - $193,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Big Data, Databases, Continuous Integration, Information Engineering, Distributed Computing Environment, Github, Python (Programming Language), Operational Databases, Software Engineering, Unstructured Data, Parquet, Datadog, Data Logging, Data Ingestion, Grafana, Apache Spark, Cloudformation, Information Technology, Machine Learning Operations, Cloudwatch, Terraform, Data Pipelines, Databricks - **Published:** August 14, 2026 - **Apply:** https://www.dice.com/job-detail/77f472de-2097-4b1d-a74a-e95f331438da ## About the Role * BS or MS in Computer Science, Engineering, or a related field, with 5+ years of software engineering experience, including production data pipeline or ML infrastructure work. * Strong Python skills, with experience building and maintaining production data or ML pipelines. * Hands-on CI/CD experience, GitHub Actions required. * Required experience with Databricks for large scale data processing and orchestration. * Required experience with AWS, including infrastructure-as-code (Terraform or CloudFormation) for provisioning distributed processing infrastructure. * Experience processing large scale time series or unstructured datasets. * Experience with observability tooling (e.g., Datadog, Grafana, CloudWatch) for production pipeline monitoring and alerting. * Experience integrating and deploying ML models into production systems - serving, monitoring, and rollback, not just training. * Strong communication skills to work across ML, perception, and simulation teams. Bonus Points! * Familiarity with auto-labeling pipelines, VLMs, or zero-shot classification for scenario extraction. * Experience with distributed compute frameworks such as Ray, Spark, or Daft. * Familiarity with robotics data formats (ROS bags, MCAP) and columnar storage formats (Parquet, Arrow). * Experience with model serving frameworks such as vLLM or SGLang. * Familiarity with scenario description standards like Pegasus layers. ## Description * Integrate and deploy automated event-tagger into production pipelines, running and monitoring tagging tasks at scale across petabytes of vehicle log data. * Build and maintain the data engineering pipelines that organize, structure, and catalog tagged scenario data into the observations database. * Own CI/CD for the Auto Tagger pipeline using GitHub Actions, keeping deployments reliable, tested, and repeatable. * Write production grade code in Python across the pipeline, from data ingestion and transformation through model integration and deployment. * Build and operate on Databricks for large scale data processing, interactive querying, and pipeline orchestration. * Design, deploy, and scale AWS infrastructure (as code) to support high-volume, distributed processing of vehicle log pipelines - working with structured/tagged outputs and metadata. * Instrument pipelines with logging, metrics, and alerting; own on-call response for tagging job failures and data quality regressions. * Partner with ML engineers on the team to take tagging and classification models from development into a scalable, monitored production pipeline. * Ensure data quality and metadata integrity as tagged events move from raw logs into the observations database used by perception, simulation, and systems teams. * Troubleshoot and improve pipeline performance, reliability, and cost as data volume and model complexity grow. ## Related Videos - [Robots are coming into the wild! Full-Stack Robotics Engineers, be ready!](https://www.wearedevelopers.com/videos/479-robots-are-coming-into-the-wild-full-stack-robotics-engineers-be-ready) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)