> Markdown version of [/jobs/ext/1985243-software-engineer-ii-autonomy-data](https://www.wearedevelopers.com/jobs/ext/1985243-software-engineer-ii-autonomy-data). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, II - Autonomy Data - **Company:** Torc Robotics, Inc. - **Location:** Blacksburg, VA, United States - **Experience:** Experienced - **Salary:** $139,000.0 - $166,800.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon S3, Build Automation, Cloud Database, Cloud Storage, Computer Engineering, Data as a Services, Data Validation, Information Engineering, Data Files, Data Infrastructure, Data Visualization, Software Debugging, Document-Oriented Databases, Python (Programming Language), Machine Learning, Open Source Technology, Software Engineering, SQL Databases, Management of Software Versions, Parquet, Data Logging, Cloudformation, Data Lakes, Infrastructure Automation Frameworks, Information Technology, Data Lineage, Terraform, Data Pipelines - **Published:** August 8, 2026 - **Apply:** https://www.careerjet.com/job/us66c95010634b15f754a20d1973390d23/eaa ## About the Role * Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, Electrical Engineering or a related field with 4+ years of data engineering experience or a Master's with 2+ years. * Strong proficiency in Python and SQL, with demonstrated ability to build production-quality data pipelines * Experience with cloud data infrastructure (AWS preferred: S3, Glue Athena, redshift, or equivalent) and infrastructure-as-code tools (Terraform, Cloud Formation). * Solid understanding of data partitioning strategies and columnar storage formats (Parquet, Orc, etc.) * Experience building and operating data pipelines that process time-series and binary data. * Proven ability to evaluate and integrate open-source tooling when appropriate versus building from scratch. * Good instincts for delivering data quality through first-class implementations of monitoring, validation and lineage tracking. Bonus points! * Experience with autonomous vehicles, robotics, or other sensor-driven autonomous systems. * Deep experience with Foxglove or Rerun beyond basic playback, e.g. building custom extensions or integrating them into a structured log review or annotation QA workflow. * Familiarity with the MCAP CLI and/or python library and experience converting MCAP data to columnar data formats for further querying and processing. * Experience with data curation for ML training, e.g. diversity sampling, pseudo-labeling, and dataset versioning. ## Description Torc is hiring an Autonomy Data Engineer Level 2 to help design, build and operate the data infrastructure that powers our autonomy program. You will build the pipelines, storage systems, and tooling that turn raw vehicle sensor logs into the curated, structured datasets that our perception, planning and simulation engineers depend on. This is a high-ownership role on a lean team. Moving large scale sensor data reliably from vehicles operating in demanding environments and making it quickly available for model training is a difficult and high-impact problem to solve. What You'll Do: * Data Lake and Ingestion Pipeline * Contribute to the design and organization of the program's data lake, including schema definitions, partitioning strategy and metadata indexing. * Build and maintain end-to-end pipelines that ingest high-bandwidth sensor logs from vehicles into cloud storage with high reliability and tolerant of ad-hoc and intermittent connectivity mechanisms. * Implement data validation and integrity checks that can detect corrupted information, missing sensors, and inconsistent calibration prior to the data being processed by downstream systems. * Implement retention, tiering and lifecycle policies for data to balance storage costs with development value. * Dataset Curation and Labeling Infrastructure * Build tooling to query raw logs to produce curated training and evaluation datasets. * Build automation to run cost-effective pseudo-labeling workflows at the scale of data ingest. * Implement data quality and model performance metrics that are used to direct labeling effort toward the highest-value examples. * Autonomy Data Visualization * Deploy and maintain data visualization tooling to support log review, annotation QA, and autonomy debugging workflows. * Build integrations between the visualization tooling and the data lake so engineers can navigate from a dataset entry or model failure directly to the origin log data * Work with autonomy engineers to define and surface custom visualization panels and implement metrics for analyzing unstructured operating environments. * Build dashboards that provide the autonomy engineers visibility into data coverage by terrain type, operating environment and geographic region. * Cross-functional Collaboration * Establish and document data contracts between the data services and model training consumers. * Partner with perception, planning and embedded engineers across the data lifecyle: from shaping the logging schemas and collection triggers to defining the dataset interfaces that supply model training and evaluation. * Follow and help evolve data engineering standards, best practices, and tooling choices for an innovative and fast-paced team. * Contribute to the data roadmap and surface findings to senior technical leadership. ## Related Videos - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) - [Tracking vehicles at scale](https://www.wearedevelopers.com/videos/1999-tracking-vehicles-at-scale) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data)