> Markdown version of [/jobs/ext/2722984-platform-engineer-data-job-in-austin](https://www.wearedevelopers.com/jobs/ext/2722984-platform-engineer-data-job-in-austin). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Platform Engineer, Data job in Austin - **Company:** Allen Control Systems - **Location:** Austin, TX, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Training Data, Artificial Intelligence, Amazon Web Services, Computer Vision, Encodings, Information Engineering, Data Infrastructure, Data Structures, Linux, Python (Programming Language), Release Management, Simple Data Format, SQL Databases, Management of Software Versions, Data Processing, Pytorch, Information Technology, Data Management, Unreal Engine, Data Generation - **Published:** September 4, 2026 - **Apply:** https://jobs.diversity.com/career/2281405/platform-engineer-data-texas-tx-austin ## About the Role * 3+ years of experience in data engineering or equivalent fields. * Solid understanding of data structures and systems design for orchestrating data-related workflows in a rapidly growing environment. * Proficient in using AWS for data management and processing. * Proficient in Python for scripting and data processing proficient with SQL and Linux. * Educational Background: Bachelor's or Master's degree in Computer Science or a related field. * Proven ability to communicate well across engineering teams, and write and maintain effective documentation. You'll Stand Out: * 5+ years of industry experience. * Experience in image/video data engineering for computer vision projects. * Experience with PyTorch DeepCore. * Experience with Unreal Engine. ## Description We are seeking a Data Platform Engineer who combines expert-level data infrastructure skills with a strong knowledge of AI & Machine Learning principles. In this role, you will go beyond simple data validation scripts you will apply your understanding of model training dynamics to design and implement existing and novel approaches to optimize our datasets. You will build and maintain large-scale image and video pipelines, but with a focus on data curation strategies-such as coreset selection, embedding-based filtering, and automated complexity scoring. You'll partner closely with our ML engineers to orchestrate ingestion, synthetic data generation, and versioned releases, ensuring that every dataset is not only high-integrity and available but strictly optimized to maximize model performance. What You'll Do: * Design and develop a scalable data infastructure, focusing on organization and curation to support continuing increases in data volume and complexity * Design and implement existing and novel approaches to optimize datasets for model training (e.g., hard example mining, class balancing, de-duplication, embedded-based filtering). * Support the data infrastructure required for optimal ingestion, transformation, and storing of datasets * Develop and use synthetic data generation workflows to create realistic synthetic training data for computer vision models. * Design and own end-to-end image and video pipelines for computer vision model training: multi-source ingestion, QA and visualization, standardization, and organization. * Coordinate collection of real-world data coordinate label creation and QA with labelers. * Develop and use data quality tooling: metrics for balance, drift, and annotation error active-learning sampling to target gaps feedback loops from production back to curation. * Implement and own dataset versioning, release management, and lineage and metadata cataloging. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [A Brief History of Data Storage](https://www.wearedevelopers.com/videos/974-a-brief-history-of-data-storage) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)