> Markdown version of [/jobs/ext/3096595-infrastructure-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3096595-infrastructure-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure Reliability Engineer - **Company:** Bright Vision Technologies - **Location:** Austin, TX, United States (Remote available) - **Experience:** Expert - **Salary:** $125,000.0 - $170,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Big Data, Code Review, Continuous Integration, Data Cleansing, Data Deduplication, Information Engineering, Data Infrastructure, Extract Transform Load (ETL), Data Systems, Distributed Systems, Document-Oriented Databases, Java Virtual Machine (JVM), Python (Programming Language), Machine Learning, Open Source Technology, Azure Machine Learning, Software Engineering, System on a Chip, Management of Software Versions, Data Processing, Apache Spark, Caching, Storage Technologies, Information Technology, Integration Frameworks, Free and Open-Source Software, Machine Learning Operations, Data Pipelines - **Published:** September 26, 2026 - **Apply:** https://www.careerjet.com/jobad/us28cc6d4754b916c3ea5a3bdf24170115 ## About the Role This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential., * Bachelor's or Master's degree in Computer Science or a related field. * Six or more years of data engineering experience, with significant work supporting ML or AI workloads. * Strong proficiency in Python and at least one JVM or systems language. * Deep experience with modern data processing frameworks such as Spark, Ray, or Beam. * Hands-on experience operating petabyte-scale storage and pipeline systems. * Strong understanding of distributed systems, data modeling, and storage formats. * Experience with dataset versioning, lineage, and reproducibility for ML workflows. * Familiarity with high-throughput data loading for accelerator-based training. * Strong software engineering practices including testing, CI/CD, and code review. * Excellent communication and cross-functional collaboration skills., * Experience with multimodal datasets at large scale. * Familiarity with data quality tooling and dataset evaluation methodology. * Exposure to privacy-preserving data systems and regulated data handling. * Open-source contributions to data infrastructure projects. * Experience supporting frontier model training pipelines. ## Description We are seeking an Infrastructure Reliability Engineer to build and operate the large-scale data systems that power modern AI training and evaluation pipelines. The role combines deep data engineering expertise with a strong understanding of AI workloads, focusing on ingestion, transformation, quality assurance, lineage, and high-throughput delivery of data to training jobs across diverse modalities. The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and clear understanding of how data infrastructure choices propagate into model quality and training efficiency., * Design and operate large-scale data pipelines supporting AI training, evaluation, and continual improvement workflows. * Build ingestion systems for diverse modalities including text, image, audio, video, and structured signals. * Implement data cleaning, deduplication, filtering, and quality assurance at petabyte scale. * Develop dataset versioning, lineage, and provenance tracking systems suitable for reproducible training. * Build high-throughput data loading systems that maximize GPU utilization during training. * Implement labeling workflows, active learning pipelines, and human-in-the-loop data improvement systems. * Design storage architectures balancing cost, throughput, and latency across data tiers. * Build evaluation dataset construction pipelines with strict integrity and contamination controls. * Implement data privacy, redaction, and consent enforcement throughout the pipeline. * Collaborate with ML researchers and engineers to align data systems with model development needs. * Drive observability of data quality, drift, and pipeline health across the AI data estate. * Optimize cost and performance through compression, format selection, and caching strategies. * Document data systems, schemas, and operational procedures for broad internal use. * Stay current with AI data infrastructure research and emerging open-source tools., Custom SoCs (System on Chip) live at the heart of AWS Machine Learning servers. As a member of the Cloud-Scale Machine Learning Acceleration team you'll be responsible for the desi… + 14 days ago