> Markdown version of [/jobs/ext/541462-sr-data-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/541462-sr-data-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Data Infrastructure Engineer - **Company:** Evolv Technology - **Location:** Waltham, MA, United States - **Experience:** Expert - **Salary:** $129,000.0 - $209,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, C++ (Programming Language), Cloud Computing, Cloud Database, Cloud Engineering, Cloud Storage, Data as a Services, Information Engineering, Data Governance, Data Infrastructure, Data Integrity, Data Security, Data Systems, Software Debugging, Distributed Computing Environment, Distributed Systems, Python (Programming Language), Machine Learning, Operational Databases, Cloud Services, Smart Devices, Software Engineering, Data Streaming, Management of Software Versions, Cloud Platform System, Data Ingestion, Model Validation, Information Technology, Data Lineage, Data Management, Functional Programming, Data Pipelines - **Published:** June 11, 2026 - **Apply:** https://www.juju.com/job/00000000g6txtm ## About the Role + Bachelor's or Master's degree in Computer Science, Data Engineering, Software Engineering, or related field. + 2-3+ years of experience building production data pipelines and data platforms that support AI/ML models. + Strong proficiency in Python, C++ and distributed data processing frameworks. + Hands-on experience with AWS services including S3, EC2, SageMaker, and Glue. + Experience designing data systems that support large-scale ML training and experimentation. + Knowledge of data governance, access control, and lifecycle management. + Experience collaborating with ML, data science, operations, and cloud teams. Preferred Qualifications: + Experience building pipelines spanning edge devices and cloud systems. + Background working with large-scale sensor, image or IoT data. + Familiarity with data labeling tools and annotation workflows. + Experience implementing dataset versioning, lineage, and reproducibility systems. + Understanding of privacy, compliance, or regulated data environments. + Experience supporting global, multi-region data platforms. Example Problems You Will Own + Design a resilient global ingestion pipeline aggregating sensor data from millions of devices. + Build ML-ready data services enabling easy discovery, versioning, and consumption of datasets. + Implement automated validation and cleaning workflows that dramatically reduce bad data. + Define and enforce lifecycle and governance policies across research and production datasets. ## Description Join Evolv as Senior Data Infrastructure Engineer in the Machine Learning & Sensors organization, responsible for building and operating the scalable, secure, and reliable data pipelines that power our AI/ML research and production systems. In this role, you will own the end-to-end data lifecycle-from collection on thousands to millions of edge devices, through cloud ingestion and processing, into a centralized data factory enabling model training, evaluation, and continuous improvement. Data is the backbone of our mission to deliver best-in-class AI-based weapon detection systems. You will ensure that data flows seamlessly across geographies, devices, and cloud systems while meeting strict requirements for quality, privacy, security, and scale. This role is ideal for someone who thrives at the intersection of distributed systems, cloud pipelines, and ML-driven data needs. Success in the Role: What performance outcomes will you work toward in the first 6-12 months? In the first 30 days: + Develop a deep understanding of existing edge-to-cloud data pipelines and deployment environments. + Review current data ingestion flows, governance policies, and cloud infrastructure. + Assess pain points in data reliability, quality, and operational scalability. + Build relationships with AI/ML, data science, field operations, and cloud engineering teams. + Design and prototype data processing pipelines (both cloud and edge) Within the first three months: + Design and implement improvements to core ingestion, validation, and processing pipelines. + Deploy scalable data pipeline with AWS-based components (S3, EC2, Lambda, Glue, Step Functions, SageMaker integrations). + Introduce automated validation workflows to detect corruption, missing metadata, or malformed data. + Design and implement automated model evaluation, model training and model improvement pipeline to speed up experiments + Partner with field operations to improve data reliability, observability, and coverage across deployments. By the end of the first year: + Own the entire lifecycle of mission-critical data pipelines supporting AI/ML research and production. + Architect next-generation edge-to-cloud data systems that scale across millions of devices. + Define and enforce data governance frameworks including retention, access control, privacy, and lineage. + Enable ML teams to rapidly experiment through high-quality, discoverable, versioned datasets. The Work: What type of work will you be doing? What assignments, requirements, or skills will you be performing on a regular basis? End-to-End Data Pipeline Ownership: + Design, build, and maintain both research and production data pipelines spanning edge devices, cloud services, and centralized data platforms. + Own the full data lifecycle: collection, ingestion, processing, obfuscation, versioning, access, retention, and retirement. + Edge-to-Cloud Data Flow: + Develop resilient ingestion pipelines capable of handling variable connectivity and device heterogeneity. + Support secure data transfer from the field to cloud storage systems. + Collaborate with field ops to enhance data coverage, observability, and operational robustness. + Data Quality, Governance & Compliance: + Implement privacy-preserving transformations and obfuscation pipelines. + Build automated cleaning/validation steps to remove duplicates, detect corruption, and validate metadata. + Establish data lineage, retention policies, and access controls ensuring compliance and traceability. Data Services for AI/ML: + Provide scalable data services for model training, evaluation, and research experimentation. + Support continuous data refresh and retraining workflows. + Integrate with data labeling services and annotation workflows. + Enable efficient access patterns for large-scale ML workloads. AWS-Based Cloud Infrastructure: + Build and optimize pipelines using AWS services (S3, EC2, SageMaker, Lambda, Glue, Step Functions). + Design for cost-efficiency, performance, and reliability at scale. Collaboration & Feedback Loops: + Partner with AI/ML engineers, scientists, and data scientists to understand data requirements. + Translate feedback into automated improvements in data collection, labeling, and consumption. + Support cross-functional teams in exploratory analysis and debugging data issues. Scaling the Data Factory: + Design and manage data schema, data versioning and data factory updates + Architect systems that scale globally across millions of devices. + Ensure the data platform remains flexible for research and reliable for production operations. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [The Data Mesh as the end of the Datalake as we know it](https://www.wearedevelopers.com/videos/156-the-data-mesh-as-the-end-of-the-datalake-as-we-know-it) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)