> Markdown version of [/jobs/ext/3077611-senior-ai-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/3077611-senior-ai-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI Infrastructure Engineer - **Company:** Anduril Industries - **Location:** Costa Mesa, CA, United States - **Experience:** Expert - **Salary:** $191,000.0 - $253,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Computer Vision, C++ (Programming Language), Concurrent Computing, Extract Transform Load (ETL), Distributed Computing Environment, Distributed Data Store, Distributed Systems, Python (Programming Language), Machine Learning, Motion Planning, Regression Testing, Azure Machine Learning, Robotic Automation Software, Software Engineering, AI Infrastructure, Reinforcement Learning, Cloud Platform System, Pytorch, Delivery Pipeline, Large Language Models, Model Validation, Generative AI, Kubernetes, Slurm, Machine Learning Operations, Data Pipelines, Docker - **Published:** September 25, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9193423/senior-ai-infrastructure-engineer ## About the Role * 5+ years of software engineering experience with demonstrated success in building and operating production-scale machine learning infrastructure or distributed systems. * Proficiency in Python, Go, or C++, with a strong grasp of software engineering fundamentals, systems design, and concurrent programming. * Hands-on experience with container orchestration (Docker, Kubernetes) and distributed training frameworks (e.g., PyTorch Distributed, Ray, Slurm, or Megatron-LM). * Experience building and maintaining distributed data pipelines handling large-scale unstructured or multi-modal datasets. * Track record of owning projects end-to-end-from technical design to production deployment and operational monitoring. * Eligible to obtain and maintain an active U.S. Top Secret security clearance., * Experience deploying ML infrastructure, model serving, or artifacts in secure, air-gapped, or regulated environments (e.g., IL5/IL6, GovCloud). * Hands-on experience profiling GPU/accelerator workloads, resolving hardware/network bottlenecks, and optimizing compute utilization. * Experience supporting workloads for Large Language Models, Generative AI, or Reinforcement Learning (RL) pipelines. * Experience with multi-tenant cluster management, including fair scheduling, GPU slicing, and quota enforcement. * Familiarity with production ML observability frameworks, including data drift detection and automated evaluation pipelines. ## Description The Air Dominance & Strike team at Anduril develops aerial and multi-domain robotic systems. The team is responsible for taking products like Fury (unmanned fighter jet) and Barracuda (air-breathing cruise missile) from concept to product. The team also develops Lattice for Mission Autonomy, Anduril's premier software platform that enables masses of Fury, Barracuda, and other first and third party robots to collaborate across various missions. We work in close coordination with specialist teams like Perception, Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems facing our customers. We are looking for software engineers and roboticists excited about creating a powerful autonomy software stack that includes computer vision, motion planning, SLAM, controls, estimation, and secure communications., We are looking for a Senior AI Infrastructure Engineer to build, scale, and optimize the end-to-end machine learning platform that powers Anduril's autonomous systems. In this role, you will own critical components of our ML platform and MLOps tooling. You will build and operate the infrastructure required to train, evaluate, host, and serve complex AI models (including LLMs, computer vision, and RL agents) across cloud environments and air-gapped, edge-deployed networks. Working closely with AI Research Scientists and Platform Engineers, you will eliminate friction in model development, optimize hardware utilization, and ensure robust delivery of models into safety-critical operational environments. WHAT YOU'LL DO * Build, optimize, and maintain scalable training, orchestration, and experimentation infrastructure to accelerate state-of-the-art model development. * Identify and resolve bottlenecks in the ML lifecycle by developing tooling for experiment tracking, automated profiling, and hyperparameter tuning. * Implement and scale robust data pipelines (ETL) to process multi-modal data (video feeds, radar, flight telemetry, and simulation logs) captured from physical assets and test sites. * Deploy high-throughput, low-latency model serving frameworks optimized for both cloud environments and resource-constrained, air-gapped tactical edge hardware. * Develop robust CI/CD pipelines for ML models, including automated regression testing, validation benchmarks, and safe rollout/rollback strategies. * Implement pipelines for model evaluation, validation, and reinforcement learning alignment loops (RLHF/DPO) to ensure predictability and safety in mission-critical deployments. * Partner with AI Researchers and Computer Vision engineers to translate modeling requirements into scalable, reusable infrastructure. * Mentor peers, conduct thorough design and code reviews, and champion engineering best practices across the team., To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind: * No Financial Requests: Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process. Our legitimate recruitment is entirely free for candidates. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)