> Markdown version of [/jobs/ext/215460-software-engineer-ml-infrastructure](https://www.wearedevelopers.com/jobs/ext/215460-software-engineer-ml-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, ML Infrastructure - **Company:** Snap Inc. - **Location:** New York, NY, United States - **Salary:** $157,000.0 - $235,000.0 - **Contract:** Permanent contract - **Skills:** Training Data, Java (Programming Language), Artificial Intelligence, Big Data, C++ (Programming Language), Computer Programming, Distributed Systems, Python (Programming Language), Machine Learning, Tensorflow, Scala (Programming Language), Software Engineering, Pytorch, Apache Spark, Scikit Learn, Information Technology, Apache Flink, Data Management, Machine Learning Operations - **Published:** May 30, 2026 - **Apply:** https://www.dice.com/job-detail/12c86e05-fc69-4d3c-867f-84be41ed25ec ## About the Role * Strong programming skills in Python, Java, Scala or C++ * Strong problem-solving skills with a focus on system performance, scalability, and efficiency * Good understanding of distributed systems and the infrastructure components of large-scale ML * Experience with big data processing frameworks such as Spark, Flink, or Ray * Ability to collaborate and work well with others * Proven track record of operating highly-available systems at significant scale * Ability to proactively learn new concepts and apply them at work * Adaptability in learning and applying evolving AI systems and tools to remain at the forefront of engineering trends and modern development practices, * Bachelor's degree in a technical field such as computer science or equivalent experience * 2+ years of post-Bachelor's software development experience; or Master's degree in a technical field + 1+ year of post-grad software development experience; or PhD in a relevant technical field * Experience building large scale production machine learning systems, distributed systems or big data processing Preferred Qualifications: * Masters/PhD in a technical field such as computer science or equivalent industry experience * Experience working with ML Training platforms or optimizing AI model inference * Familiarity with ML frameworks such as TensorFlow, PyTorch, Caffe2, Spark ML, scikit-learn, or related frameworks ## Description * Design and optimize infrastructure systems for machine learning workloads at scale and drive reliability and efficiency improvements across Snapchat's ML Infrastructure * Build and enhance feature generation and serving pipelines that power online inferencing and offline training data generation * Develop high-performance inference systems to ensure fast and efficient AI model serving * Build infrastructure to perform scalable ML model training, evaluation, and inference in the cloud * Develop high-performance inference systems to ensure fast and efficient AI model serving * Build comprehensive data management systems for scalable data collection, labeling, processing, and evaluation * Work closely with ML engineers to deploy cutting-edge models into production * Utilize AI tools and high velocity engineering workflows to design and ship scalable services while upholding rigorous standards for code correctness, security, and production ready quality code ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)