> Markdown version of [/jobs/ext/3545856-software-engineer-systems-ml](https://www.wearedevelopers.com/jobs/ext/3545856-software-engineer-systems-ml). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Systems ML - **Company:** The Meta Game, Inc. - **Location:** Menlo Park, CA, United States - **Experience:** Experienced - **Salary:** $121,992.0 - $181,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Automation of Tests, C++ (Programming Language), Profiling, Code Review, Nvidia CUDA, Computer Programming, Computer Engineering, Distributed Computing Environment, Distributed Systems, Python (Programming Language), Linux Kernel, Machine Learning, Tensorflow, Software Engineering, AI Infrastructure, Graphics Processing Unit (GPU), High Performance Computing, Pytorch, Multi-Agent Systems, Prompt Engineering, Gpu Programming, Information Technology, Low Latency, ROCm, Machine Learning Operations - **Published:** October 1, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3413454846&tx=CT3533TTI&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role 1. Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta 2. 2+ years of experience in software engineering with a focus on machine learning systems, AI infrastructure, or high-performance computing 3. Experience developing or optimizing ML training or inference pipelines using frameworks such as PyTorch, TensorFlow, or equivalent 4. Experience with distributed computing architectures and large-scale systems design for ML workloads 5. Experience programming in C++ and Python for performance-critical systems 6. Experience using profiling and performance analysis tools to identify and resolve bottlenecks in ML or compute-intensive systems, 1. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements) 2. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews) 3. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies 4. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements) 5. Experience with GPU programming using CUDA, ROCm, or equivalent hardware accelerator kernel development 6. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews) 7. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies 8. Experience optimizing large-scale ranking or recommendation model inference on AI accelerator hardware such as GPUs or TPUs 9. Experience with hardware-software co-design, including numerics optimization and SIMD or vectorization techniques ## Description Meta is seeking a Software Engineer to join our Systems ML Engineering team, focused on building and optimizing the machine learning infrastructure that powers Meta's products at massive scale. In this role, you will design and develop high-performance ML systems, working across the full stack from model training and inference pipelines to hardware-aware optimizations. You will collaborate with researchers, platform engineers, and product teams to accelerate ML workloads and improve the efficiency of AI infrastructure that serves billions of users., 1. Design, build, and optimize large-scale ML training and inference systems, including distributed computing frameworks and hardware-accelerated pipelines 2. Develop and maintain high-performance ML infrastructure components in C++ and Python, ensuring reliability, scalability, and low-latency execution 3. Identify and resolve performance bottlenecks across the ML stack using profiling, instrumentation, and benchmarking tools 4. Architect and evaluate trade-offs in ML system design, including memory bandwidth, compute utilization, and I/O throughput 5. Partner with research and product teams to translate ML model requirements into efficient infrastructure solutions 6. Write automated tests covering expected behaviors, failure modes, and error paths for ML infrastructure components, and build monitoring and alerting for production anomalies 7. Contribute to staged rollout strategies using feature flagging and experimentation frameworks to safely deploy ML system changes 8. Produce accurate technical documentation when introducing new ML infrastructure features or updating existing systems 9. Participate in on-call rotations to investigate and mitigate production incidents affecting ML serving systems, and contribute detailed retrospectives 10. Provide constructive code review feedback to other engineers and contribute to engineering programs that improve the health and quality of the ML infrastructure codebase ## Related Videos - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Agentic employees in world's most downloaded FinTech app](https://www.wearedevelopers.com/videos/100123-agentic-employees-in-world-s-most-downloaded-fintech-app) - [Enhancing Workload Security in Kubernetes](https://www.wearedevelopers.com/videos/356-enhancing-workload-security-in-kubernetes) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [ I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.](https://www.wearedevelopers.com/magazine/773-i-gave-a-video-editor-more-autonomy-than-a-trading-bot-on-purpose) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)