> Markdown version of [/jobs/ext/2421253-ai-hpc-system-performance-engineer](https://www.wearedevelopers.com/jobs/ext/2421253-ai-hpc-system-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI/HPC System Performance Engineer - **Company:** The Meta Game, Inc. - **Location:** Menlo Park, CA, United States - **Experience:** Expert - **Salary:** $154,000.0 - $217,000.0 - **Contract:** Permanent contract - **Skills:** C (Programming Language), Artificial Intelligence, Application Layers, C++ (Programming Language), Profiling, Computer Programming, Computer Engineering, Network Congestion, Software Debugging, Distributed Systems, Network Topologies, Job Scheduling, Python (Programming Language), Network Architecture, Remote Direct Memory Access, Tensorflow, Software Deployment, Software Requirements Analysis, Systems Architecture, High Performance Computing, Pytorch, Information Technology, Low Latency, Performance Monitor, Programming Languages - **Published:** August 14, 2026 - **Apply:** https://dejobs.org/x/x/6A0A6695CD2A44BBB30FCD596E7FAEF4/job/ ## About the Role 11. 5+ years of coding experience in C, C++, Python, or similar programming languages. Flexible to learn new programming languages 12. Experience profiling and optimizing distributed AI or HPC workloads, including familiarity with GPU interconnects, RDMA networking, and collective communication frameworks such as NCCL or MPI 13. Experience debugging complex, non-reproducible performance issues across multi-layer systems including network fabric, operating system, and application layers 14. Experience designing and implementing performance monitoring systems, including instrumentation, telemetry pipelines, and alerting for large-scale infrastructure 15. Experience driving cross-functional technical projects from requirements definition through production deployment, including communicating performance findings and trade-offs to diverse stakeholders 16. Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience 17. 6+ years of experience in system performance engineering, network infrastructure engineering, or a related field within large-scale distributed computing or HPC environments, 18. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements) 19. Understanding of RDMA congestion control mechanisms on IB and RoCE Networks 20. Understanding of AI training workloads and demands they exert on networks 21. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies 22. Understanding of the latest artificial intelligence (AI) technologies 23. Experience in developing systems software in languages like C++ 24. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews) 25. Experience with machine learning frameworks such as PyTorch and TensorFlow ## Description Meta is building large-scale AI and high-performance computing infrastructure to power next-generation AI research and products. As an AI/HPC System Performance Engineer on the Network Infrastructure Engineering team, you will drive end-to-end performance characterization, bottleneck analysis, and optimization of large-scale AI training and inference clusters. In this role, you will work at the intersection of network fabric design, distributed computing, and AI workload behavior to ensure Meta's HPC systems deliver maximum throughput and efficiency for frontier model development., 1. Profile and benchmark AI training and inference workloads across large-scale HPC clusters to identify network, compute, and memory bottlenecks 2. Develop and maintain performance analysis frameworks and dashboards to track system-level metrics including GPU utilization, network bandwidth, latency, and collective communication efficiency 3. Investigate and resolve performance regressions in distributed AI training environments, including issues related to RDMA fabrics, collective communication libraries, and job scheduling 4. Collaborate with network infrastructure, hardware, and AI research teams to define performance requirements and validate new HPC cluster configurations 5. Design and execute capacity and scalability experiments to inform network topology decisions for AI supercomputing infrastructure 6. Build tooling and automation to continuously monitor HPC system health, detect anomalies, and reduce mean time to mitigation during performance incidents 7. Establish service level objectives for AI cluster network performance and drive cross-functional alignment on reliability and efficiency targets 8. Lead technical design reviews for network and system architecture changes affecting AI workload performance, communicating trade-offs clearly to engineering and product stakeholders 9. Mentor other engineers on HPC performance methodologies, debugging techniques, and instrumentation best practices 10. Leverage AI-assisted workflows to accelerate root cause analysis, automate routine performance reporting, and expand coverage across the HPC stack ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [How We Built a Machine Learning-Based Recommendation System (And Survived to Tell the Tale)](https://www.wearedevelopers.com/videos/752-how-we-built-a-machine-learning-based-recommendation-system-and-survived-to-tell-the-tale) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)