> Markdown version of [/jobs/ext/3665848-software-engineer-systems-ml-tooling](https://www.wearedevelopers.com/jobs/ext/3665848-software-engineer-systems-ml-tooling). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Systems ML Tooling - **Company:** The Meta Game, Inc. - **Location:** New York, NY, United States - **Experience:** Experienced - **Salary:** $154,003.0 - $217,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Profiling, Nvidia CUDA, Computer Engineering, Software Debugging, Software Design Documents, Linux, Programming Tools, Firmware, Python (Programming Language), Linux Kernel, Open Source Technology, Software Architecture, Quick EMUlator (QEMU), Tensorflow, Memory Leaks, Software Deployment, Software Engineering, System Programming, Toolchain, Diagnostic Tools, Scripting, Pytorch, Large Language Models, Prompt Engineering, Reliability of Systems, Build Management, Perf (Linux), Information Technology, MLIR (Multi-Level Intermediate Representation) - **Published:** October 10, 2026 - **Apply:** https://dejobs.org/x/x/D411F2EFC70A4B6B8239BFC6F7C55450/job/ ## About the Role 10. Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience 11. 4+ years of experience in systems software engineering, performance engineering, developer tooling, or a closely related field 12. Experience building debugging, sanitizer, profiling, simulation, or diagnostic tools for complex software/hardware systems 13. Proficiency in C++ and Python, including low-level systems programming and scripting for tool automation 14. Experience working across multiple layers of a system stack (compiler, runtime, OS/driver, hardware)Track record of leading the technical design and delivery of tooling or infrastructure projects through to production deployment 15. Cross-stack debugging skills, with the ability to trace issues across application, runtime, and hardware boundaries, 16. 5+ years of experience in systems software, developer tooling, or accelerator software development (or equivalent with advanced degree) 17. Experience with compiler-based instrumentation and runtime shadow-memory techniques underlying sanitizers (LLVM instrumentation passes, interceptors, allocator red zones) 18. Track record of building developer tools adopted by large engineering populations, ideally with contributions to open-source projects (LLVM sanitizers, gdb, Valgrind, Triton, etc.) 19. Familiarity with ML framework internals (PyTorch graph execution, torch.compile, operator dispatch) and AI compiler stacks (MLIR, LLVM, TVM, Triton) 20. Experience with accelerator ecosystems (GPU/CUDA, TPU, custom ASICs), including memory analysis, performance profiling, and runtime debugging using their toolchains (cuda-gdb, nsight-compute, nsight-systems) 21. Experience working at the device-software boundary - accelerator runtime/driver internals, DMA and memory-mapped device behavior, on-device fault handling or firmware-assisted error reporting 22. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements) 23. Experience with Linux debugging and profiling infrastructure (gdb, perf, eBPF, ftrace, coredump analysis, hardware performance counters) and familiarity with binary formats and debugging metadata (ELF/DWARF) 24. Experience with Linux kernel and driver-level debugging, hands-on experience with sanitizer and dynamic-analysis technology, or building comparable memory checking tools 25. Experience with functional or cycle-approximate simulation techniques - instruction-set interpreters/simulators, timing and performance models - and with simulator or virtual-platform frameworks (gem5, QEMU, AModel, BModel) 26. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies 27. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews) ## Description Meta is seeking a Software Engineer to join the MTIA (Meta Training & Inference Accelerator) Software Tooling team, which develops and maintains the tooling ecosystem for Meta's in-house AI accelerator ASICs. The Tooling team provides debugging, profiling, memory analysis, and monitoring capabilities for the whole MTIA Ecosystem, redefining ML accelerator tooling by leveraging Meta's full-stack ownership from silicon specs to fleet observability.In this role, you will design and build developer tools across the MTIA tooling ecosystem, with a primary focus on debugging, sanitizer technology, and fault isolation for AI workloads running on MTIA hardware at scale. You will work at the intersection of compilers, runtime, hardware, and ML frameworks, collaborating with cross-functional partners to deliver a high-quality developer experience for Meta's custom AI accelerators.This role can focus on either debugging and sanitizing technology or simulation infrastructure for pre-silicon and post-silicon model bring-up, depending on the candidate's background and team needs., 1. Lead the design and development of MTIA's debugging and sanitizer tooling, including live debugging, core dump, and memory sanitizers for accelerator workloads 2. Build debugging capabilities spanning graph-mode debugging, kernel-level diagnostics, and multi-rank fault isolation 3. Contribute broadly to the MTIA SW tooling infrastructure - profiling, performance debugging, memory analysis, monitoring, and reliability analysis for training and inference workloads 4. Own significant components end-to-end, from technical design through implementation, production rollout, and long-term maintenance 5. Collaborate closely with the MTIA compiler, runtime, kernel, and hardware teams to instrument the software stack with the hooks, sanitizers, and debuggers required 6. Drive AI-native tooling approaches, leveraging automation and LLM-guided diagnostics to reduce time-to-root-cause 7. Partner with internal product teams across advertising, recommendations, and generative AI to understand developer pain points and prioritize tooling investments 8. Communicate architectural decisions through design documents and cross-team reviews 9. share debugging methodology and best practices with engineers across the MTIA stack