> Markdown version of [/jobs/ext/1451041-machine-learning-performance-engineer-job-in-new-york](https://www.wearedevelopers.com/jobs/ext/1451041-machine-learning-performance-engineer-job-in-new-york). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Performance Engineer job in New York - **Company:** Jane - **Location:** New York, NY, United States - **Contract:** Permanent contract - **Skills:** Computer Clusters, Nvidia CUDA, Software Debugging, InfiniBand, Machine Learning, System Programming, Computer Network Technologies, Real Time Systems, Syntactically Awesome Style Sheets (SASS) - **Published:** July 26, 2026 - **Apply:** https://jobs.diversity.com/career/1506309/Machine-Learning-Performance-Engineer-New-York-Ny-New-York ## About the Role If you've never thought about a career in finance, you're in good company. Many of us were in the same position before working here. If you have a curious mind and a passion for solving interesting problems, we have a feeling you'll fit right in. There's no fixed set of skills, but here are some of the things we're looking for: * An understanding of modern ML techniques and toolsets * The experience and systems knowledge required to debug a training run's performance end to end * Low-level GPU knowledge of PTX, SASS, warps, cooperative groups, Tensor Cores, and the memory hierarchy * Debugging and optimization experience using tools like CUDA GDB, NSight Systems, NSight Compute * Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN, and cuBLAS * Intuition about the latency and throughput characteristics of CUDA graph launch, tensor core arithmetic, warp-level synchronization, and asynchronous memory loads * Background in Infiniband, RoCE, GPUDirect, PXN, rail optimization, and NVLink, and how to use these networking technologies to link up GPU clusters * An understanding of the collective algorithms supporting distributed GPU training in NCCL or MPI * An inventive approach and the willingness to ask hard questions about whether we're taking the right approaches and using the right tools If you're a recruiting agency and want to partner with us, please reach out to agency-partnerships@janestreet.com. ## Description We are looking for an engineer with experience in low-level systems programming and optimization to join our growing ML team. Machine learning is a critical pillar of Jane Street's global business. Our ever-evolving trading environment serves as a unique, rapid-feedback platform for ML experimentation, allowing us to incorporate new ideas with relatively little friction. Your part here is optimizing the performance of our models - both training and inference. We care about efficient large-scale training, low-latency inference in real-time systems, and high-throughput inference in research. Part of this is improving straightforward CUDA, but the interesting part needs a whole-systems approach, including storage systems, networking, and host- and GPU-level considerations. Zooming in, we also want to ensure our platform makes sense even at the lowest level - is all that throughput actually goodput? Does loading that vector from the L2 cache really take that long? ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [How Machine Learning is turning the Automotive Industry upside down](https://www.wearedevelopers.com/videos/61-how-machine-learning-is-turning-the-automotive-industry-upside-down) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [The weekly developer show: Boosting Python with CUDA, CSS Updates & Navigating New Tech Stacks](https://www.wearedevelopers.com/videos/1293-the-weekly-developer-show-boosting-python-with-cuda-css-updates-navigating-new-tech-stacks) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)