> Markdown version of [/jobs/ext/3537611-machine-learning-performance-engineer](https://www.wearedevelopers.com/jobs/ext/3537611-machine-learning-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Performance Engineer - **Company:** Long Ridge Partners - **Location:** Hoboken, NJ, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Compilers, Software Code Optimization, Profiling, Data Centers, Microprocessors, Field-Programmable Gate Array (FPGA), Python (Programming Language), Linux Kernel, Machine Learning, Tensorflow, Systems Architecture, Graphics Processing Unit (GPU), Pytorch, Deep Learning, Parallel Computation, Discretization, ONNX (Open Neural Network Exchange) Format, Data Analytics, SAP Ariba, Machine Learning Operations, TensorRT, CUTLASS - **Published:** October 2, 2026 - **Apply:** https://www.disabledperson.com/jobs/75658063-machine-learning-performance-engineer ## About the Role * 2+ years optimizing deep learning inference in latency-sensitive or high-throughput production environments, in any domain * ML frameworks: deep expertise in lower-level ML framework development (PyTorch/JAX), paired with strong Python/C++ skills and a thorough understanding of mixed-precision computation * Kernel development and tooling: proven experience building custom GPU kernels, with deep familiarity with optimization libraries and compilers (Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS) and profiling tools (Nsight Systems, Nsight Compute) * GPU architecture: deep expertise in GPU microarchitecture, including SM execution, warp scheduling, and full memory hierarchy optimization from registers to HBM * Cross-architecture benchmarking: a rigorous, data-driven track record evaluating inference performance across heterogeneous compute architectures * Prior experience in financial trading is not required. Nice to Have * Practical experience targeting and optimizing inference workloads on specialized hardware ecosystems, including FPGAs and ASICs ## Description A leading high frequency trading firm is hiring a Machine Learning Performance Engineer to sit at the intersection of quantitative research and high-performance production systems. In his role, you'll architect inference pipelines that operate at the physical limits of hardware, driving the speed, efficiency, and reliability of ML inference so predictive models consistently achieve microsecond-level latency. GPU usage across the firm's trading teams has grown roughly 100x in the past year as deep learning has moved from a supporting signal to the core of how strategies are built. That growth has outpaced the decision-making around it. Strategies get pushed onto GPUs by default, without anyone systematically asking whether GPU is the right target at all. This role owns that question end to end: benchmark the workload across CPU, GPU, and FPGA, decide the architecture on evidence, then optimize and deploy against it. You will also have the chance to revisit existing models that never reached production, some of which stalled for hardware or deployment reasons, and run them through different environments to determine where they belong. What You'll Do Benchmarking & Strategy * Lead the technical evaluation of inference platforms across CPUs, GPUs, and FPGAs to guide infrastructure deployment decisions * Benchmark trading workloads across architectures before compute is committed, and identify where performance gains actually come from - code-level or hardware-level System Architecture Optimization * Analyze and enhance execution across deep memory hierarchies to maximize resource utilization and parallel processing * Assess and resolve memory subsystem and interconnect bottlenecks across the end-to-end inference lifecycle Infrastructure & Deployment Feasibility * Work with Infrastructure teams to understand the thermal, power, and operational constraints of hardware platforms, and design inference strategies for latency-critical trading strategies that fit within those envelopes * Consider the interaction between trading workloads, compute requirements, hardware selection, and fleet utilization GPU Kernel Development * Develop highly optimized kernels and integrate specialized performance libraries to extract maximum computational throughput from the underlying silicon Model Optimization & Deployment * Implement advanced model reduction techniques - quantization, pruning, distillation - to ensure compact memory footprints and numerical stability * Prioritize optimization for low-latency, event-level inference workloads that meet real-time trading requirements Cross-Functional Collaboration * Partner closely with ML Researchers, HPC Engineers, FPGA Engineers, and Datacenter Engineers to bring target deployments to production ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)