> Markdown version of [/jobs/ext/2054362-machine-leaning-performance-engineer-inference](https://www.wearedevelopers.com/jobs/ext/2054362-machine-leaning-performance-engineer-inference). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Leaning Performance Engineer (Inference) - **Company:** CAPITAL TOWERS II, INC. - **Location:** New York, United States - **Experience:** Experienced - **Salary:** $200,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Compilers, Program Optimization, Profiling, Data Centers, Microprocessors, Field-Programmable Gate Array (FPGA), Python (Programming Language), Linux Kernel, Tensorflow, Systems Architecture, Graphics Processing Unit (GPU), Pytorch, Deep Learning, Parallel Computation, ONNX (Open Neural Network Exchange) Format, Data Analytics, SAP Ariba, TensorRT - **Published:** August 14, 2026 - **Apply:** https://www.dice.com/job-detail/0a31c8ec-cbb6-40bd-963d-d1d0b463eaba ## About the Role * 2+ years of experience optimizing deep learning inference in latency-sensitive or high-throughput production environments, in any domain. * ML Frameworks: Deep expertise in lower-level ML framework development (PyTorch/JAX), paired with strong Python/C++ skills and a thorough understanding of mixed-precision computation. * Kernel Development & Optimization Tooling: Proven experience in custom GPU kernel development. Deep familiarity with advanced optimization libraries and compilers (e.g., Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS) as well as profiling tools (e.g., Nsight Systems, Nsight Compute). * GPU Architecture Mastery: Deep expertise in GPU microarchitecture, encompassing SM execution, warp scheduling, and full memory hierarchy optimization (registers to HBM). * Cross-Architecture Benchmarking: Proven record of rigorous, data-driven approach to evaluating inference performance across heterogeneous compute architectures. * Bonus: Practical experience targeting and optimizing inference workloads on specialized hardware ecosystems, including FPGAs and ASICs. * Prior experience in financial trading is not required. ## Description As part of Tower Research's Core Engineering team, you will bridge the gap between quantitative research and high-performance production systems, architecting inference pipelines that operate at the physical limits of hardware. Your objective will be to drive the speed, efficiency, and reliability of our ML inference pipelines to their absolute limits, ensuring our predictive models consistently achieve microsecond-level latency. Responsibilities: * Benchmarking & Strategy: * Lead the technical evaluation of diverse inference platforms - ranging across CPUs, GPUs, and FPGAs - to guide Tower's infrastructure deployment decisions. System Architecture Optimization: * Analyze and enhance execution across deep memory hierarchies to maximize resource utilization and parallel processing. You will assess and resolve memory subsystem and interconnect bottlenecks across the end-to-end inference lifecycle. Infrastructure & Deployment Feasibility: * Collaborate with Infrastructure teams to understand thermal, power, and operational constraints of hardware platforms to design inference strategies for our latency-critical trading strategies that fit within those envelopes. GPU Kernel Development: * Develop highly optimized kernels and integrate specialized performance libraries to extract maximum computational throughput from the underlying silicon. Model Optimization & Deployment: * Implement advanced model reduction techniques (quantization, pruning, distillation) to ensure compact memory footprints and numerical stability. Prioritize optimization for low-latency, event-level inference workloads to meet real-time trading requirements. Cross-Functional Collaboration: * Collaborate closely with ML Researchers, HPC Engineers, FPGA Engineers, and Datacenter Engineers to bring to fruition target deployments. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Just-in-time Compilation in JVM](https://www.wearedevelopers.com/videos/240-just-in-time-compilation-in-jvm) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)