> Markdown version of [/jobs/ext/3095693-machine-learning-performance-engineer](https://www.wearedevelopers.com/jobs/ext/3095693-machine-learning-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Performance Engineer - **Company:** TOWER RESEARCH CAPITAL, LLC - **Location:** New York, United States - **Experience:** Experienced - **Salary:** $200,000.0 - **Contract:** Permanent contract - **Skills:** Systems Engineering, C++ (Programming Language), Profiling, Nvidia CUDA, Computer Programming, Extract Transform Load (ETL), Data Transformation, Microprocessors, Distributed Computing Environment, Memory Management, Fault Tolerance, InfiniBand, Python (Programming Language), Linux Kernel, Machine Learning, Remote Direct Memory Access, Tensorflow, Graphics Processing Unit (GPU), Cloud Platform System, Data Ingestion, Pytorch, Delivery Pipeline, Kubernetes, Data Analytics, Slurm, GPT - **Published:** September 26, 2026 - **Apply:** https://www.dice.com/job-detail/bb79c1c3-2c21-433e-9a18-5041097962c0 ## About the Role * 3+ years of experience optimizing machine learning training workloads in high-performance, distributed, or large-scale computing environments. * Deep knowledge of machine learning frameworks such as PyTorch or JAX, including their execution models, compilation paths, autograd systems, and distributed-training capabilities. * Strong programming skills in Python and C++, with experience developing or optimizing performance-critical systems. * Proven experience with GPU kernel development and optimization using technologies such as CUDA, Triton, CUTLASS, cuBLAS, cuDNN, or related libraries. * Strong understanding of GPU architecture, including streaming multiprocessor execution, warp scheduling, tensor cores, and the memory hierarchy from registers through HBM. * Experience with distributed-training technologies and communication libraries such as NCCL, FSDP, DeepSpeed, Megatron-LM, XLA, or equivalent systems. * Proficiency with performance-analysis tools such as Nsight Systems, Nsight Compute, PyTorch Profiler, or comparable tracing and profiling platforms. * Understanding of high-performance networking, storage, and accelerator interconnects, including technologies such as InfiniBand, RDMA, NVLink, or NVSwitch. * Demonstrated ability to benchmark heterogeneous compute platforms and make rigorous, data-driven recommendations about performance, scalability, and cost., * Experience optimizing training workloads for transformer-based, time-series, reinforcement-learning, or other computationally intensive models. * Experience with cluster orchestration and scheduling technologies such as Kubernetes, Slurm, Ray, or similar platforms. * Familiarity with fault-tolerant distributed training, large-scale checkpointing, experiment reproducibility, and GPU-cluster observability. * Practical experience with specialized accelerators, custom hardware, or compiler technologies for machine learning. * Prior experience in financial trading is not required. ## Description You will bridge the gap between quantitative research and high-performance computing, building and optimizing the systems used to train machine learning models at scale. You will focus on accelerating the end-to-end training lifecycle-from data ingestion and distributed execution to kernel performance and hardware utilization-enabling researchers to iterate more quickly across increasingly complex models and datasets. Responsibilities * Training Performance and Benchmarking * Benchmark model-training workloads across CPUs, GPUs, and other accelerator platforms to identify bottlenecks and guide Tower's compute infrastructure decisions. * Develop performance models and standardized benchmarks for measuring throughput, utilization, scalability, and time to convergence. Distributed Training Optimization * Design and optimize distributed training strategies, including data, tensor, pipeline, and model parallelism. * Improve communication efficiency across multi-GPU and multi-node environments by optimizing collective operations, topology awareness, and computation-communication overlap. End-to-End Training Efficiency * Analyze and improve the full training pipeline, including data loading, preprocessing, memory management, forward and backward passes, optimizer execution, checkpointing, and experiment recovery. * Identify bottlenecks across compute, memory, storage, networking, and interconnects to increase accelerator utilization and researcher productivity. GPU Kernel and Framework Development * Develop and optimize GPU kernels and performance-critical framework components for quantitative machine learning workloads. * Integrate specialized libraries, compilers, and execution techniques to improve throughput, memory efficiency, and numerical performance. Model and Numerical Optimization * Apply techniques such as mixed-precision training, gradient accumulation, activation checkpointing, operator fusion, and memory-efficient optimizers. * Evaluate tradeoffs among training speed, numerical stability, reproducibility, model quality, and infrastructure cost. Training Infrastructure * Partner with HPC and infrastructure teams to optimize workload scheduling, resource allocation, observability, fault tolerance, and reproducibility across shared compute environments. * Help define the architecture and tooling required to support large-scale experimentation across on-premises and cloud-based infrastructure. Cross-Functional Collaboration * Work closely with ML Researchers, Quantitative Researchers, HPC Engineers, Systems Engineers, and hardware specialists to translate research requirements into highly efficient training systems. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [30 Golden Rules of Deep Learning Performance](https://www.wearedevelopers.com/videos/11-30-golden-rules-of-deep-learning-performance) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)