> Markdown version of [/jobs/ext/2123166-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/2123166-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - **Company:** VIRTU Financial Inc. - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $200,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Training Data, Java (Programming Language), Airflow, Algorithmic Trading, C++ (Programming Language), Profiling, Extract Transform Load (ETL), Linux, Job Scheduling, Python (Programming Language), Machine Learning, Open Source Technology, Tensorflow, Management of Software Versions, Parquet, Data Processing, High Performance Computing, Pytorch, Build Management, Slurm, Machine Learning Operations - **Published:** August 19, 2026 - **Apply:** https://job-boards.greenhouse.io/virtu/jobs/8457186002 ## About the Role * 5+ years of experience in ML engineering, research infrastructure, or HPC environments * Strong Python engineering skills - you write clean, maintainable, well-tested code that other engineers want to build on. Exposure to C++ in a performance-sensitive context is a plus * Experience building or operating distributed training infrastructure, with working knowledge of how collective communication libraries (NCCL, Horovod, or similar) behave at scale * Practical experience with experiment tracking systems and strong opinions about what good research infrastructure looks like * Comfort working across the Linux systems stack - storage, networking, job scheduling - enough to follow a problem wherever it leads * Excellent communication skills and the ability to work closely with researchers and engineers across disciplines * Intellectually curious and self-driven - you proactively identify problems worth solving, not just problems you've been asked to solve DESIRED, BUT NOT REQUIRED * Experience with on-prem compute environments and job orchestration tools such as Slurm * Familiarity with GPU profiling tools (NSight Systems, PyTorch Profiler) and hands-on experience optimizing GPU memory or compute utilization * Experience with columnar data formats and high-performance data processing tools such as Parquet, Arrow, and Polars * Familiarity with workflow orchestration tools (Prefect, Dagster, or similar) * Prior experience in environments with high-stakes, time-series data at scale. Open to Quantitative Finance, Algorithmic Trading, and Other * Experience contributing to or extending open-source ML frameworks or infrastructure tooling ## Description In this role, you will be responsible for the development of our ML research platform: the systems that manage data and compute, track experiments, and enable researchers to go from idea to result as efficiently as possible. You will work closely with quants and engineers alike and will play a central role in shaping how ML is done at the firm as we scale our capabilities. We mostly use Python, C++ and Java with a variety of open-source tools along with proprietary solutions., * Design and build experiment tracking, job orchestration, and reproducibility infrastructure so researchers can iterate quickly, compare runs reliably, and recover from failures without losing work * Create tools for all stages of the simulation lifecycle including historical back-tests and production monitoring. Add new features to our simulators * Own visibility into GPU cluster utilization - track allocation, surface bottlenecks, and ensure our compute investment is being used effectively * Diagnose and resolve performance issues across training pipelines: data loading throughput, storage I/O, GPU utilization, and inter-node communication in distributed training runs * Build and maintain data pipelines that move financial data from storage into training workflows efficiently, with strong guarantees on correctness and versioning * Develop feature storage and retrieval patterns that support fast, reproducible access to training data at scale * Work directly with researchers to understand friction in their workflows, and build solutions that reduce it - from tooling improvements to infrastructure changes * Collaborate with existing infrastructure engineers on capacity planning, cloud/on-prem tradeoffs, and tooling decisions - this is a collaborative environment, not a siloed one * Stay current with developments in ML infrastructure tooling and bring relevant ideas and tools into our stack where they create genuine value ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Data Fabric in Action - How to enhance a Stock Trading App with ML and Data Virtualization](https://www.wearedevelopers.com/videos/253-data-fabric-in-action-how-to-enhance-a-stock-trading-app-with-ml-and-data-virtualization) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)