Machine Learning Infrastructure Engineer, Technology

POINT72, L.P.
New York, NY, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$185,000.0 - $300,000.0
Working hours
Regular working hours
Job source

Tech stack

Airflow Amazon Web Services Microsoft Azure C++ (Programming Language) Profiling Software Debugging Digital Architecture Distributed Systems Python (Programming Language) Key Management Machine Learning Reinforcement Learning
+6 more
Google Cloud Kubernetes Infrastructure Automation Frameworks Information Technology Machine Learning Operations Terraform

Job description

  • Design and implement high-performance infrastructure to support large-scale generative AI and machine learning workloads, enabling faster model iteration and real business impact
  • Design and operate distributed systems for model training, hyperparameter tuning, inference, and data preprocessing pipelines to deliver reliable end-to-end machine learning (ML) workflows
  • Collaborate with ML researchers and engineers to produce models, optimizing compute utilization, training throughput, and inference latency
  • Develop and automate deployment, orchestration, and CI/CD pipelines for models and data workflows using container orchestration and infrastructure-as-code (IaC)
  • Implement observability, monitoring, and cost-management strategies for GPU and accelerator compute environments to maintain predictable performance and spend
  • Evaluate, integrate, and benchmark emerging hardware and software technologies across cloud and on-prem environments to improve scalability and throughput
  • Drive security, compliance, and operational runbooks for GenAI infrastructure including access controls, secrets management, and incident response procedures
  • Troubleshoot, profile, and optimize performance across GPU and CPU compute stacks to remove bottlenecks and increase reliability
  • Document architecture, operational practices, and mentor engineers to expand team capability and accelerate adoption of production-ready GenAI infrastructure

Requirements

  • Bachelor’s or master’s degree in computer science, electrical engineering, or a related technical field
  • 3-7 years of experience building and maintaining scalable compute or machine learning infrastructure systems
  • Deep understanding of distributed systems, container orchestration (Kubernetes), and public cloud platforms such as AWS, Google Cloud Platform, or Azure
  • Hands-on experience with machine learning operations and infrastructure tools such as MLflow, Ray, Airflow, Kubeflow, and Terraform
  • Strong understanding of reinforcement learning concepts and their infrastructure implications
  • Proficiency in Python and systems-level programming in one or more languages such as Go, C++, or Rust
  • Strong debugging, performance profiling, and optimization skills across GPU and CPU compute stacks
  • Experience implementing monitoring, observability, and cost-optimization for GPU/accelerator-based compute environments
  • Excellent collaboration and communication skills with a systems-thinking mindset
  • Commitment to the highest ethical standards

Benefits & conditions

3.63.6 out of 5 stars 55 Hudson Yards 10th FL, New York, NY 10001 $185,000 - $300,000 a year, Pulled from the full job description

  • Tuition reimbursement
  • Parental leave
  • Health insurance
  • 401(k) matching
  • Family leave
  • Matching gift program, We invest in our people, their careers, their health, and their well-being. When you work here, we provide:
  • Fully-paid health care benefits
  • Generous parental and family leave policies
  • Mental and physical wellness programs
  • Volunteer opportunities
  • Non-profit matching gift program
  • Support for employee-led affinity groups representing women, minorities and the LGBT+ community
  • Tuition assistance
  • A 401(k) savings program with an employer match and more, The annual base salary range for this role is $185,000-$300,000 (USD) , which does not include discretionary bonus compensation or our comprehensive benefits package. Actual compensation offered to the successful candidate may vary from posted hiring range based upon geographic location, work experience, education, and/or skill level, among other things.

About the company

As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications.

As a member of Point72’s Technology team, we encourage and support your professional development from day one-helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity-all while delivering real business impact for our multi-billion-dollar global business., Point72 Asset Management is a global firm led by Steven Cohen that invests in multiple asset classes and strategies worldwide. Resting on more than a quarter-century of investing experience, we seek to be the industry’s premier asset manager through delivering superior risk-adjusted returns, adhering to the highest ethical standards, and offering the greatest opportunities to the industry’s brightest talent. We’re inventing the future of finance by revolutionizing how we develop our people and how we use data to shape our thinking. For more information, visit www.Point72.com/working-here

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all