Member of Technical Staff, Capacity & Efficiency Infrastructure - MAI Superintelligence Team

Microsoft
Redmond, United States of America
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 304K

Job location

Mountain View, United States of America

Tech stack

C
Java
JavaScript
C Sharp (Programming Language)
C++
Computer Clusters
Profiling
Nvidia CUDA
Software Debugging
Distributed Computing Environment
Distributed Systems
InfiniBand
Python
Machine Learning
Azure
High Performance Computing
PyTorch
Large Language Models
Generative AI
Gpu Programming
Information Technology

Job description

  • Design, implement, test, and optimize distributed training infrastructure in Python and C++ for large-scale GPU clusters.

  • Build and evolve telemetry systems to provide visibility into infrastructure & ML model performance, utilization, and cost related metrics

  • Profile, benchmark, and debug performance bottlenecks across compute, memory, networking, and storage subsystems

  • Drive architectural improvements across various ML services which deliver measurable efficiency improvements

  • Build and evolve tools to automatically provide insights and recommendations to improve fleet-wide efficiency

  • Optimize collective communication libraries (e.g., NCCL) for emerging NVLink and InfiniBand topologies

  • Partner with ML researchers and infrastructure engineers to understand their plans and future needs and develop plans to balance growth with efficiency

  • Collaborate with hardware teams to optimize for next-generation accelerators (NVIDIA, MAIA, and beyond)

  • Embody our Culture and Values.

Requirements

  • Bachelor's Degree in Computer Science, or related technical discipline AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • OR equivalent experience, * Bachelor's Degree in Computer Science or related technical field AND 10+ years technical engineering experience with coding in languages including, but not limited to, C++ or Python OR Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C++ or Python
  • OR equivalent experience
  • Deep understanding of the fundamentals of GPU architectures and DL/LLM architectures
  • Deep experience in profiling and analyzing performance in large-scale distributed computing systems.
  • Deep experience in profiling and analyzing performance in ML models especially GenAI models
  • Experience with low-level GPU programming (CUDA, Triton, NCCL) and frameworks such as PyTorch or JAX.
  • Experience in leading technical projects and supporting architectural decisions with data.
  • Experience building infrastructure for large-scale machine learning or generative AI workloads.
  • Experience in networking (InfiniBand, NVLink), storage systems, or distributed training parallelisms.
  • Track record of contributing to high-performance computing or large-scale AI infrastructure projects.

About the company

Microsoft is a global technology company headquartered in Redmond, Washington. Our mission is to empower every person and every organization on the planet to achieve more. We develop, license, and support a wide range of software products, services, and devices that help individuals and businesses realize their full potential.

Our flagship products include the Microsoft 365 productivity cloud, Windows operating system, Azure cloud platform, and Dynamics 365 business applications. We are also a leader in areas such as artificial intelligence, cybersecurity, developer tools, and gaming through Xbox and Game Pass.

With operations in more than 190 countries and over 220,000 employees worldwide, Microsoft is committed to responsible innovation, inclusive economic growth, and sustainability. We work closely with governments, industries, and communities to ensure that technology serves the public good and helps address some of the world’s most pressing challenges.

As we celebrate our 50th anniversary in 2025, we continue to look forward—investing in AI, cloud, and quantum computing to shape the future of work, education, and society at large scale.

Apply for this position