Research Scientist, Artificial Intelligence

Facebook Inc.
Menlo Park, CA, United States
7 days ago
Apply on www.jobmonkeyjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Program Optimization Profiling Computer Engineering Software Design Documents Memory Management Linux Kernel Machine Learning Performance Tuning High Performance Computing Pytorch Delivery Pipeline
+4 more
Reliability of Systems Information Technology Data Analytics Machine Learning Operations

Job description

  • Lead the design and execution of TPU performance optimization research, including kernel development, memory optimization, and compute efficiency improvements
  • Develop and optimize Pallas kernels for large-scale model training and inference on TPU architectures
  • Drive model optimization techniques including Mixture of Experts (MoE), tensor parallelism, pipeline parallelism, and other distributed training strategies
  • Optimize first party models within Meta’s native PyTorch stack, ensuring efficient integration with XLA compilation and TPU execution
  • Identify and resolve complex technical challenges in model training efficiency, inference latency, and system reliability that require novel approaches
  • Define and drive multi-quarter research roadmaps for TPU optimization, aligning project milestones with broader organizational goals
  • Establish rigorous experimentation frameworks for performance benchmarking, including metric selection, profiling methodology, and data-driven optimization decisions
  • Translate research findings into production-ready optimizations by collaborating with engineering teams on deployment pipelines and reliability at scale
  • Communicate research findings and technical trade-offs clearly through publications, design documents, and presentations to both technical and non-technical audiences
  • Mentor other researchers and engineers on TPU optimization techniques, providing structured feedback on technical direction and experimental rigor

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 8+ years of experience in machine learning systems, model optimization, or high-performance computing research
  • Experience with TPU architecture and performance optimization, including profiling, kernel development, and memory management
  • Experience with XLA compilation, graph optimization, and low-level performance tuning for accelerator hardware
  • Experience developing and optimizing large-scale distributed training systems, including parallelism strategies such as data, tensor, and pipeline parallelism
  • Experience with PyTorch and its integration with accelerator backends
  • Experience communicating complex technical findings in writing, including technical reports, design documents, or peer-reviewed publications, * Experience developing custom kernels using Pallas or similar kernel authoring frameworks for TPU or GPU
  • Demonstrated track record of transitioning performance research into deployed systems used at significant scale
  • PhD in Computer Science, Machine Learning, Computer Architecture, or a related technical field, or equivalent depth of research experience
  • Publication record in systems for ML venues such as MLSys, OSDI, SOSP, or related AI conferences such as NeurIPS, ICML, or ICLR
  • Experience with Mixture of Experts (MoE) architectures and their optimization for efficient training and inference
  • Experience optimizing production-scale models with billions of parameters

About the company

Meta AI Research is at the forefront of advancing foundational and applied artificial intelligence, developing breakthroughs that power products used by billions of people and shape the future of human-computer interaction. We are seeking a Research Scientist at the Staff level (IC6) with deep expertise in TPU performance optimization, large-scale model training, and systems-level machine learning. In this role, you will lead high-impact research on model efficiency and optimization for first party models within Meta’s native PyTorch stack, collaborating across research and engineering teams to drive AI capabilities that define Meta’s next generation of products and platforms., Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today-beyond the constraints of screens, the limits of distance, and even the rules of physics.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobmonkeyjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

1:15 min

Overcoming the challenges of modifying Linux kernel code

Ayesha Kaleem · World Congress 2023

2:01 min

Exploring foundational expertise in traditional optimization and machine learning

Eric Enge · Coffee With Developers

2:37 min

Optimizing technical profiles for AI sourcing and recruitment

Mina Golesorkhi Mina Golesorkhi · World Congress 2026 Europe

1:55 min

Role of the Linux kernel in handling processes

Mohammed Aboullaite Mohammed Aboullaite · World Congress 2024

Videos

See all

Related articles

See all