Software Engineer, Systems - AI Training Data Infrastructure

Facebook Inc.
Bellevue, WA, United States
16 days ago
Apply on www.jobmonkeyjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon S3 Big Data C++ (Programming Language) Profiling Data Infrastructure File Systems Distributed Systems Information Lifecycle Management Python (Programming Language) Management of Software Versions
+2 more
Caching Data Pipelines

Job description

  • Own significant components of the AIRStore data path end to end - ingestion, metadata, client, and read path - from design through production operation
  • Attack throughput and latency as a first-class product concern: prefetching, parallelism, caching, and startup cost, measured in GPU utilization and training wall-clock rather than microbenchmarks
  • Build the multi-region and multi-cloud story: make dataset location invisible to the training job, across Meta data centers and third-party clouds
  • Get dataset lifecycle right - TTL, archival, expiration, and deletion - where the correctness bar is absolute in both directions: nothing a live run needs may disappear, and nothing that must be deleted may persist
  • Own a widely embedded client library responsibly: compatibility, rollout safety, and blast-radius control across thousands of callers you don’t control
  • Take real operational ownership. Join the oncall rotation, drive root-cause analysis on incidents affecting production model training, and convert each one into a structural fix rather than a mitigation
  • Partner directly with AI research and training teams, Manifold, Warm Storage, Privacy, and Crypto to land changes that cross system boundaries

Requirements

  • 5+ years of experience building and operating production distributed systems or large-scale data infrastructure
  • Proficiency in a systems language - C++, Rust, or Go - plus Python
  • Demonstrated ability to diagnose performance problems in production: profiling, tracing, and reasoning about I/O, network, and concurrency behavior at scale
  • Experience owning a service in production, including oncall, incident response, and postmortem follow-through
  • Track record of designing and delivering a substantial system component with limited direction

Preferred Qualifications:

  • Experience with data lifecycle, retention, and privacy-driven deletion at scale
  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • Evidence of influencing technical direction beyond your immediate team
  • Background in ML data pipelines - dataloading, checkpointing, dataset versioning, or throughput-bound training I/O
  • Experience with storage systems: object/blob stores, distributed filesystems, caching and prefetch layers, or dataset/columnar formats
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  • Experience running infrastructure across multiple cloud providers or hybrid environments
  • Familiarity with S3-compatible object storage APIs and the practical tradeoffs of compatibility layers

Benefits & conditions

We own the dataset layer that Meta’s largest AI training runs read from. AIRStore and the Anywhere Training substrate are what let a multi-petabyte dataset be written once and read at full throughput from any cluster, in any region, in any cloud - without a copy. Our customers are named model programs, not abstract services: when a training job’s GPUs go idle waiting on I/O, or a dataset isn’t where the scheduler put the job, that is our problem to own and fix.In 2026 this team moved Anywhere Training blob pointers to 100% Manifold residency, cut AIRStore dataset startup time by 10Γ—, drove the migration of AIRStore datasets onto a standard S3 interface, and reclaimed hundreds of petabytes through lifecycle work - all while holding the line on training reliability across dozens of production incidents.

About the company

Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today-beyond the constraints of screens, the limits of distance, and even the rules of physics.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobmonkeyjobs.com
Prepare application

Good distractions

Talks and stories from around this role β€” technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou Β· Coffee With Developers

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn Β· World Congress 2026 Europe

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy Β· LIVE

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann Chris Heilmann +3 Β· LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 Β· LIVE

2:33 min

Maintaining prompt structures for prefix caching

Douglas Reiser Douglas Reiser Β· Europe 2026 Virtual

Videos

See all

Related articles

See all