Software Engineer, Systems - AI Training Data Infrastructure
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+2 more
Job description
- Own significant components of the AIRStore data path end to end - ingestion, metadata, client, and read path - from design through production operation
- Attack throughput and latency as a first-class product concern: prefetching, parallelism, caching, and startup cost, measured in GPU utilization and training wall-clock rather than microbenchmarks
- Build the multi-region and multi-cloud story: make dataset location invisible to the training job, across Meta data centers and third-party clouds
- Get dataset lifecycle right - TTL, archival, expiration, and deletion - where the correctness bar is absolute in both directions: nothing a live run needs may disappear, and nothing that must be deleted may persist
- Own a widely embedded client library responsibly: compatibility, rollout safety, and blast-radius control across thousands of callers you donβt control
- Take real operational ownership. Join the oncall rotation, drive root-cause analysis on incidents affecting production model training, and convert each one into a structural fix rather than a mitigation
- Partner directly with AI research and training teams, Manifold, Warm Storage, Privacy, and Crypto to land changes that cross system boundaries
Requirements
- 5+ years of experience building and operating production distributed systems or large-scale data infrastructure
- Proficiency in a systems language - C++, Rust, or Go - plus Python
- Demonstrated ability to diagnose performance problems in production: profiling, tracing, and reasoning about I/O, network, and concurrency behavior at scale
- Experience owning a service in production, including oncall, incident response, and postmortem follow-through
- Track record of designing and delivering a substantial system component with limited direction
Preferred Qualifications:
- Experience with data lifecycle, retention, and privacy-driven deletion at scale
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Evidence of influencing technical direction beyond your immediate team
- Background in ML data pipelines - dataloading, checkpointing, dataset versioning, or throughput-bound training I/O
- Experience with storage systems: object/blob stores, distributed filesystems, caching and prefetch layers, or dataset/columnar formats
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
- Experience running infrastructure across multiple cloud providers or hybrid environments
- Familiarity with S3-compatible object storage APIs and the practical tradeoffs of compatibility layers
Benefits & conditions
We own the dataset layer that Metaβs largest AI training runs read from. AIRStore and the Anywhere Training substrate are what let a multi-petabyte dataset be written once and read at full throughput from any cluster, in any region, in any cloud - without a copy. Our customers are named model programs, not abstract services: when a training jobβs GPUs go idle waiting on I/O, or a dataset isnβt where the scheduler put the job, that is our problem to own and fix.In 2026 this team moved Anywhere Training blob pointers to 100% Manifold residency, cut AIRStore dataset startup time by 10Γ, drove the migration of AIRStore datasets onto a standard S3 interface, and reclaimed hundreds of petabytes through lifecycle work - all while holding the line on training reliability across dozens of production incidents.
About the company
Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today-beyond the constraints of screens, the limits of distance, and even the rules of physics.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role β technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Navigating the AI Shift
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
How Much FAANG Companies Actually Pay Software Engineers in 2025