> Markdown version of [/jobs/ext/3382983-software-engineer-systems-ai-training-data-infrastructure](https://www.wearedevelopers.com/jobs/ext/3382983-software-engineer-systems-ai-training-data-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Systems - AI Training Data Infrastructure - **Company:** Facebook Inc. - **Location:** Bellevue, WA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon S3, Big Data, C++ (Programming Language), Profiling, Data Infrastructure, File Systems, Distributed Systems, Information Lifecycle Management, Python (Programming Language), Management of Software Versions, Caching, Data Pipelines - **Published:** September 17, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/28029609/Software-Engineer-Systems-Ai-Training-Data-Infrastructure-Washington-Bellevue-7418 ## About the Role * 5+ years of experience building and operating production distributed systems or large-scale data infrastructure * Proficiency in a systems language - C++, Rust, or Go - plus Python * Demonstrated ability to diagnose performance problems in production: profiling, tracing, and reasoning about I/O, network, and concurrency behavior at scale * Experience owning a service in production, including oncall, incident response, and postmortem follow-through * Track record of designing and delivering a substantial system component with limited direction Preferred Qualifications: * Experience with data lifecycle, retention, and privacy-driven deletion at scale * Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies * Evidence of influencing technical direction beyond your immediate team * Background in ML data pipelines - dataloading, checkpointing, dataset versioning, or throughput-bound training I/O * Experience with storage systems: object/blob stores, distributed filesystems, caching and prefetch layers, or dataset/columnar formats * Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews) * Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements) * Experience running infrastructure across multiple cloud providers or hybrid environments * Familiarity with S3-compatible object storage APIs and the practical tradeoffs of compatibility layers ## Description * Own significant components of the AIRStore data path end to end - ingestion, metadata, client, and read path - from design through production operation * Attack throughput and latency as a first-class product concern: prefetching, parallelism, caching, and startup cost, measured in GPU utilization and training wall-clock rather than microbenchmarks * Build the multi-region and multi-cloud story: make dataset location invisible to the training job, across Meta data centers and third-party clouds * Get dataset lifecycle right - TTL, archival, expiration, and deletion - where the correctness bar is absolute in both directions: nothing a live run needs may disappear, and nothing that must be deleted may persist * Own a widely embedded client library responsibly: compatibility, rollout safety, and blast-radius control across thousands of callers you don't control * Take real operational ownership. Join the oncall rotation, drive root-cause analysis on incidents affecting production model training, and convert each one into a structural fix rather than a mitigation * Partner directly with AI research and training teams, Manifold, Warm Storage, Privacy, and Crypto to land changes that cross system boundaries ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Machine Learning for Software Developers (and Knitters)](https://www.wearedevelopers.com/videos/154-machine-learning-for-software-developers-and-knitters) - [In-Memory Computing - The Big Picture](https://www.wearedevelopers.com/videos/626-in-memory-computing-the-big-picture) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)