> Markdown version of [/jobs/ext/2864530-software-engineer-genai-frameworks](https://www.wearedevelopers.com/jobs/ext/2864530-software-engineer-genai-frameworks). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, GenAI Frameworks - **Company:** The Meta Game, Inc. - **Location:** Bellevue, WA, United States - **Salary:** $121,992.0 - $181,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), A/B Testing, Artificial Intelligence, Automation of Tests, C++ (Programming Language), Profiling, Software Quality, Code Review, Computer Engineering, Data Cleansing, Data Infrastructure, Extract Transform Load (ETL), Distributed Computing Environment, Distributed Systems, Python (Programming Language), Machine Learning, Azure Machine Learning, Software Engineering, Data Logging, Data Processing, Pytorch, Delivery Pipeline, Apache Spark, Parallel Computation, Information Technology, Low Latency, Apache Flink, Production Code, Apache Kafka, Machine Learning Operations, Data Pipelines - **Published:** September 12, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3387082691&tx=JP7771FFJ&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role 1. Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta 2. 2+ years of experience in software engineering with a focus on machine learning systems, distributed systems, or high-performance data infrastructure 3. Experience writing production-quality code in Python and at least one compiled language such as C++ or Java, including performance-sensitive systems code 4. Experience designing and implementing components of ML training or inference pipelines, including data preprocessing, model execution, or serving systems 5. Experience with distributed computing concepts such as data parallelism, model parallelism, or parameter server architectures as applied to ML workloads 6. Experience writing automated tests, building logging and alerting, and participating in production incident response for large-scale systems, 1. Experience optimizing ML workloads for throughput and latency, including profiling GPU or CPU utilization, memory bandwidth, and communication overhead in distributed training or inference settings 2. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews) 3. Experience with ML framework internals such as PyTorch or JAX, including custom operator development, execution graph optimization, or compiler integration 4. Experience with stream or batch data processing systems such as Apache Spark, Flink, or Kafka in the context of ML feature pipelines 5. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements) 6. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies 7. 1+ years production experience in GenAI post-training, RLHF, RLVR, comms/collectives, parallelism, accuracy evaluation, and performance profiling 8. Familiarity with model quantization, mixed-precision training, or other techniques for reducing compute and memory costs in production ML systems ## Description Meta is building the infrastructure and systems that power machine learning at scale across its family of products, including Feed, Reels, Ads ranking, and generative AI services. The Systems ML Engineering team sits at the intersection of ML and systems software, designing and optimizing the training and inference pipelines, distributed execution frameworks, and data processing systems that enable researchers and product teams to iterate quickly and deploy reliably. In this role, you will contribute to the full lifecycle of ML systems software - from designing scalable data pipelines and distributed training infrastructure to optimizing model serving latency and throughput - directly impacting the quality and speed of ML-powered experiences for billions of people., 1. Design and implement scalable systems for distributed ML training and inference, including data ingestion pipelines, feature processing, and model serving infrastructure 2. Develop and optimize ML platform components such as training orchestration, gradient communication, and checkpoint management across large-scale distributed environments 3. Profile and diagnose performance bottlenecks across the ML stack, including data loading, preprocessing, forward and backward passes, and serving latency 4. Write automated tests covering expected behaviors, failure modes, and error paths for ML systems components, and build monitoring and alerting for production anomalies 5. Collaborate with ML researchers and product engineers to translate model requirements into reliable, high-throughput system designs 6. Own technical design for features and components within ML infrastructure, evaluating trade-offs between throughput, latency, cost, and engineering maintainability 7. Participate in staged rollouts of ML system changes using feature flagging and A/B testing frameworks, monitoring key metrics and responding to regressions 8. Contribute to code quality through code reviews, clear technical documentation, and consolidation of duplicative implementations across the ML systems codebase 9. Support on-call rotations for owned ML infrastructure, investigating production incidents and contributing detailed retrospectives to prevent recurrence ## Related Videos - [The Future of Developer Experience with GenAI: Driving Engineering Excellence](https://www.wearedevelopers.com/videos/1107-the-future-of-developer-experience-with-genai-driving-engineering-excellence) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Bringing the power of AI to your application.](https://www.wearedevelopers.com/videos/1010-bringing-the-power-of-ai-to-your-application) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)