Production Systems Engineer, AI Systems

Facebook Inc.
Austin, TX, United States
2 days ago
Apply on www.jobmonkeyjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Data Analysis Systems Engineering Big Data Computer Clusters Code Coverage Data Centers Data Center Infrastructure Management (CIM) Software Debugging Linux Dynamic Random-Access Memory Firmware
+15 more
Hardware Design InfiniBand PCI Express Remote Infrastructure Management Software Deployment Subsystems System Testing Strategies of Testing AI Infrastructure Graphics Processing Unit (GPU) Application Specific Integrated Circuits AI Platforms Hardware Acceleration Server Operating Systems & Platforms Programming Languages

Job description

Meta is seeking a Hardware Systems Engineer to support the new product introduction (NPI) of next-generation AI and high-performance computing infrastructure for large-scale data center deployments. In this role, you will work at the intersection of server systems, AI applications and data center operations, partnering with hardware design, firmware, software, networking, and capacity engineering teams to validate and scale cutting-edge AI hardware systems from early bring-up through production readiness., * Lead end-to-end system validation strategies for AI and HPC hardware platforms, including AI accelerators, GPU clusters, and high-bandwidth memory subsystems in data center environments

  • Drive hands-on bring-up, characterization, and validation of AI server systems and associated components such as PCIe, NVLink, DRAM, and high-speed networking fabrics
  • Develop and maintain test specifications, validation procedures, and debug guides tailored to AI infrastructure NPI programs
  • Investigate and root-cause complex system failures spanning silicon, firmware, software, and hardware layers in collaboration with cross-functional engineering teams
  • Triage and track hardware and firmware defects through resolution while maintaining forward progress on NPI program milestones
  • Identify gaps in test coverage and drive improvements to test methodologies, tooling, and automation frameworks across the NPI lifecycle
  • Partner with AI platform and capacity engineering teams to define acceptance criteria and deployment readiness standards for new AI hardware systems
  • Guide data collection, analysis, and reporting efforts to surface systemic hardware quality trends and inform go/no-go decisions for production deployment
  • Communicate validation status, risk assessments, and technical findings to internal engineering teams and external hardware vendors
  • Collaborate with firmware and software teams to define hardware-software interface requirements for telemetry, diagnostics, and remote management of AI infrastructure

Requirements

  • 6+ years of experience in hardware systems engineering, silicon validation, firmware validation, or system-level bring-up for AI servers, GPUs, TPUs, or AI accelerator platforms
  • Experience in one or more of the following domains: ASIC bring-up and characterization, board-level debug, firmware validation, or large-scale system validation in data center environments
  • Experience developing test specifications, validation procedures, and debug methodologies for complex hardware systems
  • Experience leading root-cause analysis and troubleshooting of system-level failures across hardware, firmware, and software stacks
  • Experience with high-speed interconnects or memory subsystems such as PCIe, NVLink, DDR5, or HBM in the context of AI or HPC system validation
  • Experience analyzing system telemetry and fleet health data to identify reliability trends and drive engineering improvements, * Proficiency in scripting or programming languages such as Python for automation of infrastructure workflows and data analysis
  • Familiarity with Linux-based server environments and data center management tooling used in large-scale production operations
  • Experience defining hardware-software interface requirements for telemetry, out-of-band management, or remote diagnostics in data center AI systems
  • Experience with high-speed interconnects and memory subsystems such as PCIe, NVLink, InfiniBand, DDR5, or HBM in the context of AI or HPC infrastructure operations

About the company

Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today-beyond the constraints of screens, the limits of distance, and even the rules of physics.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobmonkeyjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · World Congress 2021

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

2:42 min

Dissecting artificial intelligence layers from compute to applications

Christian Nagel Christian Nagel +3 · World Congress 2026 Europe

Videos

See all

Related articles

See all