Data Center Infrastructure Architect

OPEN-SILICON INC
United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Systems Engineering Computational Fluid Dynamics Computer Clusters Computer Simulation Data Centers Data Center Infrastructure Management (CIM) Supervisory Control and Data Acquisition (SCADA) Python (Programming Language) MATLAB Operational Data Store Software Tools
+5 more
Systems Architecture Digital Twin Performance Testing Process Control Systems Hardware Acceleration

Job description

OpenAI’s Industrial Compute team is building and productizing infrastructure capabilities that help organizations deploy and operate advanced AI systems at scale. The team works across AI hardware, systems engineering, physical infrastructure, and customer delivery to turn emerging technologies into reliable, repeatable infrastructure solutions.

Our work sits at the intersection of technical strategy, product development, engineering, and deployment. We partner closely with customers and internal engineering teams to solve complex infrastructure challenges spanning compute, power, cooling, controls, and facility efficiency., We are seeking a senior, hands-on Data Center Infrastructure Architect to develop and optimize the physical infrastructure required for large-scale AI deployments.

This is a broad technical role spanning data center architecture, electrical and mechanical systems, high-density compute, controls, telemetry, and digital modeling. You will use simulation, operational data, and digital-twin approaches to evaluate infrastructure designs, identify system-level constraints, and improve efficiency, reliability, cost, and speed of deployment.

The ideal candidate can move fluidly between first-principles analysis, facility and equipment design, computational modeling, engineering review, and real-world implementation. You should be comfortable working across disciplines rather than operating solely within electrical, mechanical, or software boundaries., * Define system-level architectures for high-density AI data centers across power, cooling, IT equipment, controls, and facility infrastructure.

  • Develop digital twins and other computational models that represent the behavior of data center systems under changing workloads, environmental conditions, equipment configurations, and failure scenarios.
  • Use design and operational data to identify constraints, improve PUE and related efficiency metrics, and optimize capacity, reliability, water consumption, cost, and deployment schedules.
  • Evaluate tradeoffs across electrical topology, cooling architecture, rack density, redundancy, controls, maintainability, constructability, and operational complexity.
  • Translate evolving AI hardware requirements into practical facility, rack, power, and thermal architectures.
  • Establish reference architectures, modeling standards, design assumptions, performance requirements, and validation methodologies that can be reused across customer deployments.
  • Partner with software, data, controls, hardware, mechanical, electrical, construction, commissioning, and operations teams to connect digital models with real infrastructure behavior.
  • Integrate telemetry from systems such as BMS, EPMS, DCIM, SCADA, equipment controllers, and IT hardware into modeling and optimization workflows.
  • Lead technical reviews of customer and partner designs, identify material risks, and recommend changes grounded in quantitative analysis.
  • Work with customers and delivery teams to adapt reference solutions to site-specific constraints while preserving performance, reliability, and efficiency objectives.
  • Support pilots, commissioning, performance testing, and post-deployment analysis to validate models and continuously improve infrastructure designs.
  • Help shape the technical roadmap for Industrial Compute’s physical-infrastructure products and engineering services., * Establish a credible system-level model of power, cooling, compute, and facility behavior for priority Industrial Compute use cases.
  • Identify and validate meaningful opportunities to improve efficiency, capacity, reliability, or deployment cost.
  • Deliver reusable reference architectures and engineering methodologies for customer deployments.
  • Create a repeatable feedback loop connecting modeling, operational telemetry, commissioning results, and future design decisions.
  • Become a trusted technical partner to internal engineering teams, customers, and infrastructure delivery partners.

Requirements

  • Significant experience designing or optimizing hyperscale data centers, large mission-critical facilities, or comparable infrastructure systems.
  • Broad knowledge of data center electrical and mechanical systems, including power distribution, backup power, thermal management, liquid cooling, heat rejection, controls, and monitoring.
  • Experience making system-level design decisions across multiple engineering disciplines.
  • Experience developing or applying simulation, optimization, digital-twin, or physics-based modeling techniques to physical infrastructure.
  • Strong understanding of data center efficiency and performance metrics, including PUE, WUE, utilization, capacity, reliability, and total cost of ownership.
  • Experience working with operational telemetry and translating real-world system behavior into design improvements.
  • Ability to evaluate complex tradeoffs involving performance, reliability, cost, schedule, scalability, sustainability, and maintainability.
  • Demonstrated ability to lead technical work in ambiguous, rapidly changing environments.
  • Strong written and verbal communication skills, including the ability to explain complex engineering decisions to customers, executives, and cross-functional teams.
  • Bachelor’s degree in mechanical engineering, electrical engineering, systems engineering, applied physics, or a related technical discipline.

Preferred Skills

  • Experience with high-density GPU clusters and direct-to-chip liquid cooling.
  • Experience connecting facility models with workload, rack, server, or chip-level power and thermal behavior.
  • Familiarity with modeling or engineering tools such as Modelica, MATLAB/Simulink, Python, EnergyPlus, computational fluid dynamics tools, or equivalent platforms.
  • Experience with BMS, EPMS, DCIM, SCADA, PLCs, data historians, or industrial controls.
  • Experience developing reference designs or new infrastructure architectures within a hyperscaler, data center operator, advanced engineering organization, or major design consultancy.
  • Experience taking an infrastructure concept from modeling and prototype validation through deployment and operational feedback.
  • Advanced degree in an engineering or scientific discipline.

About the company

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:03 min

Optimizing energy consumption and sustainability in data centers

Markus Hacker Markus Hacker +3 · World Congress 2024

1:21 min

Building a digital twin to understand generative AI

Kai Mueller Kai Mueller +1 · World Congress 2024

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

2:40 min

Motivations for transitioning legacy MATLAB repositories to Python

Michael Niebisch Michael Niebisch · World Congress 2024

2:50 min

Adapting data center environments for direct liquid immersion cooling

Thomas Schmidt Thomas Schmidt · World Congress 2024

2:47 min

Structuring the digital twin simulation ecosystem

Ulrich Wurstbauer +1 · LIVE

Videos

See all

Related articles

See all