Principal Software Engineer, Simulation

OpenAI Inc.
San Francisco, CA, United States
25 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$347,000.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Programming Tools Python (Programming Language) Large Language Models Technical Debt Backend Api Design GPT

Job description

OpenAI’s research training infrastructure powers how our frontier models are trained and evaluated. The Simulation team sits at the intersection between the agentic harness that powers OpenAI’s products and the research infrastructure where GPT-next is trained, ensuring that our model’s training environment is as realistic as possible.

This team owns the integration layer that connects our production harness capabilities into the training stack. The work is highly cross-functional and high leverage: researchers depend on it to run experiments and evaluations reliably as well as to develop the next generation of harness capabilities. Failures in this surface can materially affect training velocity and correctness.

About the Role

We’re looking for a Principal Software Engineer to lead the architecture and evolution of the Simulation Platform. You’ll own a critical interface between research and engineering, building the systems, APIs, and operational patterns that let researchers use agentic coding infrastructure safely and effectively in training environments.

This role is ideal for a senior backend or infrastructure engineer with strong technical judgment, product sense for highly technical users, and the ability to drive execution across multiple teams. The highest-leverage work is building robust infrastructure that supports and accelerates research without compromising engineering quality.

In this role, you will

  • Design, build, and evolve the integration between the Codex harness that powers OpenAI’s products and research training infrastructure used for training GPT-next
  • Build a platform for our LLMs to train and be evaluated in simulated environments that mimic their deployment setting as closely as possible, on every axis: agentic harness, compute substrate, timing, tools, data sources, humans in the loop, and more
  • Own major integration surfaces end-to-end, from architecture and API design through rollout, operations, and long-term maintenance
  • Build reliable execution systems that can support demanding training workloads at scale
  • Partner closely with research, agent, infrastructure, and platform teams to support new training use cases and harness capabilities
  • Design clean, stable interfaces and workflows for highly technical internal users who move quickly and expect strong ergonomics
  • Prevent one-off workarounds from becoming long-term technical debt by establishing durable abstractions and clear ownership
  • Raise the bar for correctness, reliability, operational rigor, and engineering judgment across a critical research-facing system

Requirements

  • Have significant experience building and scaling backend or infrastructure systems in fast-moving environments
  • Bring deep strength in API design, systems design, and engineering fundamentals
  • Are highly detail-oriented and care deeply about correctness, reliability, and operational quality
  • Can work directly with demanding technical users while maintaining strong engineering discipline
  • Have a track record of leading cross-functional technical efforts and creating clarity across organizational boundaries
  • Bring strong product sense and user empathy for internal platforms and developer tooling
  • Are motivated by enabling researchers and accelerating their work, rather than doing research yourself
  • Are proficient in Python and have experience with backend platform engineering; Rust experience is a plus

About the company

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity., At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.localjobnetwork.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · WWC 2024

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

3:10 min

Understanding the core concepts of API design

Alen Pokos · LIVE

1:43 min

Platform engineering as the foundation for scaling AI tools

Julia Kordick Julia Kordick · WWC Europe 2026

2:14 min

Generating functioning backend code from OpenAPI specifications

Christopher Walles Christopher Walles · WWC 2024

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · WWC 2025

Videos

See all

Related articles

See all