Software Engineer, Infrastructure

OpenAI Inc.
San Francisco, CA, United States
3 days ago
Apply on diversityjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$210,000.0
Working hours
Regular working hours

Tech stack

Systems Engineering C++ (Programming Language) Cloud Computing Databases Continuous Integration Software Debugging Distributed Data Store Distributed Systems Fault Tolerance Python (Programming Language) Linux System Administration Reliability Engineering
+9 more
Software Engineering Rust (Programming Language) Cloud Platform System Reliability of Systems Kubernetes Api Design Terraform GPT Golang

Job description

We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas-including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure.

About the Role

All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability

Team Focus Areas

  • Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI
  • Systems Engineering: Work across layers of the stack-debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability.
  • Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience.
  • Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale.
  • Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely.
  • Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads.
  • Databases: Building high performance, distributed database systems that power all of OpenAI’s product stack.

In this role you will:

  • Design, build, and maintain reliable and performant systems used across engineering. Work with your team to define technical strategy, architecture, and long-term goals.

  • Collaborate with other engineers, product managers, and researchers to build infrastructure that meets evolving needs.
  • Improve internal tooling, automation, and developer experience.
  • Contribute to incident response, postmortems, and the development of best practices around system reliability and scalability.

Requirements

  • Strong software engineering skills with experience in Python, Go, C++, Rust, or similar languages.
  • Experience designing, operating, or scaling distributed systems or developer infrastructure.
  • Comfort working in Linux environments, and with tools like Kubernetes, Terraform, CI/CD pipelines, and modern observability stacks.
  • Ability to navigate complex systems and a willingness to dig deep when debugging tricky issues.
  • Excellent communication and collaboration skills, especially in cross-functional settings., * 4+ years of relevant industry experience, with 2+ years leading large scale, complex projects or teams as an engineer or tech lead
  • A passion for distributed systems at scale with a focus on reliability, scalability, security, and continuous improvement.
  • Excellent communication skills, with ability to build consensus among stakeholders both internally and externally.

About the company

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity., At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

1:38 min

Shifting software engineering skills for an AI-driven future

Patrick Schnell Patrick Schnell · Coffee With Developers

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · World Congress 2025

Videos

See all

Related articles

See all