> Markdown version of [/jobs/ext/2272188-software-engineer-data-infrastructure](https://www.wearedevelopers.com/jobs/ext/2272188-software-engineer-data-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Data Infrastructure - **Company:** OpenAI Inc. - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Salary:** $185,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Big Data, Data Infrastructure, Data Security, Software Debugging, Distributed Data Store, Distributed Systems, Machine Learning, Data Streaming, Apache Spark, Low Latency, Apache Flink, Apache Kafka, Terraform, Stream Processing, Data Pipelines - **Published:** August 27, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/18084055?backUrl=%2Fcareer%2F18084055%2FSoftware-Engineer-Data-Infrastructure-New-York-New-York ## About the Role * Have 4+ years in data infrastructure engineering OR * Have 4+ years in infrastructure engineering with a strong interest in data * Take pride in building and operating scalable, reliable, secure systems * Are comfortable with ambiguity and rapid change * Have an intrinsic desire to learn and fill in missing skills, and an equally strong talent for sharing learnings clearly and concisely with others This role is exclusively based in our San Francisco HQ. We offer relocation assistance to new employees. ## Description This role focuses on building and operating data infrastructure that supports massive compute fleets and storage systems, designed for high performance and scalability. You'll help design, build, and operate the next generation of data infrastructure at OpenAI. You will scale and harden big data compute and storage platforms, build and support high-throughput streaming systems, build and operate low latency data ingestions, enable secure and governed data access for ML and analytics, and design for reliability and performance at extreme scale. You will take full lifecycle ownership: architecture, implementation, production operations, and on-call participation. You've supported Spark, Kafka, Flink, Airflow, Trino, or Iceberg as platforms. You're well-versed in infrastructure tooling like Terraform, experienced in debugging large-scale distributed systems, and excited about solving data infrastructure problems in the AI space. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: * Design, build, and maintain data infrastructure systems such as distributed compute, data orchestration, distributed storage, streaming infrastructure, machine learning infrastructure while ensuring scalability, reliability, and security * Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient * Accelerate company productivity by empowering your fellow engineers & teammates with excellent data tooling and systems * Collaborate with product, research and analytics teams to build the technical foundations capabilities that unlock new features and experiences * Own the reliability of the systems you build, including participation in an on-call rotation for critical incidents ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)