Software Engineer, AI Platform
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+26 more
Job description
Location: San Francisco, CACompany Stage of Funding: Seed Stage AI Startup ($6M Raised)Office Type: Onsite (5 Days Per Week)Salary: $180,000-$250,000 + Competitive Equity Company Description
We’re representing a rapidly growing AI infrastructure startup building the data intelligence layer for the autonomous enterprise. Their platform captures how work happens across organizations, transforming massive volumes of enterprise activity into structured, AI-ready data that powers automation, analytics, and intelligent decision-making.
Backed by Accel and DST Global, the company works with Fortune 500 customers including CVS Health, Aon, and PVH. Their platform processes tens of terabytes of LLM inference per customer every week, creating one of the largest enterprise AI data platforms operating in production today. As the company continues to scale, they’re expanding their AI Platform team to build the infrastructure powering every AI capability across the business. What You Will Do
- Own and evolve the production data platform that powers every AI feature across the company’s products.
- Design, build, and optimize large-scale AI-enabled ETL pipelines processing terabytes of data continuously in production.
- Develop infrastructure that transforms LLM outputs into structured, queryable data used throughout the platform.
- Improve system reliability, observability, throughput, latency, and infrastructure costs across AI workloads.
- Build production tooling for monitoring, debugging, evaluation, and operational visibility of distributed AI systems.
- Partner closely with AI engineers to develop scalable platform capabilities supporting new product features.
- Design resilient orchestration workflows for distributed data processing, retries, backfills, and failure recovery.
- Contribute to cloud infrastructure, Infrastructure-as-Code, and platform automation initiatives.
- Participate in production incident response and on-call rotations while continuously improving system reliability.
- Help define the long-term architecture of a rapidly scaling AI data platform serving enterprise customers.
Requirements
- 3-8 years of experience in backend, infrastructure, platform, or data engineering.
- Proven experience owning production ETL pipelines or large-scale data platforms end-to-end.
- Strong proficiency with Python and modern backend frameworks such as FastAPI.
- Experience building distributed systems processing high-volume production workloads.
- Familiarity with AWS services including ECS, Lambda, SQS, Step Functions, RDS, and S3.
- Experience with PostgreSQL, orchestration frameworks, and Infrastructure-as-Code tools.
- Strong understanding of production reliability, observability, failure recovery, and operational excellence.
- Experience working at early-stage startups or other fast-moving engineering environments.
- Bachelor’s degree in Computer Science, Engineering, or equivalent professional experience.
Preferred
- Experience operating LLM-powered production systems with responsibility for cost, latency, and reliability.
- Familiarity with AI-enabled ETL pipelines, vector databases, and modern AI infrastructure.
- Experience with orchestration platforms such as Dagster, Airflow, Prefect, or Temporal.
- Infrastructure-as-Code experience using Terraform or similar technologies.
- Strong systems thinking with experience designing resilient distributed architectures.
- Previous experience supporting high-throughput enterprise SaaS platforms.
- Side projects, research, or technical work demonstrating exceptional engineering curiosity.
- Passion for solving large-scale infrastructure challenges in AI-native environments., AWS Lambda, Amazon Web Services (AWS), Artificial Intelligence (AI), Automation, Autoscaling, Cloud Computing, Computer Science, Continuous Improvement, Data Processing, Database Extract Transform and Load (ETL), Debugging Skills, Distributed Computing, Fortune 500 Customers, Healthcare, High Throughput, Incident Response, Machine Tool, On Call, Operational Audit, PostgreSQL, Product Support, Production Control, Production Systems, Production Volume, Python Programming/Scripting Language, Reliability Engineering, Scalable System Development, Simple Queue Service (SQS), Software Engineering, Software as a Service (SaaS), Startup, Structured Data, Systems Reliability, Technical Research
Benefits & conditions
- Base salary: $180,000-$250,000 (higher for exceptional candidates).
- Competitive equity package.
- Additional paid on-call compensation.
- $10,000 annual meal stipend.
- $12,000 annual healthcare stipend.
- Five-day onsite collaboration in San Francisco.
- E-3 visa sponsorship available for qualified Australian candidates.
- Opportunity to own the core infrastructure powering every AI product across the company.
- Work alongside a highly technical team solving some of the largest production AI data processing challenges in the industry.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
Navigating the AI Shift
Dev Digest 120 - Apple and peers
MLOps And AI Driven Development