> Markdown version of [/jobs/ext/3017488-software-engineer-ai-platform](https://www.wearedevelopers.com/jobs/ext/3017488-software-engineer-ai-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, AI Platform - **Company:** Dst Llc - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Salary:** $180,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Big Data, Software as a Service, Cloud Computing, Databases, Information Engineering, Extract Transform Load (ETL), Software Debugging, Distributed Computing Environment, Distributed Systems, Data Intelligence, Python (Programming Language), PostgreSQL, Operational Databases, Systems Development Life Cycle, Reliability Engineering, Software Engineering, AI Infrastructure, Data Processing, Scripting, Autoscaling, Large Language Models, Reliability of Systems, AWS Lambda, Backend, Fastapi, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Low Latency, Active Optical Networks, Data Management, Amazon Simple Queue Service (SQS), Terraform - **Published:** September 20, 2026 - **Apply:** https://www.careerbuilder.com/job-details/software-engineer-ai-platform-san-francisco-ca--33c61a17-cd5e-4208-9973-9e2263b7b423 ## About the Role * 3-8 years of experience in backend, infrastructure, platform, or data engineering. * Proven experience owning production ETL pipelines or large-scale data platforms end-to-end. * Strong proficiency with Python and modern backend frameworks such as FastAPI. * Experience building distributed systems processing high-volume production workloads. * Familiarity with AWS services including ECS, Lambda, SQS, Step Functions, RDS, and S3. * Experience with PostgreSQL, orchestration frameworks, and Infrastructure-as-Code tools. * Strong understanding of production reliability, observability, failure recovery, and operational excellence. * Experience working at early-stage startups or other fast-moving engineering environments. * Bachelor's degree in Computer Science, Engineering, or equivalent professional experience. Preferred * Experience operating LLM-powered production systems with responsibility for cost, latency, and reliability. * Familiarity with AI-enabled ETL pipelines, vector databases, and modern AI infrastructure. * Experience with orchestration platforms such as Dagster, Airflow, Prefect, or Temporal. * Infrastructure-as-Code experience using Terraform or similar technologies. * Strong systems thinking with experience designing resilient distributed architectures. * Previous experience supporting high-throughput enterprise SaaS platforms. * Side projects, research, or technical work demonstrating exceptional engineering curiosity. * Passion for solving large-scale infrastructure challenges in AI-native environments., AWS Lambda, Amazon Web Services (AWS), Artificial Intelligence (AI), Automation, Autoscaling, Cloud Computing, Computer Science, Continuous Improvement, Data Processing, Database Extract Transform and Load (ETL), Debugging Skills, Distributed Computing, Fortune 500 Customers, Healthcare, High Throughput, Incident Response, Machine Tool, On Call, Operational Audit, PostgreSQL, Product Support, Production Control, Production Systems, Production Volume, Python Programming/Scripting Language, Reliability Engineering, Scalable System Development, Simple Queue Service (SQS), Software Engineering, Software as a Service (SaaS), Startup, Structured Data, Systems Reliability, Technical Research ## Description Location: San Francisco, CACompany Stage of Funding: Seed Stage AI Startup ($6M Raised)Office Type: Onsite (5 Days Per Week)Salary: $180,000-$250,000 + Competitive Equity Company Description We're representing a rapidly growing AI infrastructure startup building the data intelligence layer for the autonomous enterprise. Their platform captures how work happens across organizations, transforming massive volumes of enterprise activity into structured, AI-ready data that powers automation, analytics, and intelligent decision-making. Backed by Accel and DST Global, the company works with Fortune 500 customers including CVS Health, Aon, and PVH. Their platform processes tens of terabytes of LLM inference per customer every week, creating one of the largest enterprise AI data platforms operating in production today. As the company continues to scale, they're expanding their AI Platform team to build the infrastructure powering every AI capability across the business. What You Will Do * Own and evolve the production data platform that powers every AI feature across the company's products. * Design, build, and optimize large-scale AI-enabled ETL pipelines processing terabytes of data continuously in production. * Develop infrastructure that transforms LLM outputs into structured, queryable data used throughout the platform. * Improve system reliability, observability, throughput, latency, and infrastructure costs across AI workloads. * Build production tooling for monitoring, debugging, evaluation, and operational visibility of distributed AI systems. * Partner closely with AI engineers to develop scalable platform capabilities supporting new product features. * Design resilient orchestration workflows for distributed data processing, retries, backfills, and failure recovery. * Contribute to cloud infrastructure, Infrastructure-as-Code, and platform automation initiatives. * Participate in production incident response and on-call rotations while continuously improving system reliability. * Help define the long-term architecture of a rapidly scaling AI data platform serving enterprise customers. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Intro to FastAPI](https://www.wearedevelopers.com/videos/462-intro-to-fastapi) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)