Software Engineer, Frontier Data Products
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+14 more
Job description
As AI systems become increasingly dependent on expert judgment, the company is building the orchestration platform that coordinates complex, long-running workflows between humans and AI models. This team owns the systems that sit directly between customer requests and production outputs, ensuring every stage of the data generation process is reliable, observable, and scalable. Engineers on this team will help define the architecture for one of the company’s most strategic product areas while solving distributed systems challenges at the intersection of AI, workflow orchestration, and backend infrastructure. What You Will Do
- Design and build backend services that orchestrate complex, multi-stage workflows involving AI models and human experts.
- Develop distributed state management systems capable of coordinating long-running jobs that evolve over hours or days.
- Build workflow orchestration primitives including retries, failure recovery, idempotency, auditing, and state reconciliation.
- Integrate large language model inference into production systems while maintaining reliability, observability, and human oversight.
- Design scalable APIs and internal tooling that enable Product, Operations, and Machine Learning teams to manage production workflows.
- Build highly reliable backend infrastructure where correctness and recoverability are as important as system availability.
- Debug complex distributed systems spanning asynchronous processing, model inference, and human review workflows.
- Own services throughout their lifecycle, including architecture, implementation, production operations, monitoring, and continuous improvement.
- Partner closely with engineering, product, operations, and AI research teams to translate ambiguous requirements into elegant technical systems.
- Help define the long-term architecture for one of the company’s core AI infrastructure platforms.
Requirements
- 4+ years of backend software engineering experience building production distributed systems.
- Strong experience designing scalable backend services using Python and PostgreSQL.
- Deep understanding of distributed systems concepts including asynchronous workflows, queues, retries, idempotency, and long-running processes.
- Proven experience designing reliable systems with strong observability and operational excellence.
- Strong system design skills with the ability to balance simplicity, scalability, and maintainability.
- Experience owning production services from design through deployment and ongoing operation.
- Comfortable working in highly ambiguous environments with significant ownership and autonomy.
- Excellent communication skills and ability to collaborate across engineering, product, and operations teams.
- Strong engineering judgment with an emphasis on building systems that remain maintainable as they scale.
Preferred
- Experience with workflow orchestration frameworks such as Temporal, Cadence, Airflow, or similar technologies.
- Background building orchestration platforms, distributed state machines, or event-driven backend systems.
- Experience integrating AI or machine learning inference into production applications.
- Familiarity with AWS cloud infrastructure and modern backend architecture.
- Experience building internal platforms, developer tooling, or infrastructure supporting high-throughput production systems.
- Strong understanding of fault tolerance, observability, and debugging distributed production environments.
- Startup experience building greenfield systems with significant architectural ownership.
- Passion for solving difficult infrastructure problems that directly impact customer-facing AI products., Amazon Web Services (AWS), Application Programming Interface (API), Artificial Intelligence (AI), Auditing, Building Systems, Cadence, Candidate Sourcing, Cloud Computing, Communication Skills, Continuous Improvement, Customer Relations, Data Quality, Debugging Skills, Dental Insurance, Distributed Computing, Engineering, High Throughput, Laundry, Machine Learning, Machine Tool, Machining Operations, Modeling Languages, Needs Assessment, Operations Research, PostgreSQL, Problem Solving Skills, Production Control, Production Management, Production Systems, Programming Tools, Python Programming/Scripting Language, Reconciliation, Reimbursement, Software Engineering, Startup, System Architecture, Systems Reliability, Team Lead/Manager, Team Player, Training/Teaching, Vision Plan
Benefits & conditions
- Base salary: $130,000-$400,000.
- Generous equity package vested over four years.
- Bi-annual performance bonus.
- Up to $15,000 relocation assistance.
- $10,000 housing bonus for employees living within 0.5 miles of the office.
- $1,500 monthly meal stipend.
- Complimentary Equinox membership.
- $200 monthly laundry reimbursement.
- $200 monthly wellness reimbursement.
- Comprehensive medical, dental, and vision insurance.
- Opportunity to define the architecture for one of the company’s most critical AI infrastructure platforms while building systems that directly power the world’s leading frontier AI models.
About the company
Who is Recruiting from Scratch: Recruiting from Scratch is a specialized talent firm dedicated to helping companies build exceptional teams. We partner closely with our clients to deeply understand their needs, then connect them with top-tier candidates who are not only highly skilled but also the right fit for the company’s culture and vision. Our mission is simple: place the best people in the right roles to drive long-term success for both clients and candidates. https://www.recruitingfromscratch.com, We’re representing one of the fastest-growing AI companies building the infrastructure that powers the global AI economy. Their platform connects leading AI labs and enterprises with a worldwide network of expert talent to generate the high-quality data required to train and evaluate frontier AI models.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Is Software Engineering Over-Saturated?
Why Upskilling And Reskilling is Important For Developers
Dev Digest 120 - Apple and peers