> Markdown version of [/jobs/ext/2835774-ai-engineers](https://www.wearedevelopers.com/jobs/ext/2835774-ai-engineers). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Engineers - **Company:** Platform AI LLC - **Location:** New York, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Cloud Computing, Software Debugging, Programming Tools, Distributed Systems, Python (Programming Language), Azure Machine Learning, Management of Software Versions, Reinforcement Learning, Large Language Models, Multi-Agent Systems, Caching, Backend, Rate Limiting, Build Management, AI Platforms, Kubernetes, Data Management, Machine Learning Operations, Api Design, Software Version Control, Automation Anywhere, Serverless Computing, Data Generation - **Published:** September 10, 2026 - **Apply:** https://startup.jobs/senior-staff-ai-engineer-snorkel-ai-9955123 ## About the Role * 5+ years building production software systems, with experience in AI/ML infrastructure, ML platforms, distributed systems, data platforms, or backend infrastructure * Experience operating non-deterministic AI or ML workloads in production or at significant scale - you are comfortable reasoning about behavior across models, tools, environments, and multi-step execution * Experience building infrastructure for experimentation, evaluation, model development, synthetic data, agentic workflows, training, inference, or production ML systems * Strong proficiency in Python and experience building production-quality APIs, services, and developer tooling * Strong background in distributed systems and cloud platforms (AWS preferred), including compute orchestration, storage, networking, isolation, and failure handling * Experience with workflow or distributed execution frameworks such as Prefect, Airflow, Dagster, Ray, Kubernetes, or similar systems * Strong understanding of production system fundamentals - observability, telemetry, reliability, performance, debugging, incident response, and cost management * Ability to reason about AI system quality beyond traditional service metrics, including evaluation design, experiment reproducibility, behavioral regressions, and model or agent variability * Track record of leading complex engineering initiatives, influencing stakeholders, and delivering measurable impact * Ability to work in a fast-paced environment with strong technical communication skills * Fluency with modern AI and developer tooling and a willingness to rapidly evaluate and adopt new models, frameworks, infrastructure, and techniques as the ecosystem evolves Nice to Have * Experience building or operating LLM or agent infrastructure, including model gateways, agent runtimes, tool execution, tracing, or multi-agent systems * Experience building evaluation or experimentation platforms for LLMs, agents, or other probabilistic systems * Experience with synthetic data generation, automated labeling, data refinement, or dataset quality systems * Experience building reinforcement learning environments, agent simulations, benchmarks, or other environment-based evaluation systems * Experience running large-scale distributed AI workloads across containers, Kubernetes, serverless compute, sandboxes, or heterogeneous compute environments * Experience with LLM observability, tracing, prompt/version management, token and cost attribution, rate limiting, caching, or multi-provider routing * Experience designing isolation and sandboxing infrastructure for executing model-generated code or tool calls safely * Experience building shared AI platform libraries or SDKs consumed by multiple teams, including versioning, backwards compatibility, and migration support * Experience in hyper-growth startup environments or scaling engineering organizations * Prior experience as a Tech Lead, Team Lead, or hands-on Engineering Manager ## Description We're looking for AI Engineers who combine strong software and distributed systems fundamentals with experience operating AI systems in production. You'll build the infrastructure that lets teams create, experiment with, evaluate, and operate LLM and agentic workloads at significant scale - from synthetic data and evaluation pipelines to simulation environments, orchestration systems, and LLM infrastructure. You'll work on systems where correctness is not defined by a single deterministic output. Instead, you'll build the infrastructure needed to understand behavior across models, prompts, tools, environments, and multi-step trajectories, and to continuously improve those systems through experimentation and evaluation. We are looking to grow our team of AI Engineers, and are hiring at multiple levels. What You'll Do * Design and build infrastructure for running large-scale agentic workloads, including multi-step agents interacting with tools, external services, sandboxes, and simulated environments * Build scalable synthetic data generation and automated labeling systems that allow teams to create, refine, and evaluate high-quality training and evaluation datasets * Design evaluation infrastructure for measuring AI system behavior across models, prompts, tools, environments, and multi-step trajectories - including reproducible experiments, benchmark execution, regression detection, and continuous evaluation * Build orchestration and distributed compute systems for running thousands to millions of AI experiments and simulations reliably across heterogeneous compute environments * Develop infrastructure for agent simulation environments, including environment provisioning, isolation, lifecycle management, and scalable execution * Build and operate LLM infrastructure for routing, rate limiting, retries, caching, provider failover, cost attribution, and efficient execution across multiple model providers * Instrument agent and model workloads so failures are observable and debuggable - capturing traces, model interactions, tool calls, environment state, evaluation results, latency, reliability, and cost * Design systems that make non-deterministic workloads reproducible and measurable, allowing engineers to compare experiments, diagnose behavioral regressions, and understand why an agent succeeded or failed * Improve the developer experience for AI experimentation by building APIs, SDKs, workflow abstractions, and tooling that make it easy to move workloads from local development to large-scale production execution * Collaborate with research, product, and engineering teams to turn experimental AI workflows into reliable, reusable platform capabilities ## Related Videos - [One AI API to Power Them All](https://www.wearedevelopers.com/videos/1601-one-ai-api-to-power-them-all) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Building a Multi-Agent Orchestration Engine That Actually Follows the Rules](https://www.wearedevelopers.com/videos/100159-building-a-multi-agent-orchestration-engine-that-actually-follows-the-rules) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)