> Markdown version of [/jobs/ext/2814482-senior-software-engineer-applied-ai-cloud](https://www.wearedevelopers.com/jobs/ext/2814482-senior-software-engineer-applied-ai-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer - Applied AI / Cloud - **Company:** Opentrons - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $150,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Artificial Intelligence, Amazon Web Services, Automated Storage and Retrieval Systems, Software as a Service, Cloud Computing, Software Quality, Encodings, Continuous Integration, Software Debugging, Distributed Systems, Python (Programming Language), Software Engineering, Data Streaming, Systems Integration, TypeScript, Data Logging, Large Language Models, Indexer, Fastapi, AI Platforms, Git Flow, Information Technology, Machine Learning Operations, Api Design, Restful APIs, Microservices - **Published:** September 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=726bf4d9deda4395 ## About the Role * Bachelor's degree in Computer Science or similar field, or equivalent certifications and experience Experience: * 5+ years of professional software engineering experience, including at least 3+ years spent building, deploying, and maintaining production AI/ML or LLM-powered applications. * Experience shipping and supporting a SaaS product through a real launch, including the operational work that follows. * Solid hands-on experience working within cloud microservice architectures (AWS strongly preferred) including deploying, scaling, and operating microservices in production. * Systems design experience: API design, service boundaries, data flow, and the tradeoffs involved in scaling distributed systems. * Hands-on experience building production AI or machine learning systems from prototype to production (building full systems layers beyond simply integrating third-party APIs), including NLP and classifier work. * Experience building RAG pipelines: chunking, embedding/indexing, and retrieval tuning. * Experience designing evaluation frameworks or rubrics to measure AI output quality. Knowledge, Skills & Abilities: * Practical understanding of AI safety and cost discipline: guardrails, validation, and token spend monitoring. * Proficiency in React/TypeScript and Python/FastAPI, with experience designing RESTful APIs and full-stack debugging. * Familiarity with Git workflows and CI/CD practices. * Strong communicator who can own features end-to-end, balance speed with code quality, and thrive in a fast-paced startup environment. * Solid software engineering fundamentals: testing, debugging, and writing maintainable code. * Comfort with ambiguity and iterating quickly based on evaluation results and feedback. Working Conditions and Physical Effort * Hybrid work environment based in New York with flexibility for remote days * Core collaboration hours with occasional overlap for international team coordination * Prolonged periods working at a computer for development work * Fast-paced startup environment requiring ability to deliver features iteratively while maintaining quality * Minimal physical effort required. ## Description We're looking for an Applied AI Engineer (SWE3) to help build AI-powered features as part of OpentronsAI, our platform that leverages Large Language Models to democratize laboratory automation. This role focuses on the internals of an AI-powered application: building the AWS-based microservices and asynchronous workflows that power protocol generation, integrating LLM calls and knowledge retrieval systems into production SaaS infrastructure, and making sure those systems are reliable and safe at scale for thousands of laboratories globally. You'll build the ML/AI systems layer that sits underneath the product, from parsing and understanding scientists' natural-language input to the classifiers that route and score it and the embedding infrastructure that grounds model output in real protocol knowledge. Because correctness here isn't optional, you'll treat evaluation as a first-class part of the build, designing the tests and benchmarks that catch bad output before it ships, alongside the cost and safety instrumentation (token spend tracking, guardrails, validation) that keeps the system trustworthy at scale. You'll bring the same engineering rigor to this layer that you'd bring to any production system: observability, cost control, and failure modes that fail safely rather than silently. This is a hybrid position, requiring onsite presence at our Long Island City headquarters at least 3 days per week. What You'll Do * Design, build, and operate the AWS-based microservices and asynchronous workflows that power protocol generation, including API design, service boundaries, message-driven communication, and integration points with the rest of the SaaS platform. * Apply systems design principles to how AI components fit into the broader architecture: request routing, service-to-service communication, data flow, and scaling under production load. * Build the ML/AI systems layer underneath the product: NLP pipelines for parsing scientist input, classifiers for routing and scoring requests, and embedding infrastructure that grounds model output in real protocol knowledge. * Build and iterate on knowledge retrieval and retrieval-augmented generation (RAG) pipelines, including chunking, embedding/indexing, and retrieval tuning. * Design and apply evaluation frameworks and rubrics to benchmark model and application output, and use them to drive iteration before issues reach production. * Build the cost and safety instrumentation - token spend tracking, guardrails, and validation - that keeps AI services trustworthy and sustainable at scale. * Instrument services for observability (logging, metrics, tracing) and build the monitoring needed to catch failures before customers do. * Support features post-launch: on-call debugging, incident response, and iterating on the architecture based on real production usage.