> Markdown version of [/jobs/ext/1993430-senior-ai-engineer](https://www.wearedevelopers.com/jobs/ext/1993430-senior-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI Engineer - **Company:** Z's Associates - **Location:** Princeton, NJ, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Computer Vision, Microsoft Azure, Big Data, Cloud Computing, Encodings, Computer Programming, Continuous Integration, Extract Transform Load (ETL), DevOps, Github, Python (Programming Language), Machine Learning, Modular Design, Tensorflow, Search Technologies, Systems Integration, Management of Software Versions, Feature Engineering, Pytorch, Delivery Pipeline, Large Language Models, Deep Learning, Generative AI, Backend, Keras, Rate Limiting, Fastapi, Solid Principles, Information Technology, HuggingFace, Performance Monitor, Machine Learning Operations, Celery, Data Pipelines, Docker, Key Vault, Web Api, Microservices - **Published:** August 8, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/senior-ai-engineer-princeton-nj-usa-58846589 ## About the Role systems Celery or similar for asynchronous workflows * Ensure scalable design with asynchronous Python, DI, and layered architecture * Develop real-time and batch prediction pipelines * Apply NLP and computer vision for document understanding, OCR, and layout processing * Work with HuggingFace, LangChain, LlamaIndex; manage embeddings and vector search Tasks * Master's or bachelor's degree in Computer Science or a related field * 4+ years in Machine Learning including production LLM systems * Strong ML, DL, and fine-tuning knowledge (LLMs); transformer understanding * Backend API design using FastAPI or similar; async patterns; rate limiting * Experience with vector databases (Pinecone, Weaviate, Chroma) and embedding storage * Strong Python skills; async programming; type hints and Pydantic; SOLID principles * AI/ML concepts and integrating models into backend services * MLOps experience with model tracking (MLFlow), LLM observability (Langfuse) * NLP and computer vision experience aa and document understanding, PDF extraction) * Model serving with FastAPI; feature engineering; real-time and batch inference pipelines * ML frameworks: HuggingFace, Keras/TensorFlow/PyTorch; LangChain preferred; LlamaIndex for RAG * Cloud and DevOps familiarity (Azure preferred): Azure OpenAI, Key Vault, Blob Storage; Docker; CI/CD (Azure DevOps, GitHub Actions) * Big data processing and ETL/ELT pipelines; rate limiting and cost management for LLMs * Production incident management and on-call experience * English fluency; client-first mindset; collaborative approach Key requirements * comprehensive total rewards * hybrid work model * internal mobility and career progression * professional development programs * collaborative culture * global opportunities ## Description Experteer Overview In this role you will build and operate ML engineering platforms and scalable backend systems, with a focus on MLOps and Generative AI. You will work within a collaborative, client-focused environment to deploy LLM-powered solutions, integrate Azure OpenAI, and scale models to meet strict SLAs. You'll drive end-to-end pipelines-from feature engineering to real-time inference-while maintaining cutting-edge knowledge of transformers and GenAI. This opportunity lets you shape impactful AI solutions across healthcare and consumer domains at ZS. Compensation / Benefits * Develop and maintain ML engineering platforms, backend systems, APIs, and microservices (FastAPI) * Implement MLOps: model KPI tracking, drift detection, feedback loops * Deploy ML/Deep Learning models (LLMs/GenAI) and integrate Azure OpenAI * Build and orchestrate model pipelines (feature engineering, inference, training) * Create LLM observability (Langfuse) and prompt management with versioning * Implement Celery or similar for asynchronous workflows * Ensure scalable design with asynchronous Python, DI, and layered architecture * Develop real-time and batch prediction pipelines * Apply NLP and computer vision for document understanding, OCR, and layout processing * Work with HuggingFace, LangChain, LlamaIndex; manage embeddings and vector search Tasks * Master's or bachelor's degree in Computer Science or a related field * 4+ years in Machine Learning including production LLM systems * Strong ML, DL, and fine-tuning knowledge (LLMs); transformer understanding * Backend API design using FastAPI or similar; async patterns; rate limiting * Experience with vector databases (Pinecone, Weaviate, Chroma) and embedding storage * Strong Python skills; async programming; type hints and Pydantic; SOLID principles * AI/ML concepts and integrating models into backend services * MLOps experience with model tracking (MLFlow), LLM observability (Langfuse) * NLP and computer vision experience (OCR, document understanding, PDF extraction) * Model serving with FastAPI; feature engineering; real-time and batch inference pipelines * ML frameworks: HuggingFace, Keras/TensorFlow/PyTorch; LangChain preferred; LlamaIndex for RAG * Cloud and DevOps familiarity (Azure preferred): Azure OpenAI, Key Vault, Blob Storage; Docker; CI/CD (Azure DevOps, GitHub Actions) * Big data processing and ETL/ELT pipelines; rate limiting and cost management for LLMs * Production incident management and on-call experience * English fluency; client-first mindset; collaborative approach Key requirements * comprehensive total rewards * hybrid work model * internal mobility and career progression * professional development programs * collaborative culture * global opportunities ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Celery on AWS ECS - the art of background tasks & continuous deployment](https://www.wearedevelopers.com/videos/561-celery-on-aws-ecs-the-art-of-background-tasks-continuous-deployment) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Walking into the era of Supply Chain Risks](https://www.wearedevelopers.com/videos/376-walking-into-the-era-of-supply-chain-risks) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)