Senior AI Engineer

Z's Associates
Princeton, NJ, United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Working hours
Regular working hours
Languages
English

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Computer Vision Microsoft Azure Big Data Cloud Computing Encodings Computer Programming Continuous Integration Extract Transform Load (ETL) DevOps Github
+28 more
Python (Programming Language) Machine Learning Modular Design Tensorflow Search Technologies Systems Integration Management of Software Versions Feature Engineering Pytorch Delivery Pipeline Large Language Models Deep Learning Generative AI Backend Keras Rate Limiting Fastapi Solid Principles Information Technology HuggingFace Performance Monitor Machine Learning Operations Celery Data Pipelines Docker Key Vault Web Api Microservices

Job description

Experteer Overview In this role you will build and operate ML engineering platforms and scalable backend systems, with a focus on MLOps and Generative AI. You will work within a collaborative, client-focused environment to deploy LLM-powered solutions, integrate Azure OpenAI, and scale models to meet strict SLAs. You’ll drive end-to-end pipelines-from feature engineering to real-time inference-while maintaining cutting-edge knowledge of transformers and GenAI. This opportunity lets you shape impactful AI solutions across healthcare and consumer domains at ZS. Compensation / Benefits * Develop and maintain ML engineering platforms, backend systems, APIs, and microservices (FastAPI) * Implement MLOps: model KPI tracking, drift detection, feedback loops * Deploy ML/Deep Learning models (LLMs/GenAI) and integrate Azure OpenAI * Build and orchestrate model pipelines (feature engineering, inference, training) * Create LLM observability (Langfuse) and prompt management with versioning * Implement Celery or similar for asynchronous workflows * Ensure scalable design with asynchronous Python, DI, and layered architecture * Develop real-time and batch prediction pipelines * Apply NLP and computer vision for document understanding, OCR, and layout processing * Work with HuggingFace, LangChain, LlamaIndex; manage embeddings and vector search Tasks * Master’s or bachelor’s degree in Computer Science or a related field * 4+ years in Machine Learning including production LLM systems * Strong ML, DL, and fine-tuning knowledge (LLMs); transformer understanding * Backend API design using FastAPI or similar; async patterns; rate limiting * Experience with vector databases (Pinecone, Weaviate, Chroma) and embedding storage * Strong Python skills; async programming; type hints and Pydantic; SOLID principles * AI/ML concepts and integrating models into backend services * MLOps experience with model tracking (MLFlow), LLM observability (Langfuse) * NLP and computer vision experience (OCR, document understanding, PDF extraction) * Model serving with FastAPI; feature engineering; real-time and batch inference pipelines * ML frameworks: HuggingFace, Keras/TensorFlow/PyTorch; LangChain preferred; LlamaIndex for RAG * Cloud and DevOps familiarity (Azure preferred): Azure OpenAI, Key Vault, Blob Storage; Docker; CI/CD (Azure DevOps, GitHub Actions) * Big data processing and ETL/ELT pipelines; rate limiting and cost management for LLMs * Production incident management and on-call experience * English fluency; client-first mindset; collaborative approach Key requirements * comprehensive total rewards * hybrid work model * internal mobility and career progression * professional development programs * collaborative culture * global opportunities

Requirements

systems Celery or similar for asynchronous workflows * Ensure scalable design with asynchronous Python, DI, and layered architecture * Develop real-time and batch prediction pipelines * Apply NLP and computer vision for document understanding, OCR, and layout processing * Work with HuggingFace, LangChain, LlamaIndex; manage embeddings and vector search Tasks * Master’s or bachelor’s degree in Computer Science or a related field * 4+ years in Machine Learning including production LLM systems * Strong ML, DL, and fine-tuning knowledge (LLMs); transformer understanding * Backend API design using FastAPI or similar; async patterns; rate limiting * Experience with vector databases (Pinecone, Weaviate, Chroma) and embedding storage * Strong Python skills; async programming; type hints and Pydantic; SOLID principles * AI/ML concepts and integrating models into backend services * MLOps experience with model tracking (MLFlow), LLM observability (Langfuse) * NLP and computer vision experience aa and document understanding, PDF extraction) * Model serving with FastAPI; feature engineering; real-time and batch inference pipelines * ML frameworks: HuggingFace, Keras/TensorFlow/PyTorch; LangChain preferred; LlamaIndex for RAG * Cloud and DevOps familiarity (Azure preferred): Azure OpenAI, Key Vault, Blob Storage; Docker; CI/CD (Azure DevOps, GitHub Actions) * Big data processing and ETL/ELT pipelines; rate limiting and cost management for LLMs * Production incident management and on-call experience * English fluency; client-first mindset; collaborative approach Key requirements * comprehensive total rewards * hybrid work model * internal mobility and career progression * professional development programs * collaborative culture * global opportunities

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

1:37 min

Core concepts of Celery and message broker integration

Jan Giacomelli · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

1:18 min

Converting existing Keras models to TensorFlow format

Håkan Silfvernagel · LIVE

2:10 min

Exploiting python celery dependencies for internal container access

Vandana Verma · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all