> Markdown version of [/jobs/ext/1737553-senior-ai-engineer](https://www.wearedevelopers.com/jobs/ext/1737553-senior-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI Engineer - **Company:** ZS - **Location:** Chicago, IL, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** A/B Testing, Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Computer Vision, Audit Trail, Microsoft Azure, Big Data, Encodings, Computer Programming, Databases, Continuous Integration, Information Engineering, Extract Transform Load (ETL), Data Transformation, Data Security, Dependency Injection, Software Design Patterns, DevOps, Django Web Framework, Github, Python (Programming Language), Key Management, Machine Learning, Object-Oriented Software Development, Tensorflow, Azure DevOps Pipelines, Search Technologies, SQL Databases, Strategies of Testing, Management of Software Versions, Feature Engineering, Pytorch, Flask (Web Framework), Large Language Models, Prompt Engineering, Deep Learning, Generative AI, Backend, Keras, Rate Limiting, Fastapi, Containerization, Pyspark, Solid Principles, Kubernetes, Information Technology, HuggingFace, Machine Learning Operations, Celery, Asynchronous Programming, GPT, Software Version Control, Data Pipelines, Docker, Key Vault, Microservices - **Published:** July 31, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=23b0d8efe2a85a6a ## About the Role * Master's or bachelor's degree in Computer Science or a related field from a top university. * 4+ years of hands-on experience in Machine Learning, including production LLM systems. * Strong fundamentals in machine learning, deep learning, and fine-tuning models (LLMs), including: * Understanding of transformer architectures * Prompt engineering expertise * Embeddings and vector search * Experience in backend API design using FastAPI or similar asynchronous frameworks (e.g., Flask, Django), including async patterns and rate limiting. * Experience with vector databases, including: * Pinecone, Weaviate, or Chroma * Embedding storage and similarity search * Hybrid search implementations * Strong programming expertise in Python is a must, including: * Async programming (asyncio, async/await) * Type hints and Pydantic * SOLID principles and design patterns * PySpark/Scala is optional. * Knowledge of AI/ML concepts and experience integrating AI models into backend services is mandatory. * Experience with MLOps to measure and track model performance, including: * MLFlow for model tracking * Langfuse or similar tools for LLM observability (strongly preferred) * Model versioning and A/B testing * Experience working with NLP and computer vision, including: * Text extraction and preprocessing * Document understanding (layout, tables) * OCR processing * GPT-4 Vision or similar multimodal integration * Experience implementing: * Feature engineering pipelines * Real-time inferencing systems * Batch prediction pipelines * Model serving with FastAPI * Experience with ML frameworks, including: * HuggingFace (transformers, datasets) - mandatory * Keras/TensorFlow/PyTorch * LangChain - strongly preferred * LlamaIndex for RAG * Familiarity with database technologies such as SQL. * Good problem-solving skills and the ability to work in a fast-paced, team-oriented environment. Additional Skills: * Understanding of DevOps and CI/CD, including: * Docker containerization * Azure DevOps pipelines or GitHub Actions * Kubernetes (nice to have) * Data security practices, including: * Multi-tenant data isolation * Secure key management (e.g., Azure Key Vault) * Audit trail implementation * Experience designing on cloud platforms: * Azure (strongly preferred): Azure OpenAI, Blob Storage, Key Vault, Container Registry * AWS or GCP * Experience with data engineering in Big Data systems, including large-scale data processing and ETL/ELT pipelines. * Rate limiting and quota management for high-throughput API usage. * Cost management and optimization for LLM usage at scale. * Document processing expertise (PDF extraction, OCR tooling). * Production incident management and on-call experience. * Testing strategies for non-deterministic LLM outputs (e.g., golden datasets, fuzzy matching). * Domain knowledge in regulated industries (e.g., healthcare/pharma workflows, regulatory compliance) is a plus. ## Description * Build, refine, and use ML Engineering platforms and components; develop and implement scalable backend systems, APIs, and microservices using FastAPI. * Implement MLOps including model KPI measurement, tracking, model drift detection, and model feedback loops. * Deploy and operationalize ML and Deep Learning models, with a strong focus on LLMs and Generative AI. * Integrate Azure OpenAI (GPT-4, GPT-4 Vision) and other LLM providers with proper retry logic and error handling. * Maintain up-to-date knowledge of state-of-the-art technologies such as LLMs, GenAI, and transformer architectures. * Scale machine learning algorithms to work on massive data sets under strict SLAs. * Build and orchestrate model pipelines including feature engineering, inferencing, and continuous model training. * Write backend application code in Python and SQL using strong object-oriented principles and asynchronous programming (asyncio, async/await). * Implement dependency injection patterns and layered architecture (Service, Foundation, Orchestration, DAL). * Build LLM observability (e.g., Langfuse) to track prompts, tokens, costs, and latency. * Develop prompt management systems with versioning and fallback mechanisms. * Implement Celery (or similar) workflows for asynchronous task processing and complex pipelines. ## Related Videos - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Overview of Machine Learning in Python](https://www.wearedevelopers.com/videos/840-overview-of-machine-learning-in-python) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)