> Markdown version of [/jobs/ext/3224439-senior-ai-engineer-in-chicago](https://www.wearedevelopers.com/jobs/ext/3224439-senior-ai-engineer-in-chicago). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI Engineer in Chicago - **Company:** Energy Jobline - **Location:** Chicago, IL, United States - **Experience:** Expert - **Salary:** $200,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Training Data, Application Programming Interfaces (APIs), Artificial Intelligence, Python (Programming Language), Regression Testing, Software Engineering, Large Language Models, Model Validation, Rate Limiting, Kubernetes, Low Latency, HuggingFace, Hardware Infrastructure - **Published:** September 9, 2026 - **Apply:** https://www.energyjobline.com/job/senior-ai-engineer-chicago-31579940 ## About the Role * 5+ years software engineering; strong Python * Production fine-tuning or distillation of open-weight models (not just inference API wrappers) * Experience serving LLMs on-prem (vLLM, TGI, Triton, or equivalent) * Experience managing GPU infrastructure (provisioning, scheduling, utilization monitoring) in a production environment * Model evaluation and regression testing in production * Kubernetes and GPU workload management * Strong grasp of the tradeoffs between open and closed models across cost, quality, latency, and data sensitivity : * Quantization, PEFT/LoRA, or other efficient training techniques * Model gateway or inference proxy design (routing, fallback, rate limiting) * Financial services or other regulated/sensitive-data environments * Familiarity with the open model ecosystem (Hugging Face, model cards, licensing ## Description * Build and operate a model gateway routing inference across open and closed models with cost, latency, and quality tracking * Design and run distillation pipelines: use frontier model outputs to generate training data for task-specific open models * Fine-tune and evaluate open-weight models (Llama, Qwen, Mistral, or similar) for DV-specific tasks * Deploy and maintain on-prem inference infrastructure (vLLM, TGI, or equivalent) on KubernetesBuild model evaluation frameworks for quality, cost, latency, and regression * Define criteria and tooling for model selection: when open models are production-ready vs. when to use closed APIs * Partner with the agent engineering team to ensure the model layer meets agent workload ## Related Videos - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) - [Unleash the power of 5G in your code: transform your apps](https://www.wearedevelopers.com/videos/1567-unleash-the-power-of-5g-in-your-code-transform-your-apps) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [AI That Fits Your Business, Not the Other Way Around](https://www.wearedevelopers.com/videos/100148-ai-that-fits-your-business-not-the-other-way-around) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [ I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.](https://www.wearedevelopers.com/magazine/773-i-gave-a-video-editor-more-autonomy-than-a-trading-bot-on-purpose) - [A 5-Step Open-Source Setup for Agentic Engineering](https://www.wearedevelopers.com/magazine/738-a-5-step-open-source-setup-for-agentic-engineering)