AI Solution Lead in Cary, NC (Fulltime, Onsite)

Northern Base
Cary, NC, United States
7 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Microsoft Azure Cloud Computing Continuous Integration Information Engineering Software Debugging Event-Driven Programming Python (Programming Language) Performance Tuning Search Technologies Delivery Pipeline
+10 more
Large Language Models Concurrency Caching Backend Rate Limiting Fastapi Kubernetes Front End Software Development Software Version Control Microservices

Job description

Package, serve, and monitor models for real-time and batch inference, ensuring operational readiness and performance. Build event driven, resilient integrations and containerized services, with hands-on Kubernetes debugging and Helm-based deployments. Establish observability, SLOs, CI/CD automation, testing Apply strong systems design principles (concurrency, caching, reliability, rate limiting) and robust data engineering practices. Cloud exposure preferred (Azure/AKS, managed services), with bonus experience in performance tuning, frontend collaboration, and model governance/monitoring. Roles & Responsibilities Build and productionize cloud native backend services and AI/LLM inference pipelines. Design and develop Python-based APIs and microservices (FastAPI, async patterns) and agentic AI workflows using LangChain/LangGraph. Implement and optimize LLM capabilities including embeddings, RAG, vector search, prompt/context engineering, and model versioning. Package, serve, and monitor models for real-time and batch inference, ensuring operational readiness and performance. Build event driven, resilient integrations and containerized services, with hands-on Kubernetes debugging and Helm-based deployments. Establish observability, SLOs, CI/CD automation, testing Apply strong systems design principles (concurrency, caching, reliability, rate limiting) and robust data engineering practices. Cloud exposure preferred (Azure/AKS, managed services), with bonus experience in performance tuning, frontend collaboration, and model governance/monitoring.

Requirements

Must Have Technical/Functional Skills 13+ years of experience with IT Build and productionize cloud native backend services and AI/LLM inference pipelines. Design and develop Python-based APIs and microservices (FastAPI, async patterns) and agentic AI workflows using LangChain/LangGraph. Implement and optimize LLM capabilities including embeddings, RAG, vector search, prompt/context engineering, and model versioning.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:33 min

Connecting frontends via a FastAPI proxy backend layer

Saoussen Chaabnia Saoussen Chaabnia · Europe 2026 Virtual

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

1:42 min

Introduction to the fast API web framework

Sebastián Ramírez · World Congress 2022

1:12 min

Building intelligent applications and enhancing developer productivity experiences

Alexander Wallner Alexander Wallner +3 · World Congress 2024

Videos

See all

Related articles

See all