AI Solution Lead in Cary, NC (Fulltime, Onsite)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
Package, serve, and monitor models for real-time and batch inference, ensuring operational readiness and performance. Build event driven, resilient integrations and containerized services, with hands-on Kubernetes debugging and Helm-based deployments. Establish observability, SLOs, CI/CD automation, testing Apply strong systems design principles (concurrency, caching, reliability, rate limiting) and robust data engineering practices. Cloud exposure preferred (Azure/AKS, managed services), with bonus experience in performance tuning, frontend collaboration, and model governance/monitoring. Roles & Responsibilities Build and productionize cloud native backend services and AI/LLM inference pipelines. Design and develop Python-based APIs and microservices (FastAPI, async patterns) and agentic AI workflows using LangChain/LangGraph. Implement and optimize LLM capabilities including embeddings, RAG, vector search, prompt/context engineering, and model versioning. Package, serve, and monitor models for real-time and batch inference, ensuring operational readiness and performance. Build event driven, resilient integrations and containerized services, with hands-on Kubernetes debugging and Helm-based deployments. Establish observability, SLOs, CI/CD automation, testing Apply strong systems design principles (concurrency, caching, reliability, rate limiting) and robust data engineering practices. Cloud exposure preferred (Azure/AKS, managed services), with bonus experience in performance tuning, frontend collaboration, and model governance/monitoring.
Requirements
Must Have Technical/Functional Skills 13+ years of experience with IT Build and productionize cloud native backend services and AI/LLM inference pipelines. Design and develop Python-based APIs and microservices (FastAPI, async patterns) and agentic AI workflows using LangChain/LangGraph. Implement and optimize LLM capabilities including embeddings, RAG, vector search, prompt/context engineering, and model versioning.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.careerjet.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
Navigating the AI Shift
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
What Are Large Language Models?