AI Engineer with Kubernetes

EPAM Systems, Inc.
United States
27 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Java (Programming Language) Artificial Intelligence Amazon S3 C++ (Programming Language) Encodings Databases Monitoring of Systems Python (Programming Language) Metadata Site Reliability Engineering Practices Systems Integration Cloud Platform System
+12 more
Large Language Models Multi-Agent Systems Kubernetes Helm Charts Fastapi Containerization Kubernetes HuggingFace Dataiku Api Design GPT Docker Microservices

Requirements

We are seeking a highly skilled Senior AI Engineer with strong expertise in Kubernetes and vector database technologies to join our team. In this role, you will design, build, and scale production-grade AI systems, working with cutting-edge LLM frameworks, embeddings, and cloud-native infrastructure to deliver robust and high-performance solutions. Responsibilities Deploy and manage Milvus vector databases, including schema design and index tuning (HNSW, IVF-FLAT) Build and maintain embedding and LLM pipelines using OpenAI API, Hugging Face, or Cohere Manage Kubernetes clusters, Helm charts, and containerized microservices in production Develop and maintain Docker containerization workflows, including multi-stage builds and registry management Design and deliver production-grade Python applications, integrating with Go, Java, or C++ where required Integrate object storage systems such as AWS S3, MinIO, or Google Cloud Storage Evaluate and implement alternative vector database solutions, including Qdrant, Pinecone, and Weaviate Collaborate cross-functionally with team members to deliver reliable, scalable AI services Ensure operational excellence, observability, and performance of deployed AI workloads Requirements Bachelor’s degree in Engineering with 5+ years of relevant experience Expertise in Milvus deployment, schema design, and index tuning (HNSW, IVF-FLAT) Familiarity with vector database alternatives such as Qdrant, Pinecone, Weaviate, PGVector, or Chroma Proficiency in building embedding and LLM pipelines using OpenAI API, Hugging Face, or Cohere Skills in Kubernetes cluster management, Helm charts, and containerized microservices Background in Docker containerization, multi-stage builds, and registry management Production-level Python development along with Go, Java, or C++ Knowledge of object storage integration, including AWS S3, MinIO, or Google Cloud Storage Excellent verbal and written communication skills with strong team collaboration abilities Proficiency in English at an Upper-Intermediate level (B2) or higher Nice to have Experience supporting large-scale RAG applications and multi-agent platforms Hands-on familiarity with LangChain, LlamaIndex, or custom pipelines Understanding of GPU scheduling, resource optimization, and inference acceleration Production experience with hybrid search, metadata filtering, and index tuning Implementation of LLM evaluation, governance, tracing, and monitoring tools Familiarity with CI/CD pipelines, Infrastructure-as-Code, and cloud-native deployment practices Prior work experience in the Oil and Gas industry Experience with Dataiku DSS Knowledge of SRE practices

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · World Congress 2025

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all