> Markdown version of [/jobs/ext/558764-ai-engineer](https://www.wearedevelopers.com/jobs/ext/558764-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Engineer - **Company:** AgreeYa Solutions, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Nvidia CUDA, Monitoring of Systems, Python (Programming Language), Machine Learning, Performance Tuning, Google Cloud, Retrieval-Augmented Generation, Large Language Models, Generative AI, Containerization, Kubernetes, Machine Learning Operations, TensorRT, Nim (Programming Language), Restful APIs, Docker, Microservices - **Published:** June 15, 2026 - **Apply:** https://www.dice.com/job-detail/7537ccc6-17ec-4399-a00b-1d32c80538a6 ## About the Role * Hands-on experience with NVIDIA NIM Microservices * Strong experience with NVIDIA Triton Inference Server * Experience deploying and serving Large Language Models (LLMs) * Knowledge of TensorRT-LLM and CUDA optimization * Experience with Kubernetes and Docker containerization * Strong Python programming skills * Experience building AI/ML applications in AWS, Azure, or Google Cloud Platform * Understanding of model inference, model serving, and performance tuning * Experience with REST APIs and microservices architecture Preferred Skills * Experience with NVIDIA NeMo * Experience with RAG (Retrieval-Augmented Generation) architectures * Familiarity with LangChain or LlamaIndex * Exposure to MLOps/LLMOps practices * Experience with monitoring and observability tools ## Description We are seeking a Senior AI Engineer with strong experience in NVIDIA AI technologies, specifically NVIDIA NIM Microservices and Triton Inference Server. The ideal candidate will be responsible for designing, deploying, optimizing, and scaling Generative AI and LLM-based applications in enterprise environments., * Design and deploy AI applications using NVIDIA NIM Microservices * Build and optimize model serving infrastructure using Triton Inference Server * Deploy and manage LLM workloads in Kubernetes environments * Optimize inference performance using TensorRT-LLM and CUDA * Collaborate with Data Science, MLOps, and Platform Engineering teams * Implement scalable, secure, and production-ready AI solutions * Troubleshoot and improve AI application performance and reliability * Support cloud-based AI deployments across AWS, Azure, or Google Cloud Platform ## Related Videos - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Microservices: how to get started with Spring Boot and Kubernetes](https://www.wearedevelopers.com/videos/242-microservices-how-to-get-started-with-spring-boot-and-kubernetes) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)