AI Engineer

AgreeYa Solutions, Inc.
United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Nvidia CUDA Monitoring of Systems Python (Programming Language) Machine Learning Performance Tuning Google Cloud Retrieval-Augmented Generation Large Language Models Generative AI
+8 more
Containerization Kubernetes Machine Learning Operations TensorRT Nim (Programming Language) Restful APIs Docker Microservices

Job description

We are seeking a Senior AI Engineer with strong experience in NVIDIA AI technologies, specifically NVIDIA NIM Microservices and Triton Inference Server. The ideal candidate will be responsible for designing, deploying, optimizing, and scaling Generative AI and LLM-based applications in enterprise environments., * Design and deploy AI applications using NVIDIA NIM Microservices

  • Build and optimize model serving infrastructure using Triton Inference Server
  • Deploy and manage LLM workloads in Kubernetes environments
  • Optimize inference performance using TensorRT-LLM and CUDA
  • Collaborate with Data Science, MLOps, and Platform Engineering teams
  • Implement scalable, secure, and production-ready AI solutions
  • Troubleshoot and improve AI application performance and reliability
  • Support cloud-based AI deployments across AWS, Azure, or Google Cloud Platform

Requirements

  • Hands-on experience with NVIDIA NIM Microservices
  • Strong experience with NVIDIA Triton Inference Server
  • Experience deploying and serving Large Language Models (LLMs)
  • Knowledge of TensorRT-LLM and CUDA optimization
  • Experience with Kubernetes and Docker containerization
  • Strong Python programming skills
  • Experience building AI/ML applications in AWS, Azure, or Google Cloud Platform
  • Understanding of model inference, model serving, and performance tuning
  • Experience with REST APIs and microservices architecture

Preferred Skills

  • Experience with NVIDIA NeMo
  • Experience with RAG (Retrieval-Augmented Generation) architectures
  • Familiarity with LangChain or LlamaIndex
  • Exposure to MLOps/LLMOps practices
  • Experience with monitoring and observability tools

About the company

AgreeYa is a global systems integrator delivering a competitive advantage for its customers through software, solutions, and services. Established in 1999, AgreeYa is headquartered in Folsom, California, with a global footprint and a team of more than 1,800+ professionals across offices. AgreeYa works with 550+ organizations ranging from Fortune 100 firms to small and large businesses across industries such as Telecom, Banking, Financial Services & Insurance, Healthcare, Utility & Energy, Technology, Public Sector, Pharma & Biotech, Retail, Client, and others. Please visit us at for more information.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · WWC 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

5:01 min

Leveraging large language models for code optimization and development

Stephan Gillich Stephan Gillich +3 · WWC 2024

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · WWC 2022

Videos

See all

Related articles

See all