AI Platform Engineer

Bright Vision Technologies
Raleigh, NC, United States
3 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$100,000.0 - $150,000.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Applications Architecture User Authentication Microsoft Azure C++ (Programming Language) Cloud Computing Cloud Engineering Program Optimization Nvidia CUDA Computer Programming
+36 more
Computer Engineering DevOps Distributed Systems Memory Management Python (Programming Language) Machine Learning Open Source Technology Azure Machine Learning System Programming Management of Software Versions AI Infrastructure Data Logging Google Cloud Cloud Platform System Autoscaling Istio System Availability Large Language Models Caching Cloudformation Event Driven Architecture AI Platforms Kubernetes Information Technology Deployment Automation Bicep Free and Open-Source Software Linkerd (Service Mesh) Machine Learning Operations TensorRT Cloud Optimization Api Gateway Decoding Terraform Dynatrace Docker

Job description

We are seeking an AI Platform Engineer to design, build, and operate scalable AI inference platforms for production ML workloads. The ideal candidate will have expertise in distributed systems, LLM serving, GPU optimization, autoscaling, and cloud-native infrastructure, with a strong focus on performance, reliability, and observability., * Design and maintain scalable AI model serving platforms.

  • Optimize inference performance, GPU utilization, and request routing.
  • Build autoscaling, deployment, and monitoring solutions.
  • Implement caching, security, and high-availability strategies.
  • Collaborate with ML teams to deploy and support production AI models., Bright Vision Technologies is seeking a highly experienced AI Platform Engineer with 10+ years of experience in distributed systems, cloud-native infrastructure, and AI platform engineering to design, build, and operate enterprise-scale AI inference and machine learning platforms. The ideal candidate will possess deep expertise in LLM serving, GPU optimization, Kubernetes, cloud infrastructure, distributed systems, and MLOps, with a proven ability to deliver highly scalable, reliable, secure, and cost-efficient AI platforms supporting production machine learning workloads. Key Responsibilities

  • Design, build, and maintain scalable AI inference and model-serving platforms for enterprise production environments.
  • Architect highly available, cloud-native infrastructure supporting Large Language Models (LLMs), foundation models, and machine learning services.
  • Optimize inference latency, throughput, GPU utilization, memory management, and request scheduling across distributed AI workloads.
  • Design autoscaling, workload orchestration, traffic management, and intelligent request routing strategies for AI services.
  • Implement model deployment, versioning, rollback, and lifecycle management using modern MLOps practices.
  • Develop monitoring, observability, logging, distributed tracing, and alerting solutions to ensure platform reliability and performance.
  • Implement caching strategies, API gateways, security controls, authentication, authorization, and high-availability architectures.
  • Collaborate with AI researchers, ML engineers, DevOps teams, and software engineers to deploy and support production AI models.
  • Drive cloud infrastructure optimization, resource utilization, FinOps initiatives, and operational excellence.
  • Mentor engineering teams, conduct architecture reviews, and establish best practices for AI platform engineering and cloud-native development.
  • Evaluate emerging AI infrastructure technologies, model-serving frameworks, and GPU acceleration techniques to drive continuous innovation.

Requirements

  • 6+ years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and Go, Rust, or C++.
  • Experience with LLM inference frameworks (vLLM, TensorRT-LLM), Kubernetes, cloud platforms, and GPU optimization., * Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Artificial Intelligence, or a related technical discipline.
  • 10+ years of professional experience in distributed systems, infrastructure engineering, cloud platforms, or machine learning platform engineering.
  • Strong programming skills in Python and at least one systems programming language such as Go, Rust, or C++.
  • Extensive experience with Large Language Model (LLM) serving, model inference optimization, and production AI infrastructure.
  • Hands-on experience with vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve, or similar AI serving frameworks.
  • Strong expertise in Kubernetes, container orchestration, Docker, and cloud-native application architectures.
  • Experience optimizing GPU workloads using CUDA, NVIDIA GPU technologies, distributed inference, and high-performance AI infrastructure.
  • Experience with cloud platforms including AWS, Microsoft Azure, or Google Cloud Platform (GCP).
  • Strong understanding of distributed systems, networking, scalability, observability, and security best practices.
  • Excellent analytical, communication, collaboration, and technical leadership skills., * Experience designing and operating multi-region AI platforms and globally distributed inference services.
  • Knowledge of model optimization techniques such as quantization, pruning, compression, speculative decoding, KV cache optimization, and mixed-precision inference.
  • Experience with MLOps, GitOps, Infrastructure as Code (Terraform, Bicep, CloudFormation), and CI/CD automation.
  • Familiarity with service mesh technologies such as Istio or Linkerd, API gateways, and event-driven architectures.
  • Contributions to open-source AI infrastructure projects, technical publications, patents, or conference presentations.
  • Experience implementing FinOps strategies, cloud cost optimization, and enterprise AI governance.

  • Experience with multi-region AI deployments and AI infrastructure.
  • Familiarity with model optimization techniques such as quantization or compression.
  • Open-source contributions or experience supporting large-scale AI APIs.

About the company

Learn more about Bright Vision Technologies at . We look forward to connecting with talented professionals and helping you take the next step in your career. Bright Vision Technologies is an Equal Opportunity Employer. Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees’ ability to perform their job duties may result in disciplinary action up to and including termination of employment. Powered by JazzHR, Gilead

  • Raleigh, NC
  • $146,200-189,200 per year At Gilead, we’re creating a healthier world for all people. For more than 35 years, we’ve tackled diseases such as HIV, viral hepatitis, COVID-19 and cancer - working relentlessly …

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · World Congress 2022

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

7:15 min

Installing Istio programmatically with bash scripts

Thomas Südbröcker · LIVE

Videos

See all

Related articles

See all