AI Platform Engineer
Bright Vision Technologies
United States
about 1 month ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$100,000.0 - $180,000.0
Working hours
Regular working hours
Job source
Tech stack
Application Programming Interfaces (APIs)
Artificial Intelligence
Amazon Web Services
Applications Architecture
User Authentication
Microsoft Azure
C++ (Programming Language)
Cloud Computing
Cloud Engineering
Program Optimization
Nvidia CUDA
Computer Programming
+36 more
Computer Engineering
DevOps
Distributed Systems
Memory Management
Python (Programming Language)
Machine Learning
Open Source Technology
Azure Machine Learning
System Programming
Management of Software Versions
AI Infrastructure
Data Logging
Google Cloud
Cloud Platform System
Autoscaling
Istio
System Availability
Large Language Models
Caching
Cloudformation
Event Driven Architecture
AI Platforms
Kubernetes
Information Technology
Deployment Automation
Bicep
Free and Open-Source Software
Linkerd (Service Mesh)
Machine Learning Operations
TensorRT
Cloud Optimization
Api Gateway
Decoding
Terraform
Dynatrace
Docker
Job description
- Design and maintain scalable AI model serving platforms.
- Optimize inference performance, GPU utilization, and request routing.
- Build autoscaling, deployment, and monitoring solutions.
- Implement caching, security, and high-availability strategies.
- Collaborate with ML teams to deploy and support production AI models., Bright Vision Technologies is seeking a highly experienced AI Platform Engineer with 10+ years of experience in distributed systems, cloud-native infrastructure, and AI platform engineering to design, build, and operate enterprise-scale AI inference and machine learning platforms. The ideal candidate will possess deep expertise in LLM serving, GPU optimization, Kubernetes, cloud infrastructure, distributed systems, and MLOps, with a proven ability to deliver highly scalable, reliable, secure, and cost-efficient AI platforms supporting production machine learning workloads., * Design, build, and maintain scalable AI inference and model-serving platforms for enterprise production environments.
- Architect highly available, cloud-native infrastructure supporting Large Language Models (LLMs), foundation models, and machine learning services.
- Optimize inference latency, throughput, GPU utilization, memory management, and request scheduling across distributed AI workloads.
- Design autoscaling, workload orchestration, traffic management, and intelligent request routing strategies for AI services.
- Implement model deployment, versioning, rollback, and lifecycle management using modern MLOps practices.
- Develop monitoring, observability, logging, distributed tracing, and alerting solutions to ensure platform reliability and performance.
- Implement caching strategies, API gateways, security controls, authentication, authorization, and high-availability architectures.
- Collaborate with AI researchers, ML engineers, DevOps teams, and software engineers to deploy and support production AI models.
- Drive cloud infrastructure optimization, resource utilization, FinOps initiatives, and operational excellence.
- Mentor engineering teams, conduct architecture reviews, and establish best practices for AI platform engineering and cloud-native development.
- Evaluate emerging AI infrastructure technologies, model-serving frameworks, and GPU acceleration techniques to drive continuous innovation.
Requirements
- 6+ years of experience in distributed systems, infrastructure, or ML platform engineering.
- Strong proficiency in Python and Go, Rust, or C++.
- Experience with LLM inference frameworks (vLLM, TensorRT-LLM), Kubernetes, cloud platforms, and GPU optimization., Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position., * Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Artificial Intelligence, or a related technical discipline.
- 10+ years of professional experience in distributed systems, infrastructure engineering, cloud platforms, or machine learning platform engineering.
- Strong programming skills in Python and at least one systems programming language such as Go, Rust, or C++.
- Extensive experience with Large Language Model (LLM) serving, model inference optimization, and production AI infrastructure.
- Hands-on experience with vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve, or similar AI serving frameworks.
- Strong expertise in Kubernetes, container orchestration, Docker, and cloud-native application architectures.
- Experience optimizing GPU workloads using CUDA, NVIDIA GPU technologies, distributed inference, and high-performance AI infrastructure.
- Experience with cloud platforms including AWS, Microsoft Azure, or Google Cloud Platform (GCP).
- Strong understanding of distributed systems, networking, scalability, observability, and security best practices.
- Excellent analytical, communication, collaboration, and technical leadership skills., * Experience designing and operating multi-region AI platforms and globally distributed inference services.
- Knowledge of model optimization techniques such as quantization, pruning, compression, speculative decoding, KV cache optimization, and mixed-precision inference.
- Experience with MLOps, GitOps, Infrastructure as Code (Terraform, Bicep, CloudFormation), and CI/CD automation.
- Familiarity with service mesh technologies such as Istio or Linkerd, API gateways, and event-driven architectures.
- Contributions to open-source AI infrastructure projects, technical publications, patents, or conference presentations.
-
Experience implementing FinOps strategies, cloud cost optimization, and enterprise AI governance.
- Experience with multi-region AI deployments and AI infrastructure.
- Familiarity with model optimization techniques such as quantization or compression.
- Open-source contributions or experience supporting large-scale AI APIs.
Benefits & conditions
4.2 Remote Remote $100,000 - $180,000 a year - Full-time
About the company
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
almost 3 years ago
BB
Benedikt Bischof
MLOps And AI Driven Development
over 4 years ago
DC
Daniel Cranney
Stephan Gillich - Bringing AI Everywhere
almost 2 years ago
ER
Erin Rifkin
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
about 1 year ago
BB
Benedikt Bischof
MLOps – What’s the deal behind it?
almost 4 years ago
LM
Luis Minvielle
What Are Large Language Models?
almost 3 years ago