> Markdown version of [/jobs/ext/1958484-ai-systems-performance-specialist](https://www.wearedevelopers.com/jobs/ext/1958484-ai-systems-performance-specialist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Systems Performance Specialist - **Company:** Bright Vision Technologies - **Location:** Apex, NC, United States (Remote available) - **Experience:** Expert - **Salary:** $130,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, C++ (Programming Language), Profiling, Nvidia CUDA, Information Systems, Computer Programming, Computer Engineering, Distributed Computing Environment, Distributed Systems, Memory Management, Python (Programming Language), Open Source Technology, Performance Tuning, Regression Testing, AI Infrastructure, Google Cloud, High Performance Computing, Pytorch, Large Language Models, Deep Learning, Information Technology, Low Latency, Optimization Algorithms, Hardware Acceleration, Machine Learning Operations, TensorRT, Cloud Optimization, Decoding - **Published:** August 6, 2026 - **Apply:** https://www.careerjet.com/jobad/usd9f6f5bfc3789c6f81847525d8e8b7d0 ## About the Role * Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, or a related technical discipline. * 10+ years of professional experience in performance engineering, AI infrastructure, machine learning systems, High-Performance Computing (HPC), or distributed computing. * Expert-level programming skills in Python and C++. * Extensive experience optimizing GPU-accelerated AI workloads using CUDA, distributed training frameworks, and modern deep learning libraries. * Strong knowledge of Large Language Models (LLMs), deep learning frameworks, model serving, and production AI inference. * Hands-on experience with profiling tools such as NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, TensorBoard, or similar performance analysis tools. * Experience deploying and optimizing AI workloads on AWS, Microsoft Azure, or Google Cloud Platform (GCP). * Strong understanding of distributed systems, networking, storage optimization, and AI infrastructure architecture. * Excellent analytical, troubleshooting, communication, and technical leadership skills. Preferred Qualifications * Experience optimizing production-scale LLM inference and serving large foundation models. * Hands-on experience with vLLM, TensorRT-LLM, DeepSpeed, Triton Inference Server, CUTLASS, FasterTransformer, or similar AI optimization frameworks. * Knowledge of model compression, KV cache optimization, speculative decoding, and advanced inference optimization techniques. * Experience implementing FinOps strategies for AI infrastructure cost optimization and resource management. * Contributions to AI systems research, open-source AI infrastructure projects, patents, or technical publications. * Familiarity with emerging AI accelerator technologies, including AMD ROCm, Intel oneAPI, or custom AI hardware. ## Description Bright Vision Technologies is seeking a highly experienced AI Systems Performance Specialist with 10+ years of experience in AI infrastructure, machine learning systems, High-Performance Computing (HPC), and performance engineering. The ideal candidate will optimize AI training and inference workloads for maximum performance, scalability, reliability, and cost efficiency. This role requires deep expertise in GPU optimization, distributed training, Large Language Model (LLM) inference, Python, C++, CUDA, and production AI systems, along with the ability to lead performance optimization initiatives across enterprise-scale AI platforms., * Optimize AI training and inference pipelines for maximum throughput, low latency, scalability, and infrastructure efficiency. * Analyze and improve GPU utilization, memory management, kernel execution, and multi-GPU performance across production AI workloads. * Design and implement optimization techniques including quantization, pruning, mixed precision, batching, caching, speculative decoding, and model parallelism. * Profile AI applications using industry-standard performance analysis tools and identify bottlenecks across compute, memory, networking, and storage. * Optimize distributed training and inference using NCCL, DeepSpeed, PyTorch Distributed, Ray, MPI, or similar distributed computing frameworks. * Collaborate with AI researchers, ML engineers, platform engineers, and infrastructure teams to improve model performance and production reliability. * Build automated benchmarking frameworks, performance dashboards, monitoring solutions, and regression testing pipelines. * Evaluate emerging AI hardware, GPU architectures, inference frameworks, and optimization technologies to improve enterprise AI capabilities. * Drive AI infrastructure cost optimization through efficient resource utilization, cloud optimization, and FinOps best practices. * Mentor engineering teams and provide technical leadership on AI systems architecture, GPU optimization, and performance engineering., Description: ABOUT THE POSITION Collections Specialist is responsible for proactively managing delinquent accounts by contacting Credit Union members. This role involves identi… + 1 day ago, Overview: The Pharmacy Financial Systems Specialist for procurement assists pharmacy management by maintaining information systems that support financial, operational and procureme… + 1 month ago, Job Description: Overview The Pharmacy Financial Systems Specialist for procurement assists pharmacy management by maintaining information systems that support financial, opera… + 1 month ago + ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [AI That Fits Your Business, Not the Other Way Around](https://www.wearedevelopers.com/videos/100148-ai-that-fits-your-business-not-the-other-way-around) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)