> Markdown version of [/jobs/ext/2696864-ai-performance-engineer](https://www.wearedevelopers.com/jobs/ext/2696864-ai-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Performance Engineer - **Company:** Infosys - **Location:** Phoenix, AZ, United States - **Salary:** $114,400.0 - $124,800.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Profiling, Nvidia CUDA, Computer Engineering, System Configuration, Linux, Distributed Computing Environment, General-Purpose Computing on Graphics Processing Units, InfiniBand, Python (Programming Language), Regression Testing, Network Switches, Pytorch, Large Language Models, Deep Learning, Gpu Programming, ONNX (Open Neural Network Exchange) Format, TensorRT - **Published:** September 3, 2026 - **Apply:** https://www.dice.com/job-detail/9c2c52fd-a187-4e0a-ac84-6237c27d3cfc ## About the Role BS/MS in CS, Computer Engineering, EE, or equivalent practical experience Strong Python; working proficiency in at least one systems language (C++/Rust/C) Hands-on experience with a deep-learning framework (PyTorch preferred), including model execution, export, and profiling Demonstrated experience delivering measurable performance improvements in DL training or inference Solid grounding in computer architecture: memory hierarchy, bandwidth vs. compute limits, parallelism Ability to reason quantitatively about latency, throughput, batching, memory footprint, and utilization under real workloads Fluency with Linux and GPU computing environments Preferred qualifications GPU programming (CUDA, Triton, ROCm/HIP) and low-level libraries (cuBLAS, cuDNN, CUTLASS) Inference runtimes/serving engines: TensorRT(-LLM), ONNX Runtime, vLLM, SGLang, Triton Inference Server LLM inference mechanics: attention, KV caching, prefill vs. decode, continuous batching, speculative decoding Distributed training/inference: data/tensor/pipeline parallelism, NCCL, InfiniBand/RoCE Model compression research or MLPerf-style benchmarking experience Edge/on-device deployment (Jetson, NPUs, Core ML) if your systems include e ## Description Benchmark AI models (LLMs, vision, multimodal) across hardware configurations; measure latency, throughput, utilization, memory behavior, and scaling efficiency Profile workloads end-to-end using tools such as Nsight Systems/Compute, PyTorch Profiler, and system telemetry (nvidia-smi, DCGM) to isolate bottlenecks Build roofline/performance models to quantify achieved vs. theoretical performance and prioritize the highest-impact optimizations Apply and evaluate optimizations: quantization, pruning, distillation, operator/kernel fusion, graph compilation, KV-cache management, batching strategies, speculative decoding Recommend hardware/system configurations (GPU selection, memory sizing, interconnect, storage/network I/O) for given model workloads Establish performance baselines, SLAs, and regression testing so models stay fast as they evolve Write clear analyses and recommendations for engineering and leadership audiences ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)