AI Performance Engineer

Infosys
Phoenix, AZ, United States
2 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$114,400.0 - $124,800.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence C++ (Programming Language) Profiling Nvidia CUDA Computer Engineering System Configuration Linux Distributed Computing Environment General-Purpose Computing on Graphics Processing Units InfiniBand Python (Programming Language) Regression Testing
+7 more
Network Switches Pytorch Large Language Models Deep Learning Gpu Programming ONNX (Open Neural Network Exchange) Format TensorRT

Job description

Benchmark AI models (LLMs, vision, multimodal) across hardware configurations; measure latency, throughput, utilization, memory behavior, and scaling efficiency Profile workloads end-to-end using tools such as Nsight Systems/Compute, PyTorch Profiler, and system telemetry (nvidia-smi, DCGM) to isolate bottlenecks Build roofline/performance models to quantify achieved vs. theoretical performance and prioritize the highest-impact optimizations Apply and evaluate optimizations: quantization, pruning, distillation, operator/kernel fusion, graph compilation, KV-cache management, batching strategies, speculative decoding Recommend hardware/system configurations (GPU selection, memory sizing, interconnect, storage/network I/O) for given model workloads Establish performance baselines, SLAs, and regression testing so models stay fast as they evolve Write clear analyses and recommendations for engineering and leadership audiences

Requirements

BS/MS in CS, Computer Engineering, EE, or equivalent practical experience Strong Python; working proficiency in at least one systems language (C++/Rust/C) Hands-on experience with a deep-learning framework (PyTorch preferred), including model execution, export, and profiling Demonstrated experience delivering measurable performance improvements in DL training or inference Solid grounding in computer architecture: memory hierarchy, bandwidth vs. compute limits, parallelism Ability to reason quantitatively about latency, throughput, batching, memory footprint, and utilization under real workloads Fluency with Linux and GPU computing environments Preferred qualifications GPU programming (CUDA, Triton, ROCm/HIP) and low-level libraries (cuBLAS, cuDNN, CUTLASS) Inference runtimes/serving engines: TensorRT(-LLM), ONNX Runtime, vLLM, SGLang, Triton Inference Server LLM inference mechanics: attention, KV caching, prefill vs. decode, continuous batching, speculative decoding Distributed training/inference: data/tensor/pipeline parallelism, NCCL, InfiniBand/RoCE Model compression research or MLPerf-style benchmarking experience Edge/on-device deployment (Jetson, NPUs, Core ML) if your systems include e

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:18 min

Hardware architectures tailored for specific artificial intelligence computations

Stephan Gillich Stephan Gillich · World Congress 2024

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · World Congress 2024

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

Videos

See all

Related articles

See all