AI Performance Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+7 more
Job description
Benchmark AI models (LLMs, vision, multimodal) across hardware configurations; measure latency, throughput, utilization, memory behavior, and scaling efficiency Profile workloads end-to-end using tools such as Nsight Systems/Compute, PyTorch Profiler, and system telemetry (nvidia-smi, DCGM) to isolate bottlenecks Build roofline/performance models to quantify achieved vs. theoretical performance and prioritize the highest-impact optimizations Apply and evaluate optimizations: quantization, pruning, distillation, operator/kernel fusion, graph compilation, KV-cache management, batching strategies, speculative decoding Recommend hardware/system configurations (GPU selection, memory sizing, interconnect, storage/network I/O) for given model workloads Establish performance baselines, SLAs, and regression testing so models stay fast as they evolve Write clear analyses and recommendations for engineering and leadership audiences
Requirements
BS/MS in CS, Computer Engineering, EE, or equivalent practical experience Strong Python; working proficiency in at least one systems language (C++/Rust/C) Hands-on experience with a deep-learning framework (PyTorch preferred), including model execution, export, and profiling Demonstrated experience delivering measurable performance improvements in DL training or inference Solid grounding in computer architecture: memory hierarchy, bandwidth vs. compute limits, parallelism Ability to reason quantitatively about latency, throughput, batching, memory footprint, and utilization under real workloads Fluency with Linux and GPU computing environments Preferred qualifications GPU programming (CUDA, Triton, ROCm/HIP) and low-level libraries (cuBLAS, cuDNN, CUTLASS) Inference runtimes/serving engines: TensorRT(-LLM), ONNX Runtime, vLLM, SGLang, Triton Inference Server LLM inference mechanics: attention, KV caching, prefill vs. decode, continuous batching, speculative decoding Distributed training/inference: data/tensor/pipeline parallelism, NCCL, InfiniBand/RoCE Model compression research or MLPerf-style benchmarking experience Edge/on-device deployment (Jetson, NPUs, Core ML) if your systems include e
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
Stephan Gillich - Bringing AI Everywhere
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence