> Markdown version of [/jobs/ext/2618185-ai-performance-modeling-engineer](https://www.wearedevelopers.com/jobs/ext/2618185-ai-performance-modeling-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Performance Modeling Engineer - **Company:** quadric, Inc - **Location:** Burlingame, CA, United States - **Salary:** $150,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Nvidia CUDA, Computer Engineering, Extract Transform Load (ETL), Python (Programming Language), Performance Tuning, Graphics Processing Unit (GPU), Large Language Models, Information Technology, Low Latency - **Published:** August 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=681c4f816ddda3ec ## About the Role * Python & Quantitative Modeling: Strong Python skills with experience writing, validating, and calibrating numerical or quantitative models in code. * Computer Architecture Fundamentals: Solid grasp of memory hierarchies, bandwidth/latency trade-offs, pipelining, and execution bottlenecks (via industry experience, coursework, or research). * Technical Writing: Comfort writing clear technical studies that state and defend evidence-based conclusions. * Core Technical Depth (One of the following): * + Option A: Deep understanding of NN inference operators and tensor shapes (e.g., Transformers, attention mechanisms, MoE, prefill/decode split). + Option B: Proven performance modeling experience in another quantitative/technical domain. * Education: BS, MS, or Ph.D. in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience. Preferred * Prior experience with GPUs, custom AI accelerators, CUDA, or Triton kernels. * Familiarity with roofline analysis, back-of-the-envelope estimation, or architecture simulators (e.g., gem5, Timeloop, MAESTRO, Accel-Sim). * Background in compiler internals (cost models, autotuners) or proficiency in C++. * Published performance studies or technical write-ups. ## Description As an AI Performance Modeling Engineer, you will build analytical, cycle-level performance models of AI inference workloads on our next-generation architecture in Python before silicon exists. These models directly guide team decisions on hardware lane bindings, tensor placement, and architecture trade-offs. We welcome candidates across all experience levels-from early-career engineers to seasoned experts-with direct mentorship provided to help you master mapping complex workloads (like LLMs) onto our custom hardware. What You'll DoPerformance Modeling & Architectural Analysis * Build analytical, cycle-level Python models of AI inference workloads executing on next-generation GPNPU hardware. * Derive from first principles which hardware lanes operations bind on (compute, on-chip/external memory bandwidth, interconnect) and model software pipelining overlaps. * Model tensor placement, tiling across processing elements, local memory residency, and data movement across memory tiers. * Model sharding and collective boundary communication across multi-die systems. Workload Adaptation & Technical Writing * Incorporate architectural details across vision networks and Large Language Models (LLMs), including operator mix, sparsity, routing, and quantization/low-precision numeric formats. * Calibrate performance models against an instruction-set simulator and profiling traces to meet stated accuracy targets. * Write and defend technical studies presenting empirical evidence that directly informs architecture and product decisions. * Balance single-stream latency against scaled throughput performance. What Success Looks Like Within your first 6-12 months, you'll: * Own a full workload's model end to end, calibrated against simulation and trusted by the engineering team. * Build performance models that consistently predict workload behavior within 10-15% of actual measurements. * Publish a written study whose defended conclusions directly shape an architecture or product decision. * Review and extend performance models beyond your initial starting domain. ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Practical performance tuning for Serverless Java on AWS](https://www.wearedevelopers.com/videos/2075-practical-performance-tuning-for-serverless-java-on-aws) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)