> Markdown version of [/jobs/ext/1911675-performance-modeling-architect-ai-memory-systems](https://www.wearedevelopers.com/jobs/ext/1911675-performance-modeling-architect-ai-memory-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Performance Modeling Architect - AI Memory Systems - **Company:** ASGN Incorporated - **Location:** Boston, MA, United States - **Experience:** Expert - **Salary:** $200,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Abstraction Layers, Artificial Intelligence, Nvidia CUDA, Computer Engineering, Data Distribution Service, Extract Transform Load (ETL), Network Interface Controllers, Network Protocols, Software Engineering, Information Technology, Machine Learning Operations - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a4dc9f0cd52d541d ## About the Role Requirements: AI Memory Systems, Performance Modeling, Memory Expansion (NICs, SmartNICs, CXL, IPU/DPU, NoC), Memory Systems Architecture, ML Systems, CUDA Memory Management, * Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, or a closely related field. * Ability to quickly learn new ML architectures as soon as they come out, and build performance models for them. * 5-10+ years of experience in performance modeling for data movement devices: NICs, memory expansion cards (e.g., CXL), IPU/DPU, NoC. * Ability to reason across multiple abstraction layers, from architectural details to system-level performance behavior., * PhD in Computer Science, Electrical Engineering, or a related field. * Prior experience modeling performance for networking protocols with memory semantics. * Understanding of ML systems: workload sharding, KV caching hierarchies, attention optimizations, trade-offs when deploying ML models at scale, and various assumptions. * Familiarity with shared memory systems and frameworks (e.g., CUDA VMM). * Experience with scale-up and high-bandwidth interconnects (e.g., NVLink or similar technologies). ## Description We are seeking a Member of Technical Staff, Performance Modeling to develop performance models for our fabric-attached memory expansion device for AI accelerators. You'll work closely with silicon architects and workload teams to explore design tradeoffs, validate performance assumptions, and identify bottlenecks early in the development cycle. This role is well-suited for engineers who enjoy reasoning from first principles, working with incomplete information, and co-exploring the design space as hardware and software evolve together., * Build and maintain system-level performance models for a high-bandwidth data movement device operating in the scale-up domain. * Model workload from software memory access patterns to data distribution in the network and all the way down to on-device memory channels. * Work day-to-day with silicon architects, system designers, and workload owners to align performance expectations and constraints. * Identify performance bottlenecks, scaling limits, and sensitivity points across compute, memory, and interconnects in end-to-end workload settings. * Clearly communicate modeling assumptions, limitations, and conclusions to both technical and non-specialist stakeholders. ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Cloud Vendor Lock-In - Is it just a new version of the Database Abstraction Layers?](https://www.wearedevelopers.com/videos/1185-cloud-vendor-lock-in-is-it-just-a-new-version-of-the-database-abstraction-layers) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) ## Related Articles - [Introducing Redis Agent Memory Server](https://www.wearedevelopers.com/magazine/699-introducing-redis-agent-memory-server) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)