Performance Modeling Architect - AI Memory Systems
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
We are seeking a Member of Technical Staff, Performance Modeling to develop performance models for our fabric-attached memory expansion device for AI accelerators. You’ll work closely with silicon architects and workload teams to explore design tradeoffs, validate performance assumptions, and identify bottlenecks early in the development cycle. This role is well-suited for engineers who enjoy reasoning from first principles, working with incomplete information, and co-exploring the design space as hardware and software evolve together., * Build and maintain system-level performance models for a high-bandwidth data movement device operating in the scale-up domain.
- Model workload from software memory access patterns to data distribution in the network and all the way down to on-device memory channels.
- Work day-to-day with silicon architects, system designers, and workload owners to align performance expectations and constraints.
- Identify performance bottlenecks, scaling limits, and sensitivity points across compute, memory, and interconnects in end-to-end workload settings.
- Clearly communicate modeling assumptions, limitations, and conclusions to both technical and non-specialist stakeholders.
Requirements
Requirements: AI Memory Systems, Performance Modeling, Memory Expansion (NICs, SmartNICs, CXL, IPU/DPU, NoC), Memory Systems Architecture, ML Systems, CUDA Memory Management, * Bachelor’s or Master’s degree in Electrical Engineering, Computer Engineering, or a closely related field.
- Ability to quickly learn new ML architectures as soon as they come out, and build performance models for them.
- 5-10+ years of experience in performance modeling for data movement devices: NICs, memory expansion cards (e.g., CXL), IPU/DPU, NoC.
- Ability to reason across multiple abstraction layers, from architectural details to system-level performance behavior., * PhD in Computer Science, Electrical Engineering, or a related field.
- Prior experience modeling performance for networking protocols with memory semantics.
- Understanding of ML systems: workload sharding, KV caching hierarchies, attention optimizations, trade-offs when deploying ML models at scale, and various assumptions.
- Familiarity with shared memory systems and frameworks (e.g., CUDA VMM).
- Experience with scale-up and high-bandwidth interconnects (e.g., NVLink or similar technologies).
Benefits & conditions
Pulled from the full job description Health insurance 401(k) matching Vision insurance Dental insurance Relocation assistance Life insurance Visa sponsorship, * Competitive salary commensurate with experience including base salary, performance-based bonus, and early-stage equity grant
- Comprehensive benefits including health, dental, vision, and life insurance
- Well-equipped, sunny offices in Santa Clara, CA and Boston, MA
- Relocation assistance and visa sponsorship
- Perks include a daily lunch stipend, 401k match, and more
- A collaborative, continuous-learning work environment with smart, dedicated colleagues engaged in developing the next generation of architecture for high-performance computing
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?
Everything a Developer Needs to Know About MCP with Neo4j
MLOps And AI Driven Development
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud