Performance Modeling Architect - AI Memory Systems

ASGN Incorporated
Boston, MA, United States
8 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$200,000.0 - $250,000.0
Working hours
Regular working hours
Job source

Tech stack

Abstraction Layers Artificial Intelligence Nvidia CUDA Computer Engineering Data Distribution Service Extract Transform Load (ETL) Network Interface Controllers Network Protocols Software Engineering Information Technology Machine Learning Operations

Job description

We are seeking a Member of Technical Staff, Performance Modeling to develop performance models for our fabric-attached memory expansion device for AI accelerators. You’ll work closely with silicon architects and workload teams to explore design tradeoffs, validate performance assumptions, and identify bottlenecks early in the development cycle. This role is well-suited for engineers who enjoy reasoning from first principles, working with incomplete information, and co-exploring the design space as hardware and software evolve together., * Build and maintain system-level performance models for a high-bandwidth data movement device operating in the scale-up domain.

  • Model workload from software memory access patterns to data distribution in the network and all the way down to on-device memory channels.
  • Work day-to-day with silicon architects, system designers, and workload owners to align performance expectations and constraints.
  • Identify performance bottlenecks, scaling limits, and sensitivity points across compute, memory, and interconnects in end-to-end workload settings.
  • Clearly communicate modeling assumptions, limitations, and conclusions to both technical and non-specialist stakeholders.

Requirements

Requirements: AI Memory Systems, Performance Modeling, Memory Expansion (NICs, SmartNICs, CXL, IPU/DPU, NoC), Memory Systems Architecture, ML Systems, CUDA Memory Management, * Bachelor’s or Master’s degree in Electrical Engineering, Computer Engineering, or a closely related field.

  • Ability to quickly learn new ML architectures as soon as they come out, and build performance models for them.
  • 5-10+ years of experience in performance modeling for data movement devices: NICs, memory expansion cards (e.g., CXL), IPU/DPU, NoC.
  • Ability to reason across multiple abstraction layers, from architectural details to system-level performance behavior., * PhD in Computer Science, Electrical Engineering, or a related field.
  • Prior experience modeling performance for networking protocols with memory semantics.
  • Understanding of ML systems: workload sharding, KV caching hierarchies, attention optimizations, trade-offs when deploying ML models at scale, and various assumptions.
  • Familiarity with shared memory systems and frameworks (e.g., CUDA VMM).
  • Experience with scale-up and high-bandwidth interconnects (e.g., NVLink or similar technologies).

Benefits & conditions

Pulled from the full job description Health insurance 401(k) matching Vision insurance Dental insurance Relocation assistance Life insurance Visa sponsorship, * Competitive salary commensurate with experience including base salary, performance-based bonus, and early-stage equity grant

  • Comprehensive benefits including health, dental, vision, and life insurance
  • Well-equipped, sunny offices in Santa Clara, CA and Boston, MA
  • Relocation assistance and visa sponsorship
  • Perks include a daily lunch stipend, 401k match, and more
  • A collaborative, continuous-learning work environment with smart, dedicated colleagues engaged in developing the next generation of architecture for high-performance computing

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

3:04 min

Supporting modern network protocols in foundational libraries

Daniel Stenberg · Coffee With Developers

1:23 min

Handling compatibility and abstraction layers in composable systems

Loïc Carbonne Loïc Carbonne · WWC 2024

4:54 min

NVIDIA local and edge AI hardware capabilities overview

Joerg Krall Joerg Krall · WWC Europe 2026

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · WWC Europe 2026

3:51 min

The historical cost of database abstraction layers

Björn Stahl Björn Stahl · WWC 2024

Videos

See all

Related articles

See all