CPU Cluster Performance Architect
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
- Analyze and optimize CPU cluster performance, including cache hierarchies, interconnects, coherence, and memory access behavior.
- Build and use performance models and simulation environments such as Gem5 or equivalent to evaluate architectural concepts and identify performance opportunities.
- Study real workloads to understand bottlenecks related to cache misses, coherence traffic, memory bandwidth, latency, contention, and core-to-core communication.
- Lead architectural tradeoff studies across performance, scalability, bandwidth, latency, power, and implementation complexity.
- Collaborate with CPU, cache, interconnect, memory, software, and system architects to translate performance analysis into concrete design decisions.
What You Will Learn
- How to architect, model, and optimize a multi-core CPU cluster from early performance exploration through implementation and silicon.
- How cache coherence, interconnects, memory hierarchy, and core architecture interact to determine overall system performance.
- How to model and reason about scaling across cores, including contention, bandwidth limits, latency, and workload behavior.
- How architectural decisions at the cluster level translate into real-world application performance.
- How Tenstorrent approaches scalable CPU and chiplet-based system architecture across compute, interconnect, and memory.
Requirements
- You have a strong foundation in CPU microarchitecture and performance, with an understanding of how cores, caches, interconnects, and memory systems interact.
- You enjoy using modeling, simulation, and workload analysis to understand cluster-level performance bottlenecks and scaling challenges.
- You’re comfortable moving between detailed microarchitecture and broader system-level questions around latency, bandwidth, quality of service, utilization, and scalability.
- You’re analytical and hands-on, with the ability to turn large amounts of performance data into clear architectural recommendations.
- You’re a strong communicator who enjoys working across architecture, RTL, compiler, software, and system teams.
Benefits & conditions
Compensation for all engineers at Tenstorrent ranges from $100k - $500k including base and variable compensation targets. Experience, skills, education, background and location all impact the actual offer made.
Tenstorrent offers a highly competitive compensation package and benefits, and we are an equal opportunity employer.
About the company
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities.
Tenstorrent is looking for a CPU Cluster Performance Architect to help shape the performance and scalability of our next-generation CPUs. You’ll focus on how multiple CPU cores work together as a cluster - from cache hierarchies and coherent interconnects, to memory bandwidth and overall system behavior. This is a hands-on architecture role for someone who enjoys using performance models and simulation to understand bottlenecks, explore tradeoffs, and influence decisions early in the design process. You’ll work closely with CPU architects, RTL designers, software and compiler teams, and system architects to understand workload behavior and turn performance insights into architectural improvements. If you enjoy looking at the entire cluster and how performance scales across cores, memory systems, and chiplets, this is an opportunity to have a significant impact.
This role is remote, based out of North America.
We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Why Event-Driven Architecture Isn’t About Speed (and When You Actually Need It)
Making Data Warehouses Fast: A Developer’s Story
Everything a Developer Needs to Know About MCP with Neo4j
7 Cloud Computing Trends Coming in 2025 for Developers