> Markdown version of [/jobs/ext/3000843-cpu-cluster-performance-architect](https://www.wearedevelopers.com/jobs/ext/3000843-cpu-cluster-performance-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # CPU Cluster Performance Architect - **Company:** Tenstorrent Usa, Inc. - **Location:** Santa Clara, United States (Remote available) - **Salary:** $100,000.0 - **Contract:** Permanent contract - **Skills:** Application Performance Management, Multiprocessing, Systems Architecture, Network Switches, Caching - **Published:** September 19, 2026 - **Apply:** https://startup.jobs/performance-architect-cpu-cluster-tenstorrent-company-10106983 ## About the Role * You have a strong foundation in CPU microarchitecture and performance, with an understanding of how cores, caches, interconnects, and memory systems interact. * You enjoy using modeling, simulation, and workload analysis to understand cluster-level performance bottlenecks and scaling challenges. * You're comfortable moving between detailed microarchitecture and broader system-level questions around latency, bandwidth, quality of service, utilization, and scalability. * You're analytical and hands-on, with the ability to turn large amounts of performance data into clear architectural recommendations. * You're a strong communicator who enjoys working across architecture, RTL, compiler, software, and system teams. ## Description * Analyze and optimize CPU cluster performance, including cache hierarchies, interconnects, coherence, and memory access behavior. * Build and use performance models and simulation environments such as Gem5 or equivalent to evaluate architectural concepts and identify performance opportunities. * Study real workloads to understand bottlenecks related to cache misses, coherence traffic, memory bandwidth, latency, contention, and core-to-core communication. * Lead architectural tradeoff studies across performance, scalability, bandwidth, latency, power, and implementation complexity. * Collaborate with CPU, cache, interconnect, memory, software, and system architects to translate performance analysis into concrete design decisions. What You Will Learn * How to architect, model, and optimize a multi-core CPU cluster from early performance exploration through implementation and silicon. * How cache coherence, interconnects, memory hierarchy, and core architecture interact to determine overall system performance. * How to model and reason about scaling across cores, including contention, bandwidth limits, latency, and workload behavior. * How architectural decisions at the cluster level translate into real-world application performance. * How Tenstorrent approaches scalable CPU and chiplet-based system architecture across compute, interconnect, and memory. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Challenges and Solutions for Efficient, Large-Scale Video Analysis](https://www.wearedevelopers.com/videos/2022-challenges-and-solutions-for-efficient-large-scale-video-analysis) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [In-Memory Computing - The Big Picture](https://www.wearedevelopers.com/videos/626-in-memory-computing-the-big-picture) - [How to implement convenient Python bindings to C++](https://www.wearedevelopers.com/videos/618-how-to-implement-convenient-python-bindings-to-c) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Why Event-Driven Architecture Isn’t About Speed (and When You Actually Need It)](https://www.wearedevelopers.com/magazine/745-why-event-driven-architecture-isn-t-about-speed-and-when-you-actually-need-it) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Why Attend a Developer Event in 2026?](https://www.wearedevelopers.com/magazine/688-why-attend-a-developer-event-in-2026)