> Markdown version of [/jobs/ext/1974927-senior-performance-modeling-architect-cpu-fabric](https://www.wearedevelopers.com/jobs/ext/1974927-senior-performance-modeling-architect-cpu-fabric). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Performance Modeling Architect, CPU Fabric... - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $152,000.0 - $241,500.0 - **Contract:** Permanent contract - **Skills:** Big Data, C++ (Programming Language), Computer Engineering, Microarchitecture, Data Centers, Extract Transform Load (ETL), Software Debugging, Formal Verification, Hardware Design, Python (Programming Language), SystemC, Scripting, State Machines, Information Technology, Performance Monitor, Formal Methods, Hardware Infrastructure - **Published:** August 7, 2026 - **Apply:** https://www.juju.com/job/00000000gm0crn ## About the Role To be successful in this role, you should possess a deep technical foundation in computer architecture: + A Master's or Ph.D. in Computer Engineering, Electrical Engineering, or Computer Science (or equivalent experience) with a focus on architecture with 5+ years of experience. + Strong understanding of CPU microarchitecture, memory consistency models, and cache coherency protocols. + Proven experience in C++ or SystemC for cycle-accurate or functional modeling. + Proficiency in Python or similar scripting languages for processing large datasets, generating performance visualizations, and automating simulation sweeps. + Understanding of Network-on-Chip (NoC) topologies (Mesh, Ring, Torus), credit-based flow control, and arbitration logic. Ways to stand out from the crowd: We are looking for individuals who bring a "systems-thinking" approach to hardware development. You will stand out if you have: + Cross-Domain Versatility: Practical experience managing the functional safety (ISO 26262) requirements of automotive chips alongside the power-performance-area (PPA) limitations of data center hardware. + Hardware Performance Counters: Experience defining or using PMU (Performance Monitoring Unit) events to debug performance on real silicon or emulators. + Formal Methods: A background in using formal verification or mathematical modeling to prove the correctness of complex coherency state machines. + Custom Tooling: A history of building your own internal tools or frameworks to accelerate architectural exploration rather than just using off-the-shelf simulators. + Advanced Memory Systems: Knowledge of emerging memory technologies like CXL (Compute Express Link) or HBM (High Bandwidth Memory) and how they collaborate with coherent fabrics. ## Description We are looking for a highly skilled Performance Modeling Architect to lead the architectural definition and improvement of our next-generation CPU Cache Hierarchies and interconnects. This is an outstanding chance to create scalable solutions that connect two fast-paced domains: the high-reliability, low-latency needs of Automotive and the massive efficiency, high-density demands of Data Center systems. You will build the "source of truth" models that govern data movement across our silicon, ensuring our next-level caches (L3/System Cache) and coherent fabrics achieve ambitious performance goals. What you'll be doing: As a core member of the architecture team, your daily work will involve: + Developing and maintaining high-fidelity, cycle-accurate performance models (C++/SystemC) for coherent interconnects and large-scale shared caches. + Modeling and analyzing performance bottlenecks across varying scales, from small-cluster automotive SoCs to massive, multi-mesh data center architectures. + Evaluating the performance impact of different coherency protocols (e.g., CHI, ACE, or proprietary) and snooping filters. + Running and analyzing industry-standard benchmarks (SPEC, MLPerf, Automotive-specific suites) to drive architectural trade-offs. + Collaborating with build and verification teams to correlate performance models with silicon and working with software teams to optimize drivers for the underlying hardware topology. ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Intermediate Bitcoin Script](https://www.wearedevelopers.com/videos/25-intermediate-bitcoin-script) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)