> Markdown version of [/jobs/ext/2847597-performance-engineer](https://www.wearedevelopers.com/jobs/ext/2847597-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Performance Engineer - **Company:** Astera Labs - **Location:** San Jose, CA, United States - **Experience:** Expert - **Contract:** Internship / Graduate position - **Skills:** Artificial Intelligence, Systems Engineering, Confluence, Computer Clusters, Nvidia CUDA, Computer Engineering, Data Centers, Software Debugging, Ethernet, Firmware, Python (Programming Language), Machine Learning, PCI Express, Software Engineering, Test Execution Engine, Graphics Processing Unit (GPU), Application Specific Integrated Circuits, Parallel Computation, Information Technology - **Published:** September 11, 2026 - **Apply:** https://www.thejobnetwork.com/job/de70a1d7-8940-4567-8749-bd287149f074/performance-engineer ## About the Role - MS or PhD in Computer Engineering, Computer Science, Electrical Engineering, or a related field. ## Description In this role, you will define how the world measures scale-up fabric performance. You'll build the roofline models, benchmarks, and end-to-end workload studies that quantify our performance leadership, expose bottlenecks, and drive performance fine-tuning of real AI workloads on our fabric to inform product direction. Your data will directly shape architecture, firmware, and product decisions - and fuel the marketing narrative that positions Astera Labs at the center of AI connectivity. **Key Responsibilities** - **Performance Characterization & Benchmarking** - Establish theoretical and measured roofline models for Astera Labs' scale-up fabric across key performance metrics, defining the reference for all comparative testing. - Build and maintain baseline performance benchmarks using industry-standard tools such as NVBandwidth and NCCL across a range of GPU configurations and switch topologies. - Quantify the impact of differentiated Astera Labs AI fabric features (e.g., Hypercast, In-Network Computing) against baselines using both synthetic benchmarks and real inference workloads. - **Real Workload Analysis & Fabric Scalability** - Run end-to-end inference model workloads on target hardware to capture real-world performance beyond synthetic benchmarks, supporting architecture decisions and customer-facing demonstrations. - Evaluate fabric performance as inference cluster size scales from 16 to 32 GPUs and beyond, identifying bottlenecks and building performance scaling models for state-of-the-art AI workloads. - Design and execute head-to-head performance comparisons against competing fabric switch solutions to produce data-driven differentiation evidence. - **Test Infrastructure & Automation** - Design, build, and maintain automated lab infrastructure including test execution pipelines, traffic generation tooling, and data collection and reporting systems. - Enable repeatable, high-quality, and scalable performance measurements across all hardware configurations, reducing manual effort and accelerating the test cycle. - Share infrastructure and playbooks with the Product Applications team to accelerate customer application development and issue resolution. - **Cross-Functional Impact & Innovation** - Partner closely with ASIC architecture, firmware, software, Product Definition, Product Applications, and Product Marketing teams to communicate findings, influence design decisions, and resolve performance-impacting issues. - Serve as a key technical resource in the early evaluation of new fabric architectures, interconnect technologies (UALink, PCIe Gen 6/Gen 7, Ethernet, UEC), and AI/ML communication paradigms. - Provide performance data, analysis, and live benchmark support for key customer engagements and industry events; produce clear, audience-appropriate performance reports, technical briefs, and marketing collateral, and maintain living documentation in Confluence. **Basic Qualifications** - Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, or a related technical field. We welcome both recent graduates with strong, directly relevant project, research, or internship experience and candidates with 2-5 years of industry experience in performance or systems engineering. - Hands-on experience running AI/ML workloads on GPU clusters - including benchmarking, performance analysis, and fine-tuning of workloads across clusters of GPUs or accelerators. This can come from industry, research, or substantial academic projects. - Demonstrated ability to debug and root-cause system-level performance issues across hardware, firmware, software, and network boundaries. - Excellent fundamental knowledge of compute algorithms, parallel algorithms, and AI/ML algorithms and workloads. - Strong working knowledge of computer systems, GPU systems, and datacenter networking - including PCIe and Ethernet fundamentals. - Working knowledge of GPU and CPU software stacks (e.g., CUDA, MPI, collective communication libraries, drivers, and OS-level performance tooling). - Proficiency in scripting and automation (e.g., Python) to build test pipelines and analyze large performance datasets. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [42 x 2 Canvases Later: Two Years, Two Minds, Many Lessons](https://www.wearedevelopers.com/videos/1458-42-x-2-canvases-later-two-years-two-minds-many-lessons) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)