Hardware Analytics Engineer

Cerebras Systems
Sunnyvale, CA, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$213,675.0 - $225,000.0
Working hours
Regular working hours
Job source

Tech stack

A/B Testing Artificial Intelligence Data Analysis Big Data Program Optimization Computer Engineering Extract Transform Load (ETL) Data Visualization Linux Distributed Computing Environment Dynamic Random-Access Memory Firmware
+16 more
Apache Hive Python (Programming Language) Machine Learning PCI Express Performance Tuning Reliability Engineering SQL Databases Tableau (Software) Scripting Graphics Processing Unit (GPU) Apache Spark Reliability of Systems AI Platforms Information Technology Hardware Acceleration Data Pipelines

Job description

  • Design and optimize scalable data pipeline architectures for multi-terabyte hardware telemetry, reliability analytics, and performance optimization.
  • Architect, develop, and optimize hyperscale data pipeline frameworks and ETL processes to aggregate, process, and analyze multi-terabyte hardware performance and telemetry streams, including utilization, power, thermal, acoustic, and reliability metrics across heterogeneous compute, storage, and AI server platforms, ensuring hardware performance compliance and operational reliability.
  • Design and implement hardware performance analysis and anomaly detection systems using Python, SQL, Tableau, Hive, and Spark to forecast hardware failure curves, identify performance bottlenecks, and generate prescriptive recommendations for hardware and system optimization.
  • Lead hardware characterization experiments and thermal/cooling A/B studies to evaluate operational envelopes, delivering validated strategies that reduce carbon footprint, improve water usage efficiency, and maintain or enhance system reliability.
  • Engineer telemetry ingestion, monitoring, and visualization systems to provide real-time, high-fidelity hardware health data to hardware, firmware, and datacenter operations teams, enabling data-driven decision-making at scale.
  • Define, operationalize, and maintain custom efficiency and reliability metrics; perform root cause analysis of systemic failures using large-scale statistical and machine learning methods; and deploy solutions that improve platform scalability, energy efficiency, and sustainability.
  • Collaborate with cross-functional engineering teams to troubleshoot complex failures, isolate defective components, and implement systemic fixes across CPU, GPU, DRAM, PCIe, networking, and storage subsystems.
  • Support the evolution and optimization of next-generation AI platforms and silicon products, including hardware subsystems (CPU, GPU, DRAM, PCIe, networking, and storage), to meet the performance, scalability, and efficiency demands of large language model training and inference workloads.

Requirements

Master’s degree or foreign equivalent degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field and 3 years of experience as Hardware Analytics Engineer, Hardware Engineer, Data Engineer, or a related occupation required., * Large-scale data pipeline architecture and ETL, distributed data processing (Hive, Spark), and dashboard development;

  • Python, SQL, Tableau, Linux, and automation scripting;
  • Design, training, and deployment of machine learning models for hardware performance optimization and failure prediction;
  • Predictive modeling, statistical analysis, A/B testing, anomaly detection, and data visualization in hardware reliability and performance; and
  • Hardware analytics for compute, storage, and AI servers; power and thermal optimization; GPU burn-in efficiency optimization; and reliability modeling for AI hardware systems and components including CPU, GPU, DRAM, and SSD.

About the company

Cerebras Systems builds the world’s largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.

Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference., People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:

  • Build a breakthrough AI platform beyond the constraints of the GPU.

  • Publish and open source their cutting-edge AI research.

  • Work on one of the fastest AI supercomputers in the world.

  • Enjoy job stability with startup vitality.

  • Our simple, non-corporate work culture that respects individual beliefs.

Apply today and become part of the forefront of groundbreaking advancements in AI!

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

You must create an Indeed account before continuing to the company website to apply

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · World Congress 2021

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:58 min

Verifying hardware access and exploring AI inference scaling

Piotr Zaniewski Piotr Zaniewski · World Congress 2026 Europe

Videos

See all

Related articles

See all