> Markdown version of [/jobs/ext/1504795-hardware-analytics-engineer](https://www.wearedevelopers.com/jobs/ext/1504795-hardware-analytics-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Hardware Analytics Engineer - **Company:** Cerebras Systems - **Location:** Sunnyvale, CA, United States (Remote available) - **Experience:** Experienced - **Salary:** $213,675.0 - $225,000.0 - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Data Analysis, Big Data, Program Optimization, Computer Engineering, Extract Transform Load (ETL), Data Visualization, Linux, Distributed Computing Environment, Dynamic Random-Access Memory, Firmware, Apache Hive, Python (Programming Language), Machine Learning, PCI Express, Performance Tuning, Reliability Engineering, SQL Databases, Tableau (Software), Scripting, Graphics Processing Unit (GPU), Apache Spark, Reliability of Systems, AI Platforms, Information Technology, Hardware Acceleration, Data Pipelines - **Published:** July 30, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=9c3088b477c0be75 ## About the Role Master's degree or foreign equivalent degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field and 3 years of experience as Hardware Analytics Engineer, Hardware Engineer, Data Engineer, or a related occupation required., * Large-scale data pipeline architecture and ETL, distributed data processing (Hive, Spark), and dashboard development; * Python, SQL, Tableau, Linux, and automation scripting; * Design, training, and deployment of machine learning models for hardware performance optimization and failure prediction; * Predictive modeling, statistical analysis, A/B testing, anomaly detection, and data visualization in hardware reliability and performance; and * Hardware analytics for compute, storage, and AI servers; power and thermal optimization; GPU burn-in efficiency optimization; and reliability modeling for AI hardware systems and components including CPU, GPU, DRAM, and SSD. ## Description * Design and optimize scalable data pipeline architectures for multi-terabyte hardware telemetry, reliability analytics, and performance optimization. * Architect, develop, and optimize hyperscale data pipeline frameworks and ETL processes to aggregate, process, and analyze multi-terabyte hardware performance and telemetry streams, including utilization, power, thermal, acoustic, and reliability metrics across heterogeneous compute, storage, and AI server platforms, ensuring hardware performance compliance and operational reliability. * Design and implement hardware performance analysis and anomaly detection systems using Python, SQL, Tableau, Hive, and Spark to forecast hardware failure curves, identify performance bottlenecks, and generate prescriptive recommendations for hardware and system optimization. * Lead hardware characterization experiments and thermal/cooling A/B studies to evaluate operational envelopes, delivering validated strategies that reduce carbon footprint, improve water usage efficiency, and maintain or enhance system reliability. * Engineer telemetry ingestion, monitoring, and visualization systems to provide real-time, high-fidelity hardware health data to hardware, firmware, and datacenter operations teams, enabling data-driven decision-making at scale. * Define, operationalize, and maintain custom efficiency and reliability metrics; perform root cause analysis of systemic failures using large-scale statistical and machine learning methods; and deploy solutions that improve platform scalability, energy efficiency, and sustainability. * Collaborate with cross-functional engineering teams to troubleshoot complex failures, isolate defective components, and implement systemic fixes across CPU, GPU, DRAM, PCIe, networking, and storage subsystems. * Support the evolution and optimization of next-generation AI platforms and silicon products, including hardware subsystems (CPU, GPU, DRAM, PCIe, networking, and storage), to meet the performance, scalability, and efficiency demands of large language model training and inference workloads. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Agent Smith Gets Hardware: Autonomous IoT Hacking From Debug Port to Cloud API](https://www.wearedevelopers.com/videos/100258-agent-smith-gets-hardware-autonomous-iot-hacking-from-debug-port-to-cloud-api) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)