> Markdown version of [/jobs/ext/2709255-hardware-test-engineer](https://www.wearedevelopers.com/jobs/ext/2709255-hardware-test-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Hardware Test Engineer - **Company:** Specter Incorporated - **Location:** United States - **Contract:** Permanent contract - **Skills:** Data Analysis, Databases, Data Infrastructure, Firmware, Python (Programming Language), PostgreSQL, Software Tools, Prometheus, SQL Databases, Datadog, Grafana, Hardware Testing, Terraform, Data Pipelines - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/fleet-reliability-engineer-specter-8747385 ## About the Role * Strong data and software skills - Python (or Go) and SQL - and the ability to own a data pipeline end to end. * Hands-on building and tuning observability stacks (OpenTelemetry, Grafana, Prometheus, Datadog, or similar), including their cost. * Experience operating physical or embedded device fleets at scale, and reasoning about how hardware fails in the field. * Comfortable turning messy field telemetry into trends, failure modes, and forecasts. * Fluency with databases and data modeling (PostgreSQL or equivalent); infrastructure-as-code familiarity (Terraform or similar) a plus. * Bias toward building mechanisms over doing manual work. * Nice to have: reliability/SRE fundamentals (SLOs, error budgets, proof-of-recovery) applied to a physical fleet. * Nice to have: experience across the hardware-software boundary - power, connectivity, and physical failure modes. ## Description Specter is hiring a Hardware Test Engineer to develop and own our hardware test and qualification program. Our products operate in extremely demanding outdoor environments, and our customers depend on near-100% uptime. This role is central to delivering on that expectation. You will define how we validate our hardware, execute testing, manage certifications, and make sure the data you generate gets fed back into design decisions so that reliability improves with every revision., Reliability Data Platform - Primary * Own the fleet's reliability data pipeline end to end: telemetry aggregation, storage, and instrumentation. * Drive down observability cost - own the tooling spend and cut what we pay for but don't use. * Instrument the fleet and own the health metrics that measure reliability. Proof-of-Recovery & Alert Hygiene * Verify that fixes hold fleet-wide, not just on the device that paged. * Cut alert noise at the source - separate real failures from self-resolving ones. * Turn repeat failure patterns into automated detection and recovery. Fleet Health & Failure-Mode Analytics * Turn fleet telemetry into a live picture of which cohorts, hardware revisions, and firmware versions are trending toward failure, and why. * Build the failure-mode analysis that tells engineering what to fix at the source. * Own fleet-wide trend and forecasting work, including power and solar planning. Reliability Economics & Prioritization * Score reliability work in dollars - field-trip cost, hardware-return cost, observability spend - and prioritize the most expensive problems first. * Set and track the fleet's reliability targets: uptime, offline rate, truck-rolls per sensor-year. * Give the team the data to make reliability-versus-cost tradeoffs. ## Related Videos - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [The 8 Best Code Testing Tools](https://www.wearedevelopers.com/magazine/402-the-8-best-code-testing-tools) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 139 - Soft and hard queries](https://www.wearedevelopers.com/magazine/487-dev-digest-139-soft-and-hard-queries)