> Markdown version of [/jobs/ext/2060004-fleet-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2060004-fleet-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Fleet Reliability Engineer - **Company:** Quartermaster AI Inc - **Location:** Arlington, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Databases, Firmware, Failure Mode Effects Analysis, Hardware Design, Data Intelligence, Python (Programming Language), Reliability Engineering, Satcom, Smart Devices, SQL Databases, Visual Systems, Mttr, Reliability of Systems, MIL-STD-810 Environmental Testing - **Published:** August 14, 2026 - **Apply:** https://www.dice.com/job-detail/3d194e27-fe01-4872-aec3-60a3a2ebcdcb ## About the Role * Bachelor's degree in Electrical, Mechanical, Systems, Reliability, or a related engineering discipline - or equivalent hands-on experience. * 5+ years of engineering experience with deployed electro-mechanical hardware, at least 2 of which are in reliability, sustaining/field engineering, or hardware operations for a fielded product. * Demonstrated ownership of hardware reliability outcomes for a fleet or installed base - you have been directly responsible for uptime, failure rates, or MTBF/MTTR of real hardware in the field. * Hands-on proficiency with root-cause analysis methods (8D, 5-Whys, fishbone) and reliability tools such as FMEA, fault-tree analysis, and CAPA. * Practical experience diagnosing electro-mechanical systems using telemetry/logs, bench instruments, and physical teardown. * Data fluency: able to query, analyze, and visualize fleet telemetry using SQL and Python (or equivalent) to find trends and drive decisions. * Working knowledge of electronics, power systems, and mechanical enclosures, and the failure modes of hardware operating in harsh outdoor environments. * Willingness and ability to travel periodically to field sites (vessels, ports, installation locations), including occasional international travel. * Must be legally authorized to work in the United States and able to satisfy any customer- or contract-driven eligibility requirements associated with government and maritime-security work., * Experience with hardware deployed in marine, maritime, offshore, automotive, aerospace/defense, satellite, telecom, or other remote/harsh-environment fleets. * Familiarity with IP-rated enclosures, corrosion and salt-fog effects, marine power systems, and environmental qualification (e.g., IEC 60529, MIL-STD-810, IEC 60068). * Experience with connected/IoT or edge devices: remote diagnostics, OTA firmware updates, and interpreting embedded-system and connectivity (SATCOM/cellular) telemetry. * Exposure to camera/optical systems, RF/software-defined radios, batteries, or edge-AI compute hardware. * Background building a reliability or sustaining-engineering function from scratch at a hardware startup or scaling operation. * ASQ Certified Reliability Engineer (CRE) or comparable credential. ## Description As the fleet scales, keeping every deployed SmartMast healthy, connected, and producing reliable data at sea is mission-critical. We are hiring our first dedicated Fleet Reliability Engineer to own that outcome., The Fleet Reliability Engineer owns the health, uptime, and long-term reliability of our deployed SmartMast hardware fleet. You will be the person who knows - at any moment - how many units are online, which are degrading, why units fail, and what we are doing about it. You will turn fleet telemetry into action: catching failures before they take a unit offline, driving root-cause analysis on the ones that slip through, and closing the loop with hardware design, firmware, and field-service teams so the same failure never recurs., This is a hands-on, high-ownership role suited to a senior engineer who is comfortable operating at the intersection of hardware, embedded systems, data, and harsh-environment field operations. You will build the reliability program from the ground up: the metrics, the monitoring, the failure-tracking process, and the maintenance and RMA workflows that let the fleet scale from hundreds to thousands of units. What You'll Do Fleet health & monitoring * Own fleet-wide reliability metrics - uptime, availability, MTBF, MTTR, data-yield, and failure rates by component and by deployment environment - and report them to engineering and leadership. * Build and refine dashboards, alerting, and telemetry pipelines that surface degrading units (power, thermal, connectivity, camera, radio, compute) before they go offline. * Define what "healthy" means for each subsystem and set the thresholds that trigger proactive intervention. Failure analysis & continuous improvement * Lead root-cause analysis (RCA) on field failures, from telemetry forensics through physical teardown of returned units. * Maintain the fleet failure database and drive FMEA, reliability growth tracking, and corrective/preventive action (CAPA) to closure. * Close the loop with hardware, firmware, and manufacturing teams - translating field failures into design-for-reliability, component-selection, and firmware changes. Field service, maintenance & logistics * Define preventive maintenance schedules, spares strategy, and the RMA / repair-and-return process for a globally distributed fleet. * Update installation, diagnostic, and field-repair procedures and troubleshooting guides used by internal technicians and partner crews. * Support field deployments and complex repairs directly, including periodic travel to vessels, ports, and installation sites. Reliability engineering & scale * Work with the hardware team to establish environmental and life-test protocols (vibration, salt-fog/corrosion, thermal, ingress, power) to qualify hardware and predict field life before deployment. * Feed reliability requirements and acceptance criteria into new hardware revisions and supplier qualification. * Design the reliability processes and tooling so they scale as the fleet grows into the thousands of units. ## Related Videos - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [Fleet Management - Reinvented](https://www.wearedevelopers.com/videos/473-fleet-management-reinvented) - [Agent Smith Gets Hardware: Autonomous IoT Hacking From Debug Port to Cloud API](https://www.wearedevelopers.com/videos/100258-agent-smith-gets-hardware-autonomous-iot-hacking-from-debug-port-to-cloud-api) - [Robots are coming into the wild! Full-Stack Robotics Engineers, be ready!](https://www.wearedevelopers.com/videos/479-robots-are-coming-into-the-wild-full-stack-robotics-engineers-be-ready) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)