> Markdown version of [/jobs/ext/267361-soc-firmware-engineering-manager-annapurna-labs-machine-learning-acceleration-aws](https://www.wearedevelopers.com/jobs/ext/267361-soc-firmware-engineering-manager-annapurna-labs-machine-learning-acceleration-aws). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SoC Firmware Engineering Manager, Annapurna Labs Machine Learning Acceleration, AWS - **Company:** Annapurna Labs - **Location:** Cupertino, CA, United States - **Experience:** Experienced - **Salary:** $184,900.0 - $287,700.0 - **Contract:** Permanent contract - **Skills:** Abstraction Layers, Amazon Web Services, Unit Testing, C++ (Programming Language), CMake, Code Generation, Code Review, Software Debugging, Linux on Embedded Systems, Emulators, Firmware, Field-Programmable Gate Array (FPGA), Hardware-In-The-Loop Simulation, Python (Programming Language), Machine Learning, PCI Express, Software Product Management, Quick EMUlator (QEMU), Software Engineering, Software Systems, Verification and Validation (Software), SystemVerilog, Application Specific Integrated Circuits - **Published:** May 16, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=8dfcfb1410b00574 ## About the Role * 3+ years of engineering team management experience * 7+ years of professional software development in C or C++, including embedded, firmware, or systems-level development * 4+ years of designing or architecting software systems (abstraction layers, hardware/software interfaces) * Experience developing software that interfaces directly with hardware: SoC, ASIC, FPGA, or embedded microcontrollers * Experience with register-level programming and hardware debug (waveform analysis, bus-level tracing, or similar), * Experience in recruiting, hiring, mentoring/coaching and managing teams of Software Engineers to improve their skills, and make them more effective, product software engineers * Experience with silicon bring-up or pre/post-silicon software validation * Experience shipping software across multiple target platforms (simulation, emulation, production hardware) * Familiarity with bus protocols (APB, AXI, PCIe) or memory subsystems (HBM, DDR) * Experience with C++ template metaprogramming or code generation frameworks * Experience building or maintaining hardware abstraction layers or board support packages ## Description When a new Trainium or Inferentia chip comes back from the fab, our code is the first software to touch it. We're looking for a hands-on engineering manager who lives and breathes low-level software - someone who's debugged register-level issues at 3am and wants to build a team that does it better. Our SoC HAL (Hardware Abstraction Layer) team owns the lowest layer of user-space software on AWS's custom ML accelerator chips: the firmware that boots, configures, and manages every hardware block on the SoC. Your software runs as a shared library on embedded Linux, reaching into the chip to program PCIe links, initialize HBM controllers, configure PLLs, manage interrupt controllers, and orchestrate fabric interconnects across 270+ hardware block instances per chip - all deployed across millions of servers in AWS's global fleet. Tech stack: C++17, CMake, GoogleTest, Python, SystemVerilog DPI, SPI, APB/AXI bus protocols, PCIe, UCIe, HBM, PLL, custom IPs As the SoC Firmware Manager, you will: * Manage, coach, and grow a team of 6 engineers - set technical direction, own hiring, and create an environment where strong engineers want to stay * Coordinate deliverables across chip architects, RTL designers, verification engineers, validation engineers, and platform software teams - you're the single point of accountability for HAL readiness on every new chip program * Own bring-up for new SoC tape-outs, from first-silicon power-on through production fleet deployment * Prioritize work across multiple concurrent chip programs and customer teams, balancing urgent bring-up needs against long-term architecture investments * Drive the architecture of our C++ template metaprogramming framework, BUTR (Built-in Unit Test for Registers), and HITL (Hardware-in-the-Loop) test infrastructure * Ship the same C++ codebase to three execution environments: SystemVerilog DPI for chip verification, QEMU for emulation, and Carbon OS on embedded microcontrollers for production fleet * Get into the weeds alongside your team - debug register-level HW/SW interactions, review code, and write code yourself when it matters Most firmware teams target one platform and ship to a few thousand units. We target three platforms from a single source tree and deploy across AWS's global fleet - where a single register misconfiguration can impact millions of servers. Our software must be stateless, survive live-updates on running production servers without reboots, and be correct down to individual register bits. The microcontroller can reboot at any time - including during customer workloads - and the HAL must resume managing the SoC by querying hardware state on-demand. No cached state, no assumptions. Your pre-silicon software runs in simulation and emulation months before real silicon arrives. When the chip comes back from the fab, you validate those predictions on real hardware - and when they don't match, you figure out whether it's a silicon bug or a software bug. For Trainium3, our HAL enabled a full ML training workload within 12 hours of first power-on: https://www.aboutamazon.com/news/aws/trainium-3-ultraserver-faster-ai-training-lower-cost No ML background needed. Your firmware is the foundation that enables ML training across clusters of thousands of interconnected accelerators - you'll work on components like PCIe and HBM, but won't need to understand ML itself. ## Related Videos - [Code to Road in < 12 hours](https://www.wearedevelopers.com/videos/1082-code-to-road-in-12-hours) - [Unleashing the Full Potential of the Arm Architecture – Write Once, Deploy Anywhere](https://www.wearedevelopers.com/videos/940-unleashing-the-full-potential-of-the-arm-architecture-write-once-deploy-anywhere) - [Thinking Differently - How to Make Money from Cyber Attacks & Cheats](https://www.wearedevelopers.com/videos/745-thinking-differently-how-to-make-money-from-cyber-attacks-cheats) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [More efficient software for more efficient microchips](https://www.wearedevelopers.com/videos/1164-more-efficient-software-for-more-efficient-microchips) - [WeAreDevelopers LIVE – Web Scraping, Agents, Actors and more](https://www.wearedevelopers.com/videos/1764-wearedevelopers-live-web-scraping-agents-actors-and-more) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Dev Digest 118 - not a total recall](https://www.wearedevelopers.com/magazine/452-dev-digest-118-not-a-total-recall) - [Dev Digest 162: AI careers, MCP, AWS best practices & floppy sweaters](https://www.wearedevelopers.com/magazine/571-dev-digest-162-ai-careers-mcp-aws-best-practices-floppy-sweaters)