> Markdown version of [/jobs/ext/557718-ai-ml-staff-software-engineer](https://www.wearedevelopers.com/jobs/ext/557718-ai-ml-staff-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI/ML Staff Software Engineer - **Company:** Globalfoundries U.S. Inc. - **Location:** Austin, TX, United States - **Experience:** Expert - **Salary:** $106,000.0 - $184,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Compilers, Extract Transform Load (ETL), Dynamic Random-Access Memory, Open Source Technology, Tensorflow, Reduced Instruction Set Computing, Smart Devices, Toolchain, Network Switches, Machine Learning Operations - **Published:** June 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=c54f31d9a90265d8 ## About the Role Do you have experience in Technical writing within technology?, BS or MS (preferred) in EE, CE, CS, or equivalent, with 5+ years in systems engineering, hardware architecture, ML systems, or performance engineering, and a track record of technical leadership. Deep expertise in CPU and SoC architecture - memory hierarchies, out-of-order execution, vector/SIMD pipelines, power management - and how these interact with AI/ML workloads. Strong command of system-level memory bandwidth constraints (DDR/LPDDR bandwidth, channel configuration, utilization efficiency) and the ability to reason quantitatively about memory-bound vs. compute-bound workloads. Experience with AI/ML acceleration on edge devices - NPUs, dedicated inference accelerators, DSP-based pipelines - and the HW/SW co-design challenges involved. Familiarity with model quantization, sparsity, or other efficiency techniques and their hardware interaction is a strong plus. Familiarity with AI compiler infrastructure: MLIR-based toolchains, IREE, TVM, TFLite, or equivalent. Understanding how graph representations are transformed, tiled, scheduled, and lowered to hardware will improve your ability to identify where compiler strategy and hardware architecture must be co-designed. Prior contributions to such toolchains are a significant differentiator. Effective cross-functional collaborator who can drive technical consensus without direct authority, writes clearly, and calibrates technical depth for different audiences. Preferred Qualifications * Prior implementation of CPU hardware features such as vector extensions (AVX, NEON, RVV) or matrix extensions (AMX, SME) * Experience defining or co-defining SoC architecture requirements from workload analysis * Contributions to graph lowering in MLIR/IREE or similar compiler infrastructure * Internal or external publications or contributions to technical standards * Experience mentoring junior systems engineers * Knowledge of RISC-V architecture and Vector/Matrix extensions Other Requirements * English fluency (written and verbal) * Up to 10% travel ## Description We're looking for a seasoned AI/ML Staff Software Engineer to lead workload-driven architecture strategy across hardware and software boundaries. You will define how we study, model, and optimize AI/ML workloads for current and next-generation products, drive alignment across HW and SW engineering organizations, and serve as a technical authority on performance and architecture tradeoffs. This is a senior individual contributor role with significant cross-functional scope and organizational influence., We're looking for a seasoned AI/ML Staff Software Engineer to lead workload-driven architecture strategy across hardware and software boundaries. You will define how we study, model, and optimize AI/ML workloads for current and next-generation products, drive alignment across HW and SW engineering organizations, and serve as a technical authority on performance and architecture tradeoffs. This is a senior individual contributor role with significant cross-functional scope and organizational influence. Essential Responsibilities Own workload characterization and hardware performance analysis for AI/ML systems - selecting representative workloads, defining measurement methodology, building support for MIPS products (e.g., the S8200), and projecting system-level KPIs. Your findings will directly inform SoC architecture decisions, memory subsystem design, and HW/SW co-optimization strategy. Define the software frameworks across the product portfolio: what metrics matter, how to measure them accurately, how to estimate them pre-silicon, and how to use them to make architectural bets. Leverage open-source infrastructure like MLIR and IREE to implement and validate this work. Set the standard for how the team approaches this and mentor junior engineers in applying it. Represent software in architectural discussions with hardware teams (CPU, SoC, memory, interconnect) and software teams (compilers, runtimes, ML frameworks). Identify critical bottlenecks - compute throughput, DRAM bandwidth, on-chip memory, data movement latency, or software overhead - and build the case for specific architectural changes or optimization investments. Present findings and recommendations to senior engineering leadership and product stakeholders. You should be as comfortable writing a one-page architectural recommendation as a detailed technical memo. Other Responsibilities: * Perform all activities in a safe and responsible manner and support all Environmental, Health, Safety & Security requirements and programs. ## Related Videos - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) - [Just-in-time Compilation in JVM](https://www.wearedevelopers.com/videos/240-just-in-time-compilation-in-jvm) - [Speeding up Web Apps performance with WebAssembly and Emscripten](https://www.wearedevelopers.com/videos/1985-speeding-up-web-apps-performance-with-webassembly-and-emscripten) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Unleashing the Full Potential of the Arm Architecture – Write Once, Deploy Anywhere](https://www.wearedevelopers.com/videos/940-unleashing-the-full-potential-of-the-arm-architecture-write-once-deploy-anywhere) - [Building a Compiler with C#](https://www.wearedevelopers.com/videos/116-building-a-compiler-with-c) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)