ML Software Engineer, Data Plane

Amazon.com, Inc.
Springfield, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours

Tech stack

C++ (Programming Language) Computer Literacy Extract Transform Load (ETL) Linux Memory Management Open Source Technology Software Engineering Graphics Processing Unit (GPU) Pytorch Large Language Models Parallel Computation

Job description

Experteer Overview In this role you will own and optimize the inference data plane for a custom ML accelerator, enabling efficient large-model execution across hardware, memory, and data movement. You’ll work closely with cross-functional teams to bring models from validation to production, shaping high-performance kernels and integration with serving frameworks. You’ll scale and validate end-to-end architectures, building test and profiling tools to drive latency and throughput. This ground-up role offers impact across the full stack and opportunities to influence frontier-scale ML deployments. Compensation / Benefits * Develop and optimize compute kernels for a custom ML accelerator to achieve production-level LLM inference performance * Implement and validate end-to-end LLM architectures from PyTorch models to distributed execution on custom hardware * Integrate custom accelerator backends into open-source serving frameworks (vLLM, PyTorch) including scheduler extensions, memory management, and model parallelism * Build and maintain test infrastructure for model correctness across CPU, GPU, simulator, and hardware targets * Profile and optimize inference workloads, identify bottlenecks, and drive latency/throughput improvements from simulation to hardware bring-up * Own features end-to-end from design through implementation, testing, and integration into the software stack * Contribute to CI/CD pipelines to gate model and kernel changes on correctness and performance regressions Tasks * Bachelor’s degree or equivalent * 4+ years of full software development lifecycle experience * Knowledge of computer architecture, operating systems, and parallel computing * Strong proficiency in C/C++ * Strong Linux systems knowledge * Experience developing compute kernels for GPUs, DSPs, or custom accelerators * Proven track record of owning and delivering complex software features end-to-end Key requirements *

Requirements

Experteer Overview In this role you will own and optimize the inference data plane for a custom ML accelerator, enabling efficient large-model execution across hardware, memory, and data movement. You’ll work closely with cross-functional teams to bring models from validation to production, shaping high-performance kernels and integration with serving frameworks. You’ll scale and validate end-to-end architectures, building test and profiling tools to drive latency and throughput. This ground-up role offers impact across the full stack and opportunities to influence frontier-scale ML deployments. Compensation / Benefits * Develop and optimize compute kernels for a custom ML accelerator to achieve production-level LLM inference performance * Implement and validate end-to-end LLM architectures from PyTorch models to distributed execution on custom hardware * Integrate custom accelerator backends into open-source serving frameworks (vLLM, PyTorch) including scheduler extensions, memory aaa backends and model parallelism * Build and maintain test infrastructure for model correctness across CPU, GPU, simulator, and hardware targets * Profile and optimize inference workloads, identify bottlenecks, and drive latency/throughput improvements from simulation to hardware bring-up * Own features end-to-end from design through implementation, testing, and integration into the software stack * Contribute to CI/CD pipelines to gate model and kernel changes on correctness and performance regressions Tasks * Bachelor’s degree or equivalent * 4+ years of full software development lifecycle experience * Knowledge of computer architecture, operating systems, and parallel computing * Strong proficiency in C/C++ * Strong Linux systems knowledge * Experience developing compute kernels for GPUs, DSPs, or custom accelerators * Proven track record of owning and delivering complex software features end-to-end Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:42 min

Addressing educational gaps and retaining early female technological talent

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · WWC 2025

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · WWC Europe 2026

Videos

See all

Related articles

See all