ML Software Engineer, Data Plane
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Experteer Overview In this role you will own and optimize the inference data plane for a custom ML accelerator, enabling efficient large-model execution across hardware, memory, and data movement. You’ll work closely with cross-functional teams to bring models from validation to production, shaping high-performance kernels and integration with serving frameworks. You’ll scale and validate end-to-end architectures, building test and profiling tools to drive latency and throughput. This ground-up role offers impact across the full stack and opportunities to influence frontier-scale ML deployments. Compensation / Benefits * Develop and optimize compute kernels for a custom ML accelerator to achieve production-level LLM inference performance * Implement and validate end-to-end LLM architectures from PyTorch models to distributed execution on custom hardware * Integrate custom accelerator backends into open-source serving frameworks (vLLM, PyTorch) including scheduler extensions, memory management, and model parallelism * Build and maintain test infrastructure for model correctness across CPU, GPU, simulator, and hardware targets * Profile and optimize inference workloads, identify bottlenecks, and drive latency/throughput improvements from simulation to hardware bring-up * Own features end-to-end from design through implementation, testing, and integration into the software stack * Contribute to CI/CD pipelines to gate model and kernel changes on correctness and performance regressions Tasks * Bachelor’s degree or equivalent * 4+ years of full software development lifecycle experience * Knowledge of computer architecture, operating systems, and parallel computing * Strong proficiency in C/C++ * Strong Linux systems knowledge * Experience developing compute kernels for GPUs, DSPs, or custom accelerators * Proven track record of owning and delivering complex software features end-to-end Key requirements *
Requirements
Experteer Overview In this role you will own and optimize the inference data plane for a custom ML accelerator, enabling efficient large-model execution across hardware, memory, and data movement. You’ll work closely with cross-functional teams to bring models from validation to production, shaping high-performance kernels and integration with serving frameworks. You’ll scale and validate end-to-end architectures, building test and profiling tools to drive latency and throughput. This ground-up role offers impact across the full stack and opportunities to influence frontier-scale ML deployments. Compensation / Benefits * Develop and optimize compute kernels for a custom ML accelerator to achieve production-level LLM inference performance * Implement and validate end-to-end LLM architectures from PyTorch models to distributed execution on custom hardware * Integrate custom accelerator backends into open-source serving frameworks (vLLM, PyTorch) including scheduler extensions, memory aaa backends and model parallelism * Build and maintain test infrastructure for model correctness across CPU, GPU, simulator, and hardware targets * Profile and optimize inference workloads, identify bottlenecks, and drive latency/throughput improvements from simulation to hardware bring-up * Own features end-to-end from design through implementation, testing, and integration into the software stack * Contribute to CI/CD pipelines to gate model and kernel changes on correctness and performance regressions Tasks * Bachelor’s degree or equivalent * 4+ years of full software development lifecycle experience * Knowledge of computer architecture, operating systems, and parallel computing * Strong proficiency in C/C++ * Strong Linux systems knowledge * Experience developing compute kernels for GPUs, DSPs, or custom accelerators * Proven track record of owning and delivering complex software features end-to-end Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
Highest Paying Tech Companies for Developers
How to Become an AI Engineer
MLOps – What’s the deal behind it?