AI Engineer, LLM Inference & AI Agents
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
The company is developing an agent that receives specifications and test results, writes low-level code, runs it against real hardware or simulation, analyzes failures, and keeps iterating until strict correctness and performance gates pass.
There is no room for plausible-looking output that does not work. Every component is validated against reference implementations, test suites, and hardware benchmarks.
On its first target, a proprietary accelerator with no existing inference ecosystem, the system reached working tensor-parallel matrix multiplication in approximately 10 hours and ran three frontier models end to end within 10 days. It is now serving production traffic.
What you will do
-
Build the central agent controller that writes, executes, evaluates, and improves inference code.
-
Design evaluation loops that detect incorrect output, compare implementations against references, and enforce performance gates.
-
Own the test ladder used to validate every layer before other components are built on top of it.
-
Orchestrate parallel work across kernels, runtime components, and serving infrastructure, including retry, dependency, escalation, and human-review logic.
-
Review and validate low-level code produced by the agent against real hardware and simulation results.
Requirements
This is a senior engineering role for someone whose experience combines production AI agents with either LLM inference systems or GPU kernel development., You have production experience building with LLMs or coding agents, including tool-use loops, constrained code generation, evaluation harnesses, and systems where tests catch model errors.
You also have meaningful depth in at least one of these areas:
-
LLM inference systems: vLLM, KV cache, paged attention, continuous batching, quantization, speculative decoding, FlashAttention, or low-latency serving.
-
GPU and accelerator kernels: CUDA, Triton, ROCm/HIP, Metal, attention, matrix multiplication, normalization, MoE, kernel fusion, or performance optimization.
Strong Python skills are required. You should also be comfortable reading or reviewing C++, Rust, CUDA C, or similarly low-level systems code.
Experience with LLVM, MLIR, TVM, compiler code generation, accelerator bring-up, NCCL, Megatron-LM, DeepSpeed, distributed inference, firmware, or non-GPU execution models is valuable but not required.
About the company
WorkorAI is recruiting on behalf of an early-stage AI infrastructure company building a coding agent that creates and improves an entire LLM inference stack: GPU kernels, runtime, and serving infrastructure., * $200,000-$420,000 compensation range.
-
San Francisco preferred; remote work can be discussed for a strong fit.
-
Visa sponsorship may be available, including H-1B.
-
$15M raised from investors focused on AI infrastructure and silicon.
-
Approximately 14 engineers, with engineering and product operating as one team.
-
Founded by a former Google Brain researcher.
-
Production systems with measurable correctness and performance feedback rather than demo-only agent workflows.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
How to Become an AI Engineer
Stephan Gillich - Bringing AI Everywhere
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?