> Markdown version of [/jobs/ext/2209347-ai-engineer-llm-inference-ai-agents](https://www.wearedevelopers.com/jobs/ext/2209347-ai-engineer-llm-inference-ai-agents). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Engineer, LLM Inference & AI Agents - **Company:** PROGRESSIVE PHYSICAL THERAPY, P.C. - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $200,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Code Generation, Nvidia CUDA, Firmware, Python (Programming Language), Linux Kernel, Performance Tuning, Rust (Programming Language), Large Language Models - **Published:** August 24, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/pelcsmj47t ## About the Role This is a senior engineering role for someone whose experience combines production AI agents with either LLM inference systems or GPU kernel development., You have production experience building with LLMs or coding agents, including tool-use loops, constrained code generation, evaluation harnesses, and systems where tests catch model errors. You also have meaningful depth in at least one of these areas: * LLM inference systems: vLLM, KV cache, paged attention, continuous batching, quantization, speculative decoding, FlashAttention, or low-latency serving. * GPU and accelerator kernels: CUDA, Triton, ROCm/HIP, Metal, attention, matrix multiplication, normalization, MoE, kernel fusion, or performance optimization. Strong Python skills are required. You should also be comfortable reading or reviewing C++, Rust, CUDA C, or similarly low-level systems code. Experience with LLVM, MLIR, TVM, compiler code generation, accelerator bring-up, NCCL, Megatron-LM, DeepSpeed, distributed inference, firmware, or non-GPU execution models is valuable but not required. ## Description The company is developing an agent that receives specifications and test results, writes low-level code, runs it against real hardware or simulation, analyzes failures, and keeps iterating until strict correctness and performance gates pass. There is no room for plausible-looking output that does not work. Every component is validated against reference implementations, test suites, and hardware benchmarks. On its first target, a proprietary accelerator with no existing inference ecosystem, the system reached working tensor-parallel matrix multiplication in approximately 10 hours and ran three frontier models end to end within 10 days. It is now serving production traffic. What you will do * Build the central agent controller that writes, executes, evaluates, and improves inference code. * Design evaluation loops that detect incorrect output, compare implementations against references, and enforce performance gates. * Own the test ladder used to validate every layer before other components are built on top of it. * Orchestrate parallel work across kernels, runtime components, and serving infrastructure, including retry, dependency, escalation, and human-review logic. * Review and validate low-level code produced by the agent against real hardware and simulation results. ## Related Videos - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Localized Open Models in Production: What Builders Need to Know](https://www.wearedevelopers.com/videos/100270-localized-open-models-in-production-what-builders-need-to-know) - [The weekly developer show: Boosting Python with CUDA, CSS Updates & Navigating New Tech Stacks](https://www.wearedevelopers.com/videos/1293-the-weekly-developer-show-boosting-python-with-cuda-css-updates-navigating-new-tech-stacks) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [ Dev Digest 213: Petrol Prices, Agentic Workflows, AI Skills and CODE100!](https://www.wearedevelopers.com/magazine/718-dev-digest-213-petrol-prices-agentic-workflows-ai-skills-and-code100)