AI Engineer, LLM Inference & AI Agents

PROGRESSIVE PHYSICAL THERAPY, P.C.
San Francisco, CA, United States
13 days ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$200,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence C++ (Programming Language) Code Generation Nvidia CUDA Firmware Python (Programming Language) Linux Kernel Performance Tuning Rust (Programming Language) Large Language Models

Job description

The company is developing an agent that receives specifications and test results, writes low-level code, runs it against real hardware or simulation, analyzes failures, and keeps iterating until strict correctness and performance gates pass.

There is no room for plausible-looking output that does not work. Every component is validated against reference implementations, test suites, and hardware benchmarks.

On its first target, a proprietary accelerator with no existing inference ecosystem, the system reached working tensor-parallel matrix multiplication in approximately 10 hours and ran three frontier models end to end within 10 days. It is now serving production traffic.

What you will do

  • Build the central agent controller that writes, executes, evaluates, and improves inference code.

  • Design evaluation loops that detect incorrect output, compare implementations against references, and enforce performance gates.

  • Own the test ladder used to validate every layer before other components are built on top of it.

  • Orchestrate parallel work across kernels, runtime components, and serving infrastructure, including retry, dependency, escalation, and human-review logic.

  • Review and validate low-level code produced by the agent against real hardware and simulation results.

Requirements

This is a senior engineering role for someone whose experience combines production AI agents with either LLM inference systems or GPU kernel development., You have production experience building with LLMs or coding agents, including tool-use loops, constrained code generation, evaluation harnesses, and systems where tests catch model errors.

You also have meaningful depth in at least one of these areas:

  • LLM inference systems: vLLM, KV cache, paged attention, continuous batching, quantization, speculative decoding, FlashAttention, or low-latency serving.

  • GPU and accelerator kernels: CUDA, Triton, ROCm/HIP, Metal, attention, matrix multiplication, normalization, MoE, kernel fusion, or performance optimization.

Strong Python skills are required. You should also be comfortable reading or reviewing C++, Rust, CUDA C, or similarly low-level systems code.

Experience with LLVM, MLIR, TVM, compiler code generation, accelerator bring-up, NCCL, Megatron-LM, DeepSpeed, distributed inference, firmware, or non-GPU execution models is valuable but not required.

About the company

WorkorAI is recruiting on behalf of an early-stage AI infrastructure company building a coding agent that creates and improves an entire LLM inference stack: GPU kernels, runtime, and serving infrastructure., * $200,000-$420,000 compensation range.

  • San Francisco preferred; remote work can be discussed for a strong fit.

  • Visa sponsorship may be available, including H-1B.

  • $15M raised from investors focused on AI infrastructure and silicon.

  • Approximately 14 engineers, with engineering and product operating as one team.

  • Founded by a former Google Brain researcher.

  • Production systems with measurable correctness and performance feedback rather than demo-only agent workflows.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:01 min

Leveraging large language models for code optimization and development

Stephan Gillich Stephan Gillich +3 · World Congress 2024

1:15 min

Overcoming the challenges of modifying Linux kernel code

Ayesha Kaleem · World Congress 2023

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · World Congress 2021

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

1:55 min

Role of the Linux kernel in handling processes

Mohammed Aboullaite Mohammed Aboullaite · World Congress 2024

Videos

See all

Related articles

See all