AI Engineer, Model Training, Inference & Infra

INTERNATIONAL RECRUITING LLC
Bellevue, WA, United States
3 days ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence C++ (Programming Language) Computer Clusters Nvidia CUDA Computer Programming Computer Engineering Learning Management Systems Software Debugging Microprocessors Distributed Computing Environment Electronic Design Automation Fault Tolerance
+29 more
Systems Theories InfiniBand Job Scheduling Python (Programming Language) Linux Kernel Open Source Technology Performance Tuning Inference Optimization Tensorflow Software Deployment Software Engineering Reinforcement Learning Graphics Processing Unit (GPU) Pytorch DeepSpeed Large Language Models Gpu Programming Discretization Kubernetes Information Technology SGLang Slurm Machine Learning Operations TensorRT Decoding CUTLASS Megatron Data Pipelines Data Generation

Job description

We are looking for an exceptional AI Engineer to own the model training, inference, and infrastructure that power its agentic design workforce.

You will drive the full model lifecycle: data pipelines, pretraining and post-training, reinforcement learning, evaluation, and high-performance serving. You will build and operate the GPU training and serving stack - on the compute substrate our infrastructure team provides - that keeps large-scale training reliable and low-latency inference efficient at production scale. You will work on evaluation systems, model training, and the self-improving, self-evolving learning loops that let our models and agents get better over time from real execution feedback.

Your work directly determines how capable, fast, and cost-effective our agents are. This role is ideal for someone who combines strong research ability with exceptional systems and performance engineering skills, and who wants the models they train and serve deployed in real semiconductor design environments-not left in notebooks or benchmarks. What You’ll Do

  • Train, post-train, and fine-tune large language models for agentic engineering workflows, including supervised fine-tuning, RLHF/RLAIF, reinforcement learning, and distillation.
  • Build scalable data pipelines for pretraining, post-training, and evaluation, including sparse, private, and domain-specific engineering data.
  • Design and operate distributed training on multi-node GPU clusters, using data, tensor, pipeline, and sequence parallelism (for example FSDP, DeepSpeed, or Megatron-style approaches).
  • Build high-throughput, low-latency inference systems with continuous batching, KV-cache management, paged attention, quantization, and speculative decoding.
  • Write and optimize custom GPU kernels (CUDA, Triton) and profile end-to-end performance across CPUs and GPUs.
  • Build core model infrastructure: cluster orchestration, ML job scheduling, checkpointing, fault tolerance, reproducibility, model and environment management, observability, and cost and utilization tracking.
  • Build automated evaluation systems, benchmarks, and reward models that measure agent capability, reliability, and regression across complex engineering tasks, including problems where design data is private or customer-specific.
  • Design self-improving and self-evolving algorithms and learning loops, where models and agents learn from execution feedback, outcomes, and new data to improve continuously over time.
  • Integrate models with agent runtimes, tool use, retrieval, and the production serving stack.
  • Improve reliability, throughput, and cost efficiency across the training and inference platform.
  • Translate promising research ideas into reliable, scalable product capabilities.
  • Collaborate with research, product, platform, and solutions teams across San Jose, Austin, and Taiwan.
  • Contribute to patents, publications, technical presentations, and the broader development of Agentic Design Automation.

Requirements

  • PhD or master’s degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field, or equivalent practical experience.
  • Strong programming skills in Python and proficiency in at least one systems language such as C++ or Rust.
  • Deep experience with machine learning frameworks such as PyTorch or JAX.
  • Hands-on experience with one or more of the following:
  • Large-scale or distributed model training
  • High-performance model inference and serving
  • GPU programming and performance optimization
  • ML infrastructure and platform engineering
  • Automated evaluation, reward modeling, or self-improving and continuous-learning systems
  • Ability to take a model from data and problem formulation through training, evaluation, and production deployment.
  • Strong analytical, software engineering, performance-optimization, and debugging skills.
  • High ownership, intellectual curiosity, and willingness to work across research and product boundaries.
  • Clear written and verbal communication skills.

Particularly Valuable Experience

  • Pretraining or post-training large language models at scale.
  • Distributed training frameworks such as Megatron-LM, DeepSpeed, FSDP, or Ray.
  • Production inference engines such as vLLM, TensorRT-LLM, SGLang, or TGI.
  • Custom kernel development with CUDA, Triton, or CUTLASS.
  • Inference optimization techniques such as quantization (FP8, GPTQ, AWQ), speculative decoding, or KV-cache optimization.
  • Reinforcement learning, RLHF, or reward-model training for LLMs.
  • Automated evaluation, benchmarking, or LLM-as-judge systems for agents.
  • Self-improving, self-evolving, or continuous-learning systems, including learning from execution feedback, automated curricula, or synthetic data generation.
  • GPU cluster infrastructure with Kubernetes, Slurm, or Ray, and high-performance networking such as NCCL or InfiniBand.
  • Data pipelines and MLOps for training and continuous learning.
  • Experience deploying AI systems in enterprise or security-sensitive environments.
  • A strong record of implementation through research systems, open-source projects, production software, or technical competitions.

About the company

Our client is building the next generation of design automation for the semiconductor industry.

Our mission is to enable every engineering organization to build its own self-improving agentic design workforce. It combines AI agents, engineering knowledge, agent-native tools, advanced models, and continuous learning to automate complex chip-design workflows.

The team brings deep experience in artificial intelligence, electronic design automation, semiconductor design, GPU-accelerated computing, and production software systems. We work closely with leading semiconductor companies to turn advanced research into technology that improves engineering productivity, design quality, and time to market., At this company, you will have the opportunity to:

  • Help define a new category of semiconductor design technology.
  • Build the training and inference stack that powers autonomous engineering agents.
  • Develop GPU-accelerated systems that make large-scale training and low-latency serving practical and cost-effective.
  • Build the evaluation and self-improvement loops that let agents learn and get better from real engineering work.
  • Build AI systems that perform complex, consequential engineering work-not just generate recommendations.
  • Work with real semiconductor workflows, tools, and private engineering knowledge.
  • See your models deployed directly with leading chip-design organizations.
  • Work in a small, highly technical team where individual contributions can shape the product and company.
  • Collaborate with colleagues across San Jose, Austin, and Taiwan.
  • Change how chips are designed, rather than focus on only one design or one point tool.

Our client is an equal opportunity employer. We welcome candidates from diverse backgrounds who are excited to combine ambitious research with meaningful engineering impact.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:50 min

Enhancing hardware utilization with continuous batching library frameworks

Christin Pohl Christin Pohl · World Congress 2026 Europe

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

5:01 min

Leveraging large language models for code optimization and development

Stephan Gillich Stephan Gillich +3 · World Congress 2024

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

2:37 min

Optimizing technical profiles for AI sourcing and recruitment

Mina Golesorkhi Mina Golesorkhi · World Congress 2026 Europe

Videos

See all

Related articles

See all