Software Engineer, LLM Platform

Handshake
San Francisco, CA, United States
4 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Airflow BigQuery Cloud Computing Continuous Integration Data Sharing Programming Tools Distributed Systems Data Flow Control Python (Programming Language) PostgreSQL
+19 more
Machine Learning Redis Azure Machine Learning Runbook Data Streaming TypeScript Workflow Management Systems Datadog Google Cloud Pytorch Large Language Models Backend Fastapi Kubernetes Cloudflare Machine Learning Operations Api Gateway Terraform Golang

Job description

Handshake is hiring a Senior LLM Platform Engineer to join our Data and ML Platform team. This team supports Handshake’s core career marketplace and Handshake AI (HAI) by building the shared data, ML, and LLM infrastructure behind production workflows.

This infrastructure-heavy role primarily owns our shared LLM control plane: LiteLLM gateways, provider integrations, shared clients, access controls, observability, cost attribution, capacity management, and hosted or self-hosted inference. You’ll also contribute to adjacent platform systems for workflow orchestration, model serving, shared cloud infrastructure, and developer enablement-working closely with Backend Platform, HAI engineering, data science, and FDEs.

What You’ll Do

  • Build and operate our LiteLLM-based AI gateways and shared LLM clients.
  • Own provider and model onboarding, routing, failover, rate limits, and capacity planning.
  • Build self-service virtual-key, model-access, budget, and credential-management workflows.
  • Establish SLOs, observability, cost attribution, and alerts for production LLM traffic.
  • Safely qualify and roll out new models, providers, SDKs, and gateway configurations.
  • Support hosted and self-hosted inference through a consistent platform interface.
  • Partner with product, AI, and FDE teams to turn recurring delivery problems into paved-platform capabilities.
  • Contribute to the broader Data and ML Platform, including workflow orchestration, model serving, shared cloud infrastructure, and developer tooling.
  • Participate in team on-call and support, improving runbooks, automation, and reliability across owned platform services.

Requirements

  • Strong production software engineering experience in Python, TypeScript, Go, or a similar language.
  • Experience building or operating high-throughput API gateways, proxies, or multi-tenant platform services.
  • Hands-on experience with Kubernetes, Terraform, CI/CD, and production service ownership.
  • Practical experience with LLM provider APIs, including streaming, long-running requests, retries, timeouts, cancellation, and rate limits.
  • Experience with authentication, quotas, credential management, tenant isolation, and auditability.
  • Experience building observability for distributed systems and leading production incident response.
  • Experience with usage metering, cost attribution, capacity planning, or FinOps.
  • Experience with data or ML platform systems such as BigQuery, Airflow, streaming pipelines, model serving, or ML observability.
  • Strong judgment in ambiguous environments and a bias toward simple, reliable systems that scale.

Bonus Experience

  • LiteLLM, Portkey, or a comparable multi-provider AI gateway.
  • vLLM, Modal, Ray, Triton, PyTorch, or GPU-backed serving.
  • Temporal or another durable-execution system for long-running LLM requests.
  • Evaluation, batch inference, fine-tuning, RL, or other post-training infrastructure.
  • Agent runtimes, sandbox infrastructure, MCP, tool use, or coding-agent infrastructure.

Our Stack

Python, TypeScript, Go, LiteLLM, FastAPI, PostgreSQL, Redis, GCP, Kubernetes, Terraform, Spacelift, BigQuery, Airflow, Dataflow/Beam, Datastream, OpenAI, Anthropic, Gemini, OpenRouter, vLLM, Modal, Anyscale/Ray, Datadog, Arize, Temporal, Cloudflare, and Tailscale

Benefits & conditions

Financial Wellness: 401(k) match, competitive compensation, financial coaching

Family Support: Paid parental leave, fertility benefits, parental coaching

Wellbeing: Medical, dental, and vision, mental health support, $500 wellness stipend

About the company

Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.

In 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We’ve grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals every month.

Why join Handshake now:

  • Shape how every career evolves in the AI economy, at global scale, with impact your friends, family and peers can see and feel
  • Partner hand-in-hand with world-class AI labs, Fortune 500 partners and the world’s top educational institutions
  • Work together with engineers, scientists, operators, and more from Palantir, Meta, Scale AI, and former YC founders
  • Build a massive, fast-growing business with billions in revenue

About Handshake AI

Human data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data-intensive post-training techniques. We believe that data spend for AI training will increase by 3-5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:08 min

Applying large language models to infrastructure tasks

Alfonso Sandoval Rosas Alfonso Sandoval Rosas · Europe 2026 Virtual

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

3:33 min

Connecting frontends via a FastAPI proxy backend layer

Saoussen Chaabnia Saoussen Chaabnia · Europe 2026 Virtual

5:30 min

Building components of a real-world LLM lifecycle

Maxim Salnikov Maxim Salnikov · LIVE

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · World Congress 2022

Videos

See all

Related articles

See all