> Markdown version of [/jobs/ext/3080474-software-engineer-llm-platform](https://www.wearedevelopers.com/jobs/ext/3080474-software-engineer-llm-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, LLM Platform - **Company:** Handshake - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, BigQuery, Cloud Computing, Continuous Integration, Data Sharing, Programming Tools, Distributed Systems, Data Flow Control, Python (Programming Language), PostgreSQL, Machine Learning, Redis, Azure Machine Learning, Runbook, Data Streaming, TypeScript, Workflow Management Systems, Datadog, Google Cloud, Pytorch, Large Language Models, Backend, Fastapi, Kubernetes, Cloudflare, Machine Learning Operations, Api Gateway, Terraform, Golang - **Published:** September 25, 2026 - **Apply:** https://startup.jobs/senior-software-engineer-llm-platform-joinhandshake-8737618 ## About the Role * Strong production software engineering experience in Python, TypeScript, Go, or a similar language. * Experience building or operating high-throughput API gateways, proxies, or multi-tenant platform services. * Hands-on experience with Kubernetes, Terraform, CI/CD, and production service ownership. * Practical experience with LLM provider APIs, including streaming, long-running requests, retries, timeouts, cancellation, and rate limits. * Experience with authentication, quotas, credential management, tenant isolation, and auditability. * Experience building observability for distributed systems and leading production incident response. * Experience with usage metering, cost attribution, capacity planning, or FinOps. * Experience with data or ML platform systems such as BigQuery, Airflow, streaming pipelines, model serving, or ML observability. * Strong judgment in ambiguous environments and a bias toward simple, reliable systems that scale. Bonus Experience * LiteLLM, Portkey, or a comparable multi-provider AI gateway. * vLLM, Modal, Ray, Triton, PyTorch, or GPU-backed serving. * Temporal or another durable-execution system for long-running LLM requests. * Evaluation, batch inference, fine-tuning, RL, or other post-training infrastructure. * Agent runtimes, sandbox infrastructure, MCP, tool use, or coding-agent infrastructure. Our Stack Python, TypeScript, Go, LiteLLM, FastAPI, PostgreSQL, Redis, GCP, Kubernetes, Terraform, Spacelift, BigQuery, Airflow, Dataflow/Beam, Datastream, OpenAI, Anthropic, Gemini, OpenRouter, vLLM, Modal, Anyscale/Ray, Datadog, Arize, Temporal, Cloudflare, and Tailscale ## Description Handshake is hiring a Senior LLM Platform Engineer to join our Data and ML Platform team. This team supports Handshake's core career marketplace and Handshake AI (HAI) by building the shared data, ML, and LLM infrastructure behind production workflows. This infrastructure-heavy role primarily owns our shared LLM control plane: LiteLLM gateways, provider integrations, shared clients, access controls, observability, cost attribution, capacity management, and hosted or self-hosted inference. You'll also contribute to adjacent platform systems for workflow orchestration, model serving, shared cloud infrastructure, and developer enablement-working closely with Backend Platform, HAI engineering, data science, and FDEs. What You'll Do * Build and operate our LiteLLM-based AI gateways and shared LLM clients. * Own provider and model onboarding, routing, failover, rate limits, and capacity planning. * Build self-service virtual-key, model-access, budget, and credential-management workflows. * Establish SLOs, observability, cost attribution, and alerts for production LLM traffic. * Safely qualify and roll out new models, providers, SDKs, and gateway configurations. * Support hosted and self-hosted inference through a consistent platform interface. * Partner with product, AI, and FDE teams to turn recurring delivery problems into paved-platform capabilities. * Contribute to the broader Data and ML Platform, including workflow orchestration, model serving, shared cloud infrastructure, and developer tooling. * Participate in team on-call and support, improving runbooks, automation, and reliability across owned platform services. ## Related Videos - [LLMs in the wild: Building an AI agent that survives production](https://www.wearedevelopers.com/videos/100319-llms-in-the-wild-building-an-ai-agent-that-survives-production) - [Intro to FastAPI](https://www.wearedevelopers.com/videos/462-intro-to-fastapi) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Unveiling the Magic: Scaling Large Language Models to Serve Millions](https://www.wearedevelopers.com/videos/1619-unveiling-the-magic-scaling-large-language-models-to-serve-millions) - [Building and Deploying Multi-Agent Systems with ADK and Vertex AI](https://www.wearedevelopers.com/videos/1918-building-and-deploying-multi-agent-systems-with-adk-and-vertex-ai) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [13 AI Tools You Have to Try](https://www.wearedevelopers.com/magazine/219-13-ai-tools-you-have-to-try)