Lead Software Engineer - LLM Ops Platform Reliability
Hackajob Ltd
Milton, UK
6 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.collegerecruiter.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source
Tech stack
Application Programming Interfaces (APIs)
Artificial Intelligence
Amazon Web Services
Computer Clusters
Continuous Delivery
Python (Programming Language)
Open Source Technology
Reliability Engineering
Site Reliability Engineering Practices
Prometheus
Runbook
Search Technologies
+16 more
Software Engineering
Systems Integration
Management of Software Versions
AI Infrastructure
Data Logging
Autoscaling
Retrieval-Augmented Generation
System Availability
Large Language Models
Grafana
Multi-Agent Systems
Parallel Computation
Backend
Kubernetes
Hardware Infrastructure
Terraform
Job description
- Design, develop, troubleshoot, and deliver secure, high-quality production software and services for AI infrastructure
- Build backend services and APIs that enable reliable operation of AI infrastructure in production
- Operate and scale LLM serving infrastructure (such as vLLM and llm-d), including model hosting, request routing, continuous batching, and KV-cache optimization
- Deploy, host, and lifecycle-manage open-source and proprietary LLMs on Amazon EKS and Amazon SageMaker, as well as on-prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines
- Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads
- Tune GPU and accelerator capacity, autoscaling, and cost efficiency for LLM inference workloads using performance and optimization techniques (e.g., quantization, parallelism, speculative decoding)
- Lead reliability engineering for LLM endpoints through capacity planning, load/soak testing, safe rollouts (blue/green, canary), failover, and incident response for outages and model-quality regressions
- Participate in an on-call rotation, lead incident triage and mitigation, and produce clear post-incident root-cause analyses and follow-ups
- Identify recurring operational issues and automate remediation to improve platform stability and developer experience
- Build and maintain multi-agent systems with strong orchestration (planning, coordination, tool-calling, state/memory, and workflow control) where appropriate
- Contribute to an inclusive team culture grounded in diversity, opportunity, inclusion, and respect, and help drive adoption of leading-edge technologies through communities of practice, Our professionals in our Corporate Functions cover a diverse range of areas from finance and risk to human resources and marketing. Our corporate teams are an essential part of our company, ensuring that we’re setting our businesses, clients, customers and employees up for success.
Requirements
- Formal training, certification, or equivalent practical experience in software engineering concepts
- Hands-on experience with system design, application development, testing, and operational stability in production environments
- Advanced proficiency in Python for building production-grade services and tooling
- Proficiency with automation and continuous delivery methods
- Hands-on experience with AWS and Terraform for infrastructure delivery and lifecycle management
- Strong understanding of site reliability engineering practices, including incident management, root-cause analysis, runbooks, and reliability patterns
- Practical knowledge of observability and instrumentation across metrics, logs, and traces
- Comfort with on-call operations and production troubleshooting
- Hands-on production experience operating LLM inference servers such as vLLM and llm-d (or directly equivalent serving stacks)
- Hands-on experience hosting and serving LLMs on Amazon EKS and/or Amazon SageMaker, and on local GPU infrastructure
- Knowledge of LLM reliability and risk considerations, including latency/throughput trade-offs, model and weight versioning, prompt/response logging, and safe rollout patterns, * Experience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patterns
- Experience building AI agents using frameworks such as LangChain, CrewAI, LangGraph, or similar orchestration platforms
- Experience operating or integrating serving platforms such as KServe, Ray Serve, NVIDIA Triton Inference Server, Text Generation Inference (TGI), alongside vLLM/llm-d
- Familiarity with Amazon SageMaker JumpStart, SageMaker Endpoints, and Amazon Bedrock for managed model hosting
- Experience with online LLM quality monitoring (e.g., hallucination, toxicity, drift detection) and tracing via OpenTelemetry conventions
- Contributions to open-source LLM serving or inference projects (e.g., vLLM, llm-d, Ray, KServe, Triton)
About the company
J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world’s most prominent corporations, governments, wealthy individuals and institutional investors. Our first-class business in a first-class way approach to serving clients drives everything we do. We strive to build trusted, long-term partnerships to help our clients achieve their business objectives.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.collegerecruiter.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
almost 3 years ago
CH
Chris Heilmann
Dev Digest 121 - AI goes offline
over 2 years ago
BB
Benedikt Bischof
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
about 4 years ago
BB
Benedikt Bischof
MLOps And AI Driven Development
over 4 years ago
LM
Luis Minvielle
How to Become an AI Engineer
almost 3 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
over 2 years ago