Senior SRE / DevOps Engineer - AI Platform Engineering
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+19 more
Job description
We are looking for a Senior SRE / DevOps Engineer to help design and operate a shared AWS-based AI platform supporting multiple AI-driven products and services. A core objective is enabling teams to switch between AI providers and models (e.g. OpenAI, Anthropic, AWS Bedrock, hosted/open-source models) without impacting downstream applications or customer experience. You will help define reusable infrastructure patterns, deployment standards, observability tooling, and operational best practices that allow teams to rapidly deliver AI workloads while maintaining reliability, security, scalability, and cost control. This is a hands-on role suited to someone comfortable making architectural decisions, driving standardisation, and operating cloud-native AI systems in production., * Design and implement AI-agnostic platform patterns enabling interchangeable LLM/model providers
- Build routing and integration patterns for hosted and third-party AI services
- Support AI inference workloads and model-serving infrastructure in production
- Implement observability, monitoring, usage tracking, latency analysis, and cost visibility for AI systems
- Contribute to agentic workflow and orchestration capabilities
- Define reusable AWS architecture patterns and infrastructure standards
- Build and maintain Terraform modules and CI/CD deployment pipelines
- Standardise deployment and operational practices across teams through automation and shared tooling
- Embed SRE best practices including monitoring, logging, tracing, SLIs/SLOs, and alerting
- Operate and troubleshoot distributed systems, API gateways, and container platforms
Requirements
Do you have experience in Terraform?, Do you have a Master’s degree?, This role requires commercial experience operating AI/LLM workloads in production. If your experience is primarily traditional DevOps/Cloud Engineering without hands-on exposure to:
- LLMs / AI inference workloads
- OpenAI, Bedrock, HuggingFace or hosted models
- AI routing/orchestration
- LLM observability and monitoring
…then this role is unlikely to be suitable.
We are specifically looking for engineers who have helped build or operate AI-enabled platforms at scale., AWS & Platform Engineering
- Strong commercial AWS production experience
- Hands-on experience with ECS, EC2, Lambda, CloudFront, IAM, VPC networking, and multi-account AWS environments
- Strong troubleshooting and operational support experience
- Experience supporting multiple teams/projects on shared platforms
Terraform & DevOps
- Strong Terraform experience including reusable module design and CI/CD integration
- Container platform experience (ECS preferred)
- Experience building shared/internal engineering platforms and reusable infrastructure patterns
AI / LLM Experience (ESSENTIAL) You must have commercial experience managing, routing, hosting, or supporting AI/LLM workloads in production. Examples include:
- OpenAI integrations
- AWS Bedrock
- HuggingFace
- vLLM
- Hosted inference services
- AI orchestration/routing layers
- Model-serving infrastructure
You should also have familiarity with:
- AI workload scaling and cost optimisation
- GPU/compute-heavy workloads
- LLM observability and monitoring
- Prompt/API routing strategies
- Agentic workflows and orchestration patterns
Candidates without meaningful production AI/LLM experience will not be considered.
Desirable Experience
- GPU or specialised compute environments
- Vector databases / embedding pipelines
- Event-driven architectures
- Platform engineering / Internal Developer Platforms
- Security best practices (IAM, secrets management, Zero Trust)
- Kubernetes/EKS exposure
- AI gateways or AI proxy technologies
Wider Skills
- Comfortable making and defending architectural decisions
- Strong communication and stakeholder management
- Detail-oriented with a bias toward automation and standardisation
- Comfortable operating across multiple concurrent workstreams
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on indeed.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 121 - AI goes offline
Find a Developer Job: 12 Best Job Sites For Developers
From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path
What Are Large Language Models?