Senior Site Reliability Engineer
Role details
Job location
Tech stack
Job description
Experteer Overview In this Senior SRE role you will design, build, and operate large-scale, distributed infrastructure that enables AI-powered products. You'll own patching, image lifecycle, and security guardrails while collaborating with software, ML, and platform teams to bake reliability in from the start. You'll work across cloud-native environments, ML inference pipelines, and on-call rotations to keep systems resilient. This is a hands-on opportunity to shape scalable cloud and AI platforms at Adobe. Compensation / Benefits * Design, build, and operate large-scale, distributed systems and the tooling behind them (IaC, CI/CD, automation) * Own patch, vulnerability, and golden-image lifecycle management at fleet scale * Contribute to multi-quarter initiatives (compute rationalization, cloud re-platforming, ML platform migration) * Build and operate ML inference infrastructure (model serving, GPU workloads, language model gateway/routing) * Help set infrastructure standards and security guardrails for agentic AI; develop agent tooling for operations * Partner with software, ML, and platform teams to bake reliability in from the start * Share on-call duties and address issues across web services, databases, and pipelines Tasks * BSc in Computer Science or equivalent * Python (plus familiarity with PHP, Node.js, or Ruby) * Experience deploying/operating ML inference pipelines in production (SageMaker, OpenAI, Bedrock, or equivalent) * Experience with cloud-native compute at scale (AWS EC2 Auto Scaling Groups, Kubernetes; Azure/GCP a plus) * Experience with vulnerability/patch management and AMI/golden-image automation * Strong debugging skills on distributed systems * Willingness to participate in on-call rotation Key requirements * remote work * growth opportunities in AI * collaborative culture * security and guardrails focus * infrastructure modernization * competitive pay range
Requirements
Experteer Overview In this Senior SRE role you will design, build, and operate large-scale, distributed infrastructure that enables AI-powered products. You'll own patching, image lifecycle, and security guardrails while collaborating with software, ML, and platform teams to bake reliability in from the start. You'll work across cloud-native environments, ML inference pipelines, and on-call rotations to keep systems resilient. This is a hands-on opportunity to shape scalable cloud and AI platforms at Adobe. Compensation / Benefits * Design, build, and operate large-scale, distributed systems and the tooling behind them (IaC, CI/CD, automation) * Own patch, vulnerability, and golden-image lifecycle management at fleet scale * Contribute to multi-quarter initiatives (compute rationalization, cloud re-platforming, ML platform migration) * Build and operate ML inference infrastructure (model serving, GPU workloads, language model gateway/routing) * Help set infrastructure standards and aa behind guardrails for agentic AI; develop agent tooling for operations * Partner with software, ML, and platform teams to bake reliability in from the start * Share on-call duties and address issues across web services, databases, and pipelines Tasks * BSc in Computer Science or equivalent * Python (plus familiarity with PHP, Node.js, or Ruby) * Experience deploying/operating ML inference pipelines in production (SageMaker, OpenAI, Bedrock, or equivalent) * Experience with cloud-native compute at scale (AWS EC2 Auto Scaling Groups, Kubernetes; Azure/GCP a plus) * Experience with vulnerability/patch management and AMI/golden-image automation * Strong debugging skills on distributed systems * Willingness to participate in on-call rotation Key requirements * remote work * growth opportunities in AI * collaborative culture * security and guardrails focus * infrastructure modernization * competitive pay range