Senior Site Reliability Engineer

Adobe Systems
New York, NY, United States
4 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

PHP (Programming Language) Artificial Intelligence Amazon Elastic Compute Cloud Microsoft Azure Cloud Computing Cloud Engineering Databases Continuous Integration Software Debugging Distributed Systems Python (Programming Language) Node.Js
+12 more
Reliability Engineering Ruby Azure Machine Learning Web Services Autoscaling Delivery Pipeline Adobe AI Platforms Kubernetes Information Technology Patch Management Machine Learning Operations

Job description

Experteer Overview In this Senior SRE role you will design, build, and operate large-scale, distributed infrastructure that enables AI-powered products. You’ll own patching, image lifecycle, and security guardrails while collaborating with software, ML, and platform teams to bake reliability in from the start. You’ll work across cloud-native environments, ML inference pipelines, and on-call rotations to keep systems resilient. This is a hands-on opportunity to shape scalable cloud and AI platforms at Adobe. Compensation / Benefits * Design, build, and operate large-scale, distributed systems and the tooling behind them (IaC, CI/CD, automation) * Own patch, vulnerability, and golden-image lifecycle management at fleet scale * Contribute to multi-quarter initiatives (compute rationalization, cloud re-platforming, ML platform migration) * Build and operate ML inference infrastructure (model serving, GPU workloads, language model gateway/routing) * Help set infrastructure standards and security guardrails for agentic AI; develop agent tooling for operations * Partner with software, ML, and platform teams to bake reliability in from the start * Share on-call duties and address issues across web services, databases, and pipelines Tasks * BSc in Computer Science or equivalent * Python (plus familiarity with PHP, Node.js, or Ruby) * Experience deploying/operating ML inference pipelines in production (SageMaker, OpenAI, Bedrock, or equivalent) * Experience with cloud-native compute at scale (AWS EC2 Auto Scaling Groups, Kubernetes; Azure/GCP a plus) * Experience with vulnerability/patch management and AMI/golden-image automation * Strong debugging skills on distributed systems * Willingness to participate in on-call rotation Key requirements * remote work * growth opportunities in AI * collaborative culture * security and guardrails focus * infrastructure modernization * competitive pay range

Requirements

Experteer Overview In this Senior SRE role you will design, build, and operate large-scale, distributed infrastructure that enables AI-powered products. You’ll own patching, image lifecycle, and security guardrails while collaborating with software, ML, and platform teams to bake reliability in from the start. You’ll work across cloud-native environments, ML inference pipelines, and on-call rotations to keep systems resilient. This is a hands-on opportunity to shape scalable cloud and AI platforms at Adobe. Compensation / Benefits * Design, build, and operate large-scale, distributed systems and the tooling behind them (IaC, CI/CD, automation) * Own patch, vulnerability, and golden-image lifecycle management at fleet scale * Contribute to multi-quarter initiatives (compute rationalization, cloud re-platforming, ML platform migration) * Build and operate ML inference infrastructure (model serving, GPU workloads, language model gateway/routing) * Help set infrastructure standards and aa behind guardrails for agentic AI; develop agent tooling for operations * Partner with software, ML, and platform teams to bake reliability in from the start * Share on-call duties and address issues across web services, databases, and pipelines Tasks * BSc in Computer Science or equivalent * Python (plus familiarity with PHP, Node.js, or Ruby) * Experience deploying/operating ML inference pipelines in production (SageMaker, OpenAI, Bedrock, or equivalent) * Experience with cloud-native compute at scale (AWS EC2 Auto Scaling Groups, Kubernetes; Azure/GCP a plus) * Experience with vulnerability/patch management and AMI/golden-image automation * Strong debugging skills on distributed systems * Willingness to participate in on-call rotation Key requirements * remote work * growth opportunities in AI * collaborative culture * security and guardrails focus * infrastructure modernization * competitive pay range

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

50 sec

Why developer happiness matters in web frameworks

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:17 min

Eco-friendly factory operations and generative AI image errors

Chris Heilmann +1 · LIVE

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · WWC 2023

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

3:55 min

Identifying underlying Node.js runtime vulnerabilities using fuzzing tools

Sonya Moisset · WWC 2023

59 sec

Evaluating intrusive system modifications by enterprise software

Chris Heilmann +2 · LIVE

Videos

See all

Related articles

See all