Job Posting Title AI/ DevOps Engineer

Adobe Systems
San Jose, CA, United States
25 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$228,600.0 - $331,050.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Build Automation Microsoft Azure Customer Data Management Data Stores Software Debugging DevOps Distributed Systems Amazon DynamoDB PostgreSQL Reliability Engineering
+12 more
Prometheus Azure Machine Learning Datadog Aerospike Grafana Adobe Containerization Kubernetes Information Technology Low Latency Machine Learning Operations Data Pipelines

Job description

Adobe’s Real-Time Customer Data Platform (RTCDP) powers personalized experiences for some of the world’s largest brands. As a Senior SRE on this team, you’ll be central to keeping RTCDP reliable, scalable, and operationally excellent at global scale. This is a hands-on, high-ownership role at the intersection of production operations (Day 2 ownership) and core datastore engineering, with a growing surface area in operationalizing AI/ML services and workflows.

What you’ll do Own production reliability

  • Own day-to-day reliability for RTCDP services - availability, performance, and durability against SLOs
  • Participate in on-call rotations and incident response, driving mitigation and recovery through SEV3-SEV1 events
  • Lead post-incident reviews and follow-up work
  • Strengthen operational readiness, playbooks, and on-call health
  • Partner with product and platform teams on production-ready launches and regional expansions

Operate and evolve core datastores You’ll work across RTCDP’s distributed datastore ecosystem: Aerospike, FoundationDB, Postgres, and CosmosDB/DynamoDB.

  • Drive reliability, scaling, and operational excellence across these platforms
  • Own upgrades, capacity management, backup/restore, and DR testing
  • Build automation for provisioning, scaling, and lifecycle management
  • Identify and ship cost optimizations (rightsizing, storage/compute efficiency)

Drive automation and observability

  • Build automation-first solutions that reduce toil and improve system safety
  • Improve monitoring, alerting, and observability - anchored to real customer impact
  • Establish standardized operational patterns across services and regions
  • Support the rollout of SLO-driven reliability practices

Contribute to AI/ML Ops (emerging area) A complementary part of the role, not the primary focus.

  • Support infrastructure and operational needs for AI/ML-powered services in RTCDP
  • Shape operational practices for model serving and data pipelines - reliability, scaling, monitoring
  • Help land core MLOps patterns where relevant: model deployment workflows, inference observability (latency, errors), data quality and pipeline reliability signals
  • Partner with ML and data teams to get AI-driven features production-ready
  • Leverage AI-assisted tools (e.g., Copilot, Claude Code, Codex, internal tooling) to accelerate debugging, incident response, and operational workflows

Technical leadership and collaboration

  • Operate as a strong IC and technical lead on cross-cutting projects
  • Mentor junior engineers and raise team practices
  • Partner closely with engineering, infrastructure, and security
  • Live the core SRE/DevOps principles: ownership, automation, error budgets, continuous improvement

Why this role is interesting

  • Deep involvement in production systems at global scale
  • Hands-on ownership spanning operations and datastore platforms
  • Exposure to next-generation work: AI/ML systems and AI-assisted engineering
  • Direct impact on customer reliability, platform scalability, and cost
  • A clear path toward architect-level influence over time

Requirements

  • 6-10 years in SRE, infrastructure, or platform engineering
  • Proven track record operating large-scale distributed systems in production
  • Strong foundation in datastores, reliability engineering, and automation
  • Hands-on experience with Kubernetes and containerized environments, a major cloud (AWS, Azure, or GCP), and modern observability tooling (Prometheus, Grafana, OpenTelemetry, or equivalents)
  • Real experience in incident response and driving operational improvements out of it
  • Working knowledge of - or genuine interest in - AI/ML systems or MLOps (expertise not required)
  • Comfortable with scale, ambiguity, and high ownership
  • Strong problem-solving instincts and a bias for action *

About Adobe

Benefits & conditions

Our compensation reflects the cost of labor across several U.S. geographic markets, and we pay differently based on those defined markets. The U.S. pay range for this position is $173,500 – $331,050 annually. Pay within this range varies by work location and may also depend on job-related knowledge, skills, and experience. Your recruiter can share more about the specific salary range for the job location during the hiring process.

In California, the pay range for this position is $228,600 - $331,050

At Adobe, for sales roles starting salaries are expressed as total target compensation (TTC = base + commission), and short-term incentives are in the form of sales commission plans. Non-sales roles starting salaries are expressed as base salary and short-term incentives are in the form of the Annual Incentive Plan (AIP).

In addition, certain roles may be eligible for long-term incentives in the form of a new hire equity award.

About the company

Adobe empowers everyone to create through innovative platforms and tools that unleash creativity, productivity and personalized customer experiences. Adobe’s industry-leading offerings including Adobe Acrobat Studio, Adobe Express, Adobe Firefly, Creative Cloud, Adobe Experience Platform, Adobe Experience Manager, and GenStudio enable people and businesses to turn ideas into impact, powered by AI and driven by human ingenuity.

Our 30,000+ employees worldwide are creating the future and raising the bar as we drive the next decade of growth. We’re on a mission to hire the very best and believe in creating a company culture where all employees are empowered to make an impact. At Adobe, we believe that great ideas can come from anywhere in the organization. The next big idea could be yours.

Let’s Adobe together

At Adobe, we believe in creating a company culture where all employees are empowered to make an impact. Learn more about Adobe life, including our values and culture, focus on people, purpose and community, Adobe for All, comprehensive benefits programs, the stories we tell, the customers we serve, and how you can help us advance our mission of empowering everyone to create., At Adobe, we empower employees to innovate with AI - and we look for candidates eager to do the same. As part of the hiring experience, we provide clear guidance on where AI is encouraged during the process and where it’s restricted during live interviews. See how we think about AI in the hiring experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on adobe.wd5.myworkdayjobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:17 min

Eco-friendly factory operations and generative AI image errors

Chris Heilmann +1 · LIVE

1:31 min

Essential AI and human skills for future teams

Alexander Weißhaupt Alexander Weißhaupt +1 · WWC 2025

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

Videos

See all

Related articles

See all