> Markdown version of [/jobs/ext/1316341-job-posting-title-ai-devops-engineer](https://www.wearedevelopers.com/jobs/ext/1316341-job-posting-title-ai-devops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Job Posting Title AI/ DevOps Engineer - **Company:** Adobe Systems - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $228,600.0 - $331,050.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Build Automation, Microsoft Azure, Customer Data Management, Data Stores, Software Debugging, DevOps, Distributed Systems, Amazon DynamoDB, PostgreSQL, Reliability Engineering, Prometheus, Azure Machine Learning, Datadog, Aerospike, Grafana, Adobe, Containerization, Kubernetes, Information Technology, Low Latency, Machine Learning Operations, Data Pipelines - **Published:** July 17, 2026 - **Apply:** https://adobe.wd5.myworkdayjobs.com/external_experienced/job/San-Jose/Job-Posting-Title-AI--DevOps-Engineer_R170387-1 ## About the Role * 6-10 years in SRE, infrastructure, or platform engineering * Proven track record operating large-scale distributed systems in production * Strong foundation in datastores, reliability engineering, and automation * Hands-on experience with Kubernetes and containerized environments, a major cloud (AWS, Azure, or GCP), and modern observability tooling (Prometheus, Grafana, OpenTelemetry, or equivalents) * Real experience in incident response and driving operational improvements out of it * Working knowledge of - or genuine interest in - AI/ML systems or MLOps (expertise not required) * Comfortable with scale, ambiguity, and high ownership * Strong problem-solving instincts and a bias for action * About Adobe ## Description Adobe's Real-Time Customer Data Platform (RTCDP) powers personalized experiences for some of the world's largest brands. As a Senior SRE on this team, you'll be central to keeping RTCDP reliable, scalable, and operationally excellent at global scale. This is a hands-on, high-ownership role at the intersection of production operations (Day 2 ownership) and core datastore engineering, with a growing surface area in operationalizing AI/ML services and workflows. What you'll do Own production reliability * Own day-to-day reliability for RTCDP services - availability, performance, and durability against SLOs * Participate in on-call rotations and incident response, driving mitigation and recovery through SEV3-SEV1 events * Lead post-incident reviews and follow-up work * Strengthen operational readiness, playbooks, and on-call health * Partner with product and platform teams on production-ready launches and regional expansions Operate and evolve core datastores You'll work across RTCDP's distributed datastore ecosystem: Aerospike, FoundationDB, Postgres, and CosmosDB/DynamoDB. * Drive reliability, scaling, and operational excellence across these platforms * Own upgrades, capacity management, backup/restore, and DR testing * Build automation for provisioning, scaling, and lifecycle management * Identify and ship cost optimizations (rightsizing, storage/compute efficiency) Drive automation and observability * Build automation-first solutions that reduce toil and improve system safety * Improve monitoring, alerting, and observability - anchored to real customer impact * Establish standardized operational patterns across services and regions * Support the rollout of SLO-driven reliability practices Contribute to AI/ML Ops (emerging area) A complementary part of the role, not the primary focus. * Support infrastructure and operational needs for AI/ML-powered services in RTCDP * Shape operational practices for model serving and data pipelines - reliability, scaling, monitoring * Help land core MLOps patterns where relevant: model deployment workflows, inference observability (latency, errors), data quality and pipeline reliability signals * Partner with ML and data teams to get AI-driven features production-ready * Leverage AI-assisted tools (e.g., Copilot, Claude Code, Codex, internal tooling) to accelerate debugging, incident response, and operational workflows Technical leadership and collaboration * Operate as a strong IC and technical lead on cross-cutting projects * Mentor junior engineers and raise team practices * Partner closely with engineering, infrastructure, and security * Live the core SRE/DevOps principles: ownership, automation, error budgets, continuous improvement Why this role is interesting * Deep involvement in production systems at global scale * Hands-on ownership spanning operations and datastore platforms * Exposure to next-generation work: AI/ML systems and AI-assisted engineering * Direct impact on customer reliability, platform scalability, and cost * A clear path toward architect-level influence over time ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [3x Performance: A Humbling Journey](https://www.wearedevelopers.com/videos/100165-3x-performance-a-humbling-journey) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)