> Markdown version of [/jobs/ext/61438-senior-aiops-engineer-incident-response-remote-us](https://www.wearedevelopers.com/jobs/ext/61438-senior-aiops-engineer-incident-response-remote-us). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AIOps Engineer, Incident Response [Remote-US] - **Company:** Quanata, Llc - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $215,000.0 - $280,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Confluence, JIRA, Cloud Computing, DevOps, Automation of Marketing, Systems Development Life Cycle, Reliability Engineering, Large Language Models, Multi-Agent Systems, Reliability of Systems, Event Driven Architecture, Kubernetes, Information Technology, Virtual Agents - **Published:** May 30, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a851f03833646d32 ## About the Role Do you have experience in Tooling?, * 6-8 years of experience in production operations, site reliability engineering, technical support engineering, or similar operational roles * Strong background in incident management, root cause analysis, and production system troubleshooting * Experience working within modern SDLC, DevOps, and change management environments * Familiarity with operational tooling such as Jira, Confluence, and observability/monitoring platforms * Strong analytical and problem-solving skills with the ability to identify trends and drive operational improvements * Comfortable working cross-functionally with engineering, product, operations, and leadership teams * Strong communication skills and ability to operate effectively in fast-moving technical environments * Bachelor's degree in Computer Science, Engineering, or equivalent relevant experience, * Experience building or working with AI/LLM-powered systems, intelligent agents, or workflow automation tools * Familiarity with cloud platforms such as AWS and modern observability ecosystems * Experience with event-driven architectures, orchestration frameworks, or operational automation platforms * Background leading operational transformation or reliability improvement initiatives * Passion for AI-native operations, automation, and improving developer/support experiences ## Description We're looking for an experienced production operations and reliability leader to help evolve Quanata's operational support model through AI-driven automation and intelligent agent workflows. This role will own production health, incident response, and operational reliability while partnering closely with engineering and AI orchestration teams to improve scalability, reduce operational toil, and accelerate issue resolution. This is a highly collaborative role for someone who enjoys solving complex production problems, improving systems at scale, and helping modernize operations through AI-native tooling and automation. Your Day-to-Day * Own production health, reliability, and operational support processes across critical systems and services * Lead incident response efforts, stakeholder communication, root cause analysis, and post-incident reviews * Identify patterns in production issues and drive improvements to reduce recurring incidents and operational overhead * Design and implement AI-driven agents and workflows that automate support and operational tasks * Partner with engineering, product, and AI orchestration teams to improve system resilience and operational efficiency * Build and maintain operational runbooks, documentation, and knowledge base content for both human and AI-assisted workflows * Support observability, monitoring, and troubleshooting efforts across cloud-based production environments * Participate in on-call rotations and continuously improve operational readiness and response processes ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [42 x 2 Canvases Later: Two Years, Two Minds, Many Lessons](https://www.wearedevelopers.com/videos/1458-42-x-2-canvases-later-two-years-two-minds-many-lessons) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [A Founder's Journey : From Startup Chaos to Purposeful Growth](https://www.wearedevelopers.com/videos/1926-a-founder-s-journey-from-startup-chaos-to-purposeful-growth) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)