Senior AI Ops & Incident/Site Reliability Engineer

Perficient
Austin, TX, United States
17 days ago

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$93,600.0 - $170,640.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Cloud Engineering DevOps Monitoring of Systems Information Technology Operations Automation of Marketing Reliability Engineering Site Reliability Engineering Practices Software Engineering
+11 more
Toolchain Workflow Management Systems Google Cloud Enterprise Software Applications Computer Network Operations Mttr Reliability of Systems Information Technology Machine Learning Operations Dynatrace Servicenow

Job description

We currently have a career opportunity for a NOC AI-Ops Engineer to join our team located in Austin, TX. This is a hybrid role, 3 days a week in office., Perficient is always looking for the best and brightest talent and we need you! We’re a quickly-growing, global digital consulting leader, and we’re transforming the world’s largest enterprises and biggest brands. You’ll work with the latest technologies, expand your skills, and become a part of our global community of talented, diverse, and knowledgeable colleagues. Responsibilities:

  • Incident & Recovery Management
  • Monitor, document, and analyze major incident response efforts and service recovery activities.
  • Serve as a senior escalation point for Tier 1 and Tier 2 operational incidents.
  • Conduct incident reviews, root cause analysis, and corrective action planning.
  • Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
  • Site Reliability Engineering
  • Implement SRE practices to improve platform reliability, scalability, and resiliency.
  • Define and monitor SLAs, SLOs, and operational KPIs.
  • Develop proactive reliability and availability strategies.
  • AIOps & Automation
  • Implement AIOps solutions to automate incident detection, diagnosis, remediation, and prevention.
  • Build and optimize AI-powered operational agents and self-healing workflows.
  • Reduce operational effort through intelligent automation.
  • Observability & Monitoring
  • Lead enterprise monitoring initiatives using Dynatrace and related observability platforms.
  • Improve visibility across cloud, infrastructure, applications, and user experiences.
  • Enable predictive monitoring and anomaly detection.
  • ITSM & Service Operations
  • Develop and enhance incident, problem, change, and event management frameworks aligned with ITIL and ITSM best practices.
  • Leverage ServiceNow workflow automation to improve service delivery.
  • Cross-Functional Leadership
  • Partner with Infrastructure, DevOps, Cloud, Security, Application Development, and NOC teams.
  • Mentor operational teams and promote an automation-first culture.

Requirements

We are seeking a Senior AIOps and Incident/Site Reliability Engineer to lead incident management, operational resilience, and intelligent automation initiatives across enterprise technology environments. This role will partner with Network Operations Center (NOC), Infrastructure Operations, Cloud Engineering, DevOps, and Application Support teams to proactively detect, respond to, and prevent technology incidents. The ideal candidate combines hands-on incident management expertise with experience implementing observability, automation, and AI-driven operational solutions to improve system reliability, reduce operational overhead, and enhance customer experience. The candidate should possess deep expertise in AIOps, ITSM, ITIL, SRE, Incident Management, Cloud Operations, and Enterprise Infrastructure., * Bachelor’s degree in Computer Science, Information Technology, Engineering, or related field (or equivalent experience).

  • 8+ years of experience in IT Operations, Site Reliability Engineering, Infrastructure Operations, Network Operations, or Production Support environments.
  • 5+ years of experience leading incident management, operational transformation, or reliability engineering initiatives.
  • Strong experience with:

  • Site Reliability Engineering (SRE)
  • IT Service Management (ITSM)
  • ITIL Framework
  • Incident, Problem, Change, and Event Management
  • Network Operations Center (NOC)
  • Infrastructure Operations
  • Service Desk Operations
  • Application Production Support
  • Cloud Platforms (AWS, Azure, or GCP)
  • DevOps Practices and Toolchains
  • Hands-on experience with Dynatrace, monitoring platforms, and observability solutions.
  • Experience using ServiceNow for ticketing, workflow automation, and service management.
  • Strong understanding of infrastructure, networking, cloud architecture, and enterprise application ecosystems.
  • Proven experience conducting root cause analysis and implementing preventive controls.
  • Experience leading enterprise AIOps implementations.
  • Experience building AI-powered operational agents and intelligent automation solutions.
  • Certifications such as:

  • ITIL Foundation or ITIL Managing Professional
  • Certified Site Reliability Engineer (SRE)
  • AWS, Azure, or Google Cloud certifications
  • ServiceNow certifications
  • Experience with workflow orchestration and enterprise automation platforms.
  • Familiarity with predictive analytics, machine learning operations, and autonomous operations frameworks.

Benefits & conditions

  • $126,500-182,000 per year

About the company

This is a hybrid role requiring on-site attendance up to three days per week. Candidates must reside within approximately one hour commuting distance of one of the following office locations: Fort Mill, SC, Austin, TX, Boston, MA, New York, NY, Tempe, AZ, or San Diego, CA., About Us: Perficient is the global AI and technology consulting firm disrupting the traditional consulting model. Powered by our 7,000+ advisors, engineers, and designers, Perficient implements AI-first solutions that break conventions and deliver outcomes that matter. Proudly serving clients that represent the world’s most innovative brands, and in collaboration with our powerful technology partner ecosystem, we bring deep industry expertise and data-driven design to redefine how businesses run and succeed. Perficient is different. For real. Learn more at .

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · WWC 2025

2:19 min

Applying code assistant capabilities to infrastructure and cloud operations

Ryan J Salva · Coffee With Developers

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

2:35 min

The future of artificial intelligence in platform engineering operations

Anna Ozor Anna Ozor · Europe 2026 Virtual

Videos

See all

Related articles

See all