AIOps / Observability Engineer

CLOUD SECURITY WEB LLC
Phoenix, AZ, United States
12 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$93,600.0 - $108,160.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Application Performance Management Microsoft Azure Bash Shell Cloud Engineering DevOps Elasticsearch Python (Programming Language) Automation of Marketing OpenShift
+26 more
Windows PowerShell Reliability Engineering Ansible Prometheus Datadog Data Logging Scripting Google Cloud Cloud Monitoring System Availability Grafana Reliability of Systems Cloudformation Kubernetes Infrastructure Automation Frameworks Performance Monitor Apache Kafka Restful APIs Terraform Splunk New Relic (SaaS) Appdynamics Dynatrace Api Management Docker Servicenow

Job description

We are seeking an experienced AIOps / Observability Engineer to join our team in Phoenix, AZ. The ideal candidate will have strong expertise in observability platforms, cloud monitoring, automation, API integrations, and Site Reliability Engineering (SRE). This role requires hands-on experience supporting production environments, improving system reliability, and implementing proactive monitoring solutions. Banking or financial services experience is highly preferred., * Design, implement, and maintain enterprise observability and monitoring solutions.

  • Develop and optimize AIOps strategies to improve operational efficiency and reduce incident resolution time.
  • Build and integrate APIs to connect monitoring, alerting, and automation platforms.
  • Configure and manage cloud-native monitoring across AWS, Azure, or Google Cloud Platform.
  • Automate operational tasks using scripting languages and infrastructure-as-code tools.
  • Monitor production environments to ensure high availability, performance, and reliability.
  • Perform incident management, root cause analysis (RCA), and problem resolution for critical production issues.
  • Collaborate with development, infrastructure, and operations teams to improve application performance and resiliency.
  • Create dashboards, alerts, and reports to provide real-time visibility into application and infrastructure health.
  • Implement best practices for observability, logging, tracing, and performance monitoring.
  • Participate in on-call production support and continuous service improvement initiatives.

Requirements

  • 8+ years of experience in AIOps, Observability, SRE, or Production Support.
  • Strong experience with observability and monitoring platforms such as Dynatrace, Splunk, AppDynamics, Datadog, Prometheus, Grafana, New Relic, or Elastic Stack.
  • Hands-on experience with REST APIs and API integrations.
  • Experience with cloud platforms including AWS, Azure, or Google Cloud Platform (GCP).
  • Strong scripting and automation experience using Python, Shell, PowerShell, Ansible, Terraform, or similar technologies.
  • Experience with incident management, troubleshooting, and root cause analysis.
  • Strong understanding of application performance monitoring (APM), distributed tracing, logging, and metrics.
  • Experience supporting mission-critical production environments.
  • Excellent analytical, communication, and problem-solving skills.

Preferred Qualifications

  • Banking or Financial Services domain experience.
  • Experience with CI/CD pipelines and DevOps practices.
  • Knowledge of Kubernetes, Docker, and container observability.
  • Familiarity with ITSM tools such as ServiceNow.
  • Experience implementing AI/ML-driven monitoring and predictive analytics.

Nice to Have

  • Kubernetes/OpenShift administration.
  • Kafka or messaging platform monitoring.
  • Infrastructure as Code (Terraform/CloudFormation).
  • Certification in AWS, Azure, GCP, Dynatrace, Splunk, or SRE.

Benefits & conditions

$45 - $52 an hour - Contract

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all