Observability & Monitoring Engineer

Infinite Computer Solutions (ICS)
United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services JIRA Bash Shell Cloud Computing Databases Distributed Systems Monitoring of Systems Python (Programming Language) Windows PowerShell Cloud Services
+14 more
Data Logging Scripting Google Cloud Enterprise Software Applications Reliability of Systems Infrastructure as Code (IaC) Performance Monitor Cloud Migration BIG-IP Access Policy Manager (APM) ArcSight Event Correlation Terraform Splunk Dynatrace Servicenow

Job description

We are seeking an experienced Senior Observability & Monitoring Engineer to design, implement, and optimize enterprise monitoring and observability solutions supporting the migration of mission-critical applications from on-premises environments to AWS and Google Cloud Platform (Google Cloud Platform). You will play a key role in improving system reliability by enabling proactive monitoring, intelligent alerting, rapid incident detection, and operational excellence., * Design and implement enterprise-wide monitoring and observability solutions.

  • Build dashboards, KPIs, SLIs, and SLOs for applications, infrastructure, databases, APIs, and cloud services.
  • Configure intelligent monitoring, alerting, event correlation, and anomaly detection using Dynatrace, Splunk, and Moogsoft.
  • Develop synthetic monitoring and automated health checks for critical business services.
  • Partner with Cloud Architects, Developers, SREs, and Operations teams to ensure production readiness during cloud migrations.
  • Automate monitoring configurations and operational processes using scripting and AI-assisted tools.

Requirements

  • 8+ years of experience in Monitoring, Observability, SRE, or Operations Engineering.
  • Strong hands-on experience with Dynatrace, Splunk, and Moogsoft.
  • Experience supporting cloud migrations to AWS and/or Google Cloud Platform.
  • Solid understanding of APM, distributed tracing, logging, metrics, and alert management.
  • Experience with Jira, ServiceNow, and incident management.
  • Scripting skills in Python, Bash, or PowerShell.
  • Strong knowledge of enterprise applications and distributed systems., * Experience in fintech, banking, or other regulated industries.
  • Knowledge of OpenTelemetry, Terraform, and Infrastructure as Code (IaC).
  • Experience with self-healing automation and cloud-native observability.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

5:47 min

Integrating user stories and test automation via Jira tools

Christoph Ruggenthaler · LIVE

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

Videos

See all

Related articles

See all