Observability & Enterprise Monitoring Engineer

DIGITAL TECHNOLOGY SOLUTIONS
Seattle, WA, United States
16 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Configuration Management Complex Networks Computer Networks Databases Database Theory Dynamic Host Configuration Protocol Domain Name System (DNS) Executive Information Systems
+24 more
Monitoring of Systems Network Topologies Python (Programming Language) Linux System Administration Machine Learning Microsoft SQL Server NetFlow Routing Windows PowerShell Simple Network Management Protocols SQL Databases Syslog TCP/IP Google Cloud Data Ingestion Cloud Monitoring System Availability HybridCloud OpenText SolarWinds (Software) Npm(Software) Api Design Restful APIs Splunk

Job description

Observability & Enterprise Monitoring Engineer with specialized expertise in SolarWinds platform administration and broader multi-tool observability ecosystems. Working knowledge of OpenText NNMi will be an added advantage. This role will be responsible for the end-to-end administration, optimization, integration, and operational maintenance of enterprise-scale implementation of monitoring solutions (SolarWinds). Responsible for ensuring platform health, automate alert workflows, manage hybrid/cloud monitoring integrations, and collaborate closely with cross-functional infrastructure teams to maintain high availability and performance., * Core Module Management: Administer and optimize SolarWinds modules including NPM, NCM, NTA, SAM, and the broader Orion / SWOSH (Hybrid Cloud Observability) platform ecosystem.

  • Upgrades & Maintenance: Perform routine and major version updates across platform components; monitor platform health using Active Diagnostics and My Deployment health checks.
  • Polling Infrastructure: Manage, scale, and load-balance Additional Polling Engines (APEs) to ensure optimal performance across enterprise environments.
  • Database & Backup Operations: Perform operational tasks on the underlying MS SQL Database, manage, schedule, and verify configuration and database backup jobs.
  1. Network & Device Observability Operations
  • Discovery & Asset Management: Execute network discoveries, manage node onboarding/offboarding, assign Universal Device Pollers (UnDP), and maintain custom custom attributes and group hierarchies.
  • Configuration Management (NCM): Build and maintain NCM command templates, automate daily startup/running config backups, archive config files, and remediate compliance/transfer failures.
  • Topology & Visualization: Create and maintain dynamic, accurate network topology maps using Network Atlas and modern visual canvases based on operational requirements.
  1. Alerting, Dashboarding & ITSM Integration
  • Signal Optimization: Design, tune, and maintain custom Alert Triggers, Actions, and Thresholds to eliminate alert noise and drive actionable alerting.
  • Ticketing & Automation: Configure bi-directional ITSM/ticketing integrations to enable automatic ticket creation, routing, and lifecycle tracking.
  • Reporting & Visibility: Build custom operational and executive Dashboards, Views, and Reports tailored to stakeholder requirements.
  • Incident Support: Monitor alert channels for operational anomalies, troubleshoot lingering telemetry issues, and collaborate with domain teams to drive root cause resolution.
  1. AIOps Operations
  • Leverage AIOps, machine learning, and pattern-recognition capabilities to identify baseline anomalies, reduce event noise, and drive predictive incident management.
  • Collaborate with cross-functional teams to integrate AI-driven event correlation models and automated remediation workflows into the central monitoring platform.
  1. Integration, Vendor Coordination
  • Manage relationships and support escalations with platform vendors.
  • Work on REST API integrations across applications/tools as per requirements.
  1. Operational Troubleshooting & Diagnostics
  • Perform deep-dive troubleshooting and root-cause analysis for platform-level performance degradations, engine polling failures, and monitoring agent corruptions.
  • Utilize Active Diagnostics and system telemetry to investigate and resolve complex network configuration transfer failures, polling sync latency, and data ingestion issues.

Requirements

  • Multi-tool expertise (SolarWinds, OpenText NNMi, Splunk, etc.)
  • Protocol & Telemetry Knowledge: In-depth understanding of SNMP (v2c/v3), WMI, WinRM, Syslog, NetFlow/sFlow, and Observability (Metrics, Logs, Traces).
  • Automation & API Integration: Good to have skills in PowerShell/Python, and API-driven automation for monitoring workflows.
  • AIOps & Intelligent Automation: Basic understanding of AIOps concepts, machine learning algorithms for anomaly detection, automated event correlation, and predictive analytics within modern observability frameworks.
  • Cloud & Hybrid Observability: Hands-on experience extending platform monitoring into AWS, Azure, or Google Cloud Platform environments.

  • Infrastructure Knowledge

o System Administration: Intermediate knowledge of Windows and Linux administration.

o Database: Understanding of SQL/Database concepts and standard query execution.

o Networking: Good understanding of networking concepts including TCP/IP, DNS, DHCP, Routing and Switching.

o ITSM: Experience in ITSM processes and operational support.

Benefits & conditions

DTS offers excellent compensation package.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

46 sec

Automating telemetry collection through robust Telegraf deployment

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:29 min

Real-world observability applications and Cilium integration patterns

Ozan Sazak Ozan Sazak · WWC 2024

2:30 min

Discovering and instrumenting services using systemd process enumeration

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

Videos

See all

Related articles

See all