Infrastructure Engineer Staff - Netcool (NOI) SME
Role details
Job location
Tech stack
Job description
Part of a larger team delivering high quality Computer, Network, Storage and End-User Infrastructure technology solutions, on-going support to the business. Independently completes and leads the largest and most complex infrastructure project assignments. Plan, research, evaluate, design, and engineer the enterprise's technology infrastructure. Provide technical support and troubleshooting, cost estimates, justifications, and recommendations. Produces technical documentation, support and configuration. Helps manage, plan, and maintain technical platforms including upgrading systems. Monitor system performance, and install and configure hardware. Responsible for collaborating with other Job Families such as Project Managers, Architects, Solution Engineers, Technicians, Business Analysts to deliver consistent, reliable technology solutions that leverage AEP's technology standards, architectures and best practices., The Senior Enterprise Monitoring & Event Management Engineer is responsible for the design, administration, integration, automation, and support of AEP's enterprise monitoring and event management platforms. This role serves as a technical subject matter expert for IBM Netcool Operations Insight (NOI), event correlation, alert automation, monitoring integrations, ServiceNow event management, and enterprise observability solutions.
The engineer will design and maintain monitoring solutions across infrastructure, applications, databases, cloud environments, and business-critical systems while driving automation, event reduction, incident response improvements, and platform modernization initiatives.
The position requires deep technical expertise across Linux administration, monitoring platforms, scripting, event management, system integrations, and production support in a large enterprise environment., Required Technical Skills
- Monitoring & Event Management Platforms
- IBM Netcool Operations Insight (NOI)
- IBM Netcool OMNIbus
- ServiceNow Event Management (EM)
- Dynatrace
- Splunk
- SolarWinds
- OpenNMS
- Enterprise monitoring and observability solutions
Operating Systems
- Linux Administration (RHEL preferred)
- Unix/Linux troubleshooting
- Process and service management
- Performance tuning and diagnostics
- Linux interprocess communication (IPC)
Programming & Scripting
- Python
- Shell Scripting
- Perl
- JavaScript
- SQL
- JSON
- XML/XSLT
Integration Technologies
- REST APIs
- SOAP APIs
- Webhooks
- SNMP
- Syslog
- SMTP
- Kafka
- Message bus technologies
Databases
- Oracle
- SQL Server
- PostgreSQL
- MySQL
- MongoDB
Cloud & Infrastructure
- AWS monitoring integrations
- Azure monitoring integrations
- Kubernetes
- Docker
- Virtualization technologies (VMware)
Automation Technologies
- Ansible
- NOI Runbooks
- ServiceNow Flows
- Event-driven automation
- Custom workflow development
Analytics & AI
- AIOps platforms
- Event correlation
- Root cause analysis
- Anomaly detection
- Log analytics
- Predictive alerting
- Machine learning-enabled monitoring
Preferred Skills
- Utility industry monitoring experience
- SOX-regulated application support
- ITIL framework knowledge
- ServiceNow ITOM
- CMDB integrations
- Event Management architecture
- Infrastructure monitoring design
- Enterprise observability practices
- Network monitoring and management systems
Roles & Responsibilities
- Event Management Engineering
- Design, develop, and maintain enterprise event management and correlation solutions.
- Build, maintain, and optimize Netcool probes, gateways, automation policies, rules, and event processing workflows.
- Develop alarm correlation, suppression, enrichment, normalization, deduplication, and root cause analytics.
- Design enterprise event reduction and noise suppression strategies.
- Identify opportunities to automate operational monitoring processes.
Platform Administration & Support
- Administer and support IBM Netcool Operations Insight (NOI) production and non-production environments.
- Maintain platform health, availability, performance, and resiliency.
- Perform platform upgrades, patching, capacity planning, and lifecycle management.
- Troubleshoot platform issues involving event processing, integrations, databases, and infrastructure components.
- Provide Level 3 support for enterprise monitoring and event management solutions.
Monitoring Architecture
- Design monitoring solutions across servers, databases, network devices, cloud platforms, applications, middleware, and infrastructure services.
- Develop monitoring standards, onboarding processes, and alerting strategies for business-critical systems.
- Partner with application teams to create meaningful actionable alerts and reduce alert fatigue.
- Improve service visibility and operational readiness across enterprise environments.
ServiceNow Integration
- Design and support integrations between monitoring platforms and ServiceNow.
- Implement automated incident creation, event enrichment, routing, escalation, and ticket lifecycle management.
- Support Event Management and AIOps initiatives within ServiceNow.
- Integrate monitoring tools with CMDB and service mapping solutions.
Integration Engineering
- Develop and support integrations using REST, SOAP, SNMP, Syslog, webhooks, messaging technologies, and custom APIs.
- Integrate monitoring platforms with infrastructure, application, database, and cloud environments.
- Support enterprise onboarding activities for new monitoring sources and technologies.
Automation & Continuous Improvement
Develop event-driven automation and remediation workflows.
Create self-healing and closed-loop operational automation capabilities.
Implement monitoring automation using scripts, workflows, runbooks, and orchestration platforms.
Drive monitoring modernization initiatives and platform improvements.
Production Operations Support
Participate in on-call rotation and critical incident response activities.
Support enterprise monitoring platforms that provide critical operational visibility to IT and business stakeholders.
Analyze monitoring failures and implement corrective actions.
Serve as a technical escalation point for monitoring and integration-related incidents.
Governance, Risk & Compliance
Support monitoring solutions that are subject to SOX, audit, and compliance requirements.
Maintain documentation, runbooks, operational procedures, and technical standards.
Participate in audit activities and evidence collection when required.
Ensure monitoring solutions align with cybersecurity and operational standards.
Requirements
- Strong customer service and stakeholder management skills
- Exceptional troubleshooting and analytical abilities
- Ability to communicate effectively with technical and non-technical audiences
- Ability to prioritize multiple projects and operational demands
- Self-directed and highly accountable
- Strong collaboration skills across infrastructure, security, cloud, and application teams
- Ability to lead technical initiatives and mentor junior engineers
- Strategic mindset with the ability to align monitoring capabilities with business outcomes, * Bachelor's Degree in Information Technology, Computer Science, Engineering, or related field.
- 5+ years administering enterprise monitoring, event management, observability, or infrastructure management platforms.
- Experience with IBM Netcool/NOI, ServiceNow Event Management, Dynatrace, Splunk, or equivalent enterprise monitoring solutions.
- Strong Linux administration and troubleshooting skills.
- Strong experience with scripting and automation.
- Experience with enterprise incident management processes.
- Experience integrating enterprise applications using APIs, SNMP, Syslog, or messaging technologies.
What We're Looking For:
Education requirements are listed below:
- Bachelor's degree in computer science, engineering, or related technical field is required.
Work Experience requirement listed below:
- 10 years of relevant work experience is required. An equivalent combination of education and experience may be considered.
Benefits & conditions
3.73.7 out of 5 stars Gahanna, OH 43230 $116,255.00 - $151,132.50 a year - Full-time, Pulled from the full job description
- Tuition reimbursement