Infrastructure Engineer Staff - Dynatrace SME
Role details
Job location
Tech stack
Job description
Part of a larger team delivering high quality Computer, Network, Storage and End-User Infrastructure technology solutions, on-going support to the business. Independently completes and leads the largest and most complex infrastructure project assignments. Plan, research, evaluate, design, and engineer the enterprise's technology infrastructure. Provide technical support and troubleshooting, cost estimates, justifications, and recommendations. Produces technical documentation, support and configuration. Helps manage, plan, and maintain technical platforms including upgrading systems. Monitor system performance, and install and configure hardware. Responsible for collaborating with other Job Families such as Project Managers, Architects, Solution Engineers, Technicians, Business Analysts to deliver consistent, reliable technology solutions that leverage AEP's technology standards, architectures and best practices., AEP is seeking an experienced Enterprise Observability and Monitoring Engineer to serve as a subject matter expert for enterprise monitoring technologies including Dynatrace and related monitoring platforms. This position is responsible for the design, implementation, administration, enhancement, and operational support of AEP's enterprise monitoring ecosystem.
The engineer will partner with application, infrastructure, cloud, cybersecurity, and operations teams to deliver proactive monitoring, performance management, automation, and observability solutions across critical business applications and infrastructure., Enterprise Monitoring & Observability
- Administer and support Dynatrace enterprise monitoring environments.
- Design and implement monitoring solutions for business-critical applications and infrastructure.
- Develop standards for monitoring, alerting, dashboards, and operational observability.
- Continuously improve enterprise monitoring coverage and accuracy.
- Tune alerting thresholds and event correlation to reduce false positives while ensuring timely incident detection.
- Support enterprise observability initiatives across on-premises, cloud, and hybrid platforms.
Application Performance Monitoring (APM)
- Deploy and administer Dynatrace OneAgent technologies.
- Onboarding large scale applications
- Monitor application health, service dependencies, user experience, and transaction performance.
- Support Real User Monitoring (RUM) and Synthetic Monitoring implementations.
- Analyze application performance issues and identify performance bottlenecks.
- Provide end-to-end visibility across application ecosystems.
Infrastructure Monitoring
- Monitor Windows, Linux, VMware, OpenShift, and cloud-hosted environments.
- Support monitoring for databases, middleware, network infrastructure, storage systems, domain services, load balancers, and enterprise applications.
- Assist infrastructure teams with capacity planning and performance optimization.
- Proactively identify monitoring gaps and service risks.
Monitoring Engineering & Automation
- Develop custom monitoring solutions and Dynatrace Extensions 2.0.
- Create and maintain custom SQL, Oracle, and enterprise application monitoring extensions.
- Build dashboards, metrics, health checks, and alerting policies.
- Automate monitoring deployment and configuration processes.
- Support infrastructure-as-code and configuration management initiatives where applicable.
Incident Response & Root Cause Analysis
- Participate in major incident response and war room activities.
- Perform root cause analysis for application, infrastructure, and monitoring-related incidents.
- Provide monitoring expertise during outage investigations.
- Develop corrective actions and preventative monitoring improvements.
ServiceNow & Event Management Integration
- Support integration between Dynatrace, ServiceNow, and enterprise event management platforms.
- Define monitoring requirements needed for automated incident creation.
- Assist with event correlation, CI alignment, and ticket automation workflows.
- Partner with operations teams to improve incident response processes.
Cloud & Modern Application Monitoring
- Support monitoring solutions for AWS, Azure, SaaS, containerized, and microservices-based applications.
- Deploy and support ActiveGate infrastructure.
- Assist application teams with onboarding modern cloud workloads into enterprise monitoring platforms.
- Support SSO, authentication, synthetic transaction monitoring, and end-user experience monitoring.
Customer Engagement & Consulting
- Consult with application owners and infrastructure teams regarding monitoring requirements.
- Lead onboarding activities for new applications and services.
- Conduct monitoring requirement workshops and technical discovery sessions.
- Provide guidance and best practices across the enterprise.
Governance, Compliance & Documentation
- Maintain monitoring architecture documentation, procedures, and standards.
- Support NERC/CIP and other regulatory monitoring requirements.
- Participate in change management, CAB reviews, and operational governance processes.
- Develop and maintain operational runbooks and support documentation.
Requirements
This role requires strong technical expertise, excellent communication skills, and the ability to translate business requirements into actionable monitoring and observability strategies., * 5+ years supporting enterprise monitoring or observability platforms.
- 5+ years supporting large-scale enterprise infrastructure environments.
- Experience administering Dynatrace or comparable monitoring solutions.
- Experience working in regulated utility, critical infrastructure, energy, or large enterprise environments preferred.
Required Technical Skills
- Dynatrace Administration and Engineering
- Application Performance Monitoring (APM)
- Observability Engineering
- Synthetic Monitoring
- Real User Monitoring (RUM)
- Windows Server Administration
- Linux Administration
- VMware Monitoring
- AWS and Azure Monitoring
- ActiveGate Deployment and Administration
- Monitoring Dashboard Development
- Alert Management and Event Correlation
- ServiceNow Integration
- Root Cause Analysis
- Enterprise Incident Management
- TCP/IP, DNS, SSL, Load Balancers, and Networking Fundamentals
- Database Monitoring (Oracle, SQL Server, PostgreSQL, etc.)
- ITIL Incident, Problem, and Change Management, * Dynatrace Professional or Associate Certification
- Splunk administration experience
- SolarWinds experience
- OpenShift or Kubernetes monitoring experience
- Experience developing Dynatrace Extensions 2.0
- PowerShell, Python, Bash, or automation scripting
- ServiceNow Event Management integration experience
- Experience supporting enterprise authentication solutions (Entra ID, SSO, LDAP, SAML, OAuth)
- Experience supporting NERC/CIP regulated environments
Leadership & Soft Skills
- Strong customer service and stakeholder management skills
- Exceptional troubleshooting and analytical abilities
- Ability to communicate effectively with technical and non-technical audiences
- Ability to prioritize multiple projects and operational demands
- Self-directed and highly accountable
- Strong collaboration skills across infrastructure, security, cloud, and application teams
- Ability to lead technical initiatives and mentor junior engineers
- Strategic mindset with the ability to align monitoring capabilities with business outcomes
What We're Looking For:
Education requirements are listed below:
- Bachelor's degree in computer science, engineering, or related technical field is required.
Work Experience requirement listed below:
- 10 years of relevant work experience is required. An equivalent combination of education and experience may be considered.
Benefits & conditions
3.73.7 out of 5 stars Gahanna, OH 43230 $116,255.00 - $151,132.50 a year - Full-time, Pulled from the full job description
- Tuition reimbursement