Infrastructure Engineer Lead - Cloud AI

American Electric Power
Columbus, OH, United States
2 days ago
Apply on www.columbusjobsite.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$20,800.0 - $41,600.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Cloud Computing Security Cloud Engineering Cyber Security Data Transmissions Data Security Domain Name System (DNS) Monitoring of Systems Identity and Access Management
+38 more
Subnetting Virtual Private Networks (VPN) Python (Programming Language) Machine Learning Networking Basics Network Planning and Design Routing Node.Js OpenShift Windows PowerShell Role-Based Access Control Cloud Services Ansible Runbook AI Infrastructure Enterprise Data Management SSL Certificate Management Data Logging Transport Layer Security Google Cloud Load Balancing Cloud Platform System Autoscaling Istio Retrieval-Augmented Generation Snowflake Multi-Cloud Firewalls (Computer Science) Amazon Virtual Private Cloud (VPC) Build Management AI Platforms Kubernetes Information Technology Hardware Infrastructure Api Gateway Terraform Oracle Cloud Infrastructure User Administration

Job description

This hands-on role builds and operates secure, scalable infrastructure for AI and machine learning workloads across AWS and approved on-premises environments. Responsibilities include cloud foundations, networking, containers, compute, identity, cost management, and observability.

The engineer partners with AI, architecture, cybersecurity, network, and application teams to deliver secure, cost-effective, production-ready infrastructure while supporting core cloud engineering services., AI Cloud Foundations

  • Design, build and maintain the AWS account structure, landing zones, and reference patterns that host AI and machine learning workloads, leveraging AWS Control Tower, Landing Zone Accelerator, and Terraform.
  • Engineer the compute, storage, and networking foundations required for AI workloads, including GPU and accelerated instance families, high-throughput storage, and data access paths to enterprise data platforms.
  • Enable and operate AI platform services such as Amazon Bedrock and SageMaker, including private connectivity, model access provisioning, logging, and quota management.
  • Build reusable infrastructure-as-code modules, templates and pipelines so that AI teams can deploy quickly within approved guardrails.
  • Design, build and maintain on-premises AI infrastructure for edge and specialized use cases.

Support AI in disconnected or intermittently connected environments, including local compute, model distribution, patching, monitoring, backup, and recovery.

Create deployment standards, automation, runbooks, and support processes for on-premises and edge AI.

Integrate on-premises AI with enterprise identity, security, networking, monitoring, and governance where feasible.

Controls, Security & Governance

  • Partner with Cybersecurity, Enterprise Architecture and Compliance to align AI infrastructure with AEP security standards, regulatory obligations, and responsible AI guardrails.
  • Review and remediate configuration drift, vulnerabilities and audit findings across the AI cloud estate; support evidence requests and control attestations.
  • Contribute to onboarding and intake processes for new AI use cases, ensuring workloads are provisioned into the right accounts with the right controls from day one.

FinOps & Cost Management

  • Establish and operate FinOps practices for AI workloads, including tagging standards, showback/chargeback, budgets, anomaly detection, and forecasting.
  • Analyze and optimize spend on GPU compute, inference and token consumption, storage, and data transfer; recommend commitment strategies such as Savings Plans and Reserved Instances.
  • Provide cost transparency and consumption reporting to business stakeholders and technology leadership, and identify optimization opportunities before they become budget issues.

Monitoring & Observability

  • Partner with monitoring team to create logging, alerting and dashboards for AI infrastructure and workloads.
  • Define service-level objectives and operational thresholds for AI platforms, including model endpoint availability, latency, throughput, and error rates.
  • Support incident response, root cause analysis and problem management for AI-related infrastructure events; drive preventive actions and automation to reduce recurrence.

Core Cloud Engineering & Platform Services

  • Perform standard cloud engineering functions across the AEP AWS environment including account provisioning, environment builds, platform upgrades, patching, automation and lifecycle management.
  • Design, deploy and operate Kubernetes environments (Amazon EKS, ROSA/OpenShift) including cluster architecture, autoscaling, node group and GPU scheduling, ingress, service mesh, RBAC, and cluster security hardening.
  • Engineer and support platform services including load balancing, traffic management, API gateways, and integration patterns across cloud, on-premises, edge, and disconnected environments.
  • Apply networking fundamentals, VPC design, subnetting, routing, Transit Gateway, Direct Connect, VPN, firewalls, TLS and certificate management to deliver secure, performant connectivity for AI and general workloads.
  • Prepare cost estimates, justifications, alternative solutions and technical recommendations; produce technical documentation, runbooks and standards.
  • Collaborate with Project Managers, Architects, Solution Engineers, Business Analysts and vendor partners to deliver consistent, reliable solutions that leverage AEP’s technology standards, architectures and best practices.
  • Adhere to and advocate for change, incident and problem management processes; participate in on-call rotation and after-hours support as required.
  • Provide training, mentoring and technical work direction to other engineers on the team., The Physical Demand Level for this job is: S - Sedentary Work: Exerting up to 10 pounds of force occasionally (Occasionally: activity or condition exists up to 1/3 of the time) and/or a negligible amount of force frequently. (Frequently: activity or condition exists from 1/3 to 2/3 of the time) to lift, carry, push, pull or otherwise move objects, including the human body. Sedentary work involves sitting most of the time but may involve walking or standing for brief periods of time. Jobs are sedentary if walking and standing are required only occasionally, and all other sedentary criteria are met.

Requirements

  • Demonstrated hands-on engineering experience in AWS, including IAM, VPC networking, compute, storage, encryption/KMS, and account/organization structure.
  • Strong Kubernetes expertise including cluster design, operations, troubleshooting and security in EKS, ROSA or OpenShift.
  • Solid networking fundamentals, including routing, DNS, load balancing, traffic management, firewalls, hybrid connectivity, and network design for disconnected or intermittently connected environments.
  • Experience with API gateways, integration patterns, and exposing and securing services across environments.
  • Proficiency with infrastructure-as-code and automation (Terraform required; Ansible, Python or PowerShell preferred) and CI/CD pipelines.
  • Working knowledge of cloud security principles, identity and access management, and compliance requirements.
  • Experience implementing monitoring and cost management for cloud environments, with the ability to establish local monitoring, logging, patching, backup, and recovery processes for on-premises and disconnected environments.
  • Strong analytical, troubleshooting and problem-solving skills, with the ability to work independently on complex assignments.
  • Effective written and verbal communication skills, including the ability to present technical recommendations clearly to management and non-technical stakeholders., * Architecture experience including designing end-to-end cloud solutions and influencing platform direction is highly desirable.
  • Direct experience building or operating AI/ML or GenAI workloads across cloud or on-premises infrastructure (Amazon Bedrock, SageMaker, vector databases, retrieval-augmented generation patterns, GPU-based training or inference, edge AI, or disconnected environments).
  • Experience with FinOps tooling and practices for high-variability workloads.
  • Experience in a regulated industry (utility, energy, financial services) or with FedRAMP/GovCloud environments.
  • Familiarity with multi-cloud environments (Azure, OCI, GCP) and enterprise data platforms such as Snowflake.
  • Certifications: AWS Solutions Architect - Associate or Professional; AWS Certified Machine Learning; Certified Kubernetes Administrator (CKA).

Other Requirements

  • Adhere to policies, procedures, standards, codes and regulations relevant to assignments.
  • Demonstrate in-depth knowledge of AEP infrastructure, environment and components to enable efficient, comprehensive responses to projects and problems.
  • Participate in on-call rotation, after-hours maintenance windows, and storm/emergency response support as required., * Bachelor’s degree in computer science, engineering, or related technical field is required., * 12 years of relevant work experience required. An equivalent combination of education and related experience may be considered.

Benefits & conditions

  • Base salary
  • Annual bonus
  • Long-term incentive
  • 401(k) match
  • AEP Pension
  • Comprehensive benefits package designed to support and enhance the overall well-being of employees.

At AEP, we’re more than just an energy company - we’re a team of dedicated professionals committed to delivering safe, reliable, and innovative energy solutions. Guided by our mission to put the customer first, we strive to exceed expectations by listening, responding, and continuously improving the way we serve our communities. If you’re passionate about making a meaningful impact and being part of a forward-thinking organization, this is the company for you!

Compensation Data

Compensation Grade:

SP20-010

Compensation Range:

$136,539.00 - $177,503.00

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.columbusjobsite.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · World Congress 2023

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

7:15 min

Installing Istio programmatically with bash scripts

Thomas Südbröcker · LIVE

Videos

See all

Related articles

See all