> Markdown version of [/jobs/ext/3055980-lead-infrastructure-engineer-storage-enterprise-technology](https://www.wearedevelopers.com/jobs/ext/3055980-lead-infrastructure-engineer-storage-enterprise-technology). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Infrastructure Engineer - Storage (Enterprise Technology) - **Company:** JPMorgan Chase & Co. - **Location:** Columbus, OH, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Computer-Aided Design, Artificial Intelligence, Audit Trail, Build Automation, Bash Shell, Cloud Storage, Configuration Management, Computer Programming, Databases, Continuous Integration, Linux, RAID, Multipath I/O, Python (Programming Language), Key Management, NetApp Applications, Networking Basics, Performance Tuning, Ansible, Prometheus, Data Streaming, Ceph (Software), Datadog, Scripting, Cloud Platform System, Large Language Models, Grafana, Caching, Cloudformation, Kubernetes, Information Technology, Low Latency, Apache Kafka, Data Management, Feature Extraction, Puppet, Isilon, Terraform, Splunk, Servicenow, Golang - **Published:** September 24, 2026 - **Apply:** https://find.jobs/jobs-near-me/apply/ats-redirect/?id=2985299305-2 ## About the Role * Formal training or certification on infrastructure engineering concepts and 5+ years applied experience * Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support infrastructure engineering workflows with strong validation habits and awareness of data sensitivity. * Ability to review and validate AI-assisted recommendations before implementation, escalating when uncertain and ensuring outcomes align to resiliency, security, and auditability expectations. * Strong knowledge of storage fundamentals (RAID/erasure coding, replication, snapshots, tiering/caching, IOPS/latency, multipathing, SAN/NAS, object semantics). * Expertise in product and Infrastructure Support. * Hands-on experience with at least one or more major storage ecosystem (e.g., NetApp, Dell EMC PowerStore/Isilon, Pure, Hitachi, Ceph, IBM, or cloud storage services). * Solid Linux fundamentals, including system performance, networking basics, and kernel/storage-stack concepts. * Strong scripting/programming in one or more of Python, Go, Bash. * Experience with observability stacks (e.g., Prometheus/Grafana, ELK/OpenSearch, Splunk, Datadog, OpenTelemetry). * Proven incident management skills and ability to operate effectively in an on-call rotation. * Practical AI/data skills for operations (anomaly detection/forecasting/correlation/classification; feature extraction and evaluation; integrating AI into production tooling/CI/CD; safe LLM use with guardrails and human-in-the-loop review). * US Citizen or Green card Holder Only. Preferred qualifications, capabilities, and skills * Kubernetes storage (CSI), stateful workloads, and container platform operations. * Infrastructure as Code (Terraform/CloudFormation) and configuration management (Ansible/Chef/Puppet). * Streaming/queue tooling for telemetry and event pipelines (e.g., Kafka). * Experience with ITSM/event management platforms (e.g., ServiceNow). * Backup/DR products and strategy design, including RPO/RTO tradeoffs. * Security controls for data platforms (KMS/HSM, secrets management, key rotation). * Experience building/operating controlled self-service platforms with guardrails to reduce toil at scale. ## Description Assume a vital position as a key member of a high-performing team that delivers infrastructure and performance excellence. Your role will be instrumental in shaping the future at one of the world's largest and most influential companies., * Uses enterprise-authorized AI capabilities within the work environment to accelerate infrastructure analysis and design documentation, validating outputs and handling operational data according to sensitivity and security requirements. * Applies reuse-first, AI-assisted practices within delivery and automation routines to identify recurring issues and validate remediation options, ensuring changes are traceable/auditable and aligned to resiliency and security expectations. * Own and continuously improve SLOs/SLIs, error budgets, on-call readiness, and operational excellence for storage services. * Lead incident response for storage outages/performance degradations; drive RCAs and implement preventative actions. * Create and maintain runbooks, escalation paths, and standardized operational procedures. * Operate and enhance block/file/object storage platforms across on-prem and/or cloud environments. * Perform performance tuning, capacity planning, lifecycle management, and resiliency testing (failover/DR validation). * Partner with infrastructure, network, OS, database, and application teams to meet workload requirements and reliability targets. * Build automation for provisioning, patching, upgrades, replication, backup/restore, and compliance checks. * Implement AI-driven observability/AIOps (telemetry correlation, anomaly/regression detection, LLM-assisted incident/runbook workflows) with accuracy, auditability, and safe rollout. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [The Memory Leak That Ate Our Cluster: A Postmortem](https://www.wearedevelopers.com/videos/2057-the-memory-leak-that-ate-our-cluster-a-postmortem) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [The Fastest-Growing Tech Sectors to Look Out for in 2025](https://www.wearedevelopers.com/magazine/373-the-fastest-growing-tech-sectors-to-look-out-for-in-2025)