Lead Infrastructure Engineer - Storage (Enterprise Technology)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+29 more
Job description
Assume a vital position as a key member of a high-performing team that delivers infrastructure and performance excellence. Your role will be instrumental in shaping the future at one of the world’s largest and most influential companies., * Uses enterprise-authorized AI capabilities within the work environment to accelerate infrastructure analysis and design documentation, validating outputs and handling operational data according to sensitivity and security requirements.
- Applies reuse-first, AI-assisted practices within delivery and automation routines to identify recurring issues and validate remediation options, ensuring changes are traceable/auditable and aligned to resiliency and security expectations.
- Own and continuously improve SLOs/SLIs, error budgets, on-call readiness, and operational excellence for storage services.
- Lead incident response for storage outages/performance degradations; drive RCAs and implement preventative actions.
- Create and maintain runbooks, escalation paths, and standardized operational procedures.
- Operate and enhance block/file/object storage platforms across on-prem and/or cloud environments.
- Perform performance tuning, capacity planning, lifecycle management, and resiliency testing (failover/DR validation).
- Partner with infrastructure, network, OS, database, and application teams to meet workload requirements and reliability targets.
- Build automation for provisioning, patching, upgrades, replication, backup/restore, and compliance checks.
- Implement AI-driven observability/AIOps (telemetry correlation, anomaly/regression detection, LLM-assisted incident/runbook workflows) with accuracy, auditability, and safe rollout.
Requirements
- Formal training or certification on infrastructure engineering concepts and 5+ years applied experience
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support infrastructure engineering workflows with strong validation habits and awareness of data sensitivity.
- Ability to review and validate AI-assisted recommendations before implementation, escalating when uncertain and ensuring outcomes align to resiliency, security, and auditability expectations.
- Strong knowledge of storage fundamentals (RAID/erasure coding, replication, snapshots, tiering/caching, IOPS/latency, multipathing, SAN/NAS, object semantics).
- Expertise in product and Infrastructure Support.
- Hands-on experience with at least one or more major storage ecosystem (e.g., NetApp, Dell EMC PowerStore/Isilon, Pure, Hitachi, Ceph, IBM, or cloud storage services).
- Solid Linux fundamentals, including system performance, networking basics, and kernel/storage-stack concepts.
- Strong scripting/programming in one or more of Python, Go, Bash.
- Experience with observability stacks (e.g., Prometheus/Grafana, ELK/OpenSearch, Splunk, Datadog, OpenTelemetry).
- Proven incident management skills and ability to operate effectively in an on-call rotation.
- Practical AI/data skills for operations (anomaly detection/forecasting/correlation/classification; feature extraction and evaluation; integrating AI into production tooling/CI/CD; safe LLM use with guardrails and human-in-the-loop review).
- US Citizen or Green card Holder Only.
Preferred qualifications, capabilities, and skills
- Kubernetes storage (CSI), stateful workloads, and container platform operations.
- Infrastructure as Code (Terraform/CloudFormation) and configuration management (Ansible/Chef/Puppet).
- Streaming/queue tooling for telemetry and event pipelines (e.g., Kafka).
- Experience with ITSM/event management platforms (e.g., ServiceNow).
- Backup/DR products and strategy design, including RPO/RTO tradeoffs.
- Security controls for data platforms (KMS/HSM, secrets management, key rotation).
- Experience building/operating controlled self-service platforms with guardrails to reduce toil at scale.
Benefits & conditions
We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.
About the company
hackajob is collaborating with J.P. Morgan to connect them with exceptional professionals for this role., JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
7 Cloud Computing Trends Coming in 2025 for Developers
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
How to Become an AI Engineer
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production