Lead Software Engineer (AI Platform) - Vice President

JPMorgan Chase & Co.
New York, NY, United States
8 days ago
Apply on jpmc.fa.oraclecloud.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Working hours
Regular working hours

Tech stack

Microsoft Access Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Backup Devices Configuration Management Cyber Security Information Systems Computer Engineering Continuous Integration Data Control Programming Tools
+25 more
Distributed Systems Python (Programming Language) Machine Learning Cloud Services Runbook Shell Script Software Engineering AI Infrastructure Autoscaling Large Language Models Model Validation Jupyter Agentic-AI Cloudformation Containerization AI Platforms Kubernetes Information Technology Enterprise Integration Slurm Machine Learning Operations Hardware Infrastructure Cloud Optimization Model Inference Terraform

Job description

As an AI Platform Engineer - Vice President at JPMorgan Chase within the Global Technology Applied Research (GTAR) center, you will engineer, operate, and evolve the shared platforms that enable GTAR researchers to develop, evaluate, and mature AI solutions in secure, scalable, firm-approved environments. You will own the cloud and compute foundation for AI research workloads, build reusable platform services and automation, and partner with researchers and enterprise technology teams to make experimentation reproducible, reliable, well-controlled, and ready for downstream integration., * Own the architecture, technical roadmap, and day-to-day engineering of the team’s AI research platform.

  • Design and operate secure, scalable cloud and CPU/GPU infrastructure, including clusters, workload scheduling, storage, networking, capacity planning, performance, and cost.
  • Build reproducible, self-service research environments and platform automation using standardized container images, GPU software stacks, dependency and artifact management, infrastructure-as-code, CI/CD, and orchestration.
  • Partner with researchers to translate requirements from LLM, agentic AI, model evaluation, retrieval, inference, and accelerator benchmarking projects into dependable platform capabilities, and prepare research prototypes for downstream integration through packaging, service interfaces, and operational-readiness reviews.
  • Assess and onboard new AI infrastructure technologies, cloud services, accelerators, and developer tooling through practical benchmarks, architecture reviews, and controlled pilots.
  • Maintain reliable, well-controlled platforms through observability, incident response, patching and lifecycle management, access and data controls, risk and control evidence, operational documentation, and runbooks.

Requirements

  • Master’s degree in computer science, software engineering, computer engineering, information systems, or a related technical field, plus at least 4 years of relevant industry experience building or operating cloud, platform, or infrastructure systems.
  • Hands-on experience operating production-grade environments, including EC2, networking, storage, monitoring, and cost management.
  • Strong knowledge of Linux systems, containers, networking, storage, and distributed systems, with experience in orchestration or workload-scheduling technologies.
  • Experience with infrastructure-as-code and delivery automation using tools such as Terraform or CloudFormation, configuration management, and CI/CD pipelines.
  • Proficiency in Python and shell scripting, with sound software engineering practices for building maintainable automation, services, and platform integrations.
  • Working knowledge of AI/ML development workflows and demonstrated ability to translate research requirements into platform designs, communicate technical trade-offs, manage delivery dependencies, and drive risk and control items to closure.
  • Experience in one or more of the following domains: cloud and accelerated compute platforms (e.g., AWS, CPU/GPU infrastructure, autoscaling, workload scheduling, resource quotas, capacity planning); research environments and developer platforms (e.g., containers, Jupyter, package and dependency management, artifact repositories, experiment tracking, self-service tooling); AI workload enablement (e.g., LLM inference and evaluation, agentic workflows, retrieval pipelines, model-serving infrastructure, distributed processing); reliability and operations (e.g., observability, service-level objectives, incident response, resilience, backup and recovery, operational readiness), * Experience supporting AI/ML research or development teams and operating platforms for GPU-intensive, LLM, agentic, or other distributed AI workloads.
  • Experience with Kubernetes or managed container platforms, GPU software stacks, cluster schedulers or distributed-compute frameworks such as Slurm or Ray, observability platforms, and cloud cost optimization.
  • Experience delivering infrastructure in a regulated enterprise and working with cybersecurity, architecture, technology risk, and audit stakeholders.
  • Familiarity with finance or financial use cases.

Benefits & conditions

We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.

About the company

JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jpmc.fa.oraclecloud.com
Prepare application

Good distractions

Loading talks and stories from around this role…