Lead Software Engineer - Applied AI ML Lead
Role details
Job location
Tech stack
Job description
As a Lead Software Engineer at JPMorganChase within the Enterprise Technology, Infrastructure Platforms team, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor, you are responsible for conducting critical technology solutions across multiple technical areas within various business functions in support of the firm's business objectives., * Executes creative software solutions, design, development, and technical troubleshooting with ability to think beyond routine or conventional approaches to build solutions or breakdown technical problem.
- Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
- Own infrastructure capacity optimization solutions and build predictive/prescriptive models to identify capacity risk, performance bottlenecks, and right-sizing opportunities.
- Design, develop, and productionize GenAI/agentic AI solutions for automation, decision support, and operational workflows, including LLM/SLM apps such as RAG and summarization/extraction.
- Engineer production-grade backend services in Python/Java (REST APIs, microservices, reusable libraries) and own cloud-native data ingestion/processing pipelines for capacity analytics and AI use cases.
- Build prompt engineering assets, routing strategies, and guardrails, and implement automated plus human-in-the-loop evaluation to improve quality.
- Identifies opportunities to eliminate or automate remediation of recurring issues to improve overall operational stability of software applications and systems
- Apply MLOps best practices across experimentation, versioning, CI/CD, deployment, monitoring, and lifecycle management; implement testing/benchmarking and observability; define success metrics/governance with stakeholders; and mentor engineers to uphold high standards.
- Own and govern the end-to-end AI/ML optimization strategy-from architecture and engineering standards (quality, lifecycle, observability, secure SDLC) through cross-functional execution with SRE/platform/business-to deliver scalable automation, risk reduction, and measurable enterprise outcomes.
- Leads evaluation sessions with external vendors, startups, and internal teams to drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture
Requirements
- Formal training or certification on software engineering concepts and 5+ years applied experience.
- Hands-on practical experience delivering system design, application development, testing, and operational stability
- Advanced in one or more programming language(s)
- Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
- Proficient in all aspects of the Software Development Life Cycle
- Strong hands-on data engineering stack: Apache Spark (batch optimization, partitioning, shuffle tuning, reliability), Apache Airflow (DAG design, backfills, alerting, operational reliability, CI patterns), and Apache Iceberg (schema evolution, partition specs, snapshots, compaction).
- Proven applied AI/ML and GenAI delivery with measurable impact (RAG, extraction, summarization, ranking/classification, copilots, evaluation) and demonstrated ability to lead across teams and influence technical standards and execution.
- Deep distributed systems + production engineering expertise across APIs/microservices, CI/CD, observability, containers/Kubernetes, security, and reliability., * Experience with MCP (Model Context Protocol), Agent Skills, and structured agentic architectures.
- Strong practical usage of AI engineering productivity tooling (for example, GitHub Copilot, Claude Code) in enterprise SDLC environments.
- Familiarity with VSI and Cloud Foundry contexts.
- Advanced Java engineering proficiency in addition to Python.
- Expert-level Python for production systems (packaging, dependency management, performance); strong Java proficiency is a plus.
Benefits & conditions
We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.