Lead Software Engineer - Data Engg. - Databricks / Snowflake

JPMorgan Chase & Co.
Plano, TX, United States
10 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Airflow Amazon Web Services Amazon S3 Automation of Tests Cloud Database Software Quality Code Review Continuous Integration Directed Acyclic Graph (Directed Graphs) Data Governance
+42 more
Extract Transform Load (ETL) Data Systems Document-Oriented Databases Identity and Access Management JSON Python (Programming Language) Automation of Marketing Performance Tuning Systems Development Life Cycle Query Optimization Software Tools Standard Sql Runbook Secure Coding Service Design Software Engineering Software Systems SONAR (Symantec) Spinnaker SQL Databases Systems Integration Strategies of Testing Toolchain Management of Software Versions Alwayson Parquet Data Processing Snowflake Apache Spark Data Strategy Cloudformation Build Management Data Lakes Pyspark Avro Data Management Amazon Simple Queue Service (SQS) Terraform Code Restructuring Data Pipelines Jenkins Databricks

Job description

As a Lead Software Engineer at JPMorgan Chase within the (IAM) Identity and Access Management Data team, you will play a crucial role in designing, developing, and maintaining scalable data processing solutions using Databricks, Python, and AWS. You will collaborate with cross-functional teams to deliver high-quality data solutions that support our business objectives., * Execute creative, data-driven software solutions end-to-end (design, development, troubleshooting), thinking beyond routine approaches to solve complex technical problems.

  • Design and build a control plane for enterprise data pipelines, standardizing pipeline definition, scheduling, deployment, governance, and run-time management (Databricks today; extensible for future engines).
  • Develop self-service APIs/SDKs, templates, and configuration-driven onboarding with consistent guardrails (standards, validation, environment promotion, approvals) and centralized pipeline metadata (ownership, SLAs/SLOs, dependencies, schema/parameter/version tracking).
  • Design, develop, and maintain scalable data pipelines and processing workflows using Python, PySpark, SQL, Databricks on AWS; develop fact/dimension models for analytics and reporting.
  • Ensure data quality, security, lineage, and operational transparency via standardized observability (logs/metrics/traces), dashboards, alerting, runbooks, and automated remediation patterns (retries/backfills, common-failure automation).
  • Lead and participate in the full SDLC (requirements, design, build, test, deploy, maintain), acting as SRE/production support for pipeline and platform services to improve stability and reliability.
  • Collaborate with stakeholders to shape data management strategy and translate requirements into scalable, compliant solutions; document data flows, logic, and transformation rules for knowledge sharing.
  • Mentor engineers and lead communities of practice to drive adoption of modern engineering practices and tools, fostering an inclusive, high-performing culture; utilize firm-approved AI-assisted development tools to accelerate delivery and testing
  • Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.

Requirements

  • Formal training or certification on software engineering concepts and 5+ years applied experience
  • Proven experience in data management and ETL/ELT for large-scale processing, including strong SQL, Python, and PySpark with performance tuning and query optimization.
  • Hands-on experience with Databricks/Spark and cloud data lake patterns, integrating compute/workflows with AWS services (e.g., S3, ECS, SNS/SQS, Lambda).
  • Proven experience building platform services/control planes (or similar orchestration/automation platforms), including API/service design, configuration-driven systems, and versioning/backward compatibility.
  • Strong understanding of data quality, security-by-design, and lineage/auditability, including IAM/least privilege and secrets management principles.
  • Strong production engineering mindset: observability (logs/metrics/traces), monitoring/alerting, incident response, and operational excellence for always-on services.
  • Proficiency in CI/CD and release engineering (quality gates, automated testing, safe deployments/rollbacks) using firm-standard tooling (e.g., Jenkins/Jules, Spinnaker, Sonar).
  • Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices

Preferred qualifications, capabilities, and skills

  • Experience with orchestration/execution frameworks (Databricks Workflows/Jobs, Airflow, Step Functions) and operational patterns such as dependency graphs (DAGs), replays, and backfills.
  • Experience with data governance integrations (e.g., Unity Catalog concepts such as cataloging, permissions, and lineage hooks), where applicable.
  • Infrastructure-as-Code experience (Terraform/CloudFormation) and developer-platform “golden path” enablement (internal CLIs, templates, paved roads, onboarding automation).
  • Experience with FinOps/cost controls for Spark/Databricks workloads (telemetry, quotas, chargeback/showback) and data formats (Parquet, JSON, CSV, Avro, Delta Lake), Knowledge of regulatory reporting and financial data aggregation techniques; Databricks and/or AWS certifications.

Benefits & conditions

We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.

About the company

JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jpmc.fa.oraclecloud.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier ¡ WWC 2023

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff ¡ WWC Europe 2026

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell ¡ LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou ¡ Coffee With Developers

57 sec

Extracting API schemas automatically during continuous integration builds

Axel Barbier ¡ WWC 2023

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters ¡ WWC 2025

Videos

See all

Related articles

See all