Software Engineer II - Databricks

JPMorgan Chase & Co.
Christchurch, UK
11 days ago
Apply on jobs.theguardian.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Airflow Amazon Web Services Amazon S3 Automation of Tests Unit Testing Microsoft Azure Cloud Storage Code Generation Software Quality Continuous Integration
+28 more
Data Validation Information Engineering Data Governance Extract Transform Load (ETL) Database Testing Python (Programming Language) Key Management Metadata Performance Tuning Software Tools Cloud Services Secure Coding Software Engineering SQL Databases Data Streaming Toolchain Azure Service Bus Azure Data Factory Apache Spark Data Lakes Pyspark Git Flow Deployment Automation Apache Kafka Software Coding Code Restructuring Data Pipelines Databricks

Job description

  • Design and implement batch and streaming data pipelines using Databricks (Spark), Delta Lake, and orchestrators (e.g., Workflows, Airflow, ADF).
  • Develop and optimize Spark jobs (PySpark/Scala) and SQL transformations for performance, reliability, and cost efficiency.
  • Build and maintain curated data models (bronze/silver/gold), data quality checks, and automated testing.
  • Implement CI/CD for notebooks and code (Git-based workflows), and automate deployments across environments.
  • Manage and tune Databricks clusters, jobs, and configurations; monitor production workloads and resolve incidents.
  • Integrate multiple data sources (cloud storage, relational DBs, APIs, event streams) and implement robust ingestion patterns.
  • Apply data governance and security best practices (access controls, secrets management, lineage/metadata, auditing).
  • Create clear documentation for pipelines, data contracts, and operational runbooks.
  • Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity (e.g., code generation/refactoring, unit test creation, documentation), while validating outputs through peer review, automated testing, and secure coding standards.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation

Requirements

  • Experience with building data pipelines on Databricks and/or Apache Spark in production.
  • Strong coding skills in Python (PySpark) and SQL (Scala a plus).
  • Hands-on experience with Delta Lake (MERGE/UPSERT patterns, schema evolution, partitioning, Z-ORDER, OPTIMIZE/VACUUM).
  • Experience with orchestration and scheduling (Databricks Workflows, Airflow, Azure Data Factory, etc.).
  • Familiarity with cloud data platforms ( AWS/Azure/GCP ) and storage (S3/ADLS/GCS).
  • Solid understanding of data engineering fundamentals: data modeling, ETL/ELT patterns, reliability, observability, and performance tuning.
  • Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, testing, troubleshooting, or documentation) with demonstrated ability to critically evaluate and validate AI-generated outputs.
  • Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations, * Experience with streaming (Structured Streaming, Kafka/Event Hubs/Kinesis).
  • Experience implementing data quality frameworks (Great Expectations, DQ) and data testing in CI.
  • Exposure to Unity Catalog (or similar) for governance and fine-grained permissions.

About the company

As a Software Engineer II at JPMorgan Chase within our Corporate Investment Bank, Payments Technology team, you’ll design, build, and operate scalable data pipelines and analytics workloads on the Databricks Lakehouse platform. We’re looking for a Databricks Engineer to design, build, and operate scalable data pipelines and analytics workloads on the Databricks Lakehouse platform. You’ll partner with data analysts, data scientists, and application teams to deliver trusted datasets, performant ETL/ELT pipelines, and well-governed data products., J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world’s most prominent corporations, governments, wealthy individuals and institutional investors. Our first-class business in a first-class way approach to serving clients drives everything we do. We strive to build trusted, long-term partnerships to help our clients achieve their business objectives., J.P. Morgan’s Commercial & Investment Bank is a global leader across banking, markets, securities services and payments. Corporations, governments and institutions throughout the world entrust us with their business in more than 100 countries. The Commercial & Investment Bank provides strategic advice, raises capital, manages risk and extends liquidity in markets around the world.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.theguardian.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:47 min

Comparing Egeria to alternative open metadata solutions

Ferd Scheepers · World Congress 2022

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:08 min

Creating standard APIs via the Egeria open metadata project

Ferd Scheepers · World Congress 2022

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all