Apache Spark Specialist

Nebul
Amsterdam, Netherlands
5 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Languages
Dutch, English
Job source

Tech stack

Airflow Cloud Computing Cloud Computing Security Continuous Integration Information Engineering Distributed Data Store Apache Hadoop Python (Programming Language) Open Source Technology Performance Tuning Prometheus Data Processing
+9 more
Grafana Apache Spark Caching Electronic Medical Records Data Lakes Kubernetes Infrastructure Automation Frameworks Machine Learning Operations Databricks

Job description

We’re looking for an Apache Spark Specialist to take a leading role in designing, deploying, and managing Nebul’s accelerated, fully-managed Spark proposition - enabling our customers to extract value from data at scale with top-tier performance and security. You will be responsible for architecting Spark environments optimized for GPU acceleration and distributed compute, ensuring seamless integration with Nebul’s sovereign cloud stack. This role combines deep technical ownership with solution innovation: from deploying high-performance clusters, to shaping the product experience and supporting early customer implementations.

This is not a maintenance role - it is an opportunity to build and evolve a flagship data capability from the ground up.

What You’ll Do

  • Architect, deploy, and operate scalable Apache Spark environments on Nebul’s sovereign AI cloud.
  • Design and optimize Spark workloads for GPU-accelerated and distributed performance.
  • Define and implement best practices for security, monitoring, governance, and data protection.
  • Partner closely with product, engineering, and customer teams to shape our managed Spark offering.
  • Evaluate and integrate complementary technologies (e.g., Delta Lake, Lakehouse components, tooling).
  • Support early customer pilots and translate feedback into roadmap improvements.
  • Develop automation and CI/CD deployment models to ensure reliability, repeatability, and efficiency.
  • Document architectures, operational procedures, and performance benchmarks.

Requirements

  • 4-7 years of experience working with Apache Spark in production environments.
  • Strong deep-dive knowledge of Spark internals: performance tuning, partition strategies, caching, and shuffle management.
  • Hands-on deployment experience in Kubernetes, cloud infrastructure, or on-prem clusters.
  • Solid understanding of distributed data platforms (e.g., Databricks, EMR, Hadoop, Lakehouse architectures).
  • Strong scripting and automation skills (Python / Scala preferred).
  • Ability to translate client needs into technical architectures and operational models.
  • Familiarity with cloud-security principles and infrastructure-as-code practices.

Bonus Points

  • Experience with GPU acceleration for Spark, RAPIDS, or ML workloads.
  • Exposure to high-security or regulated environments (government, critical industry).
  • Knowledge of observability stacks (Prometheus, Grafana) and orchestration (Airflow, Argo).
  • Contributions to open-source or performance engineering work.
  • Background in data engineering or MLOps., * Based in the Netherlands (commutable to Leiden).
  • Valid EU work permit (no sponsorship currently available).
  • Fluent in English (Dutch not required).

Benefits & conditions

  • Lead the creation of a key product capability powering AI and data innovation.
  • Collaborate with world-class engineers across cloud, security, and high-performance compute.
  • Hybrid environment near The Hague and Amsterdam.
  • Competitive compensation, growth opportunities, and equity participation in a critical market.

About the company

At Nebul, we’re building Europe’s sovereign AI cloud - trusted, secure, and purpose-built for the next generation of intelligent infrastructure.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

2:33 min

Maintaining prompt structures for prefix caching

Douglas Reiser Douglas Reiser · Europe 2026 Virtual

Videos

See all

Related articles

See all