Senior Data Generalist

AMO SERVICES
Paris, France
9 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Information Engineering Data Mining Data Systems Distributed Data Store Apache Hive Python (Programming Language) Machine Learning Query Optimization SQL Databases Parquet Data Processing Delivery Pipeline
+6 more
Data Lakes Pyspark Apache Flink Spark Streaming Stream Processing Data Pipelines

Job description

Build and maintain batch and streaming data pipelines.

  • Move production data into archival/data lake storage safely and observably.
  • Improve quality, reliability, and observability of derived datasets.
  • Design validations for freshness, completeness, schema correctness, and data quality.
  • Build reusable datasets that reduce one-off data extraction and aggregation work.
  • Optimize queries and data processing jobs for correctness, latency, and cost.
  • Prototype in the right tools, then integrate scoped production experiments with amo’s systems where needed.

Requirements

Senior-level experience building and operating data pipelines.

  • Strong Python and SQL.
  • Production experience with Python or Scala.
  • Experience with PySpark, Spark SQL, Spark Streaming, Parquet, and Iceberg.
  • Strong understanding of batch and stream processing.
  • Query optimization experience.
  • Familiarity with data engine internals and distributed data systems.
  • Experience with data lakes, table/storage formats, and derived datasets., * Experience with Flink or similar stream processing systems.
  • Experience with ML training or inference pipelines.
  • Familiarity with ranking, recommendations, embeddings, or entity resolution.
  • Experience with privacy-sensitive data systems, GDPR workflows, or data deletion/export pipelines.
  • Experience reading or making scoped changes in Rust-backed systems.

Benefits & conditions

To ensure that everyone is set up for success within our way of working, we work together onsite 5 days a week.

We wanted to make sure coming to the office was as comfortable as possible for you:

  • We chose a location in central Paris, near Opera (Metro lines 3,8,9 and RER A).
  • We have a beautiful Parisian-style office with high ceilings, balconies, and huge windows. So a lot of natural light!

Because life outside of work should also be stress-free, we cover:

  • Health care (100% coverage).
  • Maternity Leave, Paternity Leave, Second Parent Leave (salary maintained at 100%).
  • Vacation days: 8-9 weeks (total) per year - (European summer? ).
  • 5 weeks of paid time off (by the state);
  • 1-2 additional weeks (RTT, based on our company contract);
  • ~11 bank holidays. We shut down entirely twice a year, two weeks in summer and one week in winter to allow everyone to truly recharge and to avoid prolonged slowdowns (especially in the summer). These pauses are part of your total vacation time, giving everyone a real chance to unplug and recharge. No Slack, no email, no FOMO.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on fr.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

3:24 min

The governance failures of centralized data lakes

Mario Meir-Huber · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

6:24 min

Distributed data lakes and containerized computing clusters

Ulrich Wurstbauer +1 · LIVE

Videos

See all

Related articles

See all