Data Engineer

LA International
Greater London, UK
4 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Computing Platforms Big Data Code Reuse Information Engineering Extract Transform Load (ETL) Data Transformation Data Systems Distributed Systems Python (Programming Language) NumPy Performance Tuning
+10 more
DataOps SQL Databases Unstructured Data Data Processing Scripting Apache Spark Parallel Computation Pandas Pyspark Data Pipelines

Job description

  • Build and maintain Palantir Foundry ontologies, pipelines, and applications.
  • Design and develop scalable ETL/data engineering solutions using Python and PySpark.
  • Integrate and transform data from multiple enterprise data sources.
  • Optimize Spark-based data processing and workflow performance.
  • Collaborate with business stakeholders to deliver analytics and AI-ready data products.

Requirements

  • Palantir Platform Expertise
  • Working with Foundry Ontology, pipelines, and modular applications.
  • Implementing End to End data Solutions in Palantir sourcing data from
  • Programming & Scripting (Python and PySpark)
  • Data Processing & Automation: Ability to clean, transform, and process large datasets efficiently.
  • Integration with Palantir Foundry: Build custom Python functions and libraries to extend Foundry’s capabilities.
  • Pipeline Development: Write modular, reusable code for ETL workflows and data transformations.
  • Performance Optimization: Use Python libraries like Pandas, NumPy, and PySpark for scalable data operations.
  • Familiarity with SQL for querying and data modelling. Good to have
  • Data Engineering Fundamentals
  • Building ETL pipelines for ingestion and transformation.
  • Designing data models and optimizing workflows.
  • Handling structured and unstructured data.
  • Big Data & Distributed Systems
  • Experience with Spark (PySpark) for scalable data processing.
  • Understanding of parallel computing and performance tuning.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

3:33 min

Refactoring data science workflows using Rapids QDF and Pandas

Paul Graham Paul Graham · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all