Data Engineer

Thinkproject
UTRECHT, Netherlands
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Application Integration Architecture User Authentication Microsoft Azure Cloud Computing Cloud Storage Continuous Integration Information Engineering Data Governance Data Integration Data Integrity
+29 more
Extract Transform Load (ETL) Data Synchronization Relational Databases DevOps Identity and Access Management Python (Programming Language) PostgreSQL NoSQL Operational Databases Standard Sql Search Technologies SQL Databases Data Streaming Google Cloud Delivery Pipeline Backend Git Event Driven Architecture Containerization Core Data Low Latency Deployment Automation Google Cloud Functions Production Code Cloud Integration Api Design Terraform Data Pipelines Docker

Job description

We are looking for a Data Integration Engineer to own the data pipelines and integration layer that powers our AI Search Platform. You will design, build, and maintain the workflows that move data reliably from source systems into GCP services - including Vertex AI - and expose that capability through secure, well-designed APIs consumed by internal and external systems. This is a hands-on engineering role. You will write production code, own the reliability of what you ship, and work closely with DevOps, Network, and Platform Engineering teams.

Tech you will work with daily: Python | SQL | GCP (Cloud Run, Pub/Sub, Cloud Storage, Cloud Spanner, Vertex AI) | Terraform | PostgreSQL | Docker | Git | CI/CD, Data Integration & Pipeline Development

  • Design, implement, and optimise scalable data integration workflows supporting inference and data synchronisation across GCP services (Cloud Run, Pub/Sub, Cloud Storage, Cloud Spanner, Vertex AI)
  • Build and maintain event-driven pipelines and ETL/ELT workflows that deliver clean, reliable data to the AI Search Platform
  • Automate deployment, testing, and pipeline orchestration using Cloud Run, Pub/Sub triggers, and Terraform

API Development for AI Integration

  • Build and maintain APIs that expose data integration and AI inference capabilities to internal and external systems
  • Ensure secure, reliable, and performant access to the AI Search Platform - correct authentication, rate limiting, and error handling by default

Permissions & Compliance Layer

  • Integrate and enforce API and IAM policies for compliant access control across all AI Search Platform components
  • Own and evolve the permissions API layer to meet growing scalability and security requirements

Data Quality & Reliability

  • Ensure data integrity through monitoring, validation, and alerting across all integrated systems and services
  • Continuously monitor workflows for latency, reliability, and cost efficiency - implement improvements without waiting to be asked

Documentation & Standards

  • Maintain architecture documentation and runbooks
  • Contribute to best practices for data integration, reproducibility, scalability, and security, * Month 3: Core data pipelines understood and contributing to production; first reliability or latency improvement shipped
  • Month 6: Owning at least one integration area end-to-end; permissions API layer extended with evidence-backed design decisions
  • Month 12: Data integration reliability measurably improved; pipeline documentation and monitoring coverage complete; identified and closed at least one material cost or latency inefficiency

You’re probably NOT a fit if

  • Your data engineering experience is primarily batch ETL without event-driven or streaming context
  • You are not comfortable working across cloud-native GCP services in production
  • You treat IAM and access control as someone else’s concern
  • You need fully defined requirements before designing an integration

Requirements

You have 3+ years of hands-on experience in data engineering, cloud integrations, or backend development and have shipped production data pipelines on GCP. Specifically:

  • 3+ years of professional experience in data engineering, cloud integrations, or backend development

  • Strong proficiency in Python and SQL
  • Production experience with Google Cloud Platform services: Cloud Run, Pub/Sub, Cloud Storage,Cloud Spanner, and Vertex AI
  • Experience with event-driven architectures and cloud-based ETL/ELT workflows
  • Experience with relational databases (PostgreSQL, Cloud Spanner) and exposure to NoSQL
  • Proficient with Git and familiar with CI/CD workflows and containerisation (Docker)
  • Experience with Terraform or equivalent Infrastructure-as-Code tooling
  • Working knowledge of IAM, data governance, and access management principles

Nice-to-Have (Bonus Skills)

  • Azure DevOps or cross-cloud integration experience
  • API design experience (REST or gRPC)
  • Experience with AI/ML inference pipelines or Vertex AI in production
  • Prior work in construction, engineering, or real estate software domains

Soft Skills

  • Engineering rigour - you care about pipeline reliability and data correctness, not just throughput
  • Ownership mindset - you monitor what you build and fix it when it breaks
  • Clear written communication: able to document integration contracts and architecture decisions for non-specialist readers
  • Collaborative: comfortable working across DevOps, Network, and Platform teams without friction
  • Comfortable with ambiguity - you can scope and deliver integration work from incomplete upstream specs

Benefits & conditions

  • Competitive fixed salary - shared on request
  • Variable performance bonus: 5% of fixed
  • Continuous learning & certification budget Learning programmes | Career growth | International exposure

About the company

Thinkproject builds construction intelligence software for the firms that deliver Europe’s largest infrastructure, energy, and real estate projects. Our platform manages the information flow across the full lifecycle of a built asset - from design and construction through operation and eventual decommissioning.

By combining deep domain knowledge of the building, infrastructure, and energy industries with a modern SaaS architecture, Thinkproject empowers customers to digitise, connect, and control their construction workflows across their entire asset lifecycle., At Thinkproject, we run feedback cycles that are honest and frequent. We believe the best engineering cultures are built on trust, transparency, and shared ownership - not hierarchy. Our Pune team is a core part of a global organisation, collaborating across time zones with colleagues in Germany, France, the UK, UAE, Spain, New Zealand, and Australia.Lunch ‘n’ Learn Sessions I Women’s Network I LGBTQIA+ Network I Coffee Chat Roulette I Free English Lessons I Thinkproject Academy I Social Events I Volunteering Activities I Open Forum with Leadership Team (Tp Café) I Hybrid working I Unlimited learning

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on careers.thinkproject.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · WWC 2024

Videos

See all

Related articles

See all