> Markdown version of [/jobs/ext/2174084-data-engineer](https://www.wearedevelopers.com/jobs/ext/2174084-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Launchmetrics - **Location:** Girona, Spain - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Artificial Intelligence, Amazon Web Services, Amazon S3, Software Suite, JIRA, Unit Testing, Software as a Service, Software Quality, Databases, Continuous Integration, Data Architecture, Data Cleansing, Data Infrastructure, Distributed Computing Environment, Github, Data Intelligence, Python (Programming Language), MySQL, Object-Oriented Software Development, Data Streaming, TypeScript, Apache Spark, Git, Build Management, Pytest, Data Lakes, Pyspark, Information Technology, Functional Programming, Amazon Simple Queue Service (SQS), Software Version Control, Data Pipelines, Serverless Computing, Databricks - **Published:** August 22, 2026 - **Apply:** https://www.jobleads.com/es/job/e24f356f5f40c23df8fd45d2ce162f8df ## About the Role * Engineer Degree or Master Degree in Computer Science and 3+ years of relevant work experience in full-stack development in a SaaS environment * Strong Python and PySpark experience; comfort with distributed data processing at scale * Experience with a Lakehouse architecture (Databricks, Delta Lake, or comparable - e.g. Spark on EMR/Glue) * Familiarity with medallion architecture patterns (Bronze/Silver/Gold) or similar layered data design * Ability to reason about schema evolution, partitioning/clustering strategy, and pipeline reliability (retries, idempotency) * Ability to traverse logical sequences of either procedural or object-oriented code, abstracted or static - and understand it entirely * Bright, energetic, highly motivated self-starter with experience in a fast-paced, results-oriented organization * Ability to adapt, estimate workload, break down a task into logical steps, solve problems, self-improve and suggest new ways of improvement * Last, but definitely not least: you speak, read, and write English fluently ## Description We are a team of passionate engineers distributed across the world, always eager to learn new things. We are building state-of-the-art modern applications and scale it on the cloud. We innovate to solve customers problems, focusing on high-performance implementations without neglecting the user experience. We recently moved our Product and Tech teams to a pod model: A pod is a small group, 1 to 4 people, that owns a single mission end to end: from scoping and building through cleaning, enablement and watching adoption. Pods form around a mission and dissolve when it's done. Everyone is a pod owner, from the most junior person to the most senior, because ownership is the point. Your domain is your product-area home, the part of our software suite you know best. Pods change with every mission; your craft and your domain don't., The Data Platform and Data Enrichment teams are in charge of making our unified database accessible, usable, enrichable and reliable for all our teams internally. The team owns the Lakehouse that powers Launchmetrics' data products - a Databricks-based platform built on Delta Lake, S3, and following a Bronze/Silver/Gold medallion architecture. This layer stores, enriches, and serves the media intelligence data behind Discover (our client-facing platform), Genie (internal tooling), and Delta Sharing. You'll work on a batch-first system that simulates real-time behavior using primitives like Change Data Feed, S3 event triggers, serverless compute, and liquid clustering. This role plays a key part in our Tech & Product strategy, directly supporting company-wide objectives around customer retention, data trust, and the development of new AI-based insights., * Design and build data pipelines (batch and near-real-time) using PySpark and Databricks, across the Bronze/Silver/Gold medallion layers * Architect efficient Delta Lake table schemas - partitioning/liquid clustering strategy, schema evolution handling, and enrichment workflows * Work closely with product, QA, and other data engineers to translate enrichment and search requirements into reliable pipelines * Own code quality: structured PySpark jobs, unit tests (pytest), and adherence to team conventions * Continuously improve pipeline reliability and cost efficiency (OPTIMIZE scheduling, retry/backoff logic, concurrency handling) * Participate in cross-pod initiatives across the data platform Technical Stack * Languages: Python, PySpark, Typescript * Data Platform: Databricks, Delta Lake, Delta Sharing * Storage: AWS S3, MySQL * Cloud: AWS (S3, Kinesis, Lambda, SQS, ECS, Step Functions, …) * Tools: Jira, Databricks Asset Bundles, Serverless, Claude * Code Versioning: Git, GitHub * CI/CD: GitHub Actions * Testing: pytest, Jest, We're a company that puts people first, with a relaxed but genuinely dynamic atmosphere and a team of motivated, curious people who like their work. Autonomy is real here: pods make local decisions and own their outcomes, which means you'll see your fingerprints on the product quickly. You'll get a learning and development allowance, a benefits package tailored to your location, flexible working arrangements with support to set up your home office, and the room to grow along whichever path fits you. We're remote-friendly, with hubs across our twelve markets. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Branch your database like your code: How schema changes and pull requests go hand in hand](https://www.wearedevelopers.com/videos/350-branch-your-database-like-your-code-how-schema-changes-and-pull-requests-go-hand-in-hand) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)