> Markdown version of [/jobs/ext/2706877-data-platform-engineer](https://www.wearedevelopers.com/jobs/ext/2706877-data-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Platform Engineer - **Company:** Aperia Solutions, Inc. - **Location:** Dallas, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Encodings, Data Architecture, Information Engineering, Data Infrastructure, DevOps, Python (Programming Language), PostgreSQL, Load Testing, Online Analytical Processing, Online Transaction Processing, SQL Databases, Parquet, File Transfer Protocol (FTP), Azure Data Factory, Apache Spark, Pandas, Microsoft Fabric, Data Lakes, Pyspark, Vertica, Terraform, Sql Tuning - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-data-platform-engineer-aperia-9789752 ## About the Role * 5+ years data engineering experience, with at least 2 years hands-on Spark development (PySpark preferred). * Deep Delta Lake internals knowledge - transaction log format, checkpoints, retention settings (logRetentionDuration, deletedFileRetentionDuration), VACUUM, OPTIMIZE, deletion vectors, statistics collection. * Parquet format expertise - footers, statistics allowlists, row-group sizing, dictionary encoding, column indexes. * Microsoft Fabric hands-on experience - workspaces, capacities, notebooks, lakehouses, semantic models, Direct Lake, SQL analytics endpoint. * SQL performance tuning skills for both OLAP (Fabric) and OLTP (PostgreSQL) workloads. * Strong Python data engineering skills - pandas, PyArrow, delta-rs. * Ability to communicate architecture clearly in writing and to work autonomously on well-scoped features. Nice to Have * PCI-scoped data platform experience. * Financial services or payments domain background. * Power BI semantic model authoring (DAX). * Familiarity with published data-standard governance patterns - logical model vs physical profile, SCD2 dimensions, natural-key idempotency, extension attribute stores. * OneLake shortcuts and cross-workspace lakehouse patterns. * ClickHouse or similar OLAP engine familiarity. * Terraform for Azure data resources. * Prior experience with fluid or schema-flexible ingestion patterns., * Must be willing to submit to a background investigation and drug test as part of the selection process. ## Description Aperia is building a next-generation risk analysis and scoring platform on Microsoft Fabric, and this role owns its medallion data architecture end-to-end. You'll design and build Bronze / Silver / Gold Delta tables, Fabric Spark feature pipelines, Direct Lake semantic models, and physical-layout optimizations. This is our highest-utilization engineering role on the project - central to both Phase 1 platform foundation and Phase 2 risk product delivery, working alongside a Senior Technical Lead, Senior DevOps engineer, and mid-level developer. The platform ingests transaction and merchant data from clients via SFTP and API, tokenizes sensitive PAN data at the ingestion boundary using our reserved-BIN HMAC-derived format-preserving scheme, and lands data in Aperia's canonical Silver on Microsoft Fabric per our internal canonical spec. You'll build the pipeline that populates canonical Silver via MERGE, consume it via OneLake shortcuts, and materialize risk-specific Gold tables that feed both custom-authored DMN rules and pre-scored Mastercard Brighterion feeds. What You'll Do * Design and implement Bronze, canonical Silver, and Gold Delta table schemas with PAN-safe physical layout - statistics allowlist configuration, liquid clustering, deliberate retention settings, V-Order optimization. * Build Fabric Spark notebooks for the feature pipeline (rolling velocity features across 24h / 7d / 30d windows, merchant-level rolling features, hierarchical rollups, cross-merchant features via pan_token joins). * Author per-flavor-version Silver adapters - one function per flavor, unioned by name, tested against golden fixtures. * Build the canonical MERGE DAG (Bronze * canonical Silver, idempotent on _natural_key per Aperia's canonical Fabric internal spec v1.4). * Build the Gold materialization pipeline for both grains (fact_transaction_risk, fact_merchant_risk) with correct score-source stamping and idempotent merge semantics. * Design and implement Direct Lake semantic models per reporting surface; performance-test against representative data volumes. * Configure OneLake namespace layouts, workspace catalogs, and cross-workspace shortcut relationships with the residuals programme. * Establish idempotent Delta write patterns using delta-rs from Python containers (segmented parallel writes with single atomic commit). * Fabric capacity sizing (F-SKU selection, load testing, Reservation vs PAYG decisions). * Partner with the DevOps engineer on Fabric IaC (Terraform capacity provisioning plus Fabric REST workspace bootstrap). * Consult on ClickHouse contingency if Direct Lake latency proves insufficient for client-tier UIs. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know)