Data Engineer

Transflo Terminal Services, Inc.
United States
6 days ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Query Performance Application Programming Interfaces (APIs) Artificial Intelligence Airflow Amazon Web Services Amazon S3 Data Analysis Apache HTTP Server Audit Trail Automation of Tests Business Intelligence Development Software as a Service
+73 more
Databases Data as a Services Data Architecture Data Validation Data Deduplication Information Engineering Data Governance Data Infrastructure Data Integrity Extract Transform Load (ETL) Data Transformation Data Mining Data Normalization Data Sharing Data Warehousing IBM DB2 Relational Databases Database Design DevOps Dimensional Modeling Amazon DynamoDB Microsoft Exchange Server Fault Tolerance Mobile Application Software JSON Python (Programming Language) PostgreSQL Metadata Meta-Data Management Metadata Repositories MySQL NoSQL Performance Tuning Query Optimization Real-Time Operating Systems Reliability Engineering Power BI Application Data Software Engineering SQL Stored Procedures SQL Databases Data Streaming Tableau (Software) Web Analytics Parquet Transaction Processing (Computing) Data Processing Scripting Business Intelligence Development Studio Data Classification Concurrency Caching Data Layers Data Lakes Kubernetes Infrastructure Automation Frameworks Data Lineage Avro AWS Glue Data Analytics Star Schema AWS Data Analytics Real Time Data Apache Kafka Data Management Video Streaming Data Delivery Restful APIs Terraform Stream Processing Looker Analytics Data Pipelines Amazon Redshift

Job description

Transflo is seeking a Senior Data Engineer to architect and own our enterprise data platform - from raw ingestion through curated, analytics-ready data products. You will be the foundational engineer behind our data warehouse, data pipeline infrastructure, and the bronze-silver-gold medallion architecture that serves internal analytics teams, operational reporting, and our growing Data as a Service (DaaS) capability. This role demands both deep technical expertise and a strategic mindset. You will work across a wide range of source systems - APIs, relational databases, NoSQL stores, file-based feeds, and streaming data - normalizing and modeling data into reliable, governed, and high-performance analytical assets. You will build and scale systems designed for near real-time data environments supporting high-traffic, mission-critical workloads in the transportation and logistics industry. CORE AREAS OF RESPONSIBILITY:

  • Architect, build, and evolve a scalable enterprise data warehouse on Amazon Redshift, applying industry-standard concepts including star schemas, snowflake schemas, normalization, denormalization, referential integrity, and performance optimization strategies
  • Design and implement bronze, silver, and gold data layer architecture (medallion architecture): raw ingestion, cleansed and standardized intermediate layers, and curated, business-ready data products optimized for analytics consumption
  • Develop dimensional data models, fact and dimension tables, slowly changing dimensions (SCDs), and aggregate structures that support BI tooling, ad-hoc analytics, and downstream API consumption
  • Apply rigorous data modeling practices including schema design, constraint definition, indexing strategy, sort keys, distribution keys, and query plan optimization within Redshift and connected systems

  • Build, own, and maintain robust batch and streaming data pipelines that ingest data from disparate source systems including REST APIs, flat files, IBM DB2, MySQL, Amazon Aurora, Amazon DynamoDB, and PostgreSQL
  • Implement real-time and near real-time data streaming architectures using AWS-native services such as Kinesis Data Streams, Kinesis Firehose, MSK (Managed Kafka), and EventBridge to support low-latency data delivery requirements
  • Design pipeline frameworks for data extraction, transformation, and loading (ETL/ELT) using tools such as AWS Glue, dbt, Apache Airflow, or equivalent orchestration platforms
  • Ensure pipeline reliability, idempotency, fault tolerance, and automated recovery; build alerting and observability into every data workflow from day one

  • Own data quality end-to-end: design and implement automated profiling, cleansing, deduplication, standardization, and validation frameworks that enforce data integrity at each layer of the medallion architecture
  • Build and continuously evolve tooling and processes to support data governance including data cataloging, lineage tracking, metadata management, access controls, and data classification
  • Define and enforce data contracts between source systems and the warehouse, establishing clear SLAs for freshness, completeness, and accuracy
  • Partner with data consumers - Data scientists, Data analytics engineers, BI developers, product managers, and external API clients - to understand consumption patterns and ensure data products meet quality and performance expectations

  • Support the architecture and buildout of a reliable, scalable Data as a Service (DaaS) product, enabling external and internal consumers to access curated Transflo data via governed APIs and data sharing mechanisms
  • Contribute to the data platform infrastructure using infrastructure-as-code practices (Terraform), ensuring all data infrastructure is version-controlled, reproducible, and auditable
  • Design for scale: apply partitioning strategies, workload management (WLM) tuning, concurrency scaling, and caching patterns to sustain performance under high-traffic analytical and operational workloads
  • Champion security and compliance best practices across the data platform: column-level security, row-level access controls, encryption, and audit logging

  • Collaborate with software engineers, mobile platform teams, and DevOps to ensure upstream application data is well-structured, well-documented, and reliably delivered to the data platform
  • Leverage AI-assisted development practices and tooling to accelerate pipeline development, automate data quality checks, and improve engineering velocity

Requirements

  • 5+ years of professional data engineering experience with a track record of building and operating production-grade data warehouses and pipeline infrastructure
  • Expert-level experience with Amazon Redshift including cluster sizing, WLM configuration, distribution and sort key optimization, vacuuming, and query plan analysis
  • Deep proficiency in SQL for complex analytical queries, window functions, CTEs, stored procedures, and performance tuning across Redshift and ANSI-compatible engines
  • Hands-on experience ingesting data from heterogeneous source systems: REST APIs, IBM DB2, MySQL, Amazon Aurora (MySQL and PostgreSQL-compatible), Amazon DynamoDB, PostgreSQL, and file-based sources (CSV, JSON, Parquet, Avro)
  • Proven experience designing and implementing medallion (bronze/silver/gold) or equivalent layered data architectures at enterprise scale
  • Strong working knowledge of star schema and snowflake schema design, dimensional modeling theory, slowly changing dimensions, and fact table granularity decisions
  • Experience building real-time or near real-time data pipelines using streaming technologies such as Amazon Kinesis, Apache Kafka (or Amazon MSK), or equivalent
  • Proficiency with ETL/ELT orchestration tools such as AWS Glue, dbt, Apache Airflow, or AWS Step Functions
  • Demonstrated experience implementing data governance practices: data catalogs (AWS Glue Data Catalog, Apache Atlas, or equivalent), lineage, metadata tagging, and access control frameworks
  • Infrastructure-as-code experience with Terraform for provisioning and managing data infrastructure on AWS
  • Strong Python skills for pipeline development, data transformation logic, and automation scripting
  • Deep understanding of data reliability engineering: idempotency, exactly-once processing, late-arriving data handling, schema evolution, and SLA-driven pipeline design, * Experience in the transportation, logistics, trucking, or fleet management industry, or with high-volume transactional SaaS platforms processing operational telemetry data is a huge plus
  • Experience building DaaS or data product offerings including governed external data APIs, Redshift Data Sharing, or AWS Data Exchange integrations
  • Knowledge of columnar storage formats (Parquet, ORC) and lakehouse patterns using Amazon S3 as a data lake layer in conjunction with Redshift Spectrum or AWS Glue
  • Familiarity with BI and analytics consumption tools such as Tableau, Power BI, Amazon QuickSight, or Looker and how data model design decisions impact end-user query performance
  • Experience with data observability platforms such as Monte Carlo, Great Expectations, or dbt tests for automated data quality monitoring
  • Contributions to reusable data platform tooling, shared dbt packages, or internal data engineering frameworks
  • Experience working in fully remote, distributed engineering teams, Access Control, Amazon Simple Storage Service (S3), Amazon Web Services (AWS), American National Standards Institute (ANSI), Apache, Apache Avro, Apache Kafka, Application Programming Interface (API), Artificial Intelligence (AI), Automation Systems, Brokerage, Business Intelligence, Business Intelligence Software, Business Processes, Caching, Cargo/Freight, Cataloguing, Concurrency, Cost Control, Customer/Client Research, Data Analysis, Data Lake, Data Management, Data Modeling, Data Processing, Data Quality, Data Science, Data Warehousing, Database Design, Database Extract Transform and Load (ETL), Desktop as a Service (DaaS), DevOps, Dimensional Modeling, Engineering, Fleet Management, IBM DB2, Industry Standards, JSON, Logistics, Looker, Machine Tool, Metadata, Microsoft Exchange Server, Mobile Applications Development, MySQL, NoSQL, Operational Audit, Performance Analysis, Performance Tuning/Optimization, PostgreSQL, Power BI, Python Programming/Scripting Language, Quality Monitoring, Query Analysis, Query Optimization, REST (Representational State Transfer), Realtime Operating System, Relational Databases (RDBMS), Reliability Engineering, SQL (Structured Query Language), Sales Pipeline, Scripting (Scripting Languages), Service Level Agreement (SLA), Snowflake Schema, Software Engineering, Software as a Service (SaaS), Star Schema, Stored Procedures, Streaming Technology, Supply Chain, Tableau, Telemetry, Test Automation, Transaction Processing/Management, Transportation and Logistics, Trucking, Warehousing, Web Analytics

About the company

Transflo is a leading provider of mobile, telematics, and business process automation software for the transportation and logistics industry. Our solutions help freight carriers, brokers, and shippers automate and streamline their operations, reduce costs, and improve efficiency. We are on a mission to drive innovation in the industry by providing cutting-edge SaaS and AI solutions that enable seamless communication and collaboration across the supply chain.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

1:48 min

Analyzing network packets with database protocol tools

Daniël van Eeden Daniël van Eeden · World Congress 2026 Europe

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · World Congress 2025

Videos

See all

Related articles

See all