Junior Data Engineer
Role details
Job location
Tech stack
Job description
As a Junior Data Engineer at BUX, you will be a key builder of the analytics-ready data layer that powers our product decisions and business growth. You won't just write queries; you will build well-tested data ingestion scripts and transform raw data into high-quality, reliable data products. You will join the Data & AI team, collaborating closely with Analytics Engineers, Product stakeholders, and our Data Platform Engineers to leverage our scalable, self-service infrastructure safely. Your mandate focuses on execution and reliability: writing clean, testable code, owning the data quality of your pipelines, and documenting as you go. This is a role designed for growth, where curiosity, coachability, and an engineering-first mindset will serve as the foundation for your evolution into advanced analytics architecture or infrastructure-focused engineering tracks.
What you will do
- Build Reliable Ingestions: Build and maintain robust data ingestion scripts to pull data from APIs, servers, or files into raw/staging tables. You ensure correct handling of retries, incremental/idempotent loads, and graceful error management.
- Model with dbt: Grow our analytics-ready data layer by building and maintaining clean staging and mart models with clear naming, documentation, and automated tests, focusing on simplicity and readability before optimisation.
- Support Orchestration: Extend and troubleshoot our existing workflow orchestration frameworks (Airflow). You will add tasks to existing DAGs, fix failing runs, and grow into mastering dependency, scheduling, and retry behaviour.
- Champion Data Quality: Own pipeline and model health by adding and monitoring data quality checks (freshness, volume anomalies, null/duplicate validations) and investigating data discrepancies flagged by different stakeholders.
- Document & Accelerate: Write clean code, maintain useful READMEs, and leverage AI-assisted development tools to speed up drafting while taking full ownership of verifying correctness before pushing to production.
- Grow your ownership: Build a solid track record with data modelling, testing, and pipeline reliability as a foundation for growing into orchestration design, platform topics, and broader engineering responsibilities over time.
Requirements
- You have 1-2 years of experience in data engineering, analytics engineering, or a closely related discipline, with a track record of collaborative problem-solving.
- Solid fundamentals in Python and SQL, alongside comfort with version control (Git), writing basic automated tests, and interacting with HTTP APIs.
- General understanding of foundational cloud architecture and infrastructure concepts.
- Practical, hands-on experience (e.g. using dbt) for structuring data transformation layers, including staging/mart layering, sources, tests, and documentation.
- A clear grasp of workflow orchestration concepts (such as Airflow), understanding DAGs, task dependencies, scheduling, retries, and the importance of idempotency.
- Comfortable working with relational databases or data warehouses (Postgres, Snowflake, or similar) to write queries, design schemas, and handle data quality issues like nulls, duplicates, and freshness checks.
- You are curious and comfortable asking for help. You welcome constructive feedback and actively contribute to the team's code quality.
- You bring the critical eye needed to rigorously test and verify all logic, including AI-assisted code, before deployment., * Familiarity with Docker and how applications run on Kubernetes (GKE).
- Basic exposure to Terraform and understanding how cloud environments are provisioned as code.
- Initial exposure to high-throughput messaging (Kafka, Pub/Sub) or working with cloud object storage frameworks (GCS, S3).
- A basic understanding of continuous integration pipelines (like GitHub Actions) and the concept of how code changes safely transition from a local machine to a production environment.