Senior Data Engineer, AI & Agents

Appnovation Technologies
Austin, TX, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Amazon S3 Apache HTTP Server ARM Architecture Information Systems Continuous Integration Information Engineering Data Security Data Warehousing Python (Programming Language)
+14 more
Search Technologies SQL Databases YAML Enterprise Data Management Sql Optimization Snowflake Apache Spark Git Pyspark Information Technology Collibra Data Management Data Pipelines Databricks

Job description

  • Build and register domain data agents at scale over governed tables across Databricks and Snowflake.
  • Perform lakehouse migrations, including converting source tables to open formats (e.g., Apache Iceberg) to enable agent-based access.
  • Generate and curate catalogue metadata that feeds downstream automation and data-contract workflows.
  • Partner with business stakeholders and subject-matter experts through iterative build, test, and validation cycles., * Agent-Oriented Builder: You enjoy turning governed datasets into reliable, domain-grounded agents that business users can query with confidence.
  • Quality-Focused: You are rigorous about data accuracy, lineage, and observability, ensuring high standards through validation before data reaches agents or the business.
  • Collaborative Partner: You thrive working directly with stakeholders and subject-matter experts through iterative build, test, and validation cycles.
  • Governance-Minded: You understand the critical nature of data security in regulated domains and proactively apply masking and row- and column-level controls.
  • Forward-Thinking: You are interested in the “big picture” of lakehouse architecture and open formats, eager to advance agent-based access patterns and best practices.

What Appnovation Offers

  • Challenging and rewarding work with real impact
  • Direct Access to Cutting-Edge AI Platforms
  • Diverse and Inclusive Culture
  • Growth opportunities for personal and professional development
  • A collaborative and innovative work environment where your ideas are valued
  • Exposure to exciting projects and high-profile clients
  • Supportive work environment with access to company leaders

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Information Systems, Engineering, or a related field.
  • 5+ years of professional experience in data engineering, with significant hands-on experience across modern data warehousing and lakehouse platforms (Databricks and Snowflake preferred).
  • Strong data engineering background with genuine, hands-on fluency across both Databricks and Snowflake.
  • Demonstrated experience building data agents or query interfaces over governed datasets (e.g., Snowflake Cortex or Genie).
  • Advanced SQL together with Spark / PySpark, and experience with pipeline orchestration (dbt, Apache Airflow, or Databricks Workflows).
  • Experience implementing data quality, observability, and lineage, and applying governance controls such as masking and row- and column-level security.
  • Excellent stakeholder-facing skills, with a track record of translating business requirements into delivered data assets.

Preferred Skills

  • Familiarity with MCP-based data exposure and with embeddings or vector search for retrieval-augmented use cases.
  • Experience with AWS and S3, in anticipation of onboarding native cloud data sources.
  • Experience with regulated life-sciences data domains (clinical, commercial, or real-world data).

Technical Experience

  • Data platforms: Databricks, Snowflake (Cortex, Genie); AWS and S3.
  • Pipelines and modelling: SQL, PySpark, dbt, Airflow / Databricks Workflows.
  • Governance and catalogue: Unity Catalogue, Horizon, Collibra.
  • AI and agents: MCP; vector databases and embeddings.
  • Foundations: Python; YAML data contracts; Apache Iceberg; Git and CI/CD.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:35 min

Centralizing configuration logic with native YAML block references

Matthieu Vincent Matthieu Vincent · Europe 2026 Virtual

2:19 min

Introduction to Apache Airflow for advanced orchestration

Alan Mazankiewicz · LIVE

1:59 min

Evolving roles in AI driven software teams

Ignacio Riesgo Ignacio Riesgo +1 · World Congress 2024

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all