Python AI Data Engineer

Openmind Technologies
United States
2 months ago
Apply on dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Sql Data Warehouse Artificial Intelligence BigQuery Data Infrastructure Data Transformation Data Security Python (Programming Language) Marketing Information Systems Shell Script Google Cloud Large Language Models Multi-Agent Systems
+9 more
Data Layers Containerization Data Lakes Kubernetes Data Analytics Data Pipelines Docker Database Tools and Utilities Programming Languages

Job description

Client is building an internal AI data platform and is seeking a Lead AI Data Engineer to own and expand its flagship internal tool - Databot, an AI-powered data agent that enables business teams across marketing, finance, and operations to query and interact with company data through natural language.

This is a high-visibility, high-autonomy role. The engineer will be deeply embedded with internal stakeholders across the organization and must combine strong technical execution with a product mindset. The scope extends beyond maintenance: the successful candidate will build new internal data tools and, over time, contribute to productionizing these capabilities for external use.

Given the volume of internal collaboration required - all key stakeholders are US-based - this role requires US timezone availability. Night-shift coverage from India may be considered for exceptional candidates., 1. Databot Ownership & Expansion

  • Maintain, stabilize, and enhance Databot - a Python-based agentic data query tool

  • Extend Databot with new data sources and use cases as identified by business teams

  • Collaborate with the original Databot developer (peer engineer) for knowledge transfer and architectural decisions

  1. Internal Data Tool Development
  • Engage directly with internal teams (marketing, finance, product, engineering) to identify unmet data needs

  • Translate stakeholder requirements into well-scoped, production-grade internal tools

  • Design data pipelines, semantic layers, and consumption interfaces appropriate to each use case

  1. Productionization & Scalability
  • Evaluate internal tools for potential external productionization

  • Ensure tools meet deployment standards - containerized, Kubernetes-ready, monitored

  1. Stakeholder Communication
  • Run discovery sessions with internal business owners to understand data access and reporting needs

  • Translate non-technical requirements into data product specifications

  • Provide regular updates on tool roadmap and delivery status to engineering leadership, The platform operates on internal business metadata - cluster data, invoice data, marketing data, and external market data. No customer PII is involved. Data volume is in the triple-digit gigabyte range (~100-999 GB). The Databot agent currently surfaces this data to business teams via natural language queries.

Requirements

  • Python - primary development language; existing Databot codebase is Python

  • Docker - containerization; Databot is deployed via Docker files

  • Kubernetes - deployment target for all tools

  • Shell scripting - utility scripts in the existing codebase

Data & Analytics (strong preference):

  • BigQuery or equivalent cloud data warehouse

  • Semantic layer concepts - understanding of how data gets modeled for consumption

  • Experience working with dashboards, BI tools, or data reporting pipelines (tool-agnostic; mindset matters more than specific tool)

  • Monitoring - instrumentation and observability for data pipelines and agents

AI & Agentic Systems (openness to learn is acceptable):

  • MCP

  • LLM-based agentic data consumption - experience with data agents or AI-powered query tools

  • Google Cloud AI stack: Gemini, Vertex AI

  • Familiarity with vector databases a plus (Client’s context), * Full-stack engineering capability with Python as the primary language

  • Experience working directly with data - curating, wrangling, and building data-driven outputs

  • Product mindset: ability to work from ambiguous stakeholder needs to a shipped tool

  • Strong communication skills - this role is as much discovery and alignment as it is coding

  • US timezone availability for internal collaboration

  • Comfort operating as an independent contributor with minimal team structure

  • Prior exposure to agentic AI systems or LLM-based data querying

Differentiators:

  • Experience shipping internal tools that later became external products

  • Background in data lake architecture or data federation across heterogeneous sources

  • Startup or scale-up experience where scope and technology evolve rapidly

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:30 min

Leveraging BigQuery ML for scalable SQL-based segmentation experiments

Julian Joseph · LIVE

3:24 min

The governance failures of centralized data lakes

Mario Meir-Huber · LIVE

1:31 min

Essential AI and human skills for future teams

Alexander Weißhaupt Alexander Weißhaupt +1 · World Congress 2025

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

6:24 min

Distributed data lakes and containerized computing clusters

Ulrich Wurstbauer +1 · LIVE

Videos

See all

Related articles

See all