Data Engineer

TODATA ANALYTICS, LLC
Omaha, NE, United States
12 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$80,000.0 - $140,000.0
Working hours
Regular working hours
Job source

Tech stack

Unity 3d Artificial Intelligence Continuous Integration Information Engineering Data Governance Database Queries Software Debugging Software Engineering SQL Databases Scripting Large Language Models Apache Spark
+5 more
Pandas Data Lakes Pyspark Software Version Control Databricks

Job description

We’re looking for a Data Engineer who treats governance and foundational discipline as first-class work, not overhead. You’ll own the pipelines that move client data from raw ingestion through to the conformed, permissioned datasets that power our products and our conversational AI. Because we operate in regulated environments, we hire people who are meticulous about the unglamorous parts.

What You’ll Do

  • Build and maintain data ingestion and transformation pipelines across our platform.
  • Design and evolve the data models that power our products and analytics.
  • Implement data governance and access controls appropriate to a regulated environment, keeping client data correctly isolated.
  • Establish and uphold engineering fundamentals - CI/CD, consistent naming conventions, and environment separation - and hold the codebase to them.
  • Partner with product and software development to translate client requirements into well-modeled, usable data.
  • Support the infrastructure that makes governed data available to downstream consumers, including our AI products.
  • Monitor data quality and lineage and respond to issues before they reach clients.

Requirements

  • 3+ years in data engineering, with hands-on Databricks experience (Spark, Delta Lake, Unity Catalog).
  • Strong proficiency in SQL - comfortable writing complex queries, joins, window functions, and optimizing for performance, designing dimensional/star-schema data models from ambiguous requirements.

· Solid experience with Python for data processing (e.g., PySpark, pandas) and scripting

· Hands-on experience with Databricks (notebooks, Delta Lake, jobs, clusters)

· Excellent problem-solving and debugging skills - able to trace issues through logs, code, and data to find root causes

  • Demonstrated care for data governance and multi-tenant isolation - you can speak to how you’ve kept datasets correctly separated and permissioned.
  • Experience with CI/CD, version control, and disciplined naming/environment conventions in a data context.
  • Track record of independently delivering foundational infrastructure work without close supervision.
  • Clear written communication; you document what you build.

Nice to Have

  • Experience in a HIPAA, SOC 2, or otherwise regulated data environment.
  • Familiarity with healthcare/clinical research or financial/accounting data domains.
  • Exposure to enabling AI/LLM consumers of a governed semantic layer.
  • Databricks certification., * How many years of hands-on experience do you have building production pipelines in Databricks?

Benefits & conditions

You’ll work directly with leadership on a platform where your work is visible in the product and in client trust.

Comprehensive benefits package including medical, dental, vision, short and long-term disability, PTO, matching 401K, fitness center and golf simulator in building and more!

Core Values: Have Grit, Be Open, Be Curious, Collaborate, Make It Simple

Pay: $80,000.00 - $140,000.00 per year, * 401(k) matching

  • Dental insurance
  • Health insurance
  • Life insurance
  • Paid time off
  • Professional development assistance
  • Retirement plan
  • Vision insurance

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:36 min

Development tools for spatial computing and drones

Zaid Zaim Zaid Zaim · World Congress 2023

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

3:33 min

Refactoring data science workflows using Rapids QDF and Pandas

Paul Graham Paul Graham · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

Videos

See all

Related articles

See all