TELECOMMUTE Lead Data Engineer
Role details
Job location
Tech stack
Job description
Architect, build, and optimize distributed data pipelines using Apache Spark in a high-volume, mission-critical environment. Design and maintain enterprise Lakehouse architecture with Delta Lake, ensuring ACID compliance, lineage, auditability, and data governance. Develop automated ingestion frameworks (batch, streaming, and event-driven) across multiple cloud services and integration points. Enable machine-learning workflows by preparing feature-ready datasets and establishing reproducible ML deployment patterns. Lead platform-wide data quality, access control, and cataloging frameworks. Implement advanced cost-optimization, cluster tuning, and performance engineering strategies. Collaborate with Finance, BI, Operations, and ML teams to translate complex business needs into scalable data solutions. Own production reliability, troubleshooting, and root-cause analysis for data and ML pipelines.
Requirements
7+ years of experience in advanced data engineering with distributed compute technologies. Expert-level Spark engineering (performance tuning, cluster configuration, partition strategies, optimization of large datasets). Hands-on experience with Lakehouse architectures including ACID transactions, schema evolution, and governance frameworks. Deep proficiency in Python and SQL for large-scale data transformation. Experience supporting machine-learning pipelines or model operationalization. Proven experience architecting cloud-native data platforms (Azure, AWS, or Google Cloud Platform). Strong background integrating diverse, complex data sources at enterprise scale. Demonstrated ability to own mission-critical production systems.
Core Tools Databricks (Spark, Delta Lake, MLflow, Notebooks) Python & SQL Apache Spark (via Databricks) Delta Lake (for lakehouse architecture) Cloud Platforms Azure, AWS, or Google Cloud Platform Cloud Storage (ADLS, S3, GCS) Data Integration Kafka or Event Hubs (streaming) Auto Loader (Databricks file ingestion) REST APIs AI/ML MLflow (model tracking/deployment) Hugging Face Transformers LangChain / LlamaIndex (LLM integration) LLMs: Anthropic Claude, Meta LLaMA, Google Gemini DevOps Git (GitHub, GitLab, Azure Repos) Databricks Repos CI/CD: GitHub Actions, Azure DevOps Security & Governance Unity Catalog RBAC
Experience with distributed streaming frameworks (Kafka, Event Hubs, or similar). Experience building or supporting ML platforms, feature stores, or experiment-tracking systems. Background in data security, compliance controls, or audit-ready governance. Experience automating data operations with CI/CD and infrastructure-as-code.
Benefits & conditions
Our benefits package includes: (EXCLUDE on perm placements)
- Comprehensive medical benefits
- Competitive pay
- 401(k) retirement plan
- …and much more!
About, INSPYR Solutions