Databricks Architect

Tata Consultancy Services Limited
Marlborough, MA, United States
3 days ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$110,000.0 - $140,000.0
Working hours
Regular working hours

Tech stack

Unity 3d Training Data Artificial Intelligence Amazon Web Services Amazon S3 Data Analysis Business Logic Microsoft Azure Cloud Computing Configuration Management Code Review Cyber Security
+43 more
Computer Programming Databases Continuous Delivery Continuous Integration Information Engineering Data Governance Extract Transform Load (ETL) Data Masking Data Security Data Warehousing Database Development Software Debugging Software Design Patterns DevOps Dimensional Modeling Github Apache Hive Python (Programming Language) Machine Learning Performance Tuning Software Construction Software Engineering SQL Databases Data Streaming Scripting Google Cloud Data Storage Technologies Sql Optimization Apache Spark Caching Infrastructure as Code (IaC) Git Data Lakes Pyspark Data Lineage Data Management Virtual Agents Terraform Software Version Control Data Pipelines Serverless Computing Cisco Databricks

Job description

Data Pipeline Development: Design, code, and deploy robust and scalable batch and streaming data pipelines using PySpark, Spark SQL, and Lakeflow (Lakeflow Connect for ingestion and Lakeflow Declarative Pipelines) to ingest data from sources such as Point-of-Sale (POS), e-commerce platforms, loyalty systems, and marketing clouds. Data Modeling & Transformation: Implement complex data transformations and business logic within the Medallion architecture (Bronze, Silver, Gold layers). Build and optimize the final “Gold” dimension tables that will serve as the single source of truth. Define governed business KPIs on top of the Gold layer using Unity Catalog Metric Views (business semantics) so metrics are computed once and reused consistently across BI, SQL, and Genie. Data Quality: Implement data quality frameworks and cleansing routines to ensure the accuracy and trustworthiness of the Customer 360 data. Performance Optimization: Proactively monitor, debug, and tune Databricks jobs and Spark clusters for performance and cost-efficiency. Implement best practices for partitioning, caching, and data layout in Delta Lake. Infrastructure as Code (IaC) & CI/CD: Work with DevOps teams to manage Databricks environments, clusters, and job deployments using tools like Terraform and AWS DevOps/GitHub Actions. Champion and implement CI/CD best practices for data pipelines. Data Governance & Security: Implement data governance features within Databricks Unity Catalog, including data lineage tracking, fine-grained access controls (row/column-level security and ABAC policies), tags, and data masking to ensure compliance and security across BI and AI/agent workloads. C ollaboration: Partner closely with Functional Consultants, Data Scientists, and Analytics Engineers to understand their data requirements and deliver well-structured, consumption-ready datasets. Semantic Modeling & Business Semantics: Build and govern Unity Catalog Business Semantics, authoring Metric Views (measures, dimensions, joins, synonyms, materialization), Domains, Pages, and Glossary terms, and certifying trusted assets so every dashboard, SQL query, notebook, and AI agent works from the same governed definitions. Genie & Agentic Analytics Enablement: Configure and curate AI/BI Genie Spaces and the Genie Ontology (instructions, trusted assets, example queries, synonyms) to deliver accurate natural-language, conversational analytics for business users, and enable consumption through AI/BI Dashboards and Databricks One.

Requirements

Experience: 10+ years of hands-on data engineering experience, with at least 3 years focused on the Databricks/Spark Ecosystem Databricks Expertise: Deep, hands-on expertise with the Databricks Lakehouse Platform, including Delta Lake, Structured Streaming, Lakeflow Declarative Pipelines (formerly Delta Live Tables), Databricks SQL, and cluster/serverless configuration and optimization. Business Semantics & Governance: Hands-on experience with Unity Catalog Business Semantics, including Metric Views (measures, dimensions, materialization), Domains, and Pages/Glossary to define governed, reusable KPIs once and serve them consistently across SQL, BI, and AI agents. Agentic & Conversational Analytics: Working knowledge/exposure of AI/BI Genie (Genie Spaces, Genie Ontology, trusted assets), Genie Code for agentic pipeline/SQL development, Databricks One for business-user consumption, and Mosaic AI for building and serving AI/ML models and agents. Programming Mastery: Expert-level proficiency in Python and PySpark. Advanced SQL skills are essential. Data Warehousing Concepts: Strong understanding of data modeling principles, including dimensional modeling (Kimball), data warehousing concepts, and ETL/ELT design patterns. Cloud Proficiency: Proven experience working with a major cloud provider (Azure, AWS, or GCP), particularly with data storage S3 and related services. Software Engineering Mindset: Experience with software engineering best practices, including version control (Git), code reviews, testing, and CI/CD. Certification: Databricks, Access Control, Amazon Simple Storage Service (S3), Amazon Web Services (AWS), Apache Spark, Artificial Intelligence (AI), Artificial Intelligence (AI) Agents, Best Practices, Business Analysis, Business Intelligence, Business Model, Business Writing, Caching, Cisco Unity, Cloud Computing, Code Reviews, Continuous Deployment/Delivery, Continuous Integration, Data Analysis, Data Management, Data Modeling, Data Quality, Data Science, Data Storage, Data Warehousing, Database Extract Transform and Load (ETL), Debugging Skills, Design Patterns Programming Methodologies, DevOps, Dimensional Modeling, GCP (Good Clinical Practices), Git, GitHub, Information/Data Security (InfoSec), Legal, Loan Funding, Maintain Compliance, Marketing, Metrics, Microsoft Windows Azure, Ontology, Performance Metrics, Performance Tuning/Optimization, Python Programming/Scripting Language, Reporting Dashboards, SQL (Structured Query Language), Sales, Security Compliance, Software Engineering, Source Code/Configuration Management (SCM), Student Loans, Team Player, Testing, Training Data Sets, eCommerce

Benefits & conditions

Discretionary Annual Incentive. Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans. Family Support: Maternal & Parental Leaves. Insurance Options: Aut& Home Insurance, Identity Theft Protection. Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement. Time Off: Vacation, Time Off, Sick Leave & Holidays. Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:29 min

Expanding practical knowledge with community sandboxes and resources

Stuart Clark · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:59 min

Key takeaways and accessing the Databricks developer toolkit

Viktoria Semaan Viktoria Semaan · World Congress 2026 Europe

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:41 min

Visualizing the complex developer journey for JVM ecosystems

Bobur Umurzokov · LIVE

Videos

See all

Related articles

See all