> Markdown version of [/jobs/ext/2122785-big-data-architect](https://www.wearedevelopers.com/jobs/ext/2122785-big-data-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Big Data Architect - **Company:** LTIMindtree Limited - **Location:** Tampa, FL, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Data Analysis, Big Data, Cloudera Impala, Data Transformation, Data Profiling, Relational Databases, Apache Hadoop, Hadoop Distributed File System, Apache Hive, Python (Programming Language), Machine Learning, Cloudera, SQL Databases, Tableau (Software), Talend, Feature Engineering, Data Ingestion, Sql Optimization, Apache Spark, Ab Initio, Data Lakes, Pyspark, Core Data, Data Management, Artificial Intelligence Markup Language (AIML), Data Pipelines, Unsupervised Learning, Databricks - **Published:** August 19, 2026 - **Apply:** https://www.dice.com/job-detail/3ff3d705-513c-4592-8232-cc5ad1de92bf ## About the Role * Strong proficiency in Python PySpark and data ingestion tools must have * Advanced SQL and RDBMS expertise * Handson experience with Big Data technologies Spark Hive Impala HDFS S3 and data lake house architecture AWS Cloud Airflow Starburst Iceberg * Strong experience in data analysis data exploration and trend identification * Solid understanding of machine learning fundamentals like Regression classification clustering Feature engineering * Mandatory Certificate Cloudera CCA Spark and Hadoop Developer CCA175 or Databricks Certified Associate Developer for Apache Spark ## Description * Primary skill Database SQL Data Modeling skills * Secondary skill Tableau * Core Data Engineering * Design build and maintain scalable data pipelines using PySpark Spark SQL and Python * Work with largescale data platforms Hive Impala S3 HDFS * Implement ETLELT workflows using tools such as Talend or Ab Initio * Ensure data quality performance and reliability across pipelines * Optimize data models for analytics and AI use cases * Advanced Analytics ML AI Enablement * Analyse large datasets to identify trends anomalies and behavioral patterns * Apply machine learning and AI concepts supervised unsupervised learning to support predictive and exploratory analysis * Perform feature engineering and data transformations to enable ML models * Partner with analytics and data science teams to prepare validate and operationalize datasets for AIML workflows * Support trend analysis segmentation clustering and forecasting use cases * Interpret analytical results and translate them into business-friendly insights * Data Analysis Insight Generation * Conduct deep dive data analysis using SQL Python and Spark * Create analytical datasets for dashboards reporting and advanced analytics * Validate trends and findings with statistical reasoning and data profiling * Collaborate closely with stakeholders to answer complex business questions using data ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)