> Markdown version of [/jobs/ext/2697163-data-analytics-lead-engineer](https://www.wearedevelopers.com/jobs/ext/2697163-data-analytics-lead-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Analytics Lead Engineer - **Company:** Citi - **Location:** Irving, TX, United States (Remote available) - **Experience:** Expert - **Salary:** $125,760.0 - $188,640.0 - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Unit Testing, CA Workload Automation Ae, Microsoft Azure, Big Data, Cloud Engineering, Cloudera Impala, Code Generation, Databases, Couchbase Servers, Information Engineering, Data Governance, Extract Transform Load (ETL), Data Presentation, Data Warehousing, Relational Databases, DevOps, Apache Hadoop, Apache HBase, Apache Hive, IBM Cognos Business Intelligence, Python (Programming Language), MongoDB, Neo4j, NoSQL, Query Optimization, Cloud Services, Cloudera, Azure Machine Learning, Shell Script, Data Streaming, Tableau (Software), Working Model 2D, Scripting, Freeform SQL, Cloud Platform System, Data Ingestion, Snowflake, Apache Spark, Parallel Computation, Data Lakes, AI Platforms, Pyspark, Cassandra, Data Analytics, Build Tools, Software Version Control, Data Pipelines, Databricks - **Published:** September 3, 2026 - **Apply:** https://citi.wd5.myworkdayjobs.com/2/job/Irving-Texas-United-States/Data-Analytics-Lead-Engineer_26985324-1/apply ## About the Role * 6+ years of hands-on experience building and managing data pipelines, data warehouses, and data lake solutions using technologies such as Hadoop, Apache Spark, PySpark, Databricks, Delta Lake, Hive, Impala, and Iceberg. * Practical experience with cloud data platforms including Snowflake, Cloudera, used to build and automate ETL and data ingestion workflows. * Fluency in one or more scripting languages - Python, Scala, or Shell Scripting - applied actively to data engineering, pipeline development, and automation tasks. * Strong ability to design and query relational and non-relational data stores, with a clear understanding of schema design trade-offs and data modelling principles. * Hands-on experience with workflow scheduling tools such as Autosys or Apache Airflow to manage and orchestrate data pipeline execution. * Confident use of DevOps practices including version control, build tools, unit testing, monitoring, and change management to support reliable and repeatable delivery. * Experience with data visualization platforms such as Tableau, Cognos to support data presentation and reporting needs. * A Bachelor's degree or equivalent university qualification; a Master's degree is preferred., * Exposure to cloud-based AI and ML services such as Amazon SageMaker, Azure Machine Learning, or Google AI Platform, used to integrate predictive models within data pipelines. * Familiarity with NoSQL database technologies such as HBase, MongoDB, Couchbase, Cassandra, or Neo4j. * Databricks certification or cloud platform certification in AWS, Azure, or GCP. * A proactive approach to troubleshooting - able to independently investigate root causes and resolve pipeline or data issues with thoroughness and pace. ## Description Citi is looking for a Data Analytics Lead Engineer to design, build, and operate scalable data pipelines and cloud-based data architectures within our Lending business, spanning Mortgage and Personal Loans. This is a hands-on data engineering role where you will develop and maintain production-grade data systems - working across big data platforms, data lakes, and cloud infrastructure - that directly power lending analytics at scale. You will also bring an understanding of AI and ML integration as an additional capability applied within a strong data engineering foundation., * Build, deploy, and manage end-to-end data pipelines that ingest, transform, and deliver large-scale lending datasets across Mortgage and Personal Loans with high reliability and performance. * Design and implement scalable data architectures on cloud platforms, selecting the right tools and approaches across data lakes, data warehouses, and streaming environments. * Architect and implement data schemas - choosing from relational, dimensional, normalized, or partitioned models - to meet performance, scalability, and business requirements. * Write and optimize complex SQL queries against large-scale datasets, applying sound decisions around distributed and parallel processing to improve pipeline efficiency. * Monitor, diagnose, and resolve operational and data quality issues across pipelines to ensure accuracy, completeness, and timely delivery of data. * Apply generative AI tools to accelerate core engineering tasks such as code generation, query optimization, and data summarization where appropriate. * Contribute to data engineering standards and collaborate with Business Analysts, Data Engineers, and Data Governance teams to translate business requirements into robust technical solutions., At Citi, you will work as a practicing data engineer at the center of one of the world's largest financial institutions, solving complex, real-world data challenges across lending products that serve millions of customers globally. We offer the technical scale and team environment to do meaningful engineering work, alongside the flexibility and investment to support your continued growth. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)