> Markdown version of [/jobs/ext/2219579-data-engineer](https://www.wearedevelopers.com/jobs/ext/2219579-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** CGI Technologies and Solutions, Inc. - **Location:** Pittsburgh, PA, United States - **Experience:** Experienced - **Salary:** $70,800.0 - **Contract:** Permanent contract - **Skills:** Airflow, Amazon S3, Microsoft Azure, Batch Processing, Big Data, Continuous Integration, Data Architecture, Information Engineering, Data Integration, Hadoop Distributed File System, Apache Hive, Job Scheduling, Python (Programming Language), Microsoft SQL Server, MongoDB, Neo4j, Oracle (Applications), Performance Tuning, Standard Sql, Shell Script, Data Streaming, Workflow Management Systems, Data Logging, Data Ingestion, Large Language Models, Apache Spark, Model Validation, Build Management, Data Lakes, Pyspark, Information Technology, Apache Kafka, Graphql, Spark Streaming, Data Pipelines, Jenkins, Control M - **Published:** August 25, 2026 - **Apply:** https://dejobs.org/x/x/6003528E441F4DC8BEF18919C23273CF/job/ ## About the Role Bachelors degree in Computer Science or related field . 4+ years of experience in data engineering and big data processing . 3+ years of experience in building Spark Streaming and Kafka streaming . Strong expertise in Apache Spark (Spark Core, Spark SQL) and PySpark for large scale batch processing. . Experience working with structured and semi structured data, including complex transformations and performance tuning . Proficiency in data ingestion and integration from sources like Oracle, SQL Server, Hive, HDFS, and S3; transform data into 'curated data models' . Experience writing data to Hive tables, Data Lakes (Iceberg), and downstream reporting systems . Strong knowledge of SQL and data modeling concepts . Hands on experience with Apache Airflow for workflow orchestration (DAG design, scheduling expectations, monitoring) . Proficiency in shell scripting for job automation, file validation, dependency handling, and logging. Trigger Spark Jobs, perform file checks and validation; Archive & purge data; mange job dependency, logging & error handling . Strong understanding of batch processing and batch job scheduling frameworks . Experience migrating from CA7/Control M, Airflow (daily, hourly, weekly schedules) CI/CD for data pipelines . Experience ensuring data quality, reliability, and compliance in regulated environments . Good communication and documentation skills, * Apache Spark * AutoGen * Azure * GraphQL * Jenkins * Large Language Model (LLM) * Python * MongoDB * Neo4J ## Description This role will require someone at our client site 5 days a week preferably in Pittsburgh, PA, We are seeking a Data Engineer with 6+ years of experience to design and maintain scalable data pipeline supporting analytics, reporting, and operational needs. The role involves collaborating with cross functional teams to ensure data alignment with business requirements and enterprise standards. Duties and Responsibilities: . Design and build scalable data pipelines aligned with business needs . Process large dataset (batch + sometimes near Realtime) . Ensure data quality, consistency, and governance standards across systems . Support data integration and transformation efforts for analytics and reporting platforms . Maintain data dictionaries, metadata, and documentation . Participate in data architecture reviews and model validation processes . Support analytics reporting and risk platforms ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)