> Markdown version of [/jobs/ext/1959201-data-engineer](https://www.wearedevelopers.com/jobs/ext/1959201-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** BigBear.ai, Inc. - **Location:** Honolulu, HI, United States - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Artificial Intelligence, Amazon Web Services, Data Analysis, Software Applications, Big Data, Cloud Computing, Data Auditing, Data Cleansing, Information Engineering, Data Integration, Extract Transform Load (ETL), Data Normalization, Database Design, JSON, Python (Programming Language), Scala (Programming Language), SQL Databases, Unstructured Data, Jupyter Notebook, Data Processing, Scripting, Data Ingestion, Apache Spark, Containerization, Apache Kafka, Apache Nifi, Software Version Control, Data Pipelines, Docker, Databricks - **Published:** August 6, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9079796/data-engineer ## About the Role * Bachelor's Degree and 0 to 2 years of experience; 6 to 8 years with no degree * Must maintain an active TS /SCI clearance * 1+ years of Python experience including developing, running, packaging, and testing Python scripts * Experience with distributed version control systems (VCS) * Experience with the entire ETL/ELT pipeline, including data ingestion, data normalization, data preparation, and database design * Experience with conducting exploratory data analysis to communicate qualitative and quantitative findings to analysts * Experience processing and fusing structured and unstructured data * Experience with data engineering projects supporting data science and AI/ML workloads * Experience creating solutions within a collaborative, cross-functional team environment in team sprint cycles What we'd like you to have * Experience with using Palantir products for data manipulation, correlation, and visualization * Experience with AWS or other cloud computing services * Experience with Kafka and NiFi development * Experience with containerization tools, including Docker and Kubernetes * TS/SCI with Counterintelligence Polygraph ## Description * Design, develop, and implement end-to-end data pipelines, utilizing ETL processes and technologies such as Databricks, Python, Spark, Scala, JavaScript/JSON, SQL, and Jupyter Notebooks. * Create and optimize data pipelines from scratch, ensuring scalability, reliability, and high-performance processing. * Perform data cleansing, data integration, and data quality assurance activities to maintain the accuracy and integrity of large datasets. * Leverage big data technologies to efficiently process and analyze large datasets, particularly those encountered in a federal agency. * Troubleshoot data-related problems and provide innovative solutions to address complex data challenges. * Implement and enforce data governance policies and procedures, ensuring compliance with regulatory requirements and industry best practices. * Work closely with cross-functional teams to understand data requirements and design optimal data models and architectures. * Collaborate with data scientists, analysts, and stakeholders to provide timely and accurate data insights and support decision-making processes. * Maintain documentation for software applications, workflows, and processes. * Stay updated with emerging trends and advancements in data engineering and recommend suitable tools and technologies for continuous improvement. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Introducing JSON Structure](https://www.wearedevelopers.com/videos/100219-introducing-json-structure) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)