> Markdown version of [/jobs/ext/283420-aws-database-engineer-cloud-dba](https://www.wearedevelopers.com/jobs/ext/283420-aws-database-engineer-cloud-dba). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AWS Database Engineer / Cloud DBA - **Company:** OpenKyber LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Amazon Web Services, Amazon S3, Apache HTTP Server, Systems Engineering, Microsoft Azure, Big Data, Business Software, Cloud Computing, Databases, Continuous Delivery, Data Deduplication, Data Governance, Data Integrity, Extract Transform Load (ETL), Data Visualization, DevOps, Perl (Programming Language), Apache Hadoop, Monitoring of Systems, Python (Programming Language), Machine Learning, Meta-Data Management, Metadata Repositories, Oracle (Applications), Queueing Systems, Cloud Services, Ruby, SAS (Software), SQL Databases, Tableau (Software), Teradata SQL, Unstructured Data, Data Ingestion, Azure Data Factory, Apache Spark, Git, Data Lakes, Collibra, Qlikview, Apache Kafka, Apache Nifi, Spark Streaming, Data Pipelines, Databricks - **Published:** May 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=3f0923d83d24d2c8 ## About the Role Do you have experience in Metadata management?, Do you have a Bachelor's degree?, * Must be eligible for a Position of Public Trust, including U.S. citizenship or permanent residency, five years of U.S. residency, and no more than six months of international travel in the past five years (excluding travel for U.S.-based work). * Bachelor s degree and 13 years of experience. * A degree from an accredited College/University in the applicable field of services is preferred. * Four additional years of relevant experience in lieu of a college degree is required. * If Degree is not in the applicable field, then four additional years of related experience is required. * 5+ years demonstrated experience designing and implementing data ingestion pipelines using tools such as: Azure Data Factory Apache Kafka Apache NiFi Spark Structured Streaming or equivalent technologies. * 5+ years of experience applying de-duplication techniques at scale, including: Record linkage Fuzzy matching Entity resolution across structured and unstructured datasets. * 5+ years Hands-on experience with: Data tagging Metadata management Tagging schemas Data catalogs (e.g., Azure Purview, Apache Atlas) Automated classification tools to support data governance and lineage tracking. * 5+ years Demonstrated experience working with unstructured data. * 2+ years of experience using Databricks or other Spark-based platforms. * Fluency in at least one scripting language: Python Perl Ruby or equivalent. Desired Skills: * Integration of Git in continuous deployment and experience with DevOps monitoring tools. * Experience with one or more of the following products and technologies: SAS Python C++ Hadoop SQL Database/Coding Teradata Oracle Amazon S3 Apache Spark Machine Learning Natural Language Processing (NLP) Visualization tools such as: Tableau Strategy QLIK Strong skills and experience in Cloud Operations support in Azure. ## Description We are seeking a Senior Data Engineer to support our client with data ingestion, data deduplication, and data tagging for migration of a large-scale data environment into Databricks., * Design, develop, and maintain scalable data ingestion pipelines to onboard structured, semi-structured, and unstructured data from batch and streaming sources (e.g., APIs, databases, flat files, message queues) into the Azure/Databricks environment. * Implement de-duplication strategies across large-scale datasets using deterministic and probabilistic matching techniques to ensure data integrity and reduce redundancy within the Data Lake. * Develop and enforce data tagging frameworks to classify, label, and annotate datasets with appropriate metadata (e.g., sensitivity, source, domain, lineage) to support data governance, discoverability, and compliance requirements. * Assist with Operationalizing deployments and support of Cloud services for ETL Operations. This will include standardizing and automating processes and workflows, creating documentation/knowledge articles, and overall assisting Operations staff who have limited experience in Cloud. * Written and oral presentations to high-level CIO management on status of current efforts. * Possesses skills and experience related to business management, systems engineering, operations research, and management engineering. Typically has specialization in a particular technology or business application. Keeps abreast of technological developments and industry trends. * Assist with deployment, configuration, and management of Azure Cloud environment. * Assist with migration efforts of existing ETL jobs into Azure/Databricks cloud environment. * Ability to share optimization and efficiencies with the larger team and management. * Ability to automate solutions to repetitive problems/tasks. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)