> Markdown version of [/jobs/ext/2897765-data-harmonization-analyst](https://www.wearedevelopers.com/jobs/ext/2897765-data-harmonization-analyst). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Harmonization Analyst - **Company:** Nyu Langone Medical Center - **Location:** New York, NY, United States - **Experience:** Experienced - **Salary:** $84,578.0 - $126,992.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Data Analysis, Big Data, Health Informatics, Computer Networks, Databases, Data Dictionary, Data Governance, Data Hub, Extract Transform Load (ETL), Data Structures, Data Visualization, R (Programming Language), Apache Hadoop, MapReduce, Python (Programming Language), Machine Learning, Metadata, Metadata Standards, MongoDB, Software Architecture, Azure Machine Learning, SQL Databases, Data Streaming, Systems Integration, Virtualization Technology, Apache Spark, Containerization, Information Technology, Build Tools, Machine Learning Operations, Programming Languages - **Published:** September 14, 2026 - **Apply:** http://hbcu.com/cgi-bin/jobs/searchJobs.cgi?job_id=35227410 ## About the Role To qualify you must have a Masters degree in a quantitative discipline (Biomedical Informatics, Computer Science, Machine Learning, Applied Stascs, Mathematics or similar field) and 3 years of experience in machine learning/ data science. Proficiency in at least one programming language (Python, R) and machine learning tools (scikitlearn, R) Knowledge of predictive modeling and machine learning concepts, including design, development, evaluation, deployment and scaling to large datasets Familiarity with computing models for big data Hadoop / MapReduce, Spark etc. Knowledge of databases (Relational / SQL, NOSQL MongoDB etc.) Good grasp of soware engineering principles. Experience in integrating modern software architectures. Knowledge and some experience in operational aspects of soware development and deployment, including automation, testing, virtualization and container technology Knowledge of clinical and operational aspects of healthcare delivery. Excellent written and oral communication skills for a variety of audiences Preferred Qualifications: Experience with OMOP Common Data Model or other biomedical research CDMs Experience with programming languages (Python, JAVA, R) Familiarity with healthcare or life sciences data standards (UMLS, SNOMED-CT, LOINC), or sequencing data standards (FASTQ, BAM, VCF), etc. Knowledge of FAIR data principles and metadata standards Familiarity with NAMs methodologies or preclinical research data Experience with federated data networks or distributed query systems Familiarity with AI/ML tools applied to terminology matching, automated mapping recommendations, or data quality assessment Qualified candidates must be able to effectively communicate with all levels of the organization. ## Description We have an exciting opportunity to join our team as a Data Science Analyst/Engineer. As part of the Complement-ARIE program, the NYU-Sage New Approach Methodologies (NAMs) Data Hub and Coordinating Center will create a controlled access platform for researchers to share and analyze data resulting from NAMs approaches. The program will build tools to standardize and harmonize NAMs data, store it securely, and provide researchers with powerful analytical and visualization tools. The successful candidate will support implementation of integrated standards [SG1.1]tailored for NAMs data. This position focuses on analyzing source data structures, developing source-to-target mapping specifications, authoring ETL functional requirements and data quality assurance frameworks, and ensuring that harmonization workflows are well-defined, reproducible, and aligned with FAIR data principles. The role is analytical and specification-oriented: the Data Analyst designs and documents the logic that guides implementation, rather than executing engineering tasks directly., * Analyze source NAMs datasets (such as transcriptomics, proteomics, microscopy, imaging, electrophysiology, etc.) to characterize data structure, content, and quality prior to harmonization * Develop detailed source-to-CDM mapping specifications, including transformation rules, value set crosswalks, and handling of edge cases * Author functional ETL requirements and data flow documentation to guide pipeline development by engineering staff * Design data quality assurance (QA) frameworks and acceptance criteria for NAMs datasets, including completeness, conformance, and plausibility checks * Evaluate and document terminology alignment across existing Vocabularies, Metadata requirements, and source ontologies * Conduct mapping gap analyses and propose remediation strategies for non-standard or missing terminology coverage * Collaborate with the Lead Metadata and Standards Specialist to ensure mapping outputs align with metadata standards * Produce and maintain clear analytical documentation: data dictionaries, mapping catalogs, QA specification sheets, and implementation guide * Support onboarding of new data contributors by reviewing their data structures and advising on harmonization pathways * Participate in data quality review cycles, analyze QA outputs, and document findings and recommended remediation steps ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [A Data Mesh needs Open Metadata](https://www.wearedevelopers.com/videos/505-a-data-mesh-needs-open-metadata) - [40 Minutes to Build a Serverless COVID-19 REST and GraphQL APIs](https://www.wearedevelopers.com/videos/208-40-minutes-to-build-a-serverless-covid-19-rest-and-graphql-apis) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Building an AI-Ready Content Lake: Scaling RAG and Document AI Beyond Demos](https://www.wearedevelopers.com/videos/1977-building-an-ai-ready-content-lake-scaling-rag-and-document-ai-beyond-demos) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)