> Markdown version of [/jobs/ext/2454975-junior-data-scientist](https://www.wearedevelopers.com/jobs/ext/2454975-junior-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Junior Data Scientist - **Company:** Historic Records Limited - **Location:** UK (Remote available) - **Experience:** Starter - **Salary:** £24,000.0 - £28,000.0 - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Data Governance, Text Processing, Python (Programming Language), Machine Learning, Natural Language Processing, Named Entity Recognition, Regular Expressions, Standard Sql, Technical Data Management Systems, Data Processing, Pandas, Information Technology, Data Analytics - **Published:** August 29, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=e4b865b8e563f00c ## About the Role About the Role We are seeking a Junior Data Scientist with strong problem-solving skills, initially for a 6-month, part-time contract (2 days per week)., We are looking for someone with highly reliable skills in Python and Pandas who is, above all, a thoughtful problem solver. You will take ownership of transcription post-processing, working collaboratively to devise programmatic approaches to handle noisy text data, and confidently stand behind the accuracy of your results., This is an excellent opportunity for a confident, autonomous early-career Data Scientist or a freelancer looking for a regular, long-term project that deals with unique, highly unstructured real-world data., * Bachelor's or Master's degree in a quantitative or technical field (e.g. Data Science, Computer Science, or Mathematics) or equivalent professional experience including a demonstrable programming background. * 2+ years of experience working in a Data Science or adjacent role. You must have a track record of building programmatic solutions to data problems. This will be discussed during the interview and confirmed at the reference stage. * Strong hands-on experience using Python and Pandas * Advanced text data proficiency. Practical experience solving problems with unstructured text data. You should be familiar with techniques such as TF-IDF, entity extraction, string manipulation, Regular Expressions (Regex), text similarity metrics, and machine learning approaches such as clustering. This experience will be tested at interview stage * Working knowledge of SQL to extract, manipulate, and query data efficiently. * A meticulous approach to your work. In genealogy, a single misspelled name or wrong date can alter an entire family tree, so accuracy is a high concern. * Ability to communicate technical data concepts to non-technical stakeholders clearly. * You will be one member of a small, growing team. Because you will be heavily involved in the data quality process, you must be comfortable working autonomously, managing your own time, and building robust workflows. ## Description In this role, you will work directly alongside genealogist and AI engineer team-mates to enhance, clean, and validate historical and genealogical data. Historical data is inherently messy, unstructured, and complex. Your primary role will be to apply data wrangling and data science techniques to actively fix, structure, and enhance this data, making it reliable for genealogical research., * Data Wrangling & Enhancement: Take charge of manipulating, cleaning, and standardising raw, poorly transcribed datasets so they are structured and ready for downstream use. * Collaborative Problem Solving: Work closely with genealogists to understand the nuances of historical records. You will need to translate their domain expertise into programmatic data science approaches to resolve complex transcription errors. * Text Data & NLP Application: Utilize text processing and Natural Language Processing techniques to extract meaning, identify entities, and correct anomalies in messy OCR/ transcription data. * Data Ownership: Independently verify your own work. You must be able to confidently stand behind your data outputs, explain your methodology, and iterate on your approaches when results are questioned. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) - [Getting to Know Your Legacy (System) with AI-Driven Software Archeology](https://www.wearedevelopers.com/videos/1437-getting-to-know-your-legacy-system-with-ai-driven-software-archeology) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)