Data Analyst (Python / Data Quality)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
We are seeking a detail-oriented Data Analyst for a 6-month, part-time contract (2 days per week). In this role, you will work directly alongside genealogist and AI engineer team members to manage, clean, and validate historical and genealogical data recently transcribed via AI models and/or human transcribers.
Historical data can often be messy, unstructured, and complex. Your primary role will be to ensure this transcribed data is of the highest possible quality, making it reliable for genealogical research.
You will be the point person for data quality concerns on this project, taking ownership of transcription post-processing and cleaning, handling ad-hoc data questions, and confidently standing behind the accuracy of your results.
This is an excellent opportunity for a confident, autonomous early-to-mid career analyst or a freelancer looking for a regular, long-term project that deals with unique, real-world data., * Data Cleaning & Processing: Take charge of cleaning, processing, and standardising raw transcribed datasets so they are ready for analysis.
- Quality Assurance: Work closely with genealogists to identify discrepancies, anomalies, or errors in the transcriptions, ensuring a high degree of data integrity.
- Ad Hoc Querying: Respond to ad hoc data requests from genealogists to support ongoing process refinements.
- Data Ownership: Independently verify your own work. You must be able to confidently stand behind your data outputs and explain your methodology if results are questioned.
Requirements
- Bachelor’s degree in a quantitative or technical field (e.g. Data Science, Computer Science) or equivalent professional experience including demonstrable technical and programming experience.
- 2+ years experience working with data as a data analyst, data quality engineer, or equivalent role. This experience will be discussed during the interview and confirmed at reference stage.
- Strong hands-on experience using Python, Pandas and NumPy. This experience will be tested at interview stage
- Advanced text processing skills: High proficiency in string manipulation, Regular Expressions (Regex), and fuzzy matching to handle messy OCR/transcription data.
- Working knowledge of SQL to extract, manipulate, and query data efficiently.
- Proven experience in cleaning, wrangling, and standardizing messy or unstructured data.
- A meticulous approach to your work. In genealogy, a single misspelled name or wrong date can alter an entire family tree, so accuracy is a high concern.
- Ability to communicate technical data concepts to non-technical stakeholders clearly and effectively.
- You will be one member of a small, growing team. Because you will be heavily involved in the data quality process, you must be comfortable working autonomously, managing your own time, and building robust workflows.
Benefits & conditions
Pay: £24,000.00-£28,800.00 per year
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Data Engineer Salary UK
Data Analyst Salary Germany
How to Become an AI Engineer
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production