Data Migration Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+2 more
Job description
Company is building a brand-new internal contract hub to replace its legacy contract tracking tools. Ahead of the technical build, they need a data-savvy analyst to lead a time-boxed project: extracting, profiling, and migrating their entire back-catalogue of executed contracts into the new repository - cleaned, de-duplicated, and mapped against a proper metadata schema.
This is a data quality and data engineering problem as much as anything else. You’ll be working across multiple disconnected sources - eSignature platforms, a legacy SQL database, SharePoint, and old ticketing systems - to build a single, reconciled, trustworthy dataset.
What You’ll Be Doing
- Source discovery and profiling - inventory every source system, produce actual record/document counts, and assess data quality, access routes, and risk (week one deliverable)
- Scripted extraction - write and run scripts (SQL and beyond) to pull documents and structured/unstructured metadata in bulk from eSignature platforms, a legacy SQL database, SharePoint/shared drives, and a contract management system
- Data cleansing and de-duplication - consolidate everything into a single staging dataset, applying a documented, defensible rule for which record wins where duplicates exist
- Metadata schema and mapping - build and populate a metadata register against agreed core fields (contract type, counter party, entity, effective/end/renewal dates, value, governing document reference, source system), explicitly flagging gaps rather than inferring or leaving blank
- Validation via sample migration - run a sample batch into the live repository to test and refine the mapping logic and transformation rules before scaling up
- Full-scale load - execute the bulk migration into the new repository structure, correctly mapped against the agreed schema and naming convention
- Documentation and handover - produce a clear data lineage/handover note: what was migrated, what couldn’t be and why, anomalies parked for decision, and the scripts and logic used, so the pipeline is repeatable and auditable
Requirements
- Strong data extraction and scripting skills (SQL plus a scripting language such as Python) - this is a scripted, repeatable-process role first, with manual review reserved for genuine exceptions
- Experience with data profiling, cleansing, de-duplication, and metadata/schema mapping across disparate, messy sources
- Comfort reconciling structured and unstructured data from multiple systems into one clean data set
- A methodical, detail-obsessed approach - this role is about data accuracy, traceability, and repeatability, not shortcuts
- Good judgement on when to escalate rather than infer, especially with ambiguous, incomplete, or sensitive records
- Experience handling confidential or commercially sensitive data professionally and securely
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Data Analyst Salary in the UK
The Most Popular IT Jobs on the Market
Fully Remote Software Engineer Jobs
Top Big Data Technologies That You Need to Know