Remote Nlp Information Extraction Engineer

Odixcity Consulting
Cartagena, Spain
7 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
5 years minimum
Working hours
Regular working hours
Languages
English

Tech stack

Training Data Computational Linguistics Databases Information Extraction Python (Programming Language) Machine Learning Natural Language Processing NLTK (NLP Analysis) Open Source Technology Data Processing Web Content HuggingFace
+2 more
Data Analytics Spacy

Job description

Job Title: Information Extraction SpecialistLocation: Remote (Worldwide)Job Summary: An Information Extraction Specialist is responsible for identifying, extracting, structuring, and validating relevant data from unstructured and semi-structured sources such as documents, reports, web content, databases, and multimedia files.The role involves applying natural language processing (NLP), machine learning models, rule-based systems, and data processing techniques to convert raw information into structured, usable datasets.ResponsibilitiesDesign and implement information extraction pipelines for diverse document types, including legal contracts, medical records, financial reports, news articles, and technical documentation.Oversee the creation of high-quality training datasets for extraction models.This includes defining sampling strategies, managing annotation teams, conducting quality assurance, and resolving ambiguous cases.Evaluate extraction model performance using metrics such as precision, recall, and F1 score.Analyze model errors, identify root causes, and iterate on guidelines, training data, or model architecture to improve results.Evaluate and implement information extraction tools and platforms (open-source and commercial).Develop scripts and workflows to automate aspects of the extraction pipeline.Adapt extraction systems to new domains or document types, rapidly acquiring the necessary domain knowledge to create accurate guidelines.RequirementsMinimum of 5 years of experience in Information Extraction, Natural Language Processing, Computational Linguistics, or relating fields.Experience with Python for data analysis and NLP tasks.Familiarity with NLP libraries such as spaCy, NLTK, Hugging Face Transformers, or Stanford CoreNLP.Proven experience designing annotation schemas and guidelines for complex extractions tasks.Ability to anticipate edge cases and create clear, unambiguous instructions.Deep understanding of evaluation methodologies for extraction tasks.Experience calculating and interpreting precision, recall, F1, and other relevant metrics.Strong problem-solving skills with ability to analyze model errors, identify patterns, and propose data-driven solutions.Excellent written and verbal communication skills in English.Ability to document complex guidelines clearly and explain technical concepts to diverse stakeholders.#J-*****-Ljbffr

Requirements

Develop scripts and workflows to automate aspects of the extraction pipeline.Adapt extraction systems to new domains or document types, rapidly acquiring the necessary domain knowledge to create accurate guidelines.RequirementsMinimum of 5 years of experience in Information Extraction, Natural Language Processing, Computational Linguistics, or relating fields.Experience with Python for data analysis and NLP tasks. Familiarity with NLP libraries such as spaCy, NLTK, Hugging Face Transformers, or Stanford CoreNLP.Proven experience designing annotation schemas and guidelines for complex extractions tasks. Ability to anticipate edge cases and create clear, unambiguous instructions.Deep understanding of evaluation methodologies for extraction tasks. Experience calculating and interpreting precision, recall, F1, and other relevant metrics.Strong problem-solving skills with ability to analyze model errors, identify patterns, and propose data-driven solutions.Excellent written and verbal communication skills in English. Ability to document complex guidelines clearly and explain technical concepts to diverse stakeholders. #J-*****-Ljbffr

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

9:02 min

Building enduring web content against algorithmic search paywalls

Chris Heilmann +2 · LIVE

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

3:29 min

Binary and count vectorization techniques for text

Jodie Burchell · LIVE

3:19 min

Automating data extraction pipelines using natural language assistants

Vidas Bacevičius Vidas Bacevičius · World Congress 2025

2:16 min

Implementing content permissions for large language web crawlers

Farooq Sheikh Farooq Sheikh +3 · World Congress 2025

4:01 min

Managing application isolation via pluggable database models

Wei Hu Wei Hu · World Congress 2022

Videos

See all

Related articles

See all