Remote Nlp Information Extraction Engineer

Odixcity Consulting
Álava, Spain
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
5 years minimum
Working hours
Regular working hours
Languages
English

Tech stack

Training Data Computational Linguistics Databases Information Extraction Python (Programming Language) Machine Learning Natural Language Processing NLTK (NLP Analysis) Open Source Technology Data Processing Corenlp Web Content
+3 more
HuggingFace Data Analytics Spacy

Job description

Job Title:Information Extraction SpecialistLocation:Remote (Worldwide)Job Summary:AnInformation Extraction Specialistis responsible for identifying, extracting, structuring, and validating relevant data from unstructured and semi-structured sources such as documents, reports, web content, databases, and multimedia files. The role involves applying natural language processing (NLP), machine learning models, rule-based systems, and data processing techniques to convert raw information into structured, usable datasets.ResponsibilitiesDesign and implement information extraction pipelines for diverse document types, including legal contracts, medical records, financial reports, news articles, and technical documentation.Oversee the creation of high-quality training datasets for extraction models. This includes defining sampling strategies, managing annotation teams, conducting quality assurance, and resolving ambiguous cases.Evaluate extraction model performance using metrics such as precision, recall, and F1 score. Analyze model errors, identify root causes, and iterate on guidelines, training data, or model architecture to improve results.Evaluate and implement information extraction tools and platforms (open-source and commercial). Develop scripts and workflows to automate aspects of the extraction pipeline.Adapt extraction systems to new domains or document types, rapidly acquiring the necessary domain knowledge to create accurate guidelines.RequirementsMinimum of 5 years of experience in Information Extraction, Natural Language Processing, Computational Linguistics, or relating fields.Experience with Python for data analysis and NLP tasks. Familiarity with NLP libraries such as spaCy, NLTK, Hugging Face Transformers, or Stanford CoreNLP.Proven experience designing annotation schemas and guidelines for complex extractions tasks. Ability to anticipate edge cases and create clear, unambiguous instructions.Deep understanding of evaluation methodologies for extraction tasks. Experience calculating and interpreting precision, recall, F1, and other relevant metrics.Strong problem-solving skills with ability to analyze model errors, identify patterns, and propose data-driven solutions.Excellent written and verbal communication skills in English. Ability to document complex guidelines clearly and explain technical concepts to diverse stakeholders.#J-*****-Ljbffr

Requirements

Minimum of 5 years of experience in Information Extraction, Natural Language Processing, Computational Linguistics, or relating fields. Experience with Python for data analysis and NLP tasks. Familiarity with NLP libraries such as spaCy, NLTK, Hugging Face Transformers, or Stanford CoreNLP. Proven experience designing annotation schemas and guidelines for complex extractions tasks. Ability to anticipate edge cases and create clear, unambiguous instructions. Deep understanding of evaluation methodologies for extraction tasks. Experience calculating and interpreting precision, recall, F1, and other relevant metrics. Strong problem-solving skills with ability to analyze model errors, identify patterns, and propose data-driven solutions. Excellent written and verbal communication skills in English. Ability to document complex guidelines clearly and explain technical concepts to diverse stakeholders. #J-*****-Ljbffr

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:29 min

Binary and count vectorization techniques for text

Jodie Burchell · LIVE

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

3:19 min

Automating data extraction pipelines using natural language assistants

Vidas Bacevičius Vidas Bacevičius · World Congress 2025

54 sec

Open-source tooling and libraries for building hybrid NLP

Jan Schweiger · World Congress 2022

4:01 min

Managing application isolation via pluggable database models

Wei Hu Wei Hu · World Congress 2022

1:52 min

Evolution of distributed SQL database management systems

Akmal Chaudhri Akmal Chaudhri · LIVE

Videos

See all

Related articles

See all