Data Scientist

Circle Corporation
Amsterdam, United States
3 months ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$44,500.0 - $74,300.0
Working hours
Shift work
Job source

Tech stack

Artificial Intelligence Content Analysis Decision Support Systems Graph Database Information Retrieval Python (Programming Language) Machine Learning Natural Language Processing Named Entity Recognition NumPy Tensorflow SciPy
+18 more
Search Technologies Unstructured Data Feature Engineering Pytorch Large Language Models Deep Learning Model Validation Electronic Medical Records Pandas Matplotlib Build Management Question Answering Scikit Learn Information Technology Code Testing Data Analytics Data Pipelines Unsupervised Learning

Job description

Our global team support products education electronic health records that introduce students to digital charting and prepare them to document care in today’s modern clinical environment. We have a very stable product that we’ve worked to get to and strive to maintain. Our team values trust, respect, collaboration, agility, and quality., In this role, you will design and build machine learning, NLP, and generative AI solutions that support scientific discovery, knowledge extraction, decision support, and intelligent content understanding. You will work with large-scale scientific content and data, applying the right techniques to solve complex problems and deliver reliable, production-ready systems. Working closely with cross-functional partners, you will help turn ambiguous challenges into measurable outcomes that improve how researchers discover and use knowledge., * Design and build machine learning, NLP, and generative AI systems for scientific discovery, knowledge extraction, decision support, and intelligent content understanding.

  • Work with large-scale, complex, and heterogeneous data, including scientific publications, research datasets, knowledge graphs, ontologies, taxonomies, citations, metadata, and content from every scientific discipline.
  • Apply the right technique to each problem, using approaches such as classification, regression, clustering, ranking, feature engineering, deep learning, embeddings, LLMs, retrieval, and generative AI.
  • Develop capabilities for semantic search, information retrieval, entity extraction, content classification, recommendation, ranking, summarization, question answering, and evidence-grounded generation.
  • Build, evaluate, fine-tune, prompt, and integrate models into robust production systems, while continuously improving quality, relevance, reliability, and user value.
  • Write clean, tested, production-quality Python and contribute reusable data science components, packages, and scalable data pipelines for preprocessing, inference, experimentation, monitoring, and continuous improvement.
  • Support deployment, monitoring, model maintenance, drift detection, automated retraining, and ongoing optimization of data science systems.
  • Collaborate with engineering, product, UX, analytics, research, and domain experts, and communicate technical concepts, model behavior, insights, trade-offs, and recommendations clearly to technical and non-technical audiences.

Requirements

  • Experience in data science, machine learning, artificial intelligence, NLP, statistics, applied mathematics, computer science, or a related quantitative area.
  • Experience working with frontier LLMs such as OpenAI’s GPTs, Anthropic’s Claude, and Google’s Gemini, including fine-tuning LLMs and/or SLMs.
  • Strong Python skills and a habit of writing clean, maintainable, well-tested code.
  • A solid grasp of machine learning fundamentals, including supervised and unsupervised learning, feature engineering, model evaluation, model selection, and performance measurement.
  • Experience working with structured, semi-structured, or unstructured data, especially large-scale text or content datasets.
  • Familiarity with common data science and machine learning tools such as Pandas, NumPy, SciPy, Scikit-learn, PyTorch, TensorFlow, or Matplotlib.
  • The ability to translate complex and ambiguous requirements into practical, measurable, data-driven solutions, with strong analytical thinking, problem-solving skills, and attention to quality.
  • Clear communication skills, a collaborative approach to working with engineering, product, and business stakeholders, and a genuine interest in building production-ready systems that deliver real user value.

Benefits & conditions

We promote a healthy work/life balance across the organisation. We offer an appealing working prospect for our people. With numerous wellbeing initiatives, shared parental leave, study assistance, and sabbaticals, we will help you meet your immediate responsibilities and your long-term goals.

Working Pattern

Working flexible hours - flexing the times when you work in the day to help you fit everything in and work when you are the most productive

About The Business

A global leader in information and analytics, we help researchers and healthcare professionals advance science and improve health outcomes for the benefit of society. Building on our publishing heritage, we combine quality information and vast data sets with analytics to support visionary science and research, health education and interactive learning, as well as exceptional healthcare and clinical practice. At Elsevier, your work contributes to the world’s grand challenges and a more sustainable future. We harness innovative technologies to support science and healthcare to partner for a better world.

Primary Location Base Pay Range: NLD Amsterdam (Radarweg) €44,500 - €74,300. This role is covered by the Collective Labor Agreement Publishing Industry.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:54 min

Development history of scientific computation libraries and PyViz tools

Radovan Kavický · LIVE

3:23 min

Exploring specialized career paths within the data science ecosystem

Julian Joseph · LIVE

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham · World Congress 2025

Videos

See all

Related articles

See all