Data Scientist II

LexisNexis
Chicago, IL, United States
25 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$110,100.0 - $183,500.0
Working hours
Shift work
Job source

Tech stack

Mxnet Java (Programming Language) Application Programming Interfaces (APIs) Amazon Web Services Amazon Elastic Compute Cloud Cloud Computing Computer Programming Data Cleansing Elasticsearch Graph Database Information Retrieval Python (Programming Language)
+27 more
Machine Learning Natural Language Processing Named Entity Recognition NumPy Open Source Technology Recommender Systems Cloud Services Tensorflow Sentiment Analysis Apache Solr SQL Databases Lexis Latent Dirichlet Allocation Pytorch Large Language Models Apache Spark Caffe Deep Learning Model Validation Keras Pandas Scikit Learn Machine Learning Operations Gensim Opennlp Spacy GPT

Job description

The Yoda team is a small, focused division within LexisNexis responsible for managing core datasets related to people, organizations, and taxonomies. These datasets are published internally and used by teams across the company to build products and deliver customer value., As a Senior Data Scientist II on the Yoda team, you will:

  • Solve challenging problems in natural language processing, machine learning, and information retrieval, including topical classification, sentiment analysis, entity extraction, and user intent detection.

  • Research, build, train, evaluate, and deploy machine learning models using both traditional and deep learning techniques.

  • Develop robust NLP-based models over large-scale corpora, including news, financial, legal, and business data.

  • Design and improve scalable NLP and machine learning pipelines.

  • Evaluate state-of-the-art algorithms, models, APIs, and open-source tools, including BERT, ELMo, GPT-based models, and related technologies.

  • Translate complex business requirements into actionable technical stories with practical estimates.

  • Partner with product leaders, engineers, and cross-functional stakeholders to apply data science solutions to real business problems.

  • Contribute to best practices for model development, evaluation, deployment, monitoring, and maintenance.

  • Support and mentor junior team members while contributing as part of a small, collaborative team.

Requirements

Do you enjoy collaborating with others to build impactful data products that power real-world applications?, * Strong understanding of machine learning techniques, including classification, clustering, recommendation systems, regression, and statistical modeling.

  • Hands-on experience with Python machine learning and data science libraries such as scikit-learn, pandas, NumPy, and related tools.

  • Experience with NLP tools and methods such as OpenNLP, Stanford NLP, LDA, Gensim, spaCy, or similar frameworks.

  • Proficiency training large-scale models using at least one modern deep learning framework such as TensorFlow, Keras, PyTorch, MXNet, Caffe, or Caffe2.

  • Experience building and deploying cloud-based services, preferably using AWS services such as EC2 and Lambda.

  • At least 5 years of recent coding experience using Python and/or Java or Scala.

  • SQL programming experience.

  • Experience designing, working with, and reasoning complex data models.

  • Familiarity with cloud-based machine learning environments, Spark, visualization and dashboarding tools, Elasticsearch, Solr, and graph databases such as JanusGraph, Neptune, or similar technologies.

  • Strong ability to set, communicate, implement, and achieve business objectives and goals.

  • Ability to work effectively on a small team and provide technical leadership or mentorship to junior team members.

Preferred Qualifications:

  • Experience with large language models and generative AI workflows.

  • Experience with entity extraction, taxonomy management, knowledge graphs, or data enrichment.

  • Experience working with large-scale legal, news, financial, business, or professional data.

  • Familiarity with model evaluation, experimentation frameworks, MLOps practices, and production of ML monitoring.

  • Ability to quickly evaluate new approaches and determine the right tool or model for a given business problem.

About the company

LexisNexis Legal & Professional, which serves customers in more than 150 countries with 11,800 employees worldwide, is part of RELX (, a global provider of information-based analytics and decision tools for professional and business customers. Our company has been a long-time leader in deploying AI and advanced technologies to the legal market to improve productivity and transform the overall business and practice of law, deploying ethical and powerful generative AI solutions with a flexible, multi-model approach that prioritizes using the best model from today’s top model creators for each individual legal use case. The company employs over 2,000 technologists, data scientists, and experts to develop, test, and validate solutions in line with RELX Responsible AI Principles (;br>, Work in a Way That Works for You We promote a healthy work/life balance across the organisation. We offer an appealing working prospect for our people. With numerous wellbeing initiatives, shared parental leave, study assistance and sabbaticals, we will help you meet your immediate responsibilities and your long-term goals.

Working Pattern Working flexible hours - flexing the times when you work in the day to help you fit everything in and work when you are the most productive.

About the Business LexisNexis Legal & Professional provides legal, regulatory, and business information and analytics that help customers increase their productivity, improve decision-making, achieve better outcomes, and advance the rule of law around the world. As a digital pioneer, the company was the first to bring legal and business information online with its Lexis and Nexis services.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

1:18 min

Converting existing Keras models to TensorFlow format

Håkan Silfvernagel · LIVE

3:29 min

Binary and count vectorization techniques for text

Jodie Burchell · LIVE

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · WWC 2024

1:34 min

Bringing diverse skills to industrial data science roles

Katja Träumner

2:20 min

Speeding up model training cycles with transfer learning techniques

Anirudh Koul · LIVE

Videos

See all

Related articles

See all