Senior Data Scientist (AI & Synthetic Intelligence)

Cint
Greater London, UK
11 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Artificial Neural Networks Big Data Cluster Analysis Databases Data Validation Data Mining Statistical Hypothesis Testing Python (Programming Language) Machine Learning Open Source Technology Performance Tuning
+10 more
Sql Optimization Retrieval-Augmented Generation Large Language Models Prompt Engineering Apache Spark Generative AI Pyspark Synthesizing Data Databricks Data Generation

Job description

As a Senior Data Scientist at Cint, you will play a pivotal role in developing next-generation AI solutions that power our product portfolio. Collaborating closely with Product and Engineering teams, you will bridge the gap between traditional research data and synthetic intelligence, focusing on the research, validation, and delivery of models-including Large Language Models (LLMs)-that augment high-quality human signals across the Cint Exchange. This role involves advanced data mining, robust data validation, and the development of sophisticated statistical and machine learning methodologies.

The ideal candidate can independently research, develop, and maintain high-impact solutions that align Cint’s AI capabilities with market research trends, directly influencing strategic decisions for Cint’s proprietary synthetic data platform., * Lead the research, discovery, and development of machine learning models, specifically focused on synthetic row generation, open-ended text generation, and data augmentation.

  • Design and drive advanced statistical methods and complex experiments to validate LLM performance and synthetic modeling hypotheses.
  • Develop logic for on-demand and dynamic boosting capabilities, collaborating with Engineering to integrate these models into Cint Exchange fielding workflows.
  • Design and refine sophisticated profiling taxonomies, leveraging large-scale datasets to create syndicated audiences.
  • Independently manage complex project planning, development, and maintenance end-to-end with minimal supervision.
  • Partner with Product, Engineering, and Strategy teams to align technical AI delivery with business goals, ensuring a seamless transition from MVP to scaled integration.
  • Create clear, effective prototypes and deliverables that explain and defend complex Generative AI concepts to both technical and non-technical audiences.

Requirements

  • Minimum 5 years of experience in a Data Science capacity, with a track record of leading complex projects.
  • A Master’s degree (or equivalent) in Statistics, Data Science, or a related quantitative field.
  • Deep understanding of Generative AI and LLMs, particularly for applications in text generation and data synthesis.
  • Advanced knowledge of statistical techniques: hypothesis testing, sampling theory, experimental design, and causal inference.
  • Strong knowledge of a variety of ML techniques (e.g., clustering, regression, neural networks, etc.) and their real-world trade-offs.
  • Expert proficiency in Python (DS/ML stack) and experience with frameworks used for LLM development and fine-tuning.
  • Advanced SQL skills and experience working with large-scale databases.
  • Ability to research and adopt new methods while mentoring junior colleagues on their application., * Highly accountable self-starter and quick learner, consistently motivated to deliver high-quality, impactful results.
  • Strong data-driven mindset with the ability to translate abstract business requests into actionable AI initiatives and solutions.
  • Excellent written and verbal communication skills, with the ability to explain and defend complex AI concepts to non-experts., * Direct experience with Synthetic Data Generation techniques and the evaluation of synthetic data quality/utility.
  • Experience with Prompt Engineering, RAG (Retrieval-Augmented Generation), or fine-tuning open-source LLMs for open-end generation.
  • Experience with probabilistic modeling, or advanced profiling techniques.
  • Familiarity with online market research or survey exchange platforms.
  • Experience using Databricks, Spark, or PySpark for large-scale workflows.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

4:01 min

Managing application isolation via pluggable database models

Wei Hu Wei Hu · World Congress 2022

Videos

See all

Related articles

See all