Principle Data Scientist
Role details
Job location
Tech stack
Job description
- Technical direction on the analytical methods, graph libraries, and tooling we build with partnering with engineering on the underlying data architecture
- Cross-functional alignment between product goals and our data capabilities
- Data science standards, tools, and methodology across Provenir AI - how we experiment, validate, and ship
- Pragmatic use of LLMs and agentic tooling to add value, whether internally for productivity or directly in the product
Your Responsibilities
- Design and evolve a deterministic identity graph that represents key entities, relationships, and business logic
- Identify and prioritise the features, signals, and patterns in our connected data that create product value
- Build and guide analytics pipelines for link analysis, identity resolution, clustering, ranking, anomaly detection, and discovery
- Partner with data engineers to define ingestion, enrichment, validation, and publishing workflows into the identity data
- Work with product managers and stakeholders to translate ambiguous business problems into product features
- Prototype and evaluate approaches for identity resolution, identity intelligence, fraud detection, and risk intelligence
- Partner with data engineers to define the quality and coverage signals that matter for trustworthy identity intelligence - coverage, match confidence, and accuracy
- Select the right technologies and patterns, and create clear documentation and decision frameworks so the work can be reused across teams
- Mentor engineers and data scientists on graph methods, modelling, and analytical techniques and grow a team as value is proven, You'll be the most senior data scientist in the Provenir AI team, which means real independence and impact, but also means you won't have a large peer group of data scientists around you initially. If you thrive when given ownership and like building things from scratch, this could be ideal.
Much of the work is about making relationship data better, more connected, and more usable not just building models for their own sake. Credibility comes from shipping production capabilities that improve the product, working very closely with software and data engineers.
The identity graph and Living Identity® data are the priority and where you'll make your mark first. Over time, as the most senior data scientist, you'll also shape how data science is done across Provenir AI the tools, methods, and pragmatic use of LLMs and agents that add value, whether internally or in the product.
Interview Process
The Interview Process Would Typically Be Structured As Follows
Team Fit. A conversation with our CTO to explore how you work, what you're looking for, and how you'd fit with the team
Case Study. We'll share a realistic identity/graph data problem for you to work through ahead of the next round
Technical Interview. With data science and/or engineering leadership - you'll present and discuss your case study, and we'll explore your approach to graph and identity problems, production systems, and team building
Final Interview. A conversation with our CEO on vision, impact, and how you'd approach the first 90 days
Your Benefits
- Comprehensive private health cover and wellness plans
- Flexible and remote-friendly opportunities
- Maternity/paternity leave
- Retirement benefits such as pension contributions to plan for your future
- Macbook Pro
Our employees are our top priority; we offer comprehensive health and wellness plans. You will enjoy paid time off and company holidays, flexible and remote-friendly opportunities, and maternity/paternity leave.
Requirements
- Strong experience in graph data science, knowledge graphs, or graph/network analytics
- Hands-on experience building and working with graph structures using libraries or databases (e.g. Neo4j, Amazon Neptune, JanusGraph, GraphFrames, NetworkX, or similar)
- Experience with identity resolution, entity resolution, or link analysis
- Strong Python and SQL skills
- A genuine product mindset, you start from the business problem and the product outcome, not the technique, and you can prioritise accordingly
- Ability to design analytical approaches from ambiguous problems
- Track record of operating in production environments with quality, scale, and reliability requirements
- Strong communication skills and comfort working embedded with product and engineering teams
- Although not essential, it would be great if you have experience with:
- Building graph-enabled products in identity, fraud, financial crime, risk, or AI-driven decisioning
- Graph embeddings or other graph-native ML methods
- Ontology design, semantic modelling, or taxonomy development
- Large-scale distributed data processing (e.g. Spark, Dask)
- Building and growing data science teams
- Working in scale-up or early-stage environments