AI Data Scientist- RAG, SLM & Distributed Data...

Insight Global
Hartford, CT, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
1 year minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Automated Storage and Retrieval Systems Databases Distributed Data Store Distributed Systems Data Flow Control Python (Programming Language) Language Modeling Search Technologies Software Engineering SQL Databases
+9 more
Google Cloud Cloud Platform System Sql Optimization Flask (Web Framework) Large Language Models Prompt Engineering Generative AI Fastapi Api Design

Job description

We are looking for a mid-level AI Engineer with hands-on experience in Retrieval-Augmented Generation (RAG)systems, Small Language Models (SLMs), and distributed databases such as Google Cloud Spanner.

You will work closely with senior engineers and product teams to build scalable AI systems that integrate retrieval pipelines, language models, and distributed transactional infrastructure. This role is ideal for someone who has already built AI features in production and wants to deepen their expertise in applied GenAI systems.

Requirements

AI RAG pipelines, Embeddings, Prompt engineering

Models SLM/LLM integration

Database Spanner schema design, SQL optimization

Backend Python, APIs

Cloud GCP, * 3-5 years of software engineering experience.

  • 1-2 years working with LLM or RAG-based systems.

  • Strong proficiency in Python.

  • Experience with:

o Embedding models and vector search

o LangChain, LlamaIndex, or similar frameworks

o API development (FastAPI/Flask)

  • Experience working with Google Cloud Spanner or similar distributed SQL databases.

  • Solid understanding of distributed systems fundamentals.

  • Comfortable working in cloud environments (GCP preferred). * Experience fine-tuning or quantizing small language models.

  • Familiarity with evaluation metrics for retrieval systems (Recall@K, etc.).

  • Knowledge of:

o Vertex AI

o Pub/Sub

o Dataflow

  • Experience optimizing AI inference for cost and latency.

  • Exposure to CI/CD pipelines.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

3:33 min

Connecting frontends via a FastAPI proxy backend layer

Saoussen Chaabnia Saoussen Chaabnia · Europe 2026 Virtual

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

3:10 min

Understanding the core concepts of API design

Alen Pokos · LIVE

1:50 min

Simplifying generative AI deployments using the RagStack opinionated framework

David Leconte David Leconte +1 · WWC 2024

1:42 min

Introduction to the fast API web framework

Sebastián Ramírez · WWC 2022

Videos

See all

Related articles

See all