Data Scientist - AI / LLM

DIAGONAL MATRIX LLC
United States
about 1 month ago

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Compensation
$120,000.0 - $170,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Data Analysis Computer Vision Microsoft Azure Cloud Computing Computer Programming Data Cleansing Information Extraction Python (Programming Language) Machine Learning
+33 more
Natural Language Processing NumPy Recommender Systems Power BI Standard Sql Software Safety Search Technologies SQL Databases Tableau (Software) Unstructured Data Supervised Learning Google Cloud Data Storage Technologies Feature Engineering Chatbots Retrieval-Augmented Generation Large Language Models Model Validation Generative AI Git Fastapi Pandas Matplotlib Question Answering AI Platforms Scikit Learn Information Technology HuggingFace Machine Learning Operations Streamlit Framework Looker Analytics Docker Unsupervised Learning

Job description

The successful candidate will work on real-world AI and data science projects involving data analysis, predictive modelling, natural language processing, generative AI, Retrieval-Augmented Generation, and LLM-powered automation. This is an excellent opportunity for candidates on OPT, CPT, or STEM OPT who want to build strong commercial experience in the AI and data science industry., Design, develop, and evaluate machine learning models for business and customer-focused use cases.

Work with structured and unstructured datasets to identify patterns, trends, insights, and predictive signals.

Build data science solutions using Python, SQL, Pandas, NumPy, Scikit-learn, and related libraries.

Support the development of AI and LLM-based applications using tools and frameworks such as OpenAI, Azure OpenAI, LangChain, Hugging Face, and vector databases.

Assist in building Retrieval-Augmented Generation pipelines using embeddings, document chunking, vector search, and semantic retrieval.

Develop and test prompts for LLM applications, including summarisation, classification, question answering, information extraction, and chatbot workflows.

Collaborate with data engineers, software developers, and business stakeholders to understand requirements and convert them into data-driven solutions.

Perform exploratory data analysis, data cleaning, feature engineering, model training, validation, and performance evaluation.

Work with cloud platforms such as AWS, Azure, or Google Cloud for data storage, AI services, and model deployment.

Create dashboards, reports, and visualisations to communicate insights clearly to technical and non-technical audiences.

Monitor model performance and support improvements to accuracy, reliability, explainability, and responsible AI practices.

Maintain clear documentation of models, datasets, experiments, assumptions, and technical processes.

Requirements

We are seeking a motivated Data Scientist with AI and Large Language Model experience to join our growing technology team in the USA. This role is ideal for recent graduates or early-career professionals who are passionate about data science, machine learning, artificial intelligence, and LLM-based applications., Bachelor’s or Master’s degree in Data Science, Computer Science, Artificial Intelligence, Machine Learning, Statistics, Mathematics, Engineering, or a related field.

Recent graduate or early-career professional with academic, internship, project, or commercial experience in data science or AI.

Strong understanding of machine learning concepts such as supervised learning, unsupervised learning, classification, regression, clustering, and model evaluation.

Good programming skills in Python.

Good knowledge of SQL for querying and analysing data.

Experience with common data science libraries such as Pandas, NumPy, Scikit-learn, Matplotlib, or similar tools.

Understanding of data preprocessing, feature engineering, model training, and evaluation techniques.

Basic understanding of Large Language Models, Generative AI, NLP, embeddings, vector databases, or RAG pipelines.

Strong analytical thinking and problem-solving skills.

Good communication skills with the ability to explain technical findings in a simple and clear way.

Preferred Qualifications

Hands-on experience with OpenAI, Azure OpenAI, Anthropic, Hugging Face, LangChain, LlamaIndex, or similar AI/LLM tools.

Experience working with vector databases such as Pinecone, FAISS, ChromaDB, Weaviate, or Qdrant.

Experience with cloud platforms such as AWS, Azure, or Google Cloud.

Knowledge of data visualisation tools such as Power BI, Tableau, Looker, or Streamlit.

Experience with Git, APIs, Docker, FastAPI, or basic software development practices.

Academic or personal projects involving chatbots, document Q&A systems, recommendation engines, forecasting, NLP, computer vision, or predictive analytics.

Understanding of responsible AI, bias, model explainability, data privacy, and AI safety principles.

Visa / Work Authorization

We welcome applications from candidates who are currently authorised to work in the United States, including eligible F-1 students and recent graduates on CPT, OPT, or STEM OPT, subject to applicable university and immigration rules.

Candidates must have valid work authorization before starting employment. For CPT candidates, university approval and CPT authorization must be completed before the employment start date.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:59 min

Building culturally aware LLMs for global audiences

Werner Vogels Werner Vogels +1 · WWC Europe 2026

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

3:13 min

Addressing language barriers and the data science talent deficit

Markus Hacker Markus Hacker +3 · WWC 2024

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all