> Markdown version of [/jobs/ext/1183641-data-scientist-ai-llm](https://www.wearedevelopers.com/jobs/ext/1183641-data-scientist-ai-llm). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - AI / LLM - **Company:** DIAGONAL MATRIX LLC - **Location:** United States (Remote available) - **Salary:** $120,000.0 - $170,000.0 - **Contract:** Internship / Graduate position - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Data Analysis, Computer Vision, Microsoft Azure, Cloud Computing, Computer Programming, Data Cleansing, Information Extraction, Python (Programming Language), Machine Learning, Natural Language Processing, NumPy, Recommender Systems, Power BI, Standard Sql, Software Safety, Search Technologies, SQL Databases, Tableau (Software), Unstructured Data, Supervised Learning, Google Cloud, Data Storage Technologies, Feature Engineering, Chatbots, Retrieval-Augmented Generation, Large Language Models, Model Validation, Generative AI, Git, Fastapi, Pandas, Matplotlib, Question Answering, AI Platforms, Scikit Learn, Information Technology, HuggingFace, Machine Learning Operations, Streamlit Framework, Looker Analytics, Docker, Unsupervised Learning - **Published:** July 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=5a4eea41277bdff1 ## About the Role We are seeking a motivated Data Scientist with AI and Large Language Model experience to join our growing technology team in the USA. This role is ideal for recent graduates or early-career professionals who are passionate about data science, machine learning, artificial intelligence, and LLM-based applications., Bachelor's or Master's degree in Data Science, Computer Science, Artificial Intelligence, Machine Learning, Statistics, Mathematics, Engineering, or a related field. Recent graduate or early-career professional with academic, internship, project, or commercial experience in data science or AI. Strong understanding of machine learning concepts such as supervised learning, unsupervised learning, classification, regression, clustering, and model evaluation. Good programming skills in Python. Good knowledge of SQL for querying and analysing data. Experience with common data science libraries such as Pandas, NumPy, Scikit-learn, Matplotlib, or similar tools. Understanding of data preprocessing, feature engineering, model training, and evaluation techniques. Basic understanding of Large Language Models, Generative AI, NLP, embeddings, vector databases, or RAG pipelines. Strong analytical thinking and problem-solving skills. Good communication skills with the ability to explain technical findings in a simple and clear way. Preferred Qualifications Hands-on experience with OpenAI, Azure OpenAI, Anthropic, Hugging Face, LangChain, LlamaIndex, or similar AI/LLM tools. Experience working with vector databases such as Pinecone, FAISS, ChromaDB, Weaviate, or Qdrant. Experience with cloud platforms such as AWS, Azure, or Google Cloud. Knowledge of data visualisation tools such as Power BI, Tableau, Looker, or Streamlit. Experience with Git, APIs, Docker, FastAPI, or basic software development practices. Academic or personal projects involving chatbots, document Q&A systems, recommendation engines, forecasting, NLP, computer vision, or predictive analytics. Understanding of responsible AI, bias, model explainability, data privacy, and AI safety principles. Visa / Work Authorization We welcome applications from candidates who are currently authorised to work in the United States, including eligible F-1 students and recent graduates on CPT, OPT, or STEM OPT, subject to applicable university and immigration rules. Candidates must have valid work authorization before starting employment. For CPT candidates, university approval and CPT authorization must be completed before the employment start date. ## Description The successful candidate will work on real-world AI and data science projects involving data analysis, predictive modelling, natural language processing, generative AI, Retrieval-Augmented Generation, and LLM-powered automation. This is an excellent opportunity for candidates on OPT, CPT, or STEM OPT who want to build strong commercial experience in the AI and data science industry., Design, develop, and evaluate machine learning models for business and customer-focused use cases. Work with structured and unstructured datasets to identify patterns, trends, insights, and predictive signals. Build data science solutions using Python, SQL, Pandas, NumPy, Scikit-learn, and related libraries. Support the development of AI and LLM-based applications using tools and frameworks such as OpenAI, Azure OpenAI, LangChain, Hugging Face, and vector databases. Assist in building Retrieval-Augmented Generation pipelines using embeddings, document chunking, vector search, and semantic retrieval. Develop and test prompts for LLM applications, including summarisation, classification, question answering, information extraction, and chatbot workflows. Collaborate with data engineers, software developers, and business stakeholders to understand requirements and convert them into data-driven solutions. Perform exploratory data analysis, data cleaning, feature engineering, model training, validation, and performance evaluation. Work with cloud platforms such as AWS, Azure, or Google Cloud for data storage, AI services, and model deployment. Create dashboards, reports, and visualisations to communicate insights clearly to technical and non-technical audiences. Monitor model performance and support improvements to accuracy, reliability, explainability, and responsible AI practices. Maintain clear documentation of models, datasets, experiments, assumptions, and technical processes. ## Related Videos - [How to Avoid LLM Pitfalls - Mete Atamel and Guillaume Laforge](https://www.wearedevelopers.com/videos/1328-how-to-avoid-llm-pitfalls-mete-atamel-and-guillaume-laforge) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How to start an AI project for a good cause and boost your career](https://www.wearedevelopers.com/magazine/15-how-to-start-an-ai-project-for-a-good-cause-and-boost-your-career) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts)