> Markdown version of [/jobs/ext/2004896-data-scientist-ml-nlp-llms](https://www.wearedevelopers.com/jobs/ext/2004896-data-scientist-ml-nlp-llms). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist (ML / NLP / LLMs) - **Company:** Sown to Grow - **Location:** United States - **Experience:** Experienced - **Salary:** $120,000.0 - $135,000.0 - **Contract:** Permanent contract - **Skills:** A/B Testing, Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Cluster Analysis, Continuous Integration, Python (Programming Language), Machine Learning, Language Modeling, Natural Language Processing, NumPy, Performance Tuning, Recommender Systems, Tensorflow, Unstructured Data, Feature Engineering, Pytorch, Large Language Models, Apache Spark, Model Validation, Pandas, Containerization, Git Flow, Scikit Learn, Information Technology, Low Latency, Production Code, Machine Learning Operations - **Published:** August 9, 2026 - **Apply:** https://www.wayup.com/i-j-Data-Scientist-ML-NLP-LLMs-Sown-To-Grow-159951921362700/ ## About the Role + Passion for impact. You want your models to matter in the lives of real students, not just move a metric on a dashboard. + Comfort in ambiguity. In a dynamic startup, the path isn't always clear. You can shape a strategy on the fly and execute with confidence and autonomy. + Sense of urgency. You enjoy moving fast and building with purpose - and you know when to slow down and get something right, especially when it involves kids. + Aspiration for growth. You may have experience at larger organizations, but you're ready for more responsibility and the chance to shape what comes next. + Rigor and responsibility. You care about doing ML well - validating claims, understanding failure modes, and thinking hard about fairness, privacy, and unintended consequences, especially when the data belongs to children. + Range and judgment. You're fluent in both worlds - classical, feature-driven ML and modern LLMs - and you understand the fundamentals of each well enough to know where each shines. You reach for the approach that fits the problem, not the one that's trending., + Bachelor's or higher degree in Computer Science, Data Science, Machine Learning, Math, Statistics, or a related field. + 2+ years of experience as a Data Scientist, ML Engineer, or Data Engineer, solving real-world problems with machine learning. + Strong proficiency in Python and the ML stack (pandas, numpy, scikit-learn; PyTorch or TensorFlow; Spark a plus). + Experience building and deploying ML solutions that involve natural language processing of text data. + Working knowledge of core ML techniques such as classification, clustering, prediction, recommender systems, and anomaly detection. + Working knowledge of the complete machine learning lifecycle - data, training, validation, deployment, monitoring, and retraining. + Solid understanding of how modern LLMs work under the hood - transformer architecture, training and fine-tuning, tokenization, embeddings. We care more about curiosity than credentials here: you enjoy digging into why a model behaves the way it does, not just what it returns. Hands-on with at least one of prompting/evaluation, fine-tuning, retrieval-augmented generation (RAG), or agentic/tool-use patterns, with a thoughtful view of when LLMs are and aren't the right approach. + Experience writing and maintaining high-quality production code, and comfort with Git-based workflows., + Strong interest in working in education technology in an impact-driven, mission-first role. + Experience productionizing ML for real-time, low-latency inference (e.g., AWS SageMaker or comparable), including containerization and CI/CD. + Experience building data-drift detection, model monitoring, and automated retraining systems in partnership with ML engineering teams. + Experience building, training, or fine-tuning language models from the ground up - e.g., implementing transformer components, training or adapting models on domain-specific data, or working with open-weight models beyond off-the-shelf APIs. This is a longer-term direction for us, and we value candidates who can grow into it. + Experience with responsible / trustworthy AI: fairness and bias evaluation, privacy-conscious handling of sensitive data, and building guardrails for user-facing generative features. + Experience designing human-in-the-loop evaluation and running online experiments (A/B testing, feature flagging). + Familiarity with the practical, ethical, and legal considerations of working with student data. If you don't check every box, we'd still love to hear from you. Some of the strongest people on our team grew into parts of their role after they arrived. ## Description At the heart of STG is a natural language and machine learning system that turns unstructured, deeply human reflection text into structured, trustworthy insights - reading the quality of student reflection and the strength of the student-educator relationship, and surfacing insights that help educators prioritize the kids who need support most. Built on research-validated models and extended with modern AI, our work serves one guiding philosophy: make this routine lasting and sustainable for all users by making it more engaging, meaningful, and efficient. We've spent years turning these insights into real-time, user-facing features and hardening the technical backbone to serve millions of students at low latency. Now we're pushing that foundation further with the latest in AI - building richer, more contextual, AI- and expert-guided support that helps teachers respond to students more effectively, and continually expanding the range of problems we take on. We're looking for a Data Scientist who can move fluently across this whole surface - from feature engineering and model validation to production ML systems, and from classical NLP into the responsible application of modern LLMs. You'll join a small data science / machine learning team - small enough that you'll know everyone's name and see your work ship in weeks, not quarters - working closely with product and engineering. We're deliberate about our toolkit: classical, feature-driven ML is the validated core of what we do today and isn't going anywhere, while modern LLMs open new frontiers we're actively investing in. We don't reach for the trendiest technique or cling to the familiar one - we choose the right tool for each problem, and we're looking for someone who enjoys making that call with us. Your work will directly shape how educators understand and support students, at a scale that reaches historically underserved communities., + Build and refine NLP/ML models over large volumes of unstructured student and educator text, using both classical approaches (feature engineering, classification, CNNs, ensemble methods) and modern LLM-based techniques where they're the right tool. + Partner with data science, product, and engineering to identify, define, and test opportunities to improve the product through ML/NLP - from turning existing models into user-facing features to prototyping entirely new capabilities. + Extend our models beyond "proof of concept" scope: new reflection prompts, younger students etc, widening accessibility while protecting accuracy. + Design and apply LLMs responsibly for generative and assistive features - for example, contextualized teacher-response suggestions, and resource recommendations - with careful attention to prompting, retrieval, grounding, evaluation, and guardrails. + Build validation, monitoring, and retraining pipelines in partnership with the ML engineering team - including data-drift detection - so models keep performing as code and data change. + Undertake preprocessing of structured and unstructured data, and build reliable, reproducible feature and evaluation workflows. Why join us A student writes a few sentences on a Tuesday about a week that's been hard. A teacher reads it and writes back. A counselor checks in too, if more support is needed. Everything we build serves that small loop - making kids feel seen, heard, and supported by the educators around them, reliably, at the scale of hundreds of thousands of students. As a Data Scientist at Sown To Grow, you'll be shaping the DS/ML foundation the entire product rests on, not inheriting someone else's. The insights you build get used by real educators, about real kids, in the week they're written. That's a rare combination of technical depth and immediate human consequence. You'll also have unusual latitude for a team our size: room to shape the technical direction, choose the right approaches, and see your work reach real classrooms rather than sit in a notebook. ## Related Videos - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [How to implement convenient Python bindings to C++](https://www.wearedevelopers.com/videos/618-how-to-implement-convenient-python-bindings-to-c) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How to start an AI project for a good cause and boost your career](https://www.wearedevelopers.com/magazine/15-how-to-start-an-ai-project-for-a-good-cause-and-boost-your-career)