Data Scientist (ML / NLP / LLMs)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+16 more
Job description
At the heart of STG is a natural language and machine learning system that turns unstructured, deeply human reflection text into structured, trustworthy insights - reading the quality of student reflection and the strength of the student-educator relationship, and surfacing insights that help educators prioritize the kids who need support most. Built on research-validated models and extended with modern AI, our work serves one guiding philosophy: make this routine lasting and sustainable for all users by making it more engaging, meaningful, and efficient. Weâve spent years turning these insights into real-time, user-facing features and hardening the technical backbone to serve millions of students at low latency. Now weâre pushing that foundation further with the latest in AI - building richer, more contextual, AI- and expert-guided support that helps teachers respond to students more effectively, and continually expanding the range of problems we take on. Weâre looking for a Data Scientist who can move fluently across this whole surface - from feature engineering and model validation to production ML systems, and from classical NLP into the responsible application of modern LLMs. Youâll join a small data science / machine learning team - small enough that youâll know everyoneâs name and see your work ship in weeks, not quarters - working closely with product and engineering. Weâre deliberate about our toolkit: classical, feature-driven ML is the validated core of what we do today and isnât going anywhere, while modern LLMs open new frontiers weâre actively investing in. We donât reach for the trendiest technique or cling to the familiar one - we choose the right tool for each problem, and weâre looking for someone who enjoys making that call with us. Your work will directly shape how educators understand and support students, at a scale that reaches historically underserved communities., + Build and refine NLP/ML models over large volumes of unstructured student and educator text, using both classical approaches (feature engineering, classification, CNNs, ensemble methods) and modern LLM-based techniques where theyâre the right tool.
- Partner with data science, product, and engineering to identify, define, and test opportunities to improve the product through ML/NLP - from turning existing models into user-facing features to prototyping entirely new capabilities.
- Extend our models beyond âproof of conceptâ scope: new reflection prompts, younger students etc, widening accessibility while protecting accuracy.
- Design and apply LLMs responsibly for generative and assistive features - for example, contextualized teacher-response suggestions, and resource recommendations - with careful attention to prompting, retrieval, grounding, evaluation, and guardrails.
- Build validation, monitoring, and retraining pipelines in partnership with the ML engineering team - including data-drift detection - so models keep performing as code and data change.
- Undertake preprocessing of structured and unstructured data, and build reliable, reproducible feature and evaluation workflows.
Why join us A student writes a few sentences on a Tuesday about a week thatâs been hard. A teacher reads it and writes back. A counselor checks in too, if more support is needed. Everything we build serves that small loop - making kids feel seen, heard, and supported by the educators around them, reliably, at the scale of hundreds of thousands of students. As a Data Scientist at Sown To Grow, youâll be shaping the DS/ML foundation the entire product rests on, not inheriting someone elseâs. The insights you build get used by real educators, about real kids, in the week theyâre written. Thatâs a rare combination of technical depth and immediate human consequence. Youâll also have unusual latitude for a team our size: room to shape the technical direction, choose the right approaches, and see your work reach real classrooms rather than sit in a notebook.
Requirements
- Passion for impact. You want your models to matter in the lives of real students, not just move a metric on a dashboard.
- Comfort in ambiguity. In a dynamic startup, the path isnât always clear. You can shape a strategy on the fly and execute with confidence and autonomy.
- Sense of urgency. You enjoy moving fast and building with purpose - and you know when to slow down and get something right, especially when it involves kids.
- Aspiration for growth. You may have experience at larger organizations, but youâre ready for more responsibility and the chance to shape what comes next.
- Rigor and responsibility. You care about doing ML well - validating claims, understanding failure modes, and thinking hard about fairness, privacy, and unintended consequences, especially when the data belongs to children.
- Range and judgment. Youâre fluent in both worlds - classical, feature-driven ML and modern LLMs - and you understand the fundamentals of each well enough to know where each shines. You reach for the approach that fits the problem, not the one thatâs trending., + Bachelorâs or higher degree in Computer Science, Data Science, Machine Learning, Math, Statistics, or a related field.
- 2+ years of experience as a Data Scientist, ML Engineer, or Data Engineer, solving real-world problems with machine learning.
- Strong proficiency in Python and the ML stack (pandas, numpy, scikit-learn; PyTorch or TensorFlow; Spark a plus).
- Experience building and deploying ML solutions that involve natural language processing of text data.
- Working knowledge of core ML techniques such as classification, clustering, prediction, recommender systems, and anomaly detection.
- Working knowledge of the complete machine learning lifecycle - data, training, validation, deployment, monitoring, and retraining.
- Solid understanding of how modern LLMs work under the hood - transformer architecture, training and fine-tuning, tokenization, embeddings. We care more about curiosity than credentials here: you enjoy digging into why a model behaves the way it does, not just what it returns. Hands-on with at least one of prompting/evaluation, fine-tuning, retrieval-augmented generation (RAG), or agentic/tool-use patterns, with a thoughtful view of when LLMs are and arenât the right approach.
- Experience writing and maintaining high-quality production code, and comfort with Git-based workflows., + Strong interest in working in education technology in an impact-driven, mission-first role.
- Experience productionizing ML for real-time, low-latency inference (e.g., AWS SageMaker or comparable), including containerization and CI/CD.
- Experience building data-drift detection, model monitoring, and automated retraining systems in partnership with ML engineering teams.
- Experience building, training, or fine-tuning language models from the ground up - e.g., implementing transformer components, training or adapting models on domain-specific data, or working with open-weight models beyond off-the-shelf APIs. This is a longer-term direction for us, and we value candidates who can grow into it.
- Experience with responsible / trustworthy AI: fairness and bias evaluation, privacy-conscious handling of sensitive data, and building guardrails for user-facing generative features.
- Experience designing human-in-the-loop evaluation and running online experiments (A/B testing, feature flagging).
- Familiarity with the practical, ethical, and legal considerations of working with student data. If you donât check every box, weâd still love to hear from you. Some of the strongest people on our team grew into parts of their role after they arrived.
Benefits & conditions
- Competitive compensation with performance-based incentives and meaningful equity
- Comprehensive health and wellness benefits for you and your family
- Flexible work arrangements and a genuine commitment to work-life balance
- Real pathways for growth - as the platform and the data team expand, so does the scope of this role
- A collaborative, mission-driven community where every voice is heard, Full-time base salary range = $120,000 - $135,000, plus performance-based bonuses and equity compensation (incentive stock options / ISOs).
About the company
Sown To Grow (STG) is a K12 education technology platform that empowers schools to improve student social, emotional, and academic health through an easy and engaging check-in and reflection process. In a short weekly routine, students check in on how they are feeling and reflect on the strategies that are working best for them (or new ones to try), and teachers respond with support and coaching. School leaders and student support staff use real-time reporting on studentsâ emotions and reflections to proactively intervene to support student needs. The system also includes built-in screeners, supporting curriculum, and powerful artifacts of growth., We understand schools. Our founding team spent years working in schools and districts, and our team today includes educators, social workers, and experienced administrators. We work alongside schools every day, including some of the largest school districts in the country (e.g., Los Angeles Unified, Metro Nashville Public Schools, New York City, Metro Madison Public Schools, Oakland Unified, and many more).
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.wayup.comGood distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
MLOps â Whatâs the deal behind it?
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?