Machine Learning Engineer for Educational Assessment
Role details
Job location
Tech stack
Job description
You will contribute to the research, development and productionisation of AI capabilities for educational assessment. These may include automated marking, feedback generation, learner support, skill estimation, proficiency modelling and adaptive testing., · Design, develop and refine machine learning models and prototypes that support educational assessment, helping to translate assessment needs into practical AI solutions and providing evidence for future development decisions.
· Evaluate model performance against technical and assessment measures, including accuracy, fairness, bias, reliability and alignment with human marking standards, while ensuring methods and results are clearly documented and reproducible.
· Work collaboratively with AI researchers, developers, psychometricians and product teams to build, deploy and continuously improve machine learning solutions, developing robust engineering practices and end-to-end experience across the full AI lifecycle., Unlike many early-career research roles, you will have the opportunity to follow promising work beyond the prototype: learning how models are evaluated, engineered, integrated and deployed within real products and services. You will gain practical experience across the full machine learning lifecycle while working with experienced specialists in AI, software development, product development, psychometrics and educational assessment.
You will:
- Contribute to AI capabilities in areas such as automated marking, personalised feedback, item generation and learner support;
- Develop practical experience of both applied ML research and production engineering;
- Learn how responsible AI capabilities are tested, deployed, monitored and improved;
- See how your work contributes to products and services used to address meaningful educational needs.
- A 35-hour working week with flexible, hybrid working
- 25 days' annual leave (rising to 30), plus Christmas closure days
- Excellent pension (up to 11.5% employer contribution) etc, Accountable to the Head of AI for Assessment Innovation, the overall purpose of this role is to develop models and algorithms as required by new assessment products and services.
The post holder will research and develop AI capabilities that can enable new assessment products, increase the breadth of assessment services on offer and help shape long-term tech innovation and solutions. They will ideate and develop proofs of concept and prototypes and ensure they are cutting-edge, relevant and fit-for-purpose.Research, development and evaluation of AI solutions for assessment are key enablers in a range of diversification, digitisation and customer programmes. The AI for Assessment Innovation team is AQA's in-house AI for assessment lab, providing services and solutions alongside and in collaboration with contractors and partners. The team's responsibilities are: Research and development of AI features for new products or as part of contracted services EdTech partnership support through targeted evaluations and testing of third-party AI tools Providing AI for assessment expertise to the whole AQA group and advancing AQA's knowledge and know-how
The role sits within the AI for Assessment Innovations team, in the Assessment Research and Innovation business area. Reporting to the Head of AI for Assessment, the role collaborates with a team of AI researchers, developers and managers and will have line management responsibility for AI for Assessment apprentices.
Activities: AI model development for assessmentDesign, build, and refine machine learning models that support educational assessment use cases, such as automated marking (e.g., essays, short answers), feedback generation and learner support, skill estimation, proficiency modelling, and adaptive testing. Select appropriate modelling approaches (e.g., NLP models, classical ML, or deep learning) based on pedagogical and product requirements. Conduct rigorous experimentation, including hyperparameter tuning and ablation studies, to improve model performance and fairness.
Educational assessment researchWork with complex educational datasets (e.g., learner responses, interaction logs, assessment outcomes). Design evaluation frameworks that go beyond accuracy to include fairness and bias across learner groups, marking reliability and consistency, alignment with human marking standards and mark schemes Work closely with psychometricians, assessment experts and product teams to translate educational requirements into technical solutions. Incorporate domain knowledge (e.g., marking schemes, assessment objectives, curriculum standards) into model design.
From prototype to operationalisationDevelop scalable pipelines for data processing, model training, validation, and deployment. Collaborate and support the teams responsible for integrating models into production systems Contribute to CI/CD workflows, model versioning, and reproducibility practices.
Responsible AI and governanceIdentify, assess, and mitigate risks related to bias, fairness, and misuse in assessment AI systems. Contribute to the development of explainable and transparent AI systems suitable for high-stakes exams or classroom use. Work with the relevant AQA teams to ensure compliance with relevant regulatory and ethical standards in education.
Documentation and knowledge sharingDocument model architectures, decisions, evaluation results, and limitations. Communicate findings clearly to both technical and non-technical stakeholders. Contribute to internal best practices, reusable components, and knowledge sharing across teams.
To be successful in this role, you will need to demonstrate:
Essential MotivationA keen interest in the education or educational assessment sector, and a drive to furthering AQA's mission.
Requirements
- Strong Python skills, with practical experience using relevant data and machine learning libraries such as NumPy, Pandas and scikit-learn.
- Practical experience with at least one deep-learning framework, such as PyTorch or TensorFlow.
- A good foundation in machine learning, including supervised learning and model evaluation.
- A good understanding of the machine learning lifecycle, from data ingestion and cleaning through to model development and validation.
- An interest in education or educational assessment and motivation to apply technology in support of AQA's mission.
- Strong communication and collaboration skills, including the ability to explain technical ideas clearly and learn from colleagues across different disciplines.
Desirable
-
NLP knowledge or experience relevant to text-based assessment.
-
Familiarity with NLP libraries or frameworks such as Hugging Face Transformers or spaCy.
-
Experience of, or exposure to, building end-to-end machine learning systems, including deployment.
-
Familiarity with software-engineering practices such as version control, testing, code review and technical documentation.
-
Exposure to sequential modelling.
-
Experience working with multimodal data, including data processing and synchronisation., * Please include a link to a machine learning project you can share with us, such as a GitHub or other accessible repository, that showcases relevant technical skills for this role.
-
Stage 1: a 30-minute Teams interview where you will talk through the shared project or artefact and discuss the technical decisions behind it.
-
Stage 2: a face-to-face interview in Manchester or Milton Keynes, focused on your wider professional experience, collaboration style and motivation for educational assessment., Machine learning and NLP expertiseUnderstanding of machine learning techniques, including supervised learning, model evaluation and optimisation Natural Language Processing (NLP) for text-based assessment and some knowledge of multi-modal models Experience building end-to-end ML systems from data ingestion to deployment. Familiarity with model interpretability techniques (e.g., SHAP, LIME).
Engineering Proficient in Python and core ML/data libraries (e.g., PyTorch/TensorFlow, Scikit-learn, Pandas). Knowledge or experience with production systems such as API development, containerisation and cloud platforms Solid understanding of software engineering practices: version control, testing, modular design.
Research skillsExperience working with real-world datasets, including noisy or incomplete data. Understanding of evaluation methodologies, particularly in contexts where ground truth may be subjective (e.g. human marking).
Analytical and problem-solving skillsAbility to translate ambiguous, domain-specific problems into structured ML solutions. Strong critical thinking when interpreting model outputs in high-stakes contexts. Attention to detail, particularly when working with sensitive learner data and evaluation outcomes.
Communication and CollaborationAbility to work effectively in multidisciplinary teams. Strong communication skills, including explaining technical concepts to educators and non-technical stakeholders. Experience contributing to collaborative development environments.
Education and ExperienceBachelor's or Master's degree in Computer Science, Machine Learning, Data Science, or a related field, or equivalent experience Hands-on experience in machine learning or applied AI, acquired in a range of contexts
Desirable Interest in or experience with education technology, assessment systems, or learning analytics. Awareness of concepts relevant to assessment such as reliability and validity, computerised adaptive testing (CAT), Item response theory (IRT) or similar psychometric models Experience in education, assessment, or a related