AI Engineer

Educational Testing Service
Concord, NH, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Training Data Artificial Intelligence Python (Programming Language) Machine Learning Natural Language Processing Performance Tuning Tensorflow Strategies of Testing Feature Engineering Pytorch Large Language Models Deep Learning
+7 more
Model Validation Generative AI Kubernetes Information Technology Machine Learning Operations Document Classification Data Pipelines

Job description

The AI Model Development Engineer (open-rank) supports the TOEFL and GRE assessment programs by designing, developing, evaluating, and deploying machine learning models that power ETS’s next generation of assessment technologies. This role focuses significantly on building and scaling AI-driven scoring systems for constructed responses, including essays, spoken responses, short answers, and simulations.

Operating at the intersection of AI engineering, assessment science, and operational delivery, the role ensures models are accurate, fair, explainable, and production-ready. The position contributes to advancing ETS’s legacy in automated scoring, measurement, and assessment innovation. The ideal candidate brings strong applied machine learning expertise, along with experience in model evaluation, data pipelines, and quality controls within high-stakes or regulated environments.

Primary Responsibilities

  • Develop, train, and optimize machine learning and deep learning models for applications such as automated scoring (including text, speech, or multimodel responses), item generation, content classification, anomaly detection, and personalization.
  • Implement feature engineering, representation learning, and model architectures appropriate for scoring and classification tasks.
  • Develop hybrid scoring approaches combining AI models with rules-based or human-in-the-loop workflows.
  • Use NLP, large language models, and multimodal modeling techniques to support assessment creation, delivery, and feedback.
  • Build model pipelines, evaluation frameworks, and testing approaches to ensure models are valid, reliable, fair, and aligned with ETS’s Responsible AI guidelines.
  • Collaborate with psychometric and validity teams to ensure AI-driven systems meet technical, fairness, and measurement standards.
  • Partner with engineering teams to deploy models into scalable, production-grade systems integrated with ETS operational platforms.
  • Conduct experiments to compare model architectures, datasets, training regimes, and performance tradeoffs.
  • Implement monitoring solutions to track model drift, robustness, and security risks over time.
  • Participate in cross-functional design sessions, helping translate assessment or business needs into implementable AI solutions.
  • Document model design, training data specifications, evaluation metrics, and deployment requirements.
  • Stay current with advancements in LLMs, generative AI, responsible AI, and educational technology, bringing forward ideas for innovation., * We are passionate about hiring innovative thinkers who believe in the promise of education and lifelong learning.
  • We are energized by cultivating growth, innovation, and continuous transformation for the next generation of rising professionals as leaders. Â In support of this ETS offers multiple Business Resource Groups (BRG) for you to learn and advance your career growth!
  • As a not-for-profit organization we will encourage you to lean in to your passion for volunteering. Â At ETS you may qualify for up to an additional 8 hours of PTO for volunteer work on causes that are important to you!
  • The base salary range advertised represents the low and high end of the anticipated salary range for this position. The base pay actually offered will take into account internal equity and also may vary depending on the candidate’s geographic region, job-related knowledge, skills, and experience among other factors. The base pay is only one aspect of the Total Rewards Package that will be offered to the successful candidate.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Data Science, Engineering, or related field.
  • 3+ years of experience developing machine learning models in production environments, preferably in large scale scoring environments.
  • Hands-on experience developing and deploying machine learning or deep learning models (TensorFlow, PyTorch, JAX, or equivalent).
  • Experience with NLP techniques and models, including transformer architectures and LLM fine-tuning and/or speech processing (e.g., text classification, embeddings, ASR outputs).
  • Strong proficiency in Python and familiarity with modern MLOps tools (MLflow, Kubeflow, KServe, or equivalent).
  • Demonstrated experience evaluating model performance using appropriate metrics and validation strategies.
  • Understanding of model evaluation, bias and fairness assessment, and statistical validation.
  • Ability to work collaboratively with researchers, engineers, and product teams.

About the company

ETS is a global education and talent solutions organization enabling lifelong learners worldwide to be future-ready. For more than 75 years, we’ve been advancing the science of measurement to build benchmarks for fair and valid skill assessment across cultures and borders. Our worldwide impact extends through our renowned assessments including TOEFL®, TOEIC®, GRE® and Praxis® tests, serving millions of learners in more than 200 countries and territories. Through strategic acquisitions, we’ve expanded our global capabilities: PSI strengthens our workforce assessment solutions, while Edusoft, Kira Talent, Pipplet, Vericant, and Wheebox enhance our educational technology and assessment platforms across critical markets worldwide.

Through ETS Research Institute and ETS Solutions, we’re partnering with educational institutions, governments, and organizations globally to promote skill proficiency, empower upward mobility, and unlock opportunities for everyone, everywhere. With offices and partners across Asia, Europe, the Middle East, Africa, and the Americas, we deliver nearly 50 million tests annually. Join us in our journey of measuring progress to power human progress worldwide.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

Videos

See all

Related articles

See all