> Markdown version of [/jobs/ext/2488475-senior-ai-engineer](https://www.wearedevelopers.com/jobs/ext/2488475-senior-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior AI Engineer - **Company:** RWS - **Location:** UK (Remote available) - **Experience:** Expert - **Salary:** £42,615.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Computing, Code Review, Continuous Integration, Python (Programming Language), Machine Translation, Performance Tuning, Software Engineering, Systems Integration, Management of Software Versions, Datadog, Comet Programming, Large Language Models, Prompt Engineering, Kubernetes, Machine Learning Operations, Api Design - **Published:** August 20, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5849825468 ## About the Role * Significant software engineering experience (typically 5+ years) building and operating production systems, tools, libraries, or services that many other engineers depend on, with excellent API design, reliability, and developer experience. Also, a track record with CI/CD and cloud infrastructure. * Proficiency in Python and/or another general-purpose language, with strong testing discipline * Hands-on experience building with LLMs or other ML systems (prompt engineering, fine-tuning, retrieval, model integration), with an understanding of their failure modes and tradeoffs. * Proven experience designing and leading evaluation for AI/ML systems: defining metrics and methodology, building evaluation pipelines, managing test sets, and reasoning rigorously about model quality and regressions. * Strong command of evaluation concepts - various types of metrics (accuracy, precision, recall, F1), the distinction and tradeoffs between automated and human evaluation, statistical significance, and the limits of each approach. * Excellent written and verbal communication, and a history of influencing technical direction across teams and mentoring other engineers. * Comfort with ambiguity and the judgment to scope, prioritize, and sequence high-impact work with limited direction. Preferred * Deep experience evaluating NLP, machine translation, or content-generation systems, including metrics such as COMET, chrF++, BLEU, MetricX, and MQM-style human evaluation. * Experience with experimentation and observability tooling, data/test-set versioning, and rigorous benchmarking workflows. * Established practice in AI governance and documentation - model cards, system cards, reproducibility, and responsible-AI considerations - at an organizational level. * Broad familiarity with the modern LLM ecosystem (open and proprietary models, orchestration frameworks, vector stores) and well-formed views on the tradeoffs. * Experience supporting multilingual or localization-focused products at enterprise scale. ## Description Architecture and technical strategy * Contribute to the design and architecture of core platform components and evaluation systems, making the load-bearing technical decisions and bearing accountability for their reliability, scalability, and long-term maintainability. * Help set the technical direction for how AI capabilities are built, evaluated, and deployed across the company, and define a coherent platform vision that scales beyond your immediate team. * Design reusable abstractions, SDKs, and services for model integration, prompt management, experimentation, and deployment that establish organization-wide patterns and reduce duplicated effort. Research and delivery excellence * Help define the evaluation strategy and methodology for AI capabilities across the company - automated metrics, human-in-the-loop workflows, test set management, and benchmarking - and establish the quality standards other teams build against. * Build evaluation frameworks and developer tooling robust enough for production yet simple enough for non-specialist developers to adopt. * Establish observability standards for AI systems - quality, performance, cost, and regression signals - and build dashboards and reporting that turn those signals into actionable decisions. * Drive engineering rigor in delivery through testing discipline, reproducibility, sound experimental design, and statistically defensible measurement of model quality. Technical leadership * Provide technical leadership on the team's most ambiguous and highest-impact problems, scoping and sequencing work where direction is limited. * Mentor engineers and raise engineering standards through code review, design review, and leading by example. * Contribute to model and system governance practices including documentation (model cards, system cards), dataset and test-set versioning, reproducibility, and responsible-AI checks embedded directly into the platform. * Act as a technical multiplier - codifying best practices into tooling and standards adopted by hundreds of developers. Innovation and technology adoption * Track developments in LLMs, evaluation research, and AI tooling, and translate them into pragmatic, well-scoped improvements to the platform. * Prototype and de-risk emerging techniques and tools, shepherding the promising ones from experiment to supported capability. * Champion the adoption of new platform capabilities across teams, lowering the barrier for developers to use them well. * Keep the platform current and competitive without chasing novelty for its own sake. Cross-functional collaboration and stakeholder leadership * Partner with research, product, and localization leaders to align evaluation methodology with real-world quality and customer needs. * Influence roadmap and technical strategy beyond your immediate team, building consensus across engineering and product stakeholders. * Gather requirements from developers across the company and represent their needs in platform direction, acting as a trusted technical partner. * Communicate technical direction, trade-offs, and quality standards clearly to both technical and non-technical audiences. ## Related Videos - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [API Design - Getting Started](https://www.wearedevelopers.com/videos/33-api-design-getting-started) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [Rest API Antipatterns](https://www.wearedevelopers.com/videos/100208-rest-api-antipatterns) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)