Machine Learning and AI Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
Build and maintain ML/AI features
- Develop and improve the components that make AI agents intelligent: prompt engineering, classifier pipelines, goal evaluation logic, and post-call analysis
- Work with the LangChain/LangGraph agent framework to build, test, and refine conversation flows that handle real-world customer interactions
- Implement and evaluate data extraction pipelines - turning unstructured conversation transcripts into structured fields (names, dates, postcodes, appointment preferences) reliably
Integrate and evaluate models
- Integrate LLM providers (OpenAI, Anthropic, Groq, Google) into the platformâs agent orchestration layer, including prompt construction, response parsing, and error handling
- Run model evaluations - comparing output quality, latency, and cost across providers and model versions to inform which models the platform uses in production
- Work with the existing Langsmith tracing infrastructure to monitor model performance and identify regressions
Support the voice and classification pipeline
- Contribute to the STT (speech-to-text) and TTS (text-to-speech) integration layer - understanding how audio becomes text, how text becomes an agent response, and how that response becomes audio again
- Help build and extend the classification system that determines conversation outcomes (was the call successful? did the customer want a callback? was it a voicemail?) - including writing evaluation prompts, defining ground truth datasets, and measuring accuracy
- Assist with data preparation, feature engineering, and dataset curation for evaluation and fine-tuning tasks
Write production-quality code
- Write clean, tested Python that runs in a production FastAPI application - not throwaway scripts
- Participate in code reviews, both giving and receiving - learning from the senior developerâs feedback and contributing your own perspective
- Contribute to documentation that helps the rest of the engineering team understand how AI components work and how to use them correctly, * Youâve worked with LangChain, LangGraph, or similar agent frameworks - even in a personal project or hackathon
- Youâve built something with the OpenAI or Anthropic API that went beyond âhello worldâ - a chatbot, a classifier, a data extraction pipeline, an evaluation harness
- You understand the basics of how voice AI works: STT * LLM * TTS - even if youâve only read about it rather than built it
- Youâve worked with structured evaluation of LLM outputs - comparing model responses against expected answers, not just eyeballing whether it âlooks rightâ
- You have opinions about prompt engineering - youâve iterated on prompts and observed how small changes affect output quality
What You Wonât Be Doing
- Working in isolation on research problems - this is a product engineering role embedded in a delivery team
- Training large models from scratch - the platform uses hosted LLM APIs; your job is integration, evaluation, and orchestration, not pretraining
- Waiting to be told what to do - youâll have guidance and mentorship from the senior developer, but youâre expected to take ownership of your tasks and ask questions when youâre stuck
Requirements
Do you have experience in Python?, The platform runs on a practical AI stack: LangChain and LangGraph for agent orchestration, OpenAI and Anthropic for LLMs, Deepgram for speech-to-text, ElevenLabs for text-to-speech, and LiveKit for real-time voice infrastructure. You donât need to know all of these coming in, but you do need to be comfortable working with APIs, understanding model behaviour, and writing Python that runs in production - not just in notebooks., Youâve finished your degree or equivalent, and youâve spent some time - whether through jobs, internships, or serious personal projects - working with ML or AI in a way that went beyond coursework.
- 1-2 years of experience working with ML/AI (including internships, placement years, or substantial personal/open-source projects)
- Solid Python skills - you can write functions, classes, and tests confidently, not just Jupyter notebooks
- Familiarity with at least some of: LLMs and prompt engineering, NLP, text classification, or information extraction - you donât need depth in all of them, but you need to have worked with at least one area hands-on
- Basic understanding of how ML models are evaluated - you know what precision, recall, and F1 mean and why they matter; youâve compared model outputs against ground truth at least once
- Comfortable working with APIs and reading documentation - a significant part of this role involves integrating and configuring third-party AI services, not building models from scratch
- Familiar with Git and working in a team codebase - youâve committed code that other people have reviewed, and youâve reviewed other peopleâs code
Benefits & conditions
Weâre building something global at Narwhal, and we mean that in every sense. The work we do requires different ways of thinking - and different ways of thinking come from different people.
At Narwhal, weâre committed to building a diverse and inclusive team. We welcome applications from people of all backgrounds, identities, and experiences, and we actively work to ensure our hiring process is fair and accessible for everyone. Reasonable adjustments are available at every stage, just reach out and weâll make it happen.
Pay: ÂŁ75,000.00-ÂŁ100,000.00 per year
About the company
Narwhal Labs is the company behind DeepBlue OS - an autonomous revenue infrastructure platform that enables any business to answer every call, follow up every lead, and log every interaction across Voice, SMS, Email and WhatsApp. As an NVIDIA Inception Program Member and Google Partner, we are a 38-person team with our platform launching in May 2026. We build the infrastructure layer for serious businesses that want enterprise-grade revenue operations at a fraction of traditional cost.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on indeed.comGood distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
MLOps And AI Driven Development
MLOps â Whatâs the deal behind it?
Dev Digest 137 - AI'm not sure about this