Applied Machine Learning Engineer (Model Layer)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+1 more
Job description
As the Senior LLM / AI Model Routing Engineer, you will own the model layer from integration through production optimization. You will build the systems that determine which AI model should handle each request, evaluate model performance, and continuously improve the balance between quality, speed, reliability, and cost. This is a production engineering role, not a research-only position. The ideal candidate has experience taking LLM technology beyond prototypes and building reliable systems used by real users., LLM Integration & Model Management
- Integrate open-source and third-party LLMs through APIs and inference providers.
- Maintain the product’s model lineup and evaluate new models as they become available.
- Compare models based on quality, reliability, speed, cost, and suitability for different use cases.
- Determine when models should be added, replaced, or retired.
- Implement reliable model and provider fallbacks.
Model Routing
- Design and maintain intelligent model routing logic.
- Route requests to the most appropriate model based on factors such as task type, quality, latency, cost, and model behavior.
- Continuously refine routing decisions using production data and user feedback.
- Explore advanced approaches such as model cascades, ensembles, and multi-model systems when appropriate.
Evaluation & Optimization
- Build practical systems for evaluating LLM output quality and routing decisions.
- Establish and monitor key performance metrics, including quality, latency, reliability, and cost per request.
- Analyze production data to identify opportunities for improvement.
- Optimize prompts, model parameters, configurations, and fallback strategies.
- Balance response quality with performance and operating costs.
Performance & Reliability
- Monitor model performance, latency, usage, and inference costs.
- Identify and resolve issues with models, providers, routing, and integrations.
- Improve response speed and reliability as usage grows.
- Implement appropriate logging, monitoring, error handling, and fallback mechanisms.
- Maintain clear documentation for the model architecture and operational processes.
Collaboration
- Work closely with the application engineering team to integrate the model layer into the product.
- Collaborate on the connection between the AI, application, and policy layers.
- Communicate technical concepts and trade-offs clearly to founders and non-ML stakeholders.
- Provide data-driven recommendations on model strategy and product performance., * Curious: You stay current with new models, providers, and AI techniques.
- Proactive: You identify problems and opportunities without waiting for instructions.
- Ownership-oriented: You take responsibility for the performance and reliability of your systems.
- Collaborative: You can work effectively with engineers, founders, and non-technical stakeholders.
What Success Looks Like
- Requests are consistently routed to the best model for the task.
- Users receive high-quality, accurate, and relevant responses.
- Latency remains fast and consistent as usage increases.
- Cost per request is continuously optimized.
- Model and provider failures are handled reliably.
- New models are evaluated and integrated quickly when they add value.
- Routing and model decisions are based on measurable production data.
- The model layer is reliable, well-documented, and continuously improving.
Why Join?
- Own a critical AI layer: Take end-to-end ownership of the engine behind the product.
- Work with leading LLMs: Continuously evaluate and work with new open-weight and provider models.
- Make a direct impact: Your decisions will directly influence product quality, performance, and cost.
- Work with a lean team: Collaborate closely with founders and have meaningful technical influence.
- Remote & global: Work from anywhere with a flexible schedule.
- Build for real users: Move beyond experimentation and build production AI systems.
- Growth opportunity: Expand your technical ownership as the product and team scale.
Requirements
- Proven experience building LLM-powered applications in production.
- Strong Python development skills.
- Hands-on experience working with LLM APIs and inference platforms.
- Experience with providers such as OpenAI, Anthropic, Together, Groq, or similar.
- Strong understanding of LLM capabilities, prompt design, model evaluation, and production behavior.
- Ability to make practical trade-offs between quality, latency, reliability, and cost.
- Strong analytical and problem-solving abilities.
- Comfortable working independently in a small, fast-moving startup environment.
- Strong written and verbal English communication skills.
Preferred Qualifications
- Experience with LLM model routing, multi-model architectures, ensembles, or AI gateways.
- Familiarity with open-weight models such as Llama, Qwen, or Gemma.
- Experience with LLM evaluation frameworks or automated evaluation pipelines.
- Experience optimizing inference cost, latency, throughput, or reliability.
- Experience with model observability and production monitoring.
- Experience building privacy-focused or privacy-conscious AI products.
- Previous experience working in an early-stage or lean engineering team.
The Ideal Candidate
We are looking for a builder who enjoys ownership and execution.
You don’t need to be an academic AI researcher. You need to understand how modern LLMs work, know how to evaluate them in the real world, and be able to turn that knowledge into reliable production systems.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
MLOps And AI Driven Development
Prompt Engineering is a Job of the Past