> Markdown version of [/jobs/ext/2557477-agent-evaluation-evolution-machine-learning-engineer-graduate-aml-ark-us-2027-start](https://www.wearedevelopers.com/jobs/ext/2557477-agent-evaluation-evolution-machine-learning-engineer-graduate-aml-ark-us-2027-start). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Agent Evaluation & Evolution Machine Learning Engineer Graduate (AML-Ark-US) - 2027 Start - **Company:** BYTEDANCE INC. - **Location:** Seattle, WA, United States - **Experience:** Starter - **Salary:** $121,600.0 - $243,200.0 - **Contract:** Internship / Graduate position - **Skills:** Artificial Intelligence, Systems Engineering, Program Optimization, Python (Programming Language), Machine Learning, Open Source Technology, Large Language Models, Multi-Agent Systems, Prompt Engineering, Information Technology, Data Analytics, Virtual Agents - **Published:** August 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=67621850c47b1c0c ## About the Role * Individuals who are completing or have recently completed a Bachelor's/ Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science, or a related field. * Solid foundation in machine learning and deep learning. * Hands-on experience with LLM-based systems (e.g., agents, tool calling, retrieval, multi-agent systems) through research, internships, or projects. * Strong Python skills and experience with a mainstream ML or agent evaluation framework. * Demonstrated research or engineering ability through publications, substantial projects, internships, or open-source work., * Publications at top-tier ML/NLP venues (e.g., NeurIPS, ICML, ICLR, ACL etc.), especially in agent learning, self-improving/self-evolving/RSI, or agent evaluation. * Experience with evaluation methodology: metric design, model-based judging, or annotation and statistical analysis, etc. * Familiarity with LLM post-training, reasoning and planning methods, or continual learning. * Experience with feedback-driven optimization loops, or with large-scale log and trace analysis., Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment ## Description The Applied Machine Learning Ark team combines system engineering and machine learning to develop and operate Large Language Model (LLM) service platforms that offer businesses Model-as-a-Service (MaaS) solutions, serving both large model providers and downstream users. The US team drives the design, development, and operation of MaaS solutions across the US and international markets outside mainland China. We are building full-stack, end-to-end solutions spanning text and multimodal LLM algorithms, LLM training/fine-tuning/inference frameworks, prompt engineering, model alignment, and intelligent agent systems. Beyond model serving, we operate large-scale log analytics pipelines that process massive volumes of invocation logs from text models, multimodal models, and agent systems - extracting usage patterns, quality signals, and actionable insights to inform model improvement, system optimization, and product decisions through continuous, data-driven feedback loops. We are actively seeking talented engineers and researchers specializing in Large Language Models and AI Agent systems to join our dynamic team., * Design evaluation systems for LLM-based agents, covering task success, tool use, reasoning quality, and reliability. * Build benchmarks and automated judging pipelines, combining rule-based checks, model-based judging, and human review, etc. * Analyze agent execution traces and user feedback to identify failure patterns and turn them into concrete system improvements. * Support the closed loop from experience to capability, and work with research, platform, and product teams to bring methods into production. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Guiding Agentic AI with Vue](https://www.wearedevelopers.com/videos/2033-guiding-agentic-ai-with-vue) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Beyond the Benchmark: How to Evaluate AI Agents in the Real World](https://www.wearedevelopers.com/videos/100269-beyond-the-benchmark-how-to-evaluate-ai-agents-in-the-real-world) - [How Data is Shaping our Games](https://www.wearedevelopers.com/videos/176-how-data-is-shaping-our-games) - [Evals vs. Evil - AI and Package Security - Laurie Voss](https://www.wearedevelopers.com/videos/2131-evals-vs-evil-ai-and-package-security-laurie-voss) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)