> Markdown version of [/jobs/ext/2799745-ai-benchmark-engineer](https://www.wearedevelopers.com/jobs/ext/2799745-ai-benchmark-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Benchmark Engineer - **Company:** Lilt - **Location:** Wien, Austria (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Encodings, Text Processing, Python (Programming Language), Shell Script, Software Engineering, Toolchain, Data Processing, Epic Haiku, Large Language Models, OPUS (Software) - **Published:** September 9, 2026 - **Apply:** https://startup.jobs/ai-benchmark-engineer-native-language-specialist-german-austria-remote-lilt-production-9958626 ## About the Role We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches., * Experience: 5+ years of industry experience in software engineering. * Background: Proven track record at leading technology companies and/or graduation from top-tier engineering universities. * Language: Native or near-native fluency, with a deep understanding of its grammar, register, and phrasing rules. High English proficiency. * Technical Stack: Strong proficiency in Python, standard shell scripting, and data processing. * Workflow: Extensive experience with Terminal/CLI-based development workflows and a working familiarity with coding agents. * Domain Expertise: Deep technical understanding of multilingual text processing pitfalls, including: + Encoding/decoding robustness and Unicode normalization. + Locale-dependent conventions (collation, casing, non-Gregorian dates). + Text I/O, toolchain interoperability, and safe string operations. + (For specific languages) Bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts. ## Description * Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially, these assets must remain in the target language to genuinely measure multilingual handling. * Prompting & Translation: finding failure points where AI does not work, in your native language * Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary). * Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus). * Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity., * Get paid quickly and fairly. We respect your time and your expertise. Competitive rates, prompt payments, no chasing invoices. * Work on projects that actually matter. Contribute to cutting-edge AI and language technology that is shaping how humans and machines communicate. * Be part of something bigger. Join a global community of linguists, subject matter experts, and language professionals who are advancing human knowledge together. * Grow without limits. As a Lilt contractor you get access to diverse, innovative projects that expand your portfolio and sharpen your skills across industries and domains. * Have fun doing what you love. Bring your language skills to life on projects that are as interesting as they are impactful.We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows., * Not ideal as a full time job or primary income source. Work availability fluctuates with project demand, making this better suited as a supplemental income stream. As a 1099 contractor, you won't receive benefits such as health insurance, paid time off, or retirement contributions, and hours are not guaranteed. * Requires reliable availability and commitment. Once you accept a task, we expect quality work and on-time delivery. Most tasks require a minimum of 2 hours per day or 10 hours per week. If your schedule is unpredictable, this may not be the right fit. * Geographic restrictions may apply. We cannot engage contractors in regions subject to international embargo or sanctions. As a 1099 contractor, you are solely responsible for your own tax obligations. We recommend consulting a tax professional before engaging. How to join our expert community 1 - Submit your application including an updated copy of your CV in English 2 - Next, complete a GenAI assessment to evaluate your skills 3 - Finalize onboarding and profile set-up in our system, and become eligible for Applied AI projects. ## Related Videos - [Speeding up Web Apps performance with WebAssembly and Emscripten](https://www.wearedevelopers.com/videos/1985-speeding-up-web-apps-performance-with-webassembly-and-emscripten) - [A Brief History of Data Storage](https://www.wearedevelopers.com/videos/974-a-brief-history-of-data-storage) - [Hiring AI Native Talents](https://www.wearedevelopers.com/videos/100268-hiring-ai-native-talents) - [The Power of Developer Communities](https://www.wearedevelopers.com/videos/1109-the-power-of-developer-communities) - [JSON and Beyond](https://www.wearedevelopers.com/videos/968-json-and-beyond) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)