> Markdown version of [/jobs/ext/819912-llm-evaluation-engineer](https://www.wearedevelopers.com/jobs/ext/819912-llm-evaluation-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # LLM Evaluation Engineer - **Company:** CYNET SYSTEMS INC. - **Location:** Los Angeles, CA, United States - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Software Applications, Programming Tools, Software Engineering, Systems Integration, Large Language Models, Multi-Agent Systems - **Published:** June 11, 2026 - **Apply:** https://www.dice.com/job-detail/5adf9ecc-749d-4eaf-83a2-c38b6686c54d ## About the Role * Senior-level engineering or software development experience. * Demonstrated, hands-on experience with Claude (Anthropic) tooling and APIs. * Strong background in API integration, developer tooling, and CLI-based workflows. * Experience evaluating AI/LLM platforms in an enterprise environment., * Experience with AI developer platforms, copilots, or agent frameworks. * Familiarity with desktop application integrations and enterprise controls. ## Description * Evaluate Claude for Desktop capabilities, including Cowork, Code, CLI, and APIs. * Conduct hands-on testing, prototyping, and technical validation. * Assess integration patterns, security considerations, and scalability. * Provide clear findings, recommendations, and technical feedback. ## Related Videos - [AI Agents & Agentic AI](https://www.wearedevelopers.com/videos/2017-ai-agents-agentic-ai) - [AI Pair Programming with GitHub Copilot at SAP: Looking Back, Looking Forward!](https://www.wearedevelopers.com/videos/1546-ai-pair-programming-with-github-copilot-at-sap-looking-back-looking-forward) - [Designing and Deploying Distributed Multimodal Multi-Agent Systems with Google's AI Stac](https://www.wearedevelopers.com/videos/1976-designing-and-deploying-distributed-multimodal-multi-agent-systems-with-google-s-ai-stac) - [Three years of putting LLMs into Software - Lessons learned](https://www.wearedevelopers.com/videos/1508-three-years-of-putting-llms-into-software-lessons-learned) - [LLMs in the wild: Building an AI agent that survives production](https://www.wearedevelopers.com/videos/100319-llms-in-the-wild-building-an-ai-agent-that-survives-production) - [Exploring AI: Opportunities and Risks in Development](https://www.wearedevelopers.com/videos/1267-exploring-ai-opportunities-and-risks-in-development) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [13 AI Tools You Have to Try](https://www.wearedevelopers.com/magazine/219-13-ai-tools-you-have-to-try) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [Who Owns Your Content in the Age of LLMs?](https://www.wearedevelopers.com/magazine/610-who-owns-your-content-in-the-age-of-llms) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix)