> Markdown version of [/jobs/ext/1170301-c-senior-engineer-irc299093](https://www.wearedevelopers.com/jobs/ext/1170301-c-senior-engineer-irc299093). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # C++ Senior Engineer IRC299093 - **Company:** GlobalLogic - **Location:** Minneapolis, MN, United States (Remote available) - **Experience:** Expert - **Salary:** $150,000.0 - $180,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Data Analysis, C++ (Programming Language), Data Structures, Data Visualization, Video Game Development, Python (Programming Language), Machine Learning, Reinforcement Learning, Data Logging, Large Language Models, Multi-Agent Systems, Prompt Engineering, AngularJS - **Published:** July 3, 2026 - **Apply:** https://www.careerjet.com/job/us86c49db6fa0362bc4f3be1628970ffa6/eaa ## About the Role The ideal candidate is a hands-on AI evaluation engineer who can both build and use the platform. This person should be comfortable integrating algorithms, running experiments, defining metrics, analyzing results, and giving practical feedback to engineering teams. The role requires a blend of ML experimentation, LLM agent evaluation, Python engineering, and strong platform-user instincts. Important Note, * Hands-on reinforcement learning experience. * Experience using LLMs for agents, evaluation, reasoning, automation, or benchmark workflows. * Strong Python experience for ML, data workflows, experimentation, and analysis. * Experience designing and running experiments with statistical and analytical rigor. * Strong understanding of evaluation metrics, scoring frameworks, performance comparison, and benchmark design. * Experience analyzing structured logs, run outputs, model/agent performance, and experiment results. * Ability to work across APIs, logs, CLI/tools, data structures, and platform workflows. * Strong communication skills to translate experiment findings into platform improvement requirements. * Ability to work inside client-owned repositories, infrastructure, workflows, and security controls. Preferred Skills * Experience with game environments, simulation environments, Gym-like interfaces, RL environments, or agentic AI test harnesses. * Experience benchmarking LLM agents, RL policies, autonomous agents, or hybrid AI systems. * Experience with experiment tracking, run comparison tools, metrics dashboards, or evaluation pipelines. * Experience with prompt engineering, agent orchestration, tool use, and LLM evaluation frameworks. * Experience with data visualization and performance analytics. * Experience working with externally developed algorithms, reproducible experiments, and version-controlled evaluation workflows., Toole Design Group in Minneapolis, MN is looking to hire an experienced and talented full-time Senior Engineer. Do you have a strong background in civil engineering with roadway/co… ## Description We are looking for an AI Evaluation & Benchmarking Engineer with experience in reinforcement learning, LLM-based agents, experiment design, benchmarking, and performance evaluation. This role will support the productionization of an AI evaluation platform used to execute and evaluate algorithms within video game environments. The engineer will develop and integrate baseline algorithms, reinforcement learning approaches, LLM-based agents, and externally developed algorithms into the platform. This person will also design experiments, define evaluation metrics, run benchmarks, analyze performance, and serve as a primary power user of the platform to provide feedback to the engineering team., * Develop, adapt, and integrate reinforcement learning algorithms and baseline approaches into the shared evaluation platform. * Integrate LLM-based agents and/or evaluators for solving, interacting with, and benchmarking game environments. * Integrate external or off-the-shelf algorithms into the platform using defined execution and ingestion workflows. * Design and run benchmark experiments across games, environments, configurations, agents, and algorithm versions. * Define evaluation strategies for comparing RL, LLM-based, hybrid, and baseline approaches. * Define, extract, and validate meaningful performance metrics from logs, outputs, run results, and environment interactions. * Build comparison logic, scoring approaches, rankings, verdicts, and performance summaries. * Develop analytics and visualizations to evaluate algorithm performance across runs and environments. * Act as a primary power user of the platform, running experiments and identifying gaps in tooling, APIs, metrics, workflows, logs, and user experience. * Provide structured feedback to Platform and Full Stack engineers to improve execution, logging, evaluation, and reporting capabilities. * Validate existing game environments and support development or validation of new game environments. * Evaluate environment operability using baseline/reference frontier LLM models, harnesses, and agents. * Collaborate with client technical teams and engineering resources within 3M-owned repositories, workflows, infrastructure, and security processes. * Ensure all algorithms, experiments, notebooks/scripts, configuration, documentation, and outputs comply with 3M-defined standards and policies., Are you interested in being part of an innovative team that supports Westinghouse's mission to provide clean energy solutions? At WECTEC Staffing Services, a wholly-owned subsidiar… + 21 days ago, Senior Full Stack Angular Engineer Duration: Long Term Contract Location: Minneapolis, MN (Onsite) Rate: $42/hr. Interview mode: Video Conference/In-Person Joining: ASAP This… + 16 days ago ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Crypto-secure Data Management with In-Database Blockchain](https://www.wearedevelopers.com/videos/632-crypto-secure-data-management-with-in-database-blockchain) - [How to Stop Choosing JavaScript Frameworks and Start Living](https://www.wearedevelopers.com/videos/118-how-to-stop-choosing-javascript-frameworks-and-start-living) - [Building a Multi-Agent Orchestration Engine That Actually Follows the Rules](https://www.wearedevelopers.com/videos/100159-building-a-multi-agent-orchestration-engine-that-actually-follows-the-rules) - [Developers Communities in San Francisco - WeAreDevelopers LIVE from Corgi Cafe](https://www.wearedevelopers.com/videos/2144-developers-communities-in-san-francisco-wearedevelopers-live-from-corgi-cafe) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [The Prompt Engineer ✍️](https://www.wearedevelopers.com/magazine/216-the-prompt-engineer) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)