> Markdown version of [/jobs/ext/3629504-remote](https://www.wearedevelopers.com/jobs/ext/3629504-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Remote - **Company:** MAG 24 LLC - **Location:** New York, NY, United States (Remote available) - **Experience:** Experienced - **Contract:** Temporary contract - **Skills:** Java (Programming Language), JavaScript (Programming Language), Artificial Intelligence, C++ (Programming Language), Code Generation, Software Quality, Code Review, Software Debugging, Python (Programming Language), Software Architecture, Software Engineering, Product Software Implementation Methods, Software Systems, Rust (Programming Language), ReactJS, Model Validation, Api Design, Golang, Programming Languages - **Published:** October 8, 2026 - **Apply:** https://www.careerjet.com/jobad/us911c6c02bc2942f3581e891a899be8e1 ## About the Role * 3+ years of professional software-engineering experience * Strong full-stack development capabilities * Experience building scalable, production-grade software * Strong understanding of software architecture and system design * Deep knowledge of development, debugging, and code-quality assessment * Experience reviewing and improving complex software implementations * Proficiency in one or more of Python, JavaScript, Java, C++, Rust, or related languages * ReactJS, C, or Go experience may also be relevant to project assignments * Strong understanding of API design and production implementation * Familiarity with software monitoring and operational maintenance * Ability to reason across the complete software-engineering lifecycle * Strong analytical and problem-solving capabilities * Excellent written and verbal communication skills * Ability to provide clear, structured evaluation rationales * Comfortable collaborating remotely with research and technical teams ## Description We are sharing a specialised part-time consulting opportunity for experienced software engineers to contribute to advanced large language model evaluation, coding benchmark development, and AI-assisted software-engineering research. Selected professionals will curate and evaluate code, develop verification mechanisms, assess AI-generated software across multiple programming languages, and help research teams understand how advanced models perform throughout realistic software-development workflows., Code Curation & Solution Development * Curate high-quality code examples for model training and benchmarking * Develop precise solutions to software-engineering tasks * Correct and improve code across multiple programming languages * Work with Python, JavaScript, ReactJS, C/C++, Java, Rust, and Go * Maintain strong standards for correctness and maintainability AI-Generated Code Evaluation * Evaluate AI-generated code for technical correctness * Assess solutions for efficiency, scalability, and reliability * Identify implementation weaknesses and recurring error patterns * Review code quality against professional engineering standards * Provide structured rationales supporting evaluation decisions Verification & Automated Assessment * Build agents that assess code quality * Design mechanisms for automatically verifying software solutions * Identify recurring model-generated coding errors * Develop reliable checks for engineering tasks * Support reproducible evaluation across repeated assignments Software Engineering Lifecycle Evaluation * Evaluate model capabilities across the software-development lifecycle * Assess reasoning around prototyping and architecture design * Review API design and production implementation decisions * Evaluate launch, experimentation, monitoring, and maintenance scenarios * Identify areas where models struggle with real-world engineering workflows Research & Benchmark Collaboration * Collaborate with research and cross-functional technical teams * Contribute to datasets used for training and benchmarking * Help define engineering evaluation strategies * Compare model performance against professional engineering expectations * Support iterative improvements to coding-focused evaluation systems