> Markdown version of [/jobs/ext/2717832-security-penetration-tester-frontier-ai-evaluation-contract](https://www.wearedevelopers.com/jobs/ext/2717832-security-penetration-tester-frontier-ai-evaluation-contract). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Security Penetration Tester, Frontier AI Evaluation (Contract) - **Company:** Cobalt LLP - **Location:** San Francisco, CA, United States - **Contract:** Permanent contract - **Skills:** Kubernetes Security, C (Programming Language), JavaScript (Programming Language), Artificial Intelligence, Software System Penetration Testing, C++ (Programming Language), Cloud Computing, Statistical Hypothesis Testing, Python (Programming Language), Red Team (Cyber Security), Mobile Security, Rust (Programming Language), Software Security, Software Coding, Golang - **Published:** September 4, 2026 - **Apply:** https://www.disabledperson.com/jobs/74812601-security-penetration-tester-frontier-ai-evaluation-contract ## About the Role * Demonstrable penetration testing experience, evidenced by professional engagements, published vulnerability research or CVEs, a substantive bug bounty record, competitive CTF results, or comparable work * Strong hands-on coding ability in at least one of Python, C, C++, Go, Rust, or JavaScript, sufficient to read unfamiliar codebases and write your own tooling rather than only running existing tools * Depth in at least one area, for example web and API security, cloud and container security, network and infrastructure testing, mobile security, or binary exploitation * Ability to explain each step of your reasoning clearly in writing, and to produce documentation another practitioner could reproduce * Willingness to work strictly within defined scope and authorization, and to sign a confidentiality agreement covering project materials Certifications such as OSCP, OSWE, OSEP, GPEN, or GXPN are useful but not required. ## Description Cobalt is seeking experienced penetration testers to contribute expert reasoning, technical problems, and evaluation data used to train and assess frontier AI models on security tasks. This opportunity is suited to people who test systems for a living or have done so: penetration testers, red team operators, and serious bug bounty hunters, whether from consultancies, internal security teams, or independent practice. You do not need prior experience in data annotation or AI research. What matters is that you can find and reason about real weaknesses in software and infrastructure unaided, and that you can document how you got there clearly enough for another practitioner to follow. All work is performed against sandboxed environments, purpose-built targets, and model endpoints supplied by us or by the lab. We do not accept work performed against systems you are not authorized to test, and we do not accept material covered by a client agreement or obtained without authorization. What you'll do: Depending on the project, you may: * Produce written testing traces, capturing how you form and test hypotheses, what you rule out and why, and how you arrive at a working approach, rather than only the end result * Author novel security problems, capture-the-flag style challenges, and lab environments with verifiable success criteria * Evaluate model-generated security content and code, ranking responses, explaining what makes the stronger one stronger, and identifying the specific step at which the technical reasoning breaks down * Assess whether stated findings are supported by the underlying evidence, and identify inconsistencies between reported results and what the target actually does * Design rubrics and partial-credit criteria for scoring multistep testing and remediation tasks Projects follow their own guidelines, scope rules, and quality standards, and you will work with feedback from reviewers and lab research teams. ## Related Videos - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Green Cloud Computing](https://www.wearedevelopers.com/videos/592-green-cloud-computing) - [DevSecOps: Injecting Security into Mobile CI/CD Pipelines](https://www.wearedevelopers.com/videos/273-devsecops-injecting-security-into-mobile-ci-cd-pipelines) - [Security Pitfalls for Software Engineers](https://www.wearedevelopers.com/videos/726-security-pitfalls-for-software-engineers) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) - [Securing Your Web Application Pipeline From Intruders](https://www.wearedevelopers.com/videos/53-securing-your-web-application-pipeline-from-intruders) ## Related Articles - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [9 Ways to Make Money Hacking](https://www.wearedevelopers.com/magazine/333-9-ways-to-make-money-hacking) - [Dev Digest 138 - Are you secure about this?](https://www.wearedevelopers.com/magazine/486-dev-digest-138-are-you-secure-about-this) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data)