Senior Software Engineer, AI Benchmarking
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
SecureBio is looking for a Senior Software Engineer to build, scale, and run biosecurity evaluations on frontier AI systems. The tools you build will directly inform decisions made by frontier AI companies, policymakers, and researchers to improve the safety of AI systems., * Perform pre-release and post-release assessments of frontier and open-source AI models using SecureBio’s suite of biosecurity evaluations, which surface and measure dual-use biology capabilities in AI systems.
- Build adapters for models, identify and troubleshoot unexpected agent behaviors, and ensure that our evaluations are methodologically rigorous.
- Build, scale, maintain, and continuously improve SecureBio’s suite of biosecurity evaluations.
- Improve our evaluations to use state-of-the-art agent frameworks, jailbreaking methods, and elicitation strategies.
- Develop and improve internal tools that enable our researchers to run capability assessments at scale.
- Contribute to and maintain cloud infrastructure to run our evaluations and analyses.
- Contribute analyses to our public dashboard of trends in AI biology capabilities (https://securebio.org/benchmarks/).
Requirements
- A proven track record in software engineering, including 5+ years of experience in industry roles and/or open-source contributions.
- Fluency in Python.
- Proficiency with Docker and AWS (particularly ECS).
- Experience building software tools that use or evaluate LLMs or AI agents (for example, code that orchestrates agents for a specific application, or that evaluates model capabilities on a specific task).
- Self-directed and comfortable working on small, fast-moving teams.
- A commitment to good coding practices, including knowing when and how to use (or not use) coding agents.
- Care about code quality, readability, and maintaining human oversight of human- and AI-written code.
- Alignment with SecureBio’s mission of preventing catastrophic pandemics., * Experience with AI evaluations (agent systems, scorer and grader design, agent trajectory analysis, jailbreaking or red-teaming), whether for security, safety, or model performance applications.
- Experience with Terraform or other Infrastructure-As-Code tools (e.g. Tofu or AWS CDK).
- Experience with modern, component-based web frontends (e.g. React with TypeScript).
- Experience with our specific tech stack:
- Evaluation frameworks: Inspect AI
- LLM APIs: OpenAI-compatible, Anthropic, Together, etc.
- Agent frameworks: Codex, Claude Code, etc.
About the company
About SecureBio
In joining SecureBio, you’ll be joining a motivated, mission-driven, and international team of professionals dedicated to protecting the world from biological threats. Through in-house research & development, as well as collaborations with experts in academia, industry & government, we are working to build the tools, technologies, and institutions the world needs to be secure against future biotechnology.
Major projects by the SecureBio team include the SecureBio Detection project for reliable pandemic detection and our work to understand and reduce biological risks from advanced artificial intelligence.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Navigating the AI Shift
Dev Digest 120 - Apple and peers
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Dev Digest 121 - AI goes offline