Senior Software Engineer, AI Benchmarking

SecureBio, Inc.
United States
about 2 months ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Clean Code Principles Artificial Intelligence Component-Based Software Engineering Cloud Computing Software Quality Software Construction Software Engineering TypeScript AWS Cdk ReactJS Large Language Models Multi-Agent Systems
+3 more
Infrastructure Automation Frameworks Free and Open-Source Software Terraform

Job description

SecureBio is looking for a Senior Software Engineer to build, scale, and run biosecurity evaluations on frontier AI systems. The tools you build will directly inform decisions made by frontier AI companies, policymakers, and researchers to improve the safety of AI systems., * Perform pre-release and post-release assessments of frontier and open-source AI models using SecureBio’s suite of biosecurity evaluations, which surface and measure dual-use biology capabilities in AI systems.

  • Build adapters for models, identify and troubleshoot unexpected agent behaviors, and ensure that our evaluations are methodologically rigorous.
  • Build, scale, maintain, and continuously improve SecureBio’s suite of biosecurity evaluations.
  • Improve our evaluations to use state-of-the-art agent frameworks, jailbreaking methods, and elicitation strategies.
  • Develop and improve internal tools that enable our researchers to run capability assessments at scale.
  • Contribute to and maintain cloud infrastructure to run our evaluations and analyses.
  • Contribute analyses to our public dashboard of trends in AI biology capabilities (https://securebio.org/benchmarks/).

Requirements

  • A proven track record in software engineering, including 5+ years of experience in industry roles and/or open-source contributions.
  • Fluency in Python.
  • Proficiency with Docker and AWS (particularly ECS).
  • Experience building software tools that use or evaluate LLMs or AI agents (for example, code that orchestrates agents for a specific application, or that evaluates model capabilities on a specific task).
  • Self-directed and comfortable working on small, fast-moving teams.
  • A commitment to good coding practices, including knowing when and how to use (or not use) coding agents.
  • Care about code quality, readability, and maintaining human oversight of human- and AI-written code.
  • Alignment with SecureBio’s mission of preventing catastrophic pandemics., * Experience with AI evaluations (agent systems, scorer and grader design, agent trajectory analysis, jailbreaking or red-teaming), whether for security, safety, or model performance applications.
  • Experience with Terraform or other Infrastructure-As-Code tools (e.g. Tofu or AWS CDK).
  • Experience with modern, component-based web frontends (e.g. React with TypeScript).
  • Experience with our specific tech stack:
  • Evaluation frameworks: Inspect AI
  • LLM APIs: OpenAI-compatible, Anthropic, Together, etc.
  • Agent frameworks: Codex, Claude Code, etc.

About the company

About SecureBio

In joining SecureBio, you’ll be joining a motivated, mission-driven, and international team of professionals dedicated to protecting the world from biological threats. Through in-house research & development, as well as collaborations with experts in academia, industry & government, we are working to build the tools, technologies, and institutions the world needs to be secure against future biotechnology.

Major projects by the SecureBio team include the SecureBio Detection project for reliable pandemic detection and our work to understand and reduce biological risks from advanced artificial intelligence.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:35 min

Defining a serverless architecture using AWS CDK

Raphael Manke Raphael Manke · World Congress 2023

1:21 min

Exploring the target application for front end tests

Anna Mcdougall · JS Congress

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:44 min

Balancing artisanal coding skills with automated agent oversight

Jyoti Bansal Jyoti Bansal +1 · World Congress 2026 Europe

1:46 min

Choosing between AWS CDK and Terraform

Alexander Bubeck · World Congress 2023

1:36 min

Managing infrastructure as code with AWS CDK

Markus Ziller Markus Ziller · World Congress 2024

Videos

See all

Related articles

See all