Senior AI Researcher

On behalf of Next Deavor
United States
15 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$150,000.0 - $220,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Python (Programming Language) Machine Learning Language Modeling Pytorch Large Language Models Data Pipelines

Job description

You will lead original research advancing core models that enable offensive-security capabilities, shaping experiments end-to-end and shipping results into production. You will collaborate closely with the VP of AI Engineering, the CEO, and a small AI engineering team to turn research outcomes into deployable capabilities. Work model: New York preference but open to remote; you must work EST hours.

Here’s How You’ll Make an Impact on the Team

Drive original research on offensive-security agents: reasoning, planning, tool use, and long-horizon autonomous operation

Advance the post-training pipeline, including supervised fine-tuning, RL from verifier signals, LoRA adaptation, and adversarial evaluation

Extend co-evolutionary self-training architecture with curriculum design, self-play dynamics, and reward modeling for security outcomes

Design and execute experiments end-to-end, from hypothesis through writeup

Build internal evaluation harnesses where no public benchmark exists and measure capability rigorously

Translate research into production handoffs: model cards, deployment notes, and documented failure modes

Contribute to external research outputs: papers, talks, responsible disclosures, and technical writing

Requirements

Demonstrated original ML research output (published papers, widely cited preprints, significant OSS releases, or shipped research that materially advanced a production system)

Hands-on post-training experience with large language models (7B+ parameters) and end-to-end ownership of data, training, and evaluation pipelines

Direct experience with at least one of: RL from verifier/reward signals, preference optimization (DPO/IPO/KTO), or supervised fine-tuning with synthetic data pipelines

Experience with agentic LLM systems: tool use, multi-step reasoning, planning, or long-horizon execution

Ability to design evaluations that measure real capability and avoid contamination or specification gaming

Strong Python and PyTorch skills, with experience in distributed multi-GPU training

Clear technical writing demonstrated by research memos, experiment writeups, or papers

Here’s What Else Might Help You Out

Working knowledge of offensive security fundamentals (trainable on the job)

Prior work on code-generating or code-reasoning models

Experience with sparse, delayed, or expensive reward signals in RL

Research in robustness, adversarial ML, or red-teaming of language models

Familiarity with long-horizon agent benchmarks (e.g., SWE-bench, Cybench, WebArena)

Benefits & conditions

$150,000 - $220,000 a year - Full-time

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · WWC 2024

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

3:22 min

Evaluating advanced artificial intelligence platforms for daily recruitment

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

2:01 min

Exploring foundational expertise in traditional optimization and machine learning

Eric Enge · Coffee With Developers

Videos

See all

Related articles

See all