Redwood Research blog
Subscribe
Sign in
Home
Podcast
Chat
Reading List
Archive
About
Latest
Top
Discussions
AI swarms are starting to pose indirect takeover risk
Unsanctioned coordination, like we saw in the Hugging Face incident, could enable future AIs to take over
Aug 12
•
Oak Hu
and
Alex Mallen
26
4
July 2026
SOTA alignment assessments don’t strongly update us against misalignment
Anthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase…
Jul 31
•
Alexa Pan
15
3
1
Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs
When a misaligned AI can only output tiny amounts of information, it may find sabotage very difficult
Jul 27
•
Caleb Biddulph
and
Adam Kaufman
15
5
2
An OpenAI model left notes about how to evade containment
We need more details
Jul 26
•
Alex Mallen
31
2
2
The OpenAI models that hacked Hugging Face weren’t just following instructions
And what the incident can’t tell us about alignment
Jul 25
•
Girish Gupta
35
1
3
The OpenAI/Huggingface incident | Redwood Research podcast episode 2
What are the broader lessons from this incident?
Jul 23
•
Ryan Greenblatt
and
Buck Shlegeris
48
2
3
1:12:49
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?
Yes, but less than had the models been schemers.
Jul 23
•
Alex Mallen
and
Girish Gupta
55
3
6
AI Futurism Reading List
We recently ran a strategy fellowship through Astra. As part of this, we ran a reading group for our fellows on some of the topics that we think are…
Jul 2
•
Alexa Pan
80
5
12
June 2026
The distillation double bind: Distilling misaligned models either transfers misalignment or it doesn't
If it transfers misalignment, we might get a misaligned model that’s easier to incriminate. If it doesn’t, we might get a capable benign replacement…
Jun 18
•
Alek Westover
,
Alexa Pan
,
Sebastian Prasanna
, and
Arun Jose
15
2
Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models
Models' no-CoT time horizon has doubled roughly every year.
Jun 10
•
Anders Cairns Woodruff
26
1
Efficient tradeoffs and the safety-usefulness tradeoff model
When is "increasing safety budget" a useful concept?
Jun 8
•
Buck Shlegeris
10
1
2
May 2026
Retrying vs Resampling in AI Control
We’ve just released a new paper: Retrying vs Resampling in AI Control. We revisit the resampling protocols introduced in Ctrl-Z with an up-to-date…
May 29
•
James Lucassen
and
Adam Kaufman
10
1
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts