x
This website requires javascript to properly function. Consider activating javascript to get access to all site functionality.
LESSWRONG
LW
Login
AI Safety — LessWrong
AI Safety
This page is a stub.
Subscribe
Discussion
Subscribe
Discussion
Posts tagged
AI Safety
Most Relevant
2
96
Announcing AIXI Labs
Cole Wyeth
,
Aram Ebtekar
,
michaelcohen
,
Matthias Dellago
,
Marcus Hutter
1mo
9
2
20
Synthetic Scalable Oversight
Alexander Heckett
,
Shivansh Gour
,
Riyaz Ahuja
,
Tate Rowney
,
ishinshah
1mo
0
2
19
Open Problems in Mechanistic Intepretability of Biological AIs
Ihor Kendiukhov
6d
0
2
11
The Case for Physical AI Safety
kaylene_stocking
,
Bear Häon
1mo
2
2
10
Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment
Roland Pihlakas
,
lenz
,
Three Laws
1mo
0
1
164
Existential AI safety needs an effective social movement. PauseAI is building it
Maxime Fournes
,
Espedair Street
2mo
54
1
122
Synthetic Persona Pretraining: Alignment from Token Zero
Julian Minder
,
Raghav Singhal
,
Viktor Moskvoretskii
,
Stefan Krsteski
,
ashtonanderson
,
rolandaydin
,
Robert West
3mo
27
1
78
Door's Locked, Try the Window
Prakrat Agrawal
,
Jérémy Scheurer
,
Ariel_
2mo
0
1
77
Research update: RL on Debate Games shows Proposal Accuracy uplift alongside Judge Hacking
lennie
,
joanv
,
Shi
,
Jacob Pfau
2mo
4
1
58
Do LLMs Have Desires?
Christopher Ackerman
2mo
12
1
51
Human-Guided Agentic Research: A Research Agenda
fastfedora
2mo
7
1
43
Persona Cartography: Charting Language Model Personality Traits in Weight Space
antonghawthorne
,
Mariia Koroliuk
,
Irakli Shalibashvili
,
sidbaines
,
Clément Dumas
,
Konstantinos Voudouris
,
David Africa
1mo
0
1
39
Don't normalize a permanent underclass (even a rich one)
hadad
1mo
7
1
39
Help us launch AI safety university groups by referring potential founders
thomasrodskog
,
Jason Chin
1mo
1
1
37
AI Safety Can't Afford a Second Cause
atlasaligned
1mo
9