x
This website requires javascript to properly function. Consider activating javascript to get access to all site functionality.
LESSWRONG
LW
Login
AI Safety — LessWrong
AI Safety
This page is a stub.
Subscribe
Discussion
Subscribe
Discussion
Posts tagged
AI Safety
Most Relevant
2
8
Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment
Roland Pihlakas
,
lenz
,
Three Laws
6d
0
1
163
Existential AI safety needs an effective social movement. PauseAI is building it
Maxime Fournes
,
Espedair Street
17d
54
1
117
Synthetic Persona Pretraining: Alignment from Token Zero
Julian Minder
,
Raghav Singhal
,
Viktor Moskvoretskii
,
Stefan Krsteski
,
ashtonanderson
,
rolandaydin
,
Robert West
2mo
26
1
77
Research update: RL on Debate Games shows Proposal Accuracy uplift alongside Judge Hacking
lennie
,
joanv
,
Shi
,
Jacob Pfau
11d
3
1
68
Door's Locked, Try the Window
Prakrat Agrawal
,
Jérémy Scheurer
19d
0
1
57
Do LLMs Have Desires?
Christopher Ackerman
16d
10
1
48
Human-Guided Agentic Research: A Research Agenda
fastfedora
14d
7
1
37
Persona Cartography: Charting Language Model Personality Traits in Weight Space
antonghawthorne
,
Mariia Koroliuk
,
Irakli Shalibashvili
,
sidbaines
,
Clément Dumas
,
Konstantinos Voudouris
,
David Africa
3d
0
1
36
The case for fine-grained tracking of compute for AI
Farhan
,
Katherine Biewer
2mo
17
1
34
AI Safety Can't Afford a Second Cause
atlasaligned
6d
8
1
32
[paper] Training on Documents About Monitoring Leads to CoT Obfuscation
Reilly Haskins
,
bilalchughtai
,
Josh Engels
2mo
1
1
31
A brief list of ways AI safety efforts could be net negative
Elias Schmied
24d
4
1
27
Don't normalize a permanent underclass (even a rich one)
hadad
3d
6
1
21
Hannibal Mistral: the Mistral family has a problem with persona-conditioned elicitation
vigji
1mo
0
1
21
Scheming Evals Mislead in Both Directions
Chijioke Ugwuanyi
,
eric-z
,
TerryJCZhang
10d
0