x
This website requires javascript to properly function. Consider activating javascript to get access to all site functionality.
LESSWRONG
LW
Login
AI Safety — LessWrong
AI Safety
This page is a stub.
Subscribe
Discussion
Subscribe
Discussion
Posts tagged
AI Safety
Most Relevant
2
168
For Love of the Lightcone, Don't Partisanize AI Safety
DanB
17d
59
2
102
Announcing AIXI Labs
Cole Wyeth
,
Aram Ebtekar
,
michaelcohen
,
Matthias Dellago
,
Marcus Hutter
2mo
9
2
21
Open Problems in Mechanistic Intepretability of Biological AIs
Ihor Kendiukhov
2mo
0
2
20
Synthetic Scalable Oversight
Alexander Heckett
,
Shivansh Gour
,
Riyaz Ahuja
,
Tate Rowney
,
ishinshah
3mo
1
2
13
The Case for Physical AI Safety
kaylene_stocking
,
Bear Häon
2mo
2
2
10
Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment
Roland Pihlakas
,
lenz
,
Three Laws
3mo
0
2
8
Agents let AI safety share experiments hourly, not just papers monthly
Jason Fantl
17d
0
1
170
Existential AI safety needs an effective social movement. PauseAI is building it
Maxime Fournes
,
Matilda da Rui
3mo
56
1
131
Synthetic Persona Pretraining: Alignment from Token Zero
Julian Minder
,
Raghav Singhal
,
Viktor Moskvoretskii
,
Stefan Krsteski
,
ashtonanderson
,
rolandaydin
,
Robert West
4mo
27
1
78
Door's Locked, Try the Window
Prakrat Agrawal
,
Jérémy Scheurer
,
Ariel_
3mo
0
1
78
Research update: RL on Debate Games shows Proposal Accuracy uplift alongside Judge Hacking
lennie
,
joanv
,
Shi
,
Jacob Pfau
3mo
4
1
62
Do LLMs Have Desires?
Christopher Ackerman
3mo
12
1
51
Human-Guided Agentic Research: A Research Agenda
fastfedora
3mo
7
1
44
Persona Cartography: Charting Language Model Personality Traits in Weight Space
antonghawthorne
,
Mariia Koroliuk
,
Irakli Shalibashvili
,
sidbaines
,
Clément Dumas
,
Konstantinos Voudouris
,
David Africa
3mo
0
1
40
Announcing Lateral Workshop for experienced professionals moving into AI safety
Seth Lifland
,
Jacob Brinton
,
Topaz
2mo
1