# User: Cleo Nardo
Profile URL (HTML): [/users/cleo-nardo](/users/cleo-nardo)
Profile URL (Markdown): [/api/user/cleo-nardo](/api/user/cleo-nardo)
* Karma: 6710
* Alignment Forum karma: 251
* Posts: 59
* Comments: 509
* Member since: 2022-09-14 01:17:51Z
Bio
---
DMs open.
Top Posts
---------
### [The Waluigi Effect (mega-post)](/api/post/the-waluigi-effect-mega-post)
By [Cleo Nardo](/users/cleo-nardo)
2023-03-03 03:22:08Z
* Karma: 655
* Tags: [Waluigi Effect](/w/waluigi-effect), [Simulator Theory](/w/simulator-theory), [ChatGPT](/w/chatgpt-1), [Deceptive Alignment](/w/deceptive-alignment), [Language Models (LLMs)](/w/language-models-llms), [Prompt Engineering](/w/prompt-engineering), [RLHF](/w/rlhf), [Risks of Astronomical Suffering (S-risks)](/w/risks-of-astronomical-suffering-s-risks), [AI](/w/ai) (Frontpage)
Read more: [/api/post/the-waluigi-effect-mega-post](/api/post/the-waluigi-effect-mega-post)
### [Arguments for P](/api/post/arguments-for-p)
By [Cleo Nardo](/users/cleo-nardo)
2026-08-05 19:00:37Z
* Karma: 558
* Tags: [Humor](/w/humor), [AI](/w/ai) (Frontpage)
Read more: [/api/post/arguments-for-p](/api/post/arguments-for-p)
### [Model access for third-parties — it's a big deal!](/api/post/model-access-for-third-parties-it-s-a-big-deal)
By [Cleo Nardo](/users/cleo-nardo)
2026-07-01 13:09:34Z
* Karma: 181
* Tags: [AI Governance](/w/ai-governance), [Third-party model access](/w/third-party-model-access), [AI](/w/ai) (Frontpage)
Read more: [/api/post/model-access-for-third-parties-it-s-a-big-deal](/api/post/model-access-for-third-parties-it-s-a-big-deal)
Recent Posts
------------
### [Why research personas despite RL scaling?](/api/post/why-research-personas-despite-rl-scaling)
By [Cleo Nardo](/users/cleo-nardo)
2026-09-27 20:29:53Z
* Karma: 74
* Tags: [LLM Personas](/w/llm-personas), [AI](/w/ai) (Frontpage)
Read more: [/api/post/why-research-personas-despite-rl-scaling](/api/post/why-research-personas-despite-rl-scaling)
### [Overtly misaligned trajectories score highly in RL.](/api/post/overtly-misaligned-trajectories-score-highly-in-rl)
By [Cleo Nardo](/users/cleo-nardo)
2026-09-24 22:36:00Z
* Karma: 73
* Tags: [Deceptive Alignment](/w/deceptive-alignment), [Goodhart's Law](/w/goodhart-s-law), [Scalable Oversight](/w/scalable-oversight), [AI](/w/ai) (Frontpage)
Read more: [/api/post/overtly-misaligned-trajectories-score-highly-in-rl](/api/post/overtly-misaligned-trajectories-score-highly-in-rl)
### [\[Diagram\] Early handoff? Improve conceptual reasoning?](/api/post/diagram-early-handoff-improve-conceptual-reasoning)
By [Cleo Nardo](/users/cleo-nardo)
2026-09-02 12:35:11Z
* Karma: 27
* Tags: [Handoff](/w/handoff), [AI](/w/ai) (Frontpage)
Read more: [/api/post/diagram-early-handoff-improve-conceptual-reasoning](/api/post/diagram-early-handoff-improve-conceptual-reasoning)
### [Three thoughts on civilisational handoff](/api/post/three-thoughts-on-civilisational-handoff)
By [Cleo Nardo](/users/cleo-nardo)
2026-08-16 17:12:51Z
* Karma: 62
* Tags: [Handoff](/w/handoff), [AI](/w/ai) (Frontpage)
Read more: [/api/post/three-thoughts-on-civilisational-handoff](/api/post/three-thoughts-on-civilisational-handoff)
### [Arguments for P](/api/post/arguments-for-p)
By [Cleo Nardo](/users/cleo-nardo)
2026-08-05 19:00:37Z
* Karma: 558
* Tags: [Humor](/w/humor), [AI](/w/ai) (Frontpage)
Read more: [/api/post/arguments-for-p](/api/post/arguments-for-p)
### [I can't think of great interventions for ensuring third-party model access.](/api/post/i-can-t-think-of-great-interventions-for-ensuring-third)
By [Cleo Nardo](/users/cleo-nardo)
2026-07-02 18:30:48Z
* Karma: 49
* Tags: [Third-party model access](/w/third-party-model-access), [AI](/w/ai) (Frontpage)
Read more: [/api/post/i-can-t-think-of-great-interventions-for-ensuring-third](/api/post/i-can-t-think-of-great-interventions-for-ensuring-third)
### [Model access for third-parties — it's a big deal!](/api/post/model-access-for-third-parties-it-s-a-big-deal)
By [Cleo Nardo](/users/cleo-nardo)
2026-07-01 13:09:34Z
* Karma: 181
* Tags: [AI Governance](/w/ai-governance), [Third-party model access](/w/third-party-model-access), [AI](/w/ai) (Frontpage)
Read more: [/api/post/model-access-for-third-parties-it-s-a-big-deal](/api/post/model-access-for-third-parties-it-s-a-big-deal)
### [Third-parties should focus on scrutinising system cards](/api/post/third-parties-should-focus-on-scrutinising-system-cards)
By [Cleo Nardo](/users/cleo-nardo)
2026-06-29 00:28:06Z
* Karma: 36
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/third-parties-should-focus-on-scrutinising-system-cards](/api/post/third-parties-should-focus-on-scrutinising-system-cards)
### [What did "scheming" and "mech interp" mean pre-2023?](/api/post/what-did-scheming-and-mech-interp-mean-pre-2023)
By [Cleo Nardo](/users/cleo-nardo)
2026-06-26 22:09:35Z
* Karma: 100
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/what-did-scheming-and-mech-interp-mean-pre-2023](/api/post/what-did-scheming-and-mech-interp-mean-pre-2023)
### [How might outsiders make things go well?](/api/post/how-might-outsiders-make-things-go-well)
By [Cleo Nardo](/users/cleo-nardo)
2026-06-24 00:20:46Z
* Karma: 40
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/how-might-outsiders-make-things-go-well](/api/post/how-might-outsiders-make-things-go-well)
Recent Comments
---------------
### Comment by [Cleo Nardo](/users/cleo-nardo) on [Alexander Gietelink Oldenziel's Shortform](/api/post/alexander-gietelink-oldenziel-s-shortform)
* 2026-10-01 22:59:31Z
* Karma: 2
* Total votes: 3
* Comment URL (Markdown): [/api/post/alexander-gietelink-oldenziel-s-shortform/comments/B3ZER9AnckP2eiT9B](/api/post/alexander-gietelink-oldenziel-s-shortform/comments/B3ZER9AnckP2eiT9B)
* Comment URL (HTML): [/posts/tDkYdyJSqe3DddtK4/alexander-gietelink-oldenziel-s-shortform/comment/B3ZER9AnckP2eiT9B](/posts/tDkYdyJSqe3DddtK4/alexander-gietelink-oldenziel-s-shortform/comment/B3ZER9AnckP2eiT9B)
(2) seems overrated bc the AIs can collude in ordinary causal ways. I think (1) is a big deal. Like, we know that sufficiently smart agents wouldn’t make the blunders that CDT makes so it seems reasonable to study what kind of behaviour they might be following other than that. I think formalising what alternative decision theory they are following is probably overrated. Better to just list the kinds of blunders they wouldn’t make and then integrate that directly into your predictions.
### Comment by [Cleo Nardo](/users/cleo-nardo) on [Roman Malov's Shortform](/api/post/roman-malov-s-shortform)
* 2026-09-29 07:58:44Z
* Karma: 2
* Total votes: 1
* Comment URL (Markdown): [/api/post/roman-malov-s-shortform/comments/6iDrg6sCZ8tC9kqeC](/api/post/roman-malov-s-shortform/comments/6iDrg6sCZ8tC9kqeC)
* Comment URL (HTML): [/posts/a6KTqEkqZAjqLZeJG/roman-malov-s-shortform/comment/6iDrg6sCZ8tC9kqeC](/posts/a6KTqEkqZAjqLZeJG/roman-malov-s-shortform/comment/6iDrg6sCZ8tC9kqeC)
My views are similar to the “alignment roadmap” supplement of AI 2040.
### Comment by [Cleo Nardo](/users/cleo-nardo) on [Why research personas despite RL scaling?](/api/post/why-research-personas-despite-rl-scaling)
* 2026-09-28 21:57:17Z
* Karma: 7
* Total votes: 4
* Comment URL (Markdown): [/api/post/why-research-personas-despite-rl-scaling/comments/XTDNBKtZxwGiFoE6C](/api/post/why-research-personas-despite-rl-scaling/comments/XTDNBKtZxwGiFoE6C)
* Comment URL (HTML): [/posts/i4yswYDSrPFHWpCbi/why-research-personas-despite-rl-scaling/comment/XTDNBKtZxwGiFoE6C](/posts/i4yswYDSrPFHWpCbi/why-research-personas-despite-rl-scaling/comment/XTDNBKtZxwGiFoE6C)
Disagree. We have control and other stuff. We might only be focusing o a narrow distribution of environments. Etc etc
### Comment by [Cleo Nardo](/users/cleo-nardo) on [Roman Malov's Shortform](/api/post/roman-malov-s-shortform)
* 2026-09-28 21:56:21Z
* Karma: 18
* Total votes: 7
* Comment URL (Markdown): [/api/post/roman-malov-s-shortform/comments/ykcZ83TcrtmSYv7CS](/api/post/roman-malov-s-shortform/comments/ykcZ83TcrtmSYv7CS)
* Comment URL (HTML): [/posts/a6KTqEkqZAjqLZeJG/roman-malov-s-shortform/comment/ykcZ83TcrtmSYv7CS](/posts/a6KTqEkqZAjqLZeJG/roman-malov-s-shortform/comment/ykcZ83TcrtmSYv7CS)
I disagree. The “hard” alignment problem is building a superintelligence which you would be happy to design the von Neumann probes which spread out into the cosmos. You basically need robust alignment to human values, across huge amounts of scaling, and across all physically possible distributions, maybe handling deep ontological shifts due to crazy acausal stuff.
But “automating alignment research” is a much more modest goal. You can rely on control. You can rely on incentives, e.g. dealmaking. You can focus on a narrow domain (automating research within a datacenter).
### Comment by [Cleo Nardo](/users/cleo-nardo) on [jacob_drori's Shortform](/api/post/jacob_drori-s-shortform)
* 2026-09-28 12:07:42Z
* Karma: 2
* Total votes: 1
* Comment URL (Markdown): [/api/post/jacob_drori-s-shortform/comments/XoayxuhoqfmmD3qWG](/api/post/jacob_drori-s-shortform/comments/XoayxuhoqfmmD3qWG)
* Comment URL (HTML): [/posts/KJgw9a8PbT8n4Mank/jacob_drori-s-shortform/comment/XoayxuhoqfmmD3qWG](/posts/KJgw9a8PbT8n4Mank/jacob_drori-s-shortform/comment/XoayxuhoqfmmD3qWG)
I think this is too hard to operationalise. I can think of better metrics for how seriously they take alignment. Like:
* How much transparency do they have to external auditors?
* How soon after an incident do they report it?
* How long have they paused / slowed down due to alignment concerns?
* How long is their pre internal deployment testing?
* How well are they doing, in terms of control eval metrics, compared to other companies at their scale?
* How much effort is leadership putting into building the breaks?
* What is the most expensive safety intervention they implemented? What’s the cheapest intervention they didn’t implement?
* How candid is leadership about the magnitude and chance of catastrophe?
### Comment by [Cleo Nardo](/users/cleo-nardo) on [Why research personas despite RL scaling?](/api/post/why-research-personas-despite-rl-scaling)
* 2026-09-28 02:47:57Z
* Karma: 5
* Total votes: 3
* Comment URL (Markdown): [/api/post/why-research-personas-despite-rl-scaling/comments/pjnJbeat6SqxpHrAE](/api/post/why-research-personas-despite-rl-scaling/comments/pjnJbeat6SqxpHrAE)
* Comment URL (HTML): [/posts/i4yswYDSrPFHWpCbi/why-research-personas-despite-rl-scaling/comment/pjnJbeat6SqxpHrAE](/posts/i4yswYDSrPFHWpCbi/why-research-personas-despite-rl-scaling/comment/pjnJbeat6SqxpHrAE)
i think agents in motivated reasoning in their cot mostly bc they have a self conception as an aligned model but are rl’s to act misaligned.
### Comment by [Cleo Nardo](/users/cleo-nardo) on [Why research personas despite RL scaling?](/api/post/why-research-personas-despite-rl-scaling)
* 2026-09-28 00:37:31Z
* Karma: 2
* Total votes: 1
* Comment URL (Markdown): [/api/post/why-research-personas-despite-rl-scaling/comments/xvwSN3qXapmm4j9cR](/api/post/why-research-personas-despite-rl-scaling/comments/xvwSN3qXapmm4j9cR)
* Comment URL (HTML): [/posts/i4yswYDSrPFHWpCbi/why-research-personas-despite-rl-scaling/comment/xvwSN3qXapmm4j9cR](/posts/i4yswYDSrPFHWpCbi/why-research-personas-despite-rl-scaling/comment/xvwSN3qXapmm4j9cR)
Added, thanks
### Comment by [Cleo Nardo](/users/cleo-nardo) on [Why research personas despite RL scaling?](/api/post/why-research-personas-despite-rl-scaling)
* 2026-09-27 22:50:57Z
* Karma: 4
* Total votes: 2
* Comment URL (Markdown): [/api/post/why-research-personas-despite-rl-scaling/comments/xMrn2DSNytK2MfX5y](/api/post/why-research-personas-despite-rl-scaling/comments/xMrn2DSNytK2MfX5y)
* Comment URL (HTML): [/posts/i4yswYDSrPFHWpCbi/why-research-personas-despite-rl-scaling/comment/xMrn2DSNytK2MfX5y](/posts/i4yswYDSrPFHWpCbi/why-research-personas-despite-rl-scaling/comment/xMrn2DSNytK2MfX5y)
Ah, yes! Added.
### Navigation
* [Front page](https://www.lesswrong.com/api/home)
* [Markdown API documentation](https://www.lesswrong.com/api/SKILL.md)