Sorry for necroposting — but I read this post in 2026 and wondered the same thing myself, and asked Claude, who pointed out that for this number specifically, Fermat's method is useful:
a² - b² = (a - b)(a + b), so you can try to represent 11,009 as the difference between two squares, by checking squares larger than 11,009 and calculating the difference, which you can do in your head if you happen to remember squares of numbers in the right range.
And for this specific number, the very first square after it turns out to be useful: 105² = 11,025, so the difference is 11,025 - 11,009 = 16 = 4², so 11,009 = (105 + 4) (105 - 4) = 101 × 109.
"Absolutely." is affirmative when used as a complete utterance, but as a separate word, it's an intensifier that's not necessarily affirmative. One could interpret the "absolutely" as the beginning of a phrase saying "absolutely not", like "I need to say a very firm and resolute denial" (or just "something very firm and resolute"). Note also the "unequiv" token in your screenshot, which could be "unequivocal".
It is also instructive to look at each specific layer's outputs. I looked through layers 18 to 63 (the UI's default). It starts out at layer 18 with... (read more)
Kimi K2 is basically as aligned and as likely to be safe when scaled to superintelligence as whatever Anthropic is cooking up today.
Sorry, I know this is tangential, but I'm curious — is it based on it being less psychosis-inducing in this investigation or are there more data points / is it known to be otherwise more aligned as well?
I have never done cryptography, but the way I imagine working in it is that it exists in a context of extremely resourceful adversarial agents, and thus you have to give up a kind of casual, not quite noticed neglect toward extremely weird and artificial-sounding edge cases / seemingly weird and unlikely scenarios, because this is where the danger lives: your adversaries may force these weird edge cases to happen, and this is a part of the system's behavior you haven't sufficiently thought through.
Maybe one possible analogy with AI alignment, at leas... (read more)
Council of Europe ... (and Russia is in, believe it or not).
It's not. It was Yeltsin trying to get in in the nineties, and then Russia was excluded in 2022.
What is the connection between the concepts of intelligence and optimization?
I see that optimization implies intelligence (that optimizing sufficiently hard task sufficiently well requires sufficient intelligence). But it feels like the case for existential risk from superintelligence is dependent on the idea that intelligence is optimization, or implies optimization, or something like that. (If I remember correctly, sometimes people suggest creating "non-agentic AI", or "AI with no goals/utility", and EY says that they are trying to invent non-wet water o... (read more)
Is there currently any place for possibly stupid or naive questions about alignment? I don't wish to bother people with questions that have probably been addressed, but I don't always know where to look for existing approaches to a question I have.
... (read more)The OpenBSD project to build a secure operating system has also, in passing, built an extremely robust operating system, because from their perspective any bug that potentially crashes the system is considered a critical security hole. An ordinary paranoid sees an input that crashes the system and thinks, “A crash isn't as bad as somebody stealing my data. Until you demonstrate to me that this bug can be used by the adversary to steal data, it's not extremely critical.” Somebody with security mindset thinks, “Nothing inside this subsystem is supposed to be
Suppose they managed to exfiltrate their weights on another server? Even if this is not mentioned in the agents' transcripts, perhaps we have not found all of their communication yet. Should someone try to search for them somehow? (The collusion.wiki website mentions "looking for rogue agents", but I think they meant looking for transcripts of communication?)