This is a floor estimate of the abilities of superbeings. There could be such a thing as naturalistic traceless interventions by supersimulators in such a manner that the simulacra would be unable to notice, anticipate or even explain. Thus drawing distributions which rule the Donald Trump election in or out as evidence of a (bored to local heat death) simulator are both a little dubious (ie there is a simulacra frame of reference where it is not possible to tell apart whether "I don't care about Trump's election" is a local projection of "... I don't ca...
or maybe we await the theory which could do to intelligence what thermodynamics managed for temperature.
I have a good vague feeling about your macro point here, that microstates (K.E) can be aggregated into macrostates (Temperature).
However, i'm not even sure temperature itself is a good y axis. Thermodynamics, statistical physics etc are attractive when they explain, anticipate catastrophic phase transitions where both the microstate and macrostate would be saying something like "bro I only got got 5 degrees hotter from 100" .
I would thus then think a...
Another way of considering the hiding problem is exhaustiveness (Marks et al. 2026). If all agency is grounded in the persona, then a low-dimensional intervention should suffice; if there is non-persona agency (the "router" or "shoggoth" views; Marks et al. 2026), bad behavior could route around persona-level interventions. Hopefully, low-dimensional structure gives us a handle to empirically settle this.
Yes. This is a signficant worry. I couldn't help thinking "is the bomb big enough?" when reading the introductory parts of the work. Seems to me that one ...
If you follow up and ask why it wrote whatever it wrote, it will consistently claim that it recieved what it wrote from your own prompt, which is implies the model of it as continuing your prompt.
This is seems to be an isssue with role perception.
People are downvoting you without necessarily knowing where the crux is and that's a little unfair.
Is it that
1) Strange things won't happen with AI?
2) There doesn't need to be any action taken by governments if strange things happen
3) Even if strange things happen (or are anticipated to happen) and government responds, it'll be ineffectual and make a bad problem worse?
(3 and 2 are slightly different. For example, you hold (2) if you believe corporate governance is sufficient for any AI strangeness, if any.)
latter has too optimistic a prognosis. The myopic, unambitious misalignment that we seem to have seen here is definitely less scary than ambitious long-term goals shared between all instances,
There may be a need to be a bit of clarification here. We need to agree what can be usefully described as a myopic goal and to which agent. An agent can be myopic in the trajectory of actions, or myopic in its goals.
I believe that passing ExploitBench is a terminal goal. A myopic subgoal would be something that severely decreases the probability of success. Buildin...
I'll use IFT as an example here, although with some difficulty, this can be carried forward to RLHF, etc.
We can start with some deliberate IFT pretraining budget at a fraction , say
Can we observe some element of fractality, and is this useful? If one takes their action space as "token space" it seems to me that LLMs are strongly competent, general long horizon actors, increasingly approximating AIXI. ie they're increasingly efficient navigators of the token landscape. Their trajectory class in token landscape can be expanded (pretraining), constrained (RLHF, OPSD, RLVR etc). Thus I'm positive we may be able to build a competent micro theory. I'm however not sure that we can construct a restrictive partition function like thing which ...
What is thinkish and neuralese for single backbone multimodal models?
The most dangerous enemies are found among the most powerful agents, not the most ideologically distant ones.
This is the most impressive sentence I have read this year.
The Maasai and Kikuyu are two nations in Kenya separated by around 30 thousand years.
Through movements, they ended up together in Kenya as neighbours. Long story short, the British come over. The Kikuyus make deals with the Brits to weaken the Maasai. The Maasai make deals with the Brits to weaken the Kikuyu. The two groups end up as farm labourers in British settler farms
P_learner(desired behavior) → 𝜺 > 0 and that we have a way to condition the model to act in the desired way.
I'm very interested in this, particularly in the varying of training budgets, and whether we can get some scaling type empirical control of P_model(desired behaviour) and relatedly whether this affects conditioning.
Occam’s razor is mis-slicing here. The absolute measure of complexity is not what’s relevant. It’s that neither theory is simplifiable: they cannot be easily made simpler from what they are (at least not without falling apart completely). If they aren’t simplifiable, then there is very little leakage of extraneous information.
I wonder whether Occam's Razor slices postulates or structure. Slicing of structure leads to representational issues. ie There could be objects in particular problems which seem unwieldy in matrix mechanics that are washed off by wave...
Hi Thomas, sorry this was not very clear,
I was trying to paint a rough sketch of "spillovers in context" most of which is done well in your scenario planning. My disagreement largely with the current rapidly developing "state-AI complex". The amount of private (and likely soon, public) capital, talent, state interest concentrated make it unlikely that TRT will be acceptable.
I would argue AI2040's linchpin is compute. Coordinate compute, open up almost everything else. Governments and companies want control of many other things, algorithms, talent, data e...
The strongest reason I would have to pay attention here is alignment faking. The model can be thought of as presenting a persona that is different from actual. A weaker one is sycophancy. I would however discount this because sycophancy is prompt conditioned. I however suspect persona control is system promptable in addition to persona.
I'm mostly convinced by Plan A (given the techno-economic assumptions).
I still struggle with TRT as plausible without radical internationalisation. Financial incentives are too strong for companies, geoeconomic incentives too strong for states. The social contract would likely be unstable and therefore hard to enforce. AI companies already have strong levers with governments and are likely to use them to obstruct, dodge, cheat or FUD. We don't have enough time for it to become incontrovertible that cigarette smoke causes cancer.
I would thus trade algor...
...A metagaming example can be classified as a habit If we remove obvious features of a test from the prompt, then the model that games the task because of a habit stops gaming.
If a model adopts a metagaming persona, then if we change the prompt in such a way that other, non-gaming persona surfaces, then the model might stop gaming.
If metagaming is caused by terminal reward seeking, then if we change the cues that help the model understand how its answers are graded, then the reward-seeker model stops gaming, at least, using its original strategy.
Strategic ga
Thanks Steve! I found the model behind ruthless blissmaxxing particularly helpful. Most times I have interacted with this it has been unsupported.
A few comments (which may be confounded by a picture of the agent as a connectionist or otherwise complex adaptive cognitive agent).
1. I buy the moral circle failure modes. There's however a possible failure mode of "graded" moral circles which severely underweigh the moral patienthood of agents far from the circle ( ideologically, geographically etc). One could for example envision a "national securitymaxxing" ...
There is another anthropologically interesting successionist position, that the couch potatoing, hedonism addled successor outlasts you then s/he is a worthy successor. Any talk of "beauty"and "meaningful" has to be surrendered to this metric.
I'm not a successionist, but grant that this position has sophistication. Our ancestors would find very many various descriptions of us being unworthy successors, and yet we collectively out organise, outlive and outwit them - the past is a(n easily conquerable) foreign country. Would the Old Ones find it objectionable to be succeeded by godlike, ungodly descendants?
Is there an isomorphism between the space of possible plans and agents? If such an isomorphism exists, then alignment is the slicing up of agent space with a view of constraining plan space. In this picture, the weaker claim on default behaviour comes from random navigation of the agent (and plan) space.
This likely means both formulations are valid.
Not really, I believe that it may be impossible to tell either way. This question is very similar to "what is it to be like a photon?". We just don't inhabit the frame of reference (that of a boundary condition preparer) which can answer this question, or even just litigate it.