Thanks for your answer!
The main philosophically tangled part is pointing to latent structure for which no good definition is available, for which we want to say something like "The presence of a diamond is the property that causes such-and-such observable regularities."
The distinction makes me think the missing object may be a validation layer rather than another kind of mechanistic explanation.
I’ve been approaching a connected problem through the distinction between verification (is the thing right?) and validation (is it the right thing?). A mechanistic ...
Happy to hear you're back with ARC! I'm a huge fan of the thought space.
Assuming a bounded, well specified alignment property, what ensures that the object connecting a hypothesis to a mechanistic explanation to a loss function preserves the intended meaning?
Also, why would training-gaming not move up one level then? from gaming behavioral evaluations to gaming the explanatory ontology or the mechanistic explanator?
Cunningham’s law: I make assertions about what the (emerging) field of AI verification should aim for, and people with experience in international policy, cybersecurity and any relevant field of engineering can point out what this draft gets wrong.
I used to like this law pre-LLMs, but it presumes verification is cheap relative to generation and lately that doesn't fully stand.
A frame I kept applying while reading is on a separation between security and safety. I had it described as:
- security asks how we stop the world from breaking into a system;
- safety...
Thanks for gathering these thoughts here! I recently stumbled on them and got pulled into the post.
Small objection I'd like to make: You write that spec tooling most likely needs to be solved before any of the other problems become solvable, but I think it's worse than that, because better tooling conserves the problem and relocates it.
Formalization splits one hard question (is this true of the world) into two:
This seems to want f...
For example, suppose a mechanism reliably detects something labelled "diamond". We still need to know whether it tracks: - actual diamonds, - images or descriptions of diamonds, - the training data’s annotation conventions, - some accidental correlate, etc.
I'm quite eager to see ARC's theory of how internal computational strucctures become grounded in external properties we mean to ask about.