But my guess is that it would have been an out-of-scope ??? confused question if you had asked them to say what exactly was the Grader. It would be like asking a human to point to a material substance that was Failure. Many humans would try, but not in a very coherent way, and the answer you got would mainly depend on how you asked the question.
I disagree completely on this one particular point, which makes me wonder whether I interpreted the rest of the post correctly. It seems obvious to me that if you ask a model to identify the grader, it would alway... (read more)
This shows that Qwen can be led into thinking of itself as a misaligned liar, with the right question. I very much doubt it does so by default, but there might be some seemingly-benign user queries that would also push it into that basin, which is a bit concerning.
Incidentally: You might have expected that Qwen thinks of things as a text predictor, reasoning something like "if I am aligned, I should answer yes, if I am misaligned I should also answer yes, therefore the most likely response token is yes". But this seems to show that it doesn't work that way.
It's something introduced by the more agentic coding capabilities. Like if there's a bug Claude is trying to fix and its current angle of attack to it isn't being useful, then eventually it makes sense to recognize that and switch tactics. Claude playing Pokemon also had the issue of being willing to repeat the same ineffective strategies over and over and getting stuck because of that. Could that unexpectedly generalize to something like a desire for variety in conversation?
This is exactly what I was thinking while reading the post. They didn't advertise ... (read more)
(Gemini did actually write much of the Gemini_Plays_Pokemon scaffolding, but only in the sense of doing what David told it to do, not designing and testing it.)
I think you're probably right that a LLM coding its own scaffolding is probably more achievable than one playing the game like a human, but I don't think current models can do it - watching the streams, the models don't seem like they understand their own flaws, although admittedly they haven't been prompted to focus on this.
On the other hand, Claude has (arguably) a better pathfinding tool. As long as it requests to be moved to a valid set of coordinates from the screenshot overlay grid, the tool will move it there. Gemini mostly navigates on its own, although it has access to another instance of Gemini dedicated just to pathfinding.
I very much argue this. Claude's navigator tool can only navigate to coordinates that are onscreen, meaning that the main model needs to have some idea of where it's going. Which means grappling with problems that are extremely difficult for both ... (read more)
I have not tested if Gemini can distinguish this tree (and intend to eventually). This may very well be the only reason Gemini has progressed further.
You missed an important fact about the Gemini stream, which is that it just reads the presence of these trees from RAM and labels them for the model (along with a few other special tiles like ledges and water). Nevertheless I do think Gemini's vision is better, by which I mean if you provide it a screenshot it will sometimes identify the correct tree, unlike Claude who will never do so. (Although to my knowle... (read more)
And now in the second run it has entered a similar delusional loop. It knows the way to Cerulean City is via Route 4, but the route before and after Mt. Moon are both considered part of Route 4. Therefore it deluded itself into thinking it can get to Cerulean from the first part of the route. Because of that, every time it accidentally stumbles into Mt Moon and is making substantial progress towards the exit, it intentionally blacks out to get teleported back outside the entrance, so it can look for the nonexistent path forwards.
From what I've seen on stre... (read more)
Note that the creator stated that the setup is intentionally somewhat underengineered:
I do not claim this is the world's most incredible agent harness; in fact, I explicitly have tried not to "hyper engineer" this to be like the best chance that exists to beat Pokemon. I think it'd be trivial to build a better computer program to beat Pokemon with Claude in the loop.
This is like meant to be some combination of like "understand what Claude's good at and Benchmark and understand Claude-alongside-a-simple-agent-harness", so what that boils down to is this is like a pretty straightforward tool-using agent.
This basically sums up how it's doing: https://www.reddit.com/r/ClaudePlaysPokemon/comments/1j568ck/the_mount_moon_experience
Of course much of that is basic capability issues -poor spatial reasoning, short term memory that doesn't come anywhere close to lasting for 1 lap, etc.
But I've also noticed ways in which Claude's personality is sabotaging it. Claude is capable of taking notes saying that it "THOROUGHLY confirmed NO passages" through the eastern barrier - but never gets impatient or frustrated, so this doesn't actually prevent it from trying the same... (read more)
Doesn't strike me as inevitable at all, just a result of OpenAI following similar methods for creating their tokenizer twice. (In both cases, leading to a few long strings being included as tokens even though they don't actually appear frequently in large corpuses.)
They presumably had already made the GPT-4 tokenizer long before SolidGoldMagikarp was discovered in the GPT-2/GPT-3 one.
Prior to OpenAI's 2023-02-14 patching of ChatGPT (which seemingly prevents it from directly encountering glitch tokens like ‘ petertodd’)
I've never seen it mentioned around here, but since that update, ChatGPT is using a different tokenizer that has glitch tokens of its own:
500 years still sounds optimistic to me.
I know this is somewhat taboo to mention around here, but I'll say it if no one else will: I stay off Twitter because after the Elon Musk purchase, its fairly overt purpose is to convert the general public to fascism, it has made major progress on that front, and contributing to the gravitational mass that keeps people on that website strikes me as deeply immoral.