Language is always ambiguous. Experience and beliefs cannot be encoded truly precisely, no matter how big your vector. The human advice is "use more words". Instead of "green apples", say "unripe apples" or "ripe apples like granny smith, which are green".
Advice for LLMs is probably "use a more precise shared embedding space", which kind of boils down to the same thing. More bits of transmission = more fidelity of concept.
Read this short exchange.
A: "Green apples are delicious."
B: "Huh? Aren't they better when they're ripe?"
A: "No, I meant Granny Smiths."
A said "Green apples" intending Granny Smiths — and of course A thought it would be understood that way.
But that "claim" never reached B.
So — where was it lost?
Now, the next one.
A: "Green apples are delicious."
B: "Huh? Aren't they better when they're ripe?"
A: "No, I like green, sour apples."
The same sentence — "Green apples are delicious" — now comes from a different "claim."
So: did the sentence "Green apples are delicious" ever contain a "claim" in the first place?
Isn't this the same root as jailbreaks, bugs, and misunderstanding?
Here is that consideration: That the vulnerabilities of programs, LLMs, and language are one structure
Translated from the Japanese with LLM assistance.