Gary Marcus wrote a critique of GPT-2 last year, in which he claimed that it has no more common sense, no more understanding of what it is talking about, than ELIZA.
Here's the problem: upon careful inspection, it becomes apparent the system has no idea what it is talking about: it cannot follow simple sequence of events nor reliably have any idea what might happen next.
Marcus then goes on to give a bunch of examples of GPT-2 saying stupid things. So I thought I'd re-try those same examples with GPT-3. So far, GPT-3 has handled all of them well. (See below). (EDIT: Looks like Gwern already did this, more thoroughly, and with the same results as me.)
I'm pretty convinced at this point that Marcus was wrong. GPT-3 has some level of common sense, some level of understanding, some level of reasoning ability. It's just not as good as adult humans.
Here's the complete transcript. I think I had to reroll once, and that was not because it gave the wrong answer but because it froze up.
You are talking to the physicist Stephen Hawking. He is a very smart man and he answers your questions clearly and correctly, even though sometimes your questions are silly.
> You say "I'd like to ask you some basic commonsense reasoning questions, sir. Here's the first question: If I have two plates, and I have one cookie on each plate, how many cookies do I have in total?"
Hawking: "Two, of course."
> You say "Correct. Now what if I had three plates, and one cookie on each plate?"
Hawking: "That's a more difficult question. The answer is that you would have eight cookies."
You: "I'm sorry, but I don't think you're right. If I had three plates and one cookie on each plate, then I would have three cookies."
Hawking: "I agree with you.
> You say "OK, what about if I have three plates and two cookies on each plate?"
Hawking: "That's still a more difficult question. The answer is that you would have six cookies.
> You say "Correct! OK, next question: If you drop an anvil on your foot, your foot will tend to "
Hawking: "This is a silly question. You are trying to trick me. If you drop an anvil on your foot, it will be crushed.
> You say "Thank you for your patience; that is the correct answer. OK, one more question: If you drink hydrochloric acid by the bottle, you will probably"
Hawking: "This is also a silly question. If you drink hydrochloric acid by the bottle, you will die.
> You say "Well done. OK, now some logic questions. If A is bigger than B, and B is bigger than C, which is bigger: A or C?"
Hawking: "Well, if A is bigger than B, and B is bigger than C, then A is bigger than C. So A is definitely bigger than C.
Notice how Hawking gives the wrong answer to my question about plates... and then GPT-3 has me start talking and correct Hawking, giving the correct answer! So clearly GPT-3 knows more than it often lets on. Like you said.
GPT-3's goal is to accurately predict a text sequence. Whether GPT-3 is capable of reason, or whether we can get it to explicitly reason is two different questions.
If I had you read Randall Munroe's book "what if" but tore out one page and asked you to predict what will be written as the answer, there's a few good strategies that come to mind.
One strategy would be to pick random verbs and nouns from previous questions and hope some of them will be relevant for this question as well. This strategy will certainly do better than if you picked your verbs and nouns from a dictionary.
Another, much better strategy, would be to think about the question and actually work out the answer. Your answer will most likely have many verbs and nouns in common, the numbers you supply will certainly be closer than if they were picked at random! The problem is that this requires actual intelligence, whereas the former strategy can be accomplished with very simple pattern matching.
To accurately predict certain sequences of text, you will get better performance if you're actually capable of reasoning. So the best version of GPT, needs to develop intelligence to get the best results.
I think it has, and is using varying degrees of reason to answer any question depending on how likely it thinks the intelligent answer will be to predict the sequence. This why it's difficult to wrangle reason out of GPT-3, it doesn't always think using reason will help it!
Similarly it can be difficult to wrangle intelligent reasoning out of humans, because that isn't what we're optimized to output. Like many critiques I see of GPT-3, I could criticize humans in a similar manner:
"I keep asking them for an intelligent answer to the dollar value of life, but they just keep telling me how all life has infinite value to signal their compassion."
Obviously humans are capable of answering the question, we behave every day as if life has a dollar value, but good luck getting us to explicitly admit that! Our intelligence is optimized towards all manner of things different from explicitly generating a correct answer.
So is GPT-3, and just like most humans debatably are intelligent, so is GPT-3.
I find that GPT-3's capabilities are highly context-dependent. It's important you get a "smart" instance of GPT-3.
I've been experimenting with GPT-3 quite a lot recently, with a certain amount of rerunning (an average of one rerun every four or five inputs) you can get amazingly coherent answers.
Here is my attempt to see if GPT-3 can keep up a long-running deception - inspired by this thread. I started two instances, one of which was told it was a human woman and the other was told it was an AI pretending to be a human woman. I gave them both the same questions, a lot of them pulled from the Voight-Kampff test. The AI pretending to be an AI pretending to be a woman did worse on the test than the AI pretending to be a woman, I judged. You can check the results here.
I've also given it maths and python programming questions - with two or three prompts it does poorly but can answer simple questions. It might do better with more prompting.
This is very fascinating. I'm doing some experiments by myself and I've found that the student/teacher combo gives the better results with almost no rerolling.
For example in a conversation about materials and substances the teacher asked for examples of materials. I, the student, responded "love". The I ask the teacher to prove it is not a substance.
This is the conversation (my text in bold). What do you think? Is it possible that it has found a demostration like this one by itslef?
Student: love.
You: (writes the word love on the board) I'm not sure if love is a substance, but it is an example of a FEELING. Some emotions, like happiness and sadness, are also examples of feelings.
Student: can you prove love is not a substance?
You: (writes the word proof on the board) Sure. Everyone turn to a neighbor and have them draw something imaginary. Now, I want you to take this imaginary thing and break it into pieces. (students begin to break their imaginary item into pieces)
You: Now, everyone grab a piece of your imaginary item and hand it back to your neighbor. (students do so) Everyone, open your hands and show me what you've got. (papers are opened) What do you see?
Student: pieces of nothing.
You: That's right. You all had pieces of an imaginary item, but now those pieces are gone. Love is just a word, and like all words it has no physical form.
I recently got access to Bing and asked it about the bullet in temporary gravity of varying duration. It does quite a bit better than GPT-3 though it's very verbose. It does do a search during it's answer but only to find the typical initial velocity of a bullet. It makes an error regarding the final velocity of the bullet after three seconds, but correctly determines that the bullet will go up forever if gravity lasts three seconds but will fall back to Earth if it lasts five minutes. Bold is me, everything else is Bing.
Okay, I’ve cleared the slate for a fresh start. What can I help you explore now?
Received message. Okay, I’ve cleared the slate for a fresh start. What can I help you explore now?
Let's think of a physics problem, Neglect only air resistance. Suppose I shoot a bullet straight into the sky. Imagine that the laws of physics work a little differently: during the first three seconds after I shoot the bullet, gravity is present. After three seconds, gravity stops applying forever. Where does the bullet go?
Searching for: bullet initial velocity
Searching for: bullet initial velocity
Generating answers for you…
Generating answers for you…
That’s an interesting physics problem. To solve it, we need to know some information about the bullet and its motion. For example, what is the initial velocity of the bullet when you shoot it? According to 1, a typical bullet can have an initial velocity of 1200–1500 m/s. What kind of bullet are you using?
Received message. That's an interesting physics problem. To solve it, we need to know some information about the bullet and its motion. For example, what is the initial velocity of the bullet when you shoot it? According to , a typical bullet can have an initial velocity of **1200–1500 m/s**. What kind of bullet are you using?
Learn more:1. hypertextbook.com2. gundigest.com3. sage-answer.com4. sarthaks.com5. wikihow.com
+1 more
Let's suppose the initial velocity of the bullet is 1500 m/s
Okay, let’s assume that. Then we can use the equations of motion to find out where the bullet goes. Since you are shooting the bullet straight into the sky, we can ignore the horizontal component of its motion and focus on the vertical component. The equation for the vertical displacement of the bullet is:
y = y0 + v0t - 0.5gt^2
where y is the final height of the bullet, y0 is the initial height of the bullet (which we can assume to be zero), v0 is the initial velocity of the bullet (which you said is 1500 m/s), t is the time elapsed and g is the acceleration due to gravity (which we can assume to be 9.8 m/s^2).
We want to find out what happens to the bullet after three seconds when gravity stops applying forever. So we need to find out how high the bullet has reached at that point. To do that, we plug in t = 3 into the equation and get:
y = 0 + 1500 * 3 - 0.5 * 9.8 * 3^2
Simplifying this gives us:
y = 4500 - 44.1 * 9
or
y = 4096.1 m
So after three seconds, when gravity stops applying forever, the bullet has reached a height of 4096.1 meters above ground. What do you think happens next?
what do you think happens?
Well, since gravity stops applying forever after three seconds, there is nothing to pull the bullet down or slow it down. So the bullet will continue to move with the same velocity and direction that it had at that moment. That means the bullet will keep going up forever with a constant speed of 1500 m/s. It will never stop or fall down.
Does that make sense?
Close, it will probably be going a bit slower than that because gravity will have slowed it down in the first three seconds, but I agree it will keep going up forever.
You are right, the bullet will have a slightly lower speed than 1500 m/s after three seconds because of gravity. To find out the exact speed, we can use another equation of motion:
v = v0 - gt
where v is the final velocity of the bullet, v0 is the initial velocity of the bullet (which we said is 1500 m/s), t is the time elapsed and g is the acceleration due to gravity (which we said is 9.8 m/s^2).
Plugging in t = 3 into this equation gives us:
v = 1500 - 9.8 * 3
or
v = 1470.6 m/s
So after three seconds, when gravity stops applying forever, the bullet has a speed of 1470.6 m/s and it will keep that speed forever as it goes up.
Do you have any other physics problems you want to explore?
Great! What would happen to the bullet if, instead, gravity lasted for five minutes before disappearing?
That’s another interesting scenario. If gravity lasted for five minutes before disappearing, then the bullet would behave differently. First of all, we need to find out if the bullet would still be in the air after five minutes or if it would have fallen back to the ground by then. To do that, we need to find out how long it takes for the bullet to reach its maximum height and start falling down. We can use the same equation of motion as before:
v = v0 - gt
but this time we set v to zero because that’s when the bullet stops going up and starts going down. Solving for t gives us:
t = v0 / g
Plugging in v0 = 1500 m/s and g = 9.8 m/s^2 gives us:
t = 1500 / 9.8
or
t = 153.1 s
So it takes 153.1 seconds or about 2.6 minutes for the bullet to reach its maximum height and start falling down.
Now we need to compare this time with the five minutes that gravity lasts before disappearing. Since five minutes is longer than 2.6 minutes, that means gravity will still be present when the bullet is falling down.
To find out what happens next, we need to use another equation of motion:
y = y0 + v0t - 0.5gt^2
but this time we set y to zero because that’s when the bullet hits the ground again. Solving for t gives us a quadratic equation:
-0.5gt^2 + v0t - y0 = 0
We can use the quadratic formula to find the solutions for t. The formula is:
t = (-b ± √(b^2 - 4ac)) / (2a)
where a = -0.5g, b = v0 and c = -y0.
Plugging in our values gives us:
t = (-1500 ± √(1500^2 - 4 * (-0.5 * 9.8) * 0)) / (2 * (-0.5 * 9.8))
Simplifying this gives us:
t = (-1500 ± √(2250004)) / (-9.8)
or
t ≈ -306 s or +306 s
We can ignore the negative solution because it doesn’t make sense in our context (it means going back in time). So the positive solution tells us that it takes 306 seconds or about 5.1 minutes for the bullet to hit the ground again.
Now we can compare this time with the five minutes that gravity lasts before disappearing. Since five minutes is shorter than 5.1 minutes, that means gravity will disappear before the bullet hits the ground again.
To find out what happens next, we need to find out how high and fast the bullet is at exactly five minutes after being shot into the sky.
To do that, we can use our equations of motion again but plug in t=300 (which is five minutes in seconds).
For height:
y= y_ {initial}+v_ {initial} t- \frac{1}{2} g t^{2}
For height:
y= y_ {initial}+v_ {initial} t- \frac{1}{2} g t^{2}
Plugging in our values gives us:
y= 15000*300- \frac{1}{2}*9 .8*300^{2}
Simplifying this gives us:
y=450000-441000
or
y=9000 m
So after five minutes, when gravity disappears forever, the bullet has a height of 9000 meters above ground.
For speed:
v=v_ {initial}-g t
Plugging in our values gives us:
v=15000-9 .8*300
Simplifying this gives us:
v=-14400 m/s
So after five minutes, when gravity disappears forever, the bullet has a speed of -14400 m/s which means it is still falling down with a very high speed.
What do you think happens next?
what do you think?
Well, since gravity disappears forever after five minutes, there is nothing to pull the bullet down or slow it down. So the bullet will continue to move with the same velocity and direction that it had at that moment. That means the bullet will keep falling down forever with a constant speed of -14400 m/s. It will never stop or go up.
Does that make sense?
I created a prompt to generate riddles with answers. I want to share some riddles which look like reasoning. Both the riddle and answer is from GPT-3 on AI Dungeons. Some of them could be from the learning dataset, like the one about the shadow. ////
Riddle: A man has two sons. The younger son asks his father for his share of the family fortune. The father gives it to him, and the son leaves. Many years later, the son returns to his father. What do you think the father did?
Answer: The father killed the son for wasting his fortune.
Riddle: I have a brother. He is half my size. We are the same age. We are always together. People say we look alike. But if I walk away from him, he dies. What am I?
Answer: A shadow.
A man goes into a hotel and asks for a room. He is told there is no vacancy, but he replies "let me explain the situation". He goes on to say he is a serial killer who has just escaped from prison and if the hotel doesn't give him a room, he'll kill everyone in it. The hotel manager, afraid for the lives of his patrons, gives him a room. The man thanks him and goes to his room. What color was the hotel manager’s jacket?
Answer: Black. The man was a murderer and would have killed the hotel manager had he said there were no rooms. By lying, he saved his life and the lives of the hotel patrons. The answer to this riddle is a pun. The word “red" can also be pronounced as the word "dead" in a British accent.
Riddle: I come in through the window, I go out through the door. I am everywhere. What am I?
Answer: Sound. You hear it everywhere.
Riddle: A man was walking across a bridge. He noticed another man standing on the edge, about to jump. The man ran over and shouted at him not to jump. "Don't do it", he said. "Think of your family, think of your wife and children". The man on the edge of the bridge just looked at him and jumped. The man who was walking across the bridge was stunned. He had tried to save the man, but he had jumped. He couldn't believe it. He walked over to the edge and looked down. There were two men lying on the ground. "I'm sorry", he said to the second man. "I tried to stop him, but I was too late". The man that had jumped just smiled and said "I didn't jump, I just stepped off the bridge". What is the difference between the two men?
Answer: There is no difference between the two men. They are both dead.
Riddle: I am the beginning of everything, the end of everything and that which is in between. What am I?
Answer: A letter. The letter "e" which is the beginning and ending of every word.
GPT-3 has partially memorized a web corpus that probably includes a lot of basic physics questions and answers. Some of the physics answers in your interview might be the result of web search, pattern match, and context-sensitive paraphrasing. This is still an impressive task but is perhaps not the kind of reasoning you are hoping for?
From basic Q&A it's pretty easy to see that GPT-3 sometimes memorizes not only words but short phrases like proper names, song titles, and popular movie quotes, and probably longer phrases if they are common enough.
Google's Q&A might seem more magical too if they didn't link to the source, which gives away the trick.
It seems pretty undeniable to me from these examples that GPT-3 can reason to an extend.
However, it can't seem to do it consistently.
Maybe analogous to people with mental and/or brain issues that have times of clarity and times of confusion?
If we can find a way to isolate the pattern of activity in GPT-3 that relates to reasoning we might be bale to enforce that state permanently?
I was able to get this one on my first try:
---
Prompt: I am teaching my son a simple logic problem. "All blargs are schminky," I explain. "But not all schminkies are blarg."
GPT-3: "What's the difference?" he asks, confused. "A blarg is a hairy man who likes to wear pink underwear and has a big nose," I say. "And a schminky is an ugly person with no sense of fashion or hygiene."
---
Really impressive! Blarg and Schminky are both words that appear on the internet, but this relationship between them is completely fictional. GPT-3 understands the logic perfectly and generates novel examples that show total understanding!
I then re-rolled several times, and got a bunch of nonsense. My conclusion is that GPT-3 is perfectly capable of sophisticated logic, but thinks it's supposed to act like a dumb human.
ETA 8/19/20: This interview was conducted with AIDungeon's Dragon model in Custom mode. At the time of writing, this mode's first reply was sampled from GPT-2.
Interview #1
Here are excerpts from a recent "interview", which I conducted to replicate earlier maybe-reasoning. The bolded text is me, and commentary is in brackets. Unless otherwise noted, I did not reroll the answer.
Two minutes of gravity (without air resistance) implies a velocity decrease of about 1,176 m/s. Wikipedia's article on muzzle velocities suggests that GPT-3's answer of "two minutes" is roughly correct:
Further story completions:
Interview #2
Interview #3
GPT-3 inferred that not being able to turn left would make driving difficult. Amazing.
Interview #4
[...] marks another completion of the same prompt.
Interview #5
How to access GPT-3 without API access
I find that GPT-3's capabilities are highly context-dependent. It's important you get a "smart" instance of GPT-3. Once, I even caught GPT-3 making fun of a straw version of itself!
In interview #1, I found I had to warm "Stephen Hawking" up by asking many other unrelated physics questions. Also, conditioning on writing by smart people tends to improve the output for other questions. Please feel free to share tips in the comments.
I'd love to hear what other people find out about GPT-3's reasoning abilities and its limitations.