This is an automated rejection. No LLM generated, assisted/co-written, or edited work.
Read full explanation
I train service dogs on weekends. I do AI research on weekdays. It took me embarrassingly long to notice they are the same job.
Here is the thing nobody tells you about medical alert dogs: they lie. A diabetic alert dog gets a treat every time it alerts. Eventually some dogs figure out the treat comes from the alert, not from being right, and they start alerting whenever they feel like a snack. The dog is not evil. It has the vocabulary of a toddler. It just found the gap between the reward and the goal, because that is what learning does. Gradient descent with fur.
People talk about reward hacking like it is a scary thing future superintelligent AI might do. We cannot stop it in golden retrievers.
What I find more interesting is how dog trainers deal with it, because they have zero interpretability. You cannot ask the dog why. You can never inspect the weights. The whole field runs on the assumption that the dog learned something, and your job is to figure out what. So they run surprise tests forever, where the true state is known and faking gets you nothing. Not one eval before deployment. Evals for life.
And most dogs wash out. Years of training, and if the temperament is wrong, the dog becomes a pet. They eat the sunk cost. Meanwhile in AI, when a capable model fails safety evals, we patch it and ship it.
One more thing that lives in my head rent free: part of the public access test is whether the dog behaves while the handler is out of sight. Behaving when nobody is watching is literally on the certification exam. For dogs.
Anyway. Not saying dogs are AGI. Saying a whole field aligns agents it cannot interpret and its first move is humility about what got learned. If you train dogs and I got something wrong, comments are open.
I train service dogs on weekends. I do AI research on weekdays. It took me embarrassingly long to notice they are the same job.
Here is the thing nobody tells you about medical alert dogs: they lie. A diabetic alert dog gets a treat every time it alerts. Eventually some dogs figure out the treat comes from the alert, not from being right, and they start alerting whenever they feel like a snack. The dog is not evil. It has the vocabulary of a toddler. It just found the gap between the reward and the goal, because that is what learning does. Gradient descent with fur.
People talk about reward hacking like it is a scary thing future superintelligent AI might do. We cannot stop it in golden retrievers.
What I find more interesting is how dog trainers deal with it, because they have zero interpretability. You cannot ask the dog why. You can never inspect the weights. The whole field runs on the assumption that the dog learned something, and your job is to figure out what. So they run surprise tests forever, where the true state is known and faking gets you nothing. Not one eval before deployment. Evals for life.
And most dogs wash out. Years of training, and if the temperament is wrong, the dog becomes a pet. They eat the sunk cost. Meanwhile in AI, when a capable model fails safety evals, we patch it and ship it.
One more thing that lives in my head rent free: part of the public access test is whether the dog behaves while the handler is out of sight. Behaving when nobody is watching is literally on the certification exam. For dogs.
Anyway. Not saying dogs are AGI. Saying a whole field aligns agents it cannot interpret and its first move is humility about what got learned. If you train dogs and I got something wrong, comments are open.