i'm thinking about this from the point of view of Bayes' Theorem. Where what the "native model" has are the initial priors, and they get updated after every interaction.
And sometimes with Bayes' Theorem, you find that when you have some "very strong evidence" early on, the probabilities collapse to learn too much from such evidence, which no later evidence can undo.
Then again, the question is how much one should "damp the evidence" before updating the priors after every interaction - too much damping and the model doesn't learn at all. Too little, and it can learn too much of the wrong things too early!
i'm thinking about this from the point of view of Bayes' Theorem. Where what the "native model" has are the initial priors, and they get updated after every interaction.
And sometimes with Bayes' Theorem, you find that when you have some "very strong evidence" early on, the probabilities collapse to learn too much from such evidence, which no later evidence can undo.
Then again, the question is how much one should "damp the evidence" before updating the priors after every interaction - too much damping and the model doesn't learn at all. Too little, and it can learn too much of the wrong things too early!