First, I love this question.
Second, this might seem way out of left field, but I think this might help you answer it —
https://en.wikipedia.org/wiki/B%C3%BCrgerliches_Gesetzbuch#Abstract_system_of_alienation
One of the BGB's [editor: the German Civil Law Code] fundamental components is the doctrine of abstract alienation of property (German: Abstraktionsprinzip), and its corollary, the separation doctrine (Trennungsprinzip). Derived from the works of the pandectist scholar Friedrich Carl von Savigny, the Code draws a sharp distinction between obligationary agreements (BGB, Book 2), which create enforceable obligations, and "real" or alienation agreements (BGB, Book 3), which transfer property rights. In short, the two doctrines state: the owner having an obligation to transfer ownership does not make you the owner, but merely gives you the right to demand the transfer of ownership.
I have an idea of what might be going on here with your question.
It might be the case that there's two fairly-tightly-bound — yet slightly distinct — components in your conception of "theoretical evidence."
I'm having a hard time finding the precise words, but something around evidence, which behaves more-or-less similarly to how we typically use the phrase, and something around... implication, perhaps... inference, perhaps... something to do with causality or prediction... I'm having a hard time finding the right words here, but something like that.
I think it might be the case that these components are quite tightly bound together, but can be profitably broken up into two related concepts — and thus, being able to separate them BGB-style might be a sort of solution.
Maybe I'm mistaken here — my confidence isn't super high, but when I thought through this question the German Civil Law concept came to mind quickly.
It's profitable reading, anyways — BGB I think can be informative around abstract thinking, logic, and order-of-operations. Maybe intellectually fruitful towards your question or maybe not, but interesting and recommended either way.
What makes the thing you're pointing at different than just "deduction" or "logic"?
You have empirical evidence.
You use the empirical evidence to generate a theory edifice, and further evidence has so far supported it. (induction)
You use the theory to make a prediction (deduction), but that is not itself evidence, it only feels like it because we aren't logically omniscient and didn't already know what our theory implied. Whatever probability our prediction has comes from the theory, which gets its predictive value from the empirical evidence that went into creating and testing it.
The early discussions about mask effectiveness during COVID were often between people not trained in physics at all, that just wasn't part of their thinking process, so a physics-based response was new evidence because of the empirical evidence behind the relevant physics. Also, there were lots of people talking past each other because "mask," "use," and "effective" are all underspecified terms that don't allow for simple yes/no answers at the level of discourse we seem able to publicly support as a society, and institutions don't usually bother trying to make subtler points to the public for historical, legal, and psychological reasons (that we may or may not agree with in specific cases or in general).
Could be "framing conditions". I mean, it's one think to say "masks should help to not spread or receive viral particles", but it's another thing to say "masks can't not limit convection". Even if you are interested in the first, you have to separate it into the second and similar statements. Things should resemble pieces of an empirical model besides intuitive guesses, to be updateable.
I mean, it's fine to stick to the intuition, but it doesn't help with modifying the model.
There are such things as "theorem", "finding" and "understanding".
However the word evidence is heavily reserved for theory-distant pieces of data that are not prone to be negotiable. There is the sense that "evidence" is something that shifts beliefs. but this comes from the connection that a brain should be informed by the outside world. We don't call all persuasive things evidence.
If you are doing theorethical stuff and think in a way where " evidence" factors heavily you are somewhat likely to do things a bit backwards. Weighting evidence is connected to cogent argumens which are in the realm of inductive reasoning. In the realm of theory we can use proper deductive methods and definitely say stuff about things. A proof either carries or not - there is no "we can kinda say".
I'm a bit late to the game here, but you may be thinking of a facet of "logical induction". Basically, logical induction is changing your hypotheses based on putting more thought into an issue, without necessarily getting more Bayesian evidence.
The simplest example is when deciding whether a mathematical proof is true. Technically, you already have a hypothesis that perfectly predicts your data---ZFC set theory---but proving the proof is highly computationally expensive using this hypothesis, so if you want a probability estimate of whether the proof is true you need some other prediction mechanism.
See the Consequences of Logical Induction sequence for more information.
I'm basing this answer on a clarifying example from the comments section:
I believe that what I am trying to point at is indeed evidence, in the Bayesian sense of the word. For example, consider masks and COVID. Imagine that we empirically observe that they are effective 20% of the time and ineffective 80% of the time. Should we stop there and take it as our belief that there is a 20% chance that they are effective? No!
Suppose now that we know that when someone with COVID breathes, particles containing COVID remain in the air. Further suppose that our knowledge of physics would tell us that someone standing two feet away is likely to breathe in these particles at some concentration. And further suppose that our knowledge of how other diseases work tell us that when that concentration of virus is ingested, it is likely that you will get infected. When you incorporate all of this knowledge about physics and biology, it should shift your belief that masks are effective. It shouldn't stay put at 20%. We'd want to shift it upward to something like 75% maybe.
When put like this, these "evidence" sound a lot like priors. The order should be different though:
To a perfect Bayesian the order shouldn't matter, but we are not perfect Bayesians and if we try to do it the other way around and apply the theory to update the probabilities we got from the experiments, we would be able to convince ourselves the probability is 75% no matter how much empirical evidence that says otherwise we have accumulated.
I think the word you are looking for is analysis. Consider the toy scenario: You observe two pieces of evidence:
Now, without gathering any additional evidence, you can figure out (given certain assumptions about the gears level working of A, B, and C) that A = C. Because that takes finite time for your brain to realize, it feels like a new piece of information. However, it is merely the result of analyzing the existing evidence to generate additional equivalent statements. Of course, those new ways of describing the territory can be useful, but they shouldn't result in Baysean updates. Just like getting redundant evidence (eg 1. A = B 2. B = A) shouldn't move your estimate further than just getting one bit of evidence.
Another phrase for Theoretical Evidence or Instincts is No Evidence At All. What you're describing is an under-specified rationalization made in an attempt to disregard which way the evidence is pointing and let one cling to beliefs for which they don't have sufficient support. Zvi's response wrt masks in light of the evidence that they aren't effective butting up against his intuition that they are has no evidentiary weight. He was not acting as a curious inquirer, he was a clever arguer.
The point of Sabermetrics is that the "analysis" that baseball scouts used to do (and still do for the losing teams) is worthless when put up against hard statistics taken from actual games. As to your example, even the most expert basketball player's opinion can't hold a candle to the massive computational power required to test these different techniques in actual basketball games.
I mean "theoretical evidence" as something that is in contrast to empirical evidence. Alternative phrases include "inside view evidence" and "gears-level evidence".
I personally really like the phrase "gears-level evidence". What I'm trying to refer to is something like, "our knowledge of how the gears turn would imply X". However, I can't recall ever hearing someone use the phrase "gears-level evidence". On the other hand, I think I recall hearing "theoretical evidence" used before.
Here are some examples that try to illuminate what I am referring to.
Effectiveness of masks
Iirc, earlier on in the coronavirus pandemic there was empirical evidence saying that masks are not effective. However, as Zvi talked about, "belief in the physical world" would imply that they are effective.
Foxes vs hedgehogs
Foxes place more weight on empirical evidence, hedgehogs on theoretical evidence.
Harry's dark side
HPMoR chapter 10:
The Sorting Hat has empirical evidence that Harry is at risk of going dark. Harry's understanding of how the gears turn in his brain makes him think that he is not actually at risk of going dark.
Instincts vs A/B tests
Imagine that you are working on a product. A/B tests are showing that option A is better, but your instincts, based on your understanding of how the gears turn, suggest that B is better.
Posting up in basketball
Over the past 5-10 years in basketball, there has been a big push to use analytics more. Analytics people hate post-ups (an approach to scoring). The data says that they are low-efficiency.
I agree with that in a broad sense, but I believe that a specific type of posting up is very high efficiency. Namely, trying to get deep-position post seals when you have a good height-weight advantage. My knowledge of how the gears turn strongly indicates to me that this would be high efficiency offense. However, analytics people still seem to advise against this sort of offense.