"The total size of Sci-Hub database is about 100 TB."
I think there are two reasons it's not more common to retroactively analyze papers and publications for copied or closely-paraphrased segments.
First, it's not actually easy to automate. Current solutions are RIFE with false-positives and human judgement requirements to make final conclusions.
Second, and perhaps more importantly, nobody really cares, outside of graded work where the organization is basing your credentials on doing original work (but usually not even that, just semi-original presentation of other works).
It would probably be a minor scandal if any significant papers were discovered to be based on uncredited/un-footnoted other work, but unless it were egregious (in which case it probably would have already been noticed), just not that big a deal.
Automated plagerism detection software is common. But cases like the recent incident with Harvard administrator Gay have shown that egregious cases of plagerism are still being uncovered. Why would this be the case? Is it really so hard to run plagerism checks for every paper on Sci-hub? Has anyone tried?
I am curious since I am currently upskilling for the purposes of technical alignment research and it seems like an interesting project to pursue.