It would be interesting to know how Google determines, in an automated perspective, which site is the original and which is the copy - especially when copying could go both ways. For example, I could write an article, license it under the GFDL, and someone could copy it to Wikipedia. I might then copy the Wikipedia improvements back to my site. Technically, I had the content first - so would Wikipedia be penalised?
If there is a bias against smaller sites, this might make smaller sites be reluctant to license their content under licenses that let bigger sites copy them.
This is a huge problem with Wikia. Wikia has a relatively high pagerank. Sometimes, because of Wikia's often-abusive policies, communities will decide to move their wiki off Wikia and host it separately.
But Wikia will refuse to remove the original wiki even if all the contributors want to move it, in order to maximize advertising revenue. This means that it's nearly impossible to get the new wiki to rank highly on Google, even if all links across the internet are changed to point to the new wiki, because Wikia's pagerank is so high that Google deems the new one to be a "copy". This is why Wikia moved all their wikis to subdomains: in order to piggyback on the pagerank of the main site.
As a result there are a ton of long-dead wikis on Wikia that still get more search traffic than the active equivalent. Obviously this hurts users, since they get long-outdated information as a result.
In short, once a wiki is placed on Wikia, it's basically impossible to ever move it anywhere else because of Google's anti-duplicate biasing.
I assume you're saying that the Wikia admins will prevent the community from deleting or overwriting content with a pointer redirecting to the new place? This does seem extremely likely if there's significant search engine traffic and ad revenue.
Yes. Now imagine if you had a blog on Blogspot and you wanted to host it on your own site instead -- and Blogspot prevented you from deleting your posts because they brought Blogspot good ad revenue?
I don't think that Google needs to determine the originality of any one piece of content, but rather the tendency for a site to feature content verbatim from other sources.
The sites Google seem to be targeting are those who aggregate content wholesale from a number of sources. One way they could identify these would be to examine the number of different sites from which a particular site appears to have copied its content from.
Agree. What I think they should target is a set of sites that just copy content load their pages with Google Ads and mint money. Those are the ones people more likely to not want to see rather than sites like wikia.
In any case this is a bold step for Google. It shows they still care for the quality of the search and have guts to take big bold decisions to protect it.
If a site has N copies associated with it, then you could compare that to the average number of copies associated with sites of that pagerank. If the difference between N and the average is high, then that's a likely spam site. Let's call this difference σ.
This gives us a problem though, because while σ is a good indicator of spamminess, it's not foolproof. A low pagerank site could have been copied from lots of times, which would unfairly earn it a high σ.
What we can do is calculate a σ', by inversely weighing each copy with the σ of the site on the other end. A copy of a site with a low sigma will increase σ', and a copy of a site with a high sigma will have less of an effect on σ'.
So while our high σ may have initially suggested the site to be spammy, all its copies are from high σ sites, which are counted less towards σ', leading to an overall verdict of not spammy. [I wonder what would happen if you iterated this process]
As for your example, the smaller site wouldn't be penalized, because N would be low. If it were say, a full wikipedia mirror, then it would be penalized. Wikipedia would not be penalized, since it has an ungodly amount of pagerank. There's also no bias against smaller sites, since σ is calculated relative to pagerank.
It would be interesting to know how Google determines, in an automated perspective, which site is the original and which is the copy - especially when copying could go both ways. For example, I could write an article, license it under the GFDL, and someone could copy it to Wikipedia. I might then copy the Wikipedia improvements back to my site. Technically, I had the content first - so would Wikipedia be penalised?
If there is a bias against smaller sites, this might make smaller sites be reluctant to license their content under licenses that let bigger sites copy them.