Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I agree. Also interesting to see that Google defines webspam as "pages that cheat" or "violate search engine quality guidelines." By this definition, scraper sites are not spam at all. Nor are the spammy sites in my field which super-optimize for keywords in ways that make it difficult for legitimate content to rise to visibility.

If Google did not operate AdSense, it seems hard to believe the company would not have penalized this sort of behavior ages ago. A love for AdSense is probably the single largest thing spam sites have in common worldwide.



"By this definition, scraper sites are not spam at all."

Disagree. Our quality guidelines at http://www.google.com/support/webmasters/bin/answer.py?hl=en... say "Don't create multiple pages, subdomains, or domains with substantially duplicate content." Duplicate content can be content copied within the site itself or copied from other sites.

Stack Overflow is a bit of a weird case, by the way, because their content license allowed anyone to copy their content. If they didn't have that license, we could consider the clones of SO to be scraper sites that clearly violate our guidelines.


Smaller competitors can't eat their lunch in web search, because all the content that was on the web is now on Wikipedia, YouTube, or Google Maps. Personally, I search these directly from the address bar. For the past four years I've only had two use cases for web search: 1. as a spell checker for proper nouns (and before Alpha, as a calculator) 2. to circumvent paywalls on scholarly papers by doing filetype:pdf on the title (works better than Scholar most of the time).




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: