Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It might be interesting that at UIC, we do a data mining course research project on this topic - http://www.cs.uic.edu/~liub/#projects. We used resellerratings.com and apart from a handpicked, there is no way of determining a pattern. There is no luxury of mechanical turks for a course project, and we are strictly forbidden from writing any reviews. I had to use the #of reviews, variance in similarity of the review text, time interval between posts, user since date, helpful review count, average rating(user/review/store),etc. and calculated the Mahalonobis distance[http://en.wikipedia.org/wiki/Mahalanobis_distance] to separate outliers as the spammers. Then with these as labeled spammers, used graph based semi-supervised learning to classify the reviewers as spammers. Its a wild open problem - and very easy to argue on both sides of any method :D


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: