Single-Pass, Adaptive Natural Language Filtering: Measuring Value in   User Generated Comments on Large-Scale, Social Media News Forums

Manuel Amunategui

arXiv:1701.03231·cs.CL·January 13, 2017

Single-Pass, Adaptive Natural Language Filtering: Measuring Value in User Generated Comments on Large-Scale, Social Media News Forums

Manuel Amunategui

PDF

Open Access

TL;DR

This paper introduces a single-pass, adaptive natural language filtering method to efficiently remove spam, noise, and irrelevant comments from large-scale social media news forums, improving comment relevance and data quality.

Contribution

It presents a novel two-step adaptive filtering approach that dynamically updates its corpus to enhance comment filtering accuracy in social media platforms.

Findings

01

Removes over a third of irrelevant comments

02

Increases comment relevance to original articles

03

Operates efficiently in a single pass

Abstract

There are large amounts of insight and social discovery potential in mining crowd-sourced comments left on popular news forums like Reddit.com, Tumblr.com, Facebook.com and Hacker News. Unfortunately, due the overwhelming amount of participation with its varying quality of commentary, extracting value out of such data isn't always obvious nor timely. By designing efficient, single-pass and adaptive natural language filters to quickly prune spam, noise, copy-cats, marketing diversions, and out-of-context posts, we can remove over a third of entries and return the comments with a higher probability of relatedness to the original article in question. The approach presented here uses an adaptive, two-step filtering process. It first leverages the original article posted in the thread as a starting corpus to parse comments by matching intersecting words and term-ratio balance per sentence…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsWikis in Education and Collaboration · Topic Modeling · Advanced Text Analysis Techniques