Catching Unicorns with GLTR: Using AI to Detect Generated Text and Fake News

Kazuki Nakayashiki

Hatched by Kazuki Nakayashiki

Aug 04, 2023

4 min read

0

Catching Unicorns with GLTR: Using AI to Detect Generated Text and Fake News

In the age of advanced technology and artificial intelligence, it has become increasingly difficult to distinguish between genuine human-written text and generated content. The rise of deep learning models, such as GPT-2, has made it possible for machines to generate text that is almost indistinguishable from that written by humans. However, researchers have found a way to use these same models as a tool for detection, giving us the ability to catch unicorns – or in other words, identify fake text.

One such project that aims to achieve this is GLTR (short for "Generating Linguistic Tests for Rhetorical Analysis"). GLTR utilizes the same models that are used to generate fake text and repurposes them for the purpose of detection. By analyzing the patterns and probabilities of words generated by these models, GLTR can determine whether a given piece of text is likely to be generated or written by a human.

The concept behind GLTR is based on the observation that natural writing often incorporates unpredictable words that are relevant to the topic or domain. In contrast, generated text tends to lack these unpredictable elements and may exhibit high certainty and predictability. By analyzing the rankings of words generated by the model, GLTR can identify whether a text shows signs of being generated or written by a human. For example, a text that consists mostly of green and yellow words, with few purple or red words, is likely to be generated, while a text with unexpected purple and red words is more likely to be human-written.

To demonstrate the effectiveness of GLTR, researchers used the GPT-2 model to produce non-conditioned text by sampling from the top 40 predictions. The results were remarkable, with the model being able to detect its own text with high accuracy. The generated text exhibited a pattern of mostly green and yellow words, indicating its synthetic nature.

The implications of GLTR extend beyond simply detecting generated text. In the age of misinformation and fake news, it has become crucial to have tools that can help users discern fact from fiction. This is where GLTR's potential becomes even more significant. Platforms like Facebook and YouTube have already recognized the value of utilizing resources like Wikipedia to combat fake news. By integrating GLTR into their fact-checking algorithms, these platforms could further enhance their ability to detect and flag fake news articles by analyzing the linguistic patterns and probabilities of the text.

Wikipedia, with its vast collection of articles and user-generated content, could serve as a valuable resource for training and validating GLTR. Each article and user on Wikipedia has a corresponding "talk" page, which serves as a communication channel for editors to discuss and debate. This wealth of user-generated content provides ample opportunities to train GLTR to recognize the patterns and linguistic features that distinguish human-written text from generated text.

The impact of Wikipedia goes beyond its role as an information repository. It has become a trusted source of information for millions of users worldwide. In fact, in January 2007, Wikipedia ranked among the top ten most popular websites in the US, surpassing renowned platforms like The New York Times and Apple. The credibility and reliability of Wikipedia make it an ideal partner for platforms seeking to combat fake news. By suggesting fact-checking links to related Wikipedia articles, platforms like Facebook and YouTube leverage the power of crowd-sourced information to verify the accuracy of news articles.

Moreover, the social dynamics within the Wikipedia community play a significant role in maintaining the quality and accuracy of the content. Wikipedians often award each other "virtual barnstars" as tokens of appreciation for their contributions. These barnstars go beyond recognizing simple editing work and encompass various forms of support, administrative actions, and articulation work. This system of recognition fosters a sense of community and collective responsibility, ensuring that Wikipedia remains a reliable source of information.

In conclusion, the development of tools like GLTR represents a significant step forward in the fight against misinformation and fake news. By leveraging the same models used to generate fake text, GLTR provides us with the means to catch unicorns – to identify and flag generated content. Integrating GLTR into platforms like Facebook and YouTube, along with the support of resources like Wikipedia, could greatly enhance the detection of fake news and promote a more informed and trustworthy online environment.

Actionable Advice:

  1. Stay vigilant: While tools like GLTR can aid in detecting fake text, it is crucial to remain cautious and skeptical when consuming online content. Double-check the sources and cross-reference information to ensure its authenticity.
  2. Support crowd-sourced platforms: Platforms like Wikipedia rely on the collective efforts of volunteers to maintain accurate and reliable information. Consider contributing to crowd-sourced platforms or donating to support their mission.
  3. Promote media literacy: Educate yourself and others about the signs of fake news and misinformation. By developing critical thinking and media literacy skills, we can navigate the digital landscape more effectively and make informed decisions.

By employing AI to detect generated text and partnering with reputable resources like Wikipedia, we can work towards a future where the spread of fake news is mitigated, and reliable information prevails.

Sources

← Back to Library

Hatch New Ideas with Glasp AI 🐣

Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)

Start Hatching 🐣