SOTAVerified

TheEyeCorpus: Experiments in Reducing NLP Bias and Identifiability for Large LMs

2021-11-03Code Available0· sign in to hype

Anonymous

Code Available — Be the first to reproduce this paper.

Reproduce

Code

Abstract

NLP is a constantly changing and evolving subset of artificial intelligence, due to constantly increasing data requirements for SOTA (state of the art) machine learning models such as GPT-3, we often neglect the safety and privacy of our data producers. Using pure raw data often presents challenges, such as legalities and identifiability, solving these issues requires taking a machine-to-machine corpus approach, using source data to synthesize novel data. However, machine-to-machine approaches often involve a prolonged data gathering period, while also being computationally costly to produce. Rather than building from the ground up I believe publicly available datasets are needed. Building off of previous work from Gretel.AI (https://gretel.ai/), I propose TheEyeCorpus, a machine synthesized publicly available text corpus obtained from the public The-Eye Discord server and preprocessed and synthesised by Simp Labs utilizing Gretel AI’s APIS and utilities. The data is hosted at github.com/puffy310/TheEyeCorpus

Reproductions