Snorkel: Fast training set generation for information extraction

AJ Ratner, SH Bach, HR Ehrenberg, C Ré - Proceedings of the 2017 …, 2017 - dl.acm.org
Proceedings of the 2017 ACM international conference on management of data, 2017dl.acm.org
State-of-the art machine learning methods such as deep learning rely on large sets of hand-
labeled training data. Collecting training data is prohibitively slow and expensive, especially
when technical domain expertise is required; even the largest technology companies
struggle with this challenge. We address this critical bottleneck with Snorkel, a new system
for quickly creating, managing, and modeling training sets. Snorkel enables users to
generate large volumes of training data by writing labeling functions, which are simple …
State-of-the art machine learning methods such as deep learning rely on large sets of hand-labeled training data. Collecting training data is prohibitively slow and expensive, especially when technical domain expertise is required; even the largest technology companies struggle with this challenge. We address this critical bottleneck with Snorkel, a new system for quickly creating, managing, and modeling training sets. Snorkel enables users to generate large volumes of training data by writing labeling functions, which are simple functions that express heuristics and other weak supervision strategies. These user-authored labeling functions may have low accuracies and may overlap and conflict, but Snorkel automatically learns their accuracies and synthesizes their output labels. Experiments and theory show that surprisingly, by modeling the labeling process in this way, we can train high-accuracy machine learning models even using potentially lower-accuracy inputs. Snorkel is currently used in production at top technology and consulting companies, and used by researchers to extract information from electronic health records, after-action combat reports, and the scientific literature. In this demonstration, we focus on the challenging task of information extraction, a common application of Snorkel in practice. Using the task of extracting corporate employment relationships from news articles, we will demonstrate and build intuition for a radically different way of developing machine learning systems which allows us to effectively bypass the bottleneck of hand-labeling training data.
ACM Digital Library