short-paper

Snorkel: Fast Training Set Generation for Information Extraction

Authors:

Alexander J. Ratner,

Stephen H. Bach,

Henry R. Ehrenberg,

Chris RéAuthors Info & Claims

SIGMOD '17: Proceedings of the 2017 ACM International Conference on Management of Data

Pages 1683 - 1686

https://doi.org/10.1145/3035918.3056442

Published: 09 May 2017 Publication History

Get Access

Abstract

State-of-the art machine learning methods such as deep learning rely on large sets of hand-labeled training data. Collecting training data is prohibitively slow and expensive, especially when technical domain expertise is required; even the largest technology companies struggle with this challenge. We address this critical bottleneck with Snorkel, a new system for quickly creating, managing, and modeling training sets. Snorkel enables users to generate large volumes of training data by writing labeling functions, which are simple functions that express heuristics and other weak supervision strategies. These user-authored labeling functions may have low accuracies and may overlap and conflict, but Snorkel automatically learns their accuracies and synthesizes their output labels. Experiments and theory show that surprisingly, by modeling the labeling process in this way, we can train high-accuracy machine learning models even using potentially lower-accuracy inputs. Snorkel is currently used in production at top technology and consulting companies, and used by researchers to extract information from electronic health records, after-action combat reports, and the scientific literature. In this demonstration, we focus on the challenging task of information extraction, a common application of Snorkel in practice. Using the task of extracting corporate employment relationships from news articles, we will demonstrate and build intuition for a radically different way of developing machine learning systems which allows us to effectively bypass the bottleneck of hand-labeling training data.

References

[1]

S. H. Bach, B. He, A. Ratner, and C. Ré. Learning the structure of generative models without labeled data. arXiv preprint arXiv:1703.00854, 2017.

Google Scholar

[2]

A. P. Davis et al. A CTD--Pfizer collaboration: Manual curation of 88,000 scientific articles text mined for drug--disease and drug--phenotype interactions. Database, 2013.

Google Scholar

[3]

H. R. Ehrenberg, J. Shin, A. J. Ratner, J. A. Fries, and C. Ré. Data programming with DDLite: Putting humans in a different part of the loop. In HILDA@ SIGMOD, 2016.

Digital Library

Google Scholar

[4]

A. Ratner, C. De Sa, S. Wu, D. Selsam, and C. Ré. Data programming: Creating large training sets, quickly. In Neural Information Processing Systems (NIPS), 2016.

Digital Library

Google Scholar

Cited By

View all

Siqueira FPressato DPereira Fda Silva NSouza EDias Mde Carvalho A(2024)Segmenting Brazilian legislative text using weak supervision and active learningArtificial Intelligence and Law10.1007/s10506-024-09419-5Online publication date: 26-Sep-2024
https://doi.org/10.1007/s10506-024-09419-5
Liu XWang QHuang K(2023)Text-Based Technological Risk and Firm Innovation: An Empirical AnalysisSSRN Electronic Journal10.2139/ssrn.4611658Online publication date: 2023
https://doi.org/10.2139/ssrn.4611658
Domschot ERamyaa RSmith M(2023)Improving Automated Labeling for ATT&CK Tactics in Malware Threat ReportsDigital Threats: Research and Practice10.1145/35945535:1(1-16)Online publication date: 17-May-2023
https://dl.acm.org/doi/10.1145/3594553
Show More Cited By

Index Terms

Snorkel: Fast Training Set Generation for Information Extraction
1. Computing methodologies
  1. Artificial intelligence
    1. Natural language processing
      1. Information extraction
  2. Machine learning
2. Information systems

Recommendations

Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale
SIGMOD '19: Proceedings of the 2019 International Conference on Management of Data

Labeling training data is one of the most costly bottlenecks in developing machine learning-based applications. We present a first-of-its-kind study showing how existing knowledge resources from across an organization can be used as weak supervision in ...
Software 2.0 and Snorkel: Beyond Hand-Labeled Data
KDD '18: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining

In the last few years, deep learning models have simultaneously achieved high quality on conventionally challenging tasks and become easy-to-use commodity tools. These factors, combined with the ease of deployment compared to traditional software, have ...
SPL-LDP: a label distribution propagation method for semi-supervised partial label learning
Abstract
Partial label learning learns from examples represented by a single instance while associated with multiple candidate labels, among which only one valid label resides. However, in real-world applications, collecting candidate label sets for all ...

Comments

Information & Contributors

Information

Published In

SIGMOD '17: Proceedings of the 2017 ACM International Conference on Management of Data

May 2017

1810 pages

ISBN:9781450341974

DOI:10.1145/3035918

General Chairs:
Rada Chirkova
North Carolina State University, USA
,
Jun Yang
Duke University, USA
,
Program Chair:
Dan Suciu
University of Washington, USA

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]

Publisher

Association for Computing Machinery

New York, NY, United States

Publication History

Published: 09 May 2017

Permissions

Request permissions for this article.

Request Permissions

Check for updates

Author Tags

Qualifiers

Short-paper

Conference

SIGMOD/PODS'17

Sponsor:

SIGMOD

SIGMOD/PODS'17: International Conference on Management of Data

May 14 - 19, 2017

Illinois, Chicago, USA

Acceptance Rates

Overall Acceptance Rate 785 of 4,003 submissions, 20%

Contributors

Other Metrics

View Article Metrics

Bibliometrics & Citations

Bibliometrics

Article Metrics

42
Total Citations
View Citations
922
Total Downloads

Downloads (Last 12 months)38
Downloads (Last 6 weeks)4

Reflects downloads up to 03 Oct 2024

Other Metrics

View Author Metrics

Citations

Cited By

View all

Siqueira FPressato DPereira Fda Silva NSouza EDias Mde Carvalho A(2024)Segmenting Brazilian legislative text using weak supervision and active learningArtificial Intelligence and Law10.1007/s10506-024-09419-5Online publication date: 26-Sep-2024
https://doi.org/10.1007/s10506-024-09419-5
Liu XWang QHuang K(2023)Text-Based Technological Risk and Firm Innovation: An Empirical AnalysisSSRN Electronic Journal10.2139/ssrn.4611658Online publication date: 2023
https://doi.org/10.2139/ssrn.4611658
Domschot ERamyaa RSmith M(2023)Improving Automated Labeling for ATT&CK Tactics in Malware Threat ReportsDigital Threats: Research and Practice10.1145/35945535:1(1-16)Online publication date: 17-May-2023
https://dl.acm.org/doi/10.1145/3594553
Zhang HTae KPark JChu XWhang S(2023)iFlipper: Label Flipping for Individual FairnessProceedings of the ACM on Management of Data10.1145/35886881:1(1-26)Online publication date: 30-May-2023
https://dl.acm.org/doi/10.1145/3588688
Jiang MRocktäschel TGrefenstette E(2023)General intelligence requires rethinking explorationRoyal Society Open Science10.1098/rsos.23053910:6Online publication date: 21-Jun-2023
https://doi.org/10.1098/rsos.230539
Kumar ASharaff A(2023)SnorkelPlus: A Novel Approach for Identifying Relationships Among Biomedical Entities Within AbstractsThe Computer Journal10.1093/comjnl/bxad051Online publication date: 4-May-2023
https://doi.org/10.1093/comjnl/bxad051
Paul SMadan GMishra AHegde NKumar PAggarwal G(2023)Weakly Supervised Information Extraction from Inscrutable Handwritten Document ImagesDocument Analysis and Recognition - ICDAR 202310.1007/978-3-031-41685-9_28(445-463)Online publication date: 19-Aug-2023
https://doi.org/10.1007/978-3-031-41685-9_28
Kpiebaareh MWu WAgyemang BHaruna CLawrence T(2022)A Generic Graph-Based Method for Flexible Aspect-Opinion Analysis of Complex Product Customer FeedbackInformation10.3390/info1303011813:3(118)Online publication date: 28-Feb-2022
https://doi.org/10.3390/info13030118
Li BLu YKandula SIves ZBonifati AEl Abbadi A(2022)Warper: Efficiently Adapting Learned Cardinality Estimators to Data and Workload DriftsProceedings of the 2022 International Conference on Management of Data10.1145/3514221.3526179(1920-1933)Online publication date: 10-Jun-2022
https://dl.acm.org/doi/10.1145/3514221.3526179
Tabar MJung WYadav AWilson Chavez OFlores ALee DAl Hasan MXiong L(2022)WARNER: Weakly-Supervised Neural Network to Identify Eviction Filing Hotspots in the Absence of Court RecordsProceedings of the 31st ACM International Conference on Information & Knowledge Management10.1145/3511808.3557128(3514-3523)Online publication date: 17-Oct-2022
https://dl.acm.org/doi/10.1145/3511808.3557128
Show More Cited By

View Options

Get Access

Login options

Check if you have access through your login credentials or your institution to get full access on this article.

Cited By

Index Terms

Recommendations

Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale

Software 2.0 and Snorkel: Beyond Hand-Labeled Data

SPL-LDP: a label distribution propagation method for semi-supervised partial label learning