Developing a cardiovascular disease risk factor annotated corpus of Chinese electronic medical records

Su, Jia; He, Bin; Guan, Yi; Jiang, Jingchi; Yang, Jinfeng

Computer Science > Computation and Language

arXiv:1611.09020 (cs)

[Submitted on 28 Nov 2016 (v1), last revised 3 Mar 2017 (this version, v2)]

Title:Developing a cardiovascular disease risk factor annotated corpus of Chinese electronic medical records

Authors:Jia Su, Bin He, Yi Guan, Jingchi Jiang, Jinfeng Yang

View PDF

Abstract:Cardiovascular disease (CVD) has become the leading cause of death in China, and most of the cases can be prevented by controlling risk factors. The goal of this study was to build a corpus of CVD risk factor annotations based on Chinese electronic medical records (CEMRs). This corpus is intended to be used to develop a risk factor information extraction system that, in turn, can be applied as a foundation for the further study of the progress of risk factors and CVD. We designed a light annotation task to capture CVD risk factors with indicators, temporal attributes and assertions that were explicitly or implicitly displayed in the records. The task included: 1) preparing data; 2) creating guidelines for capturing annotations (these were created with the help of clinicians); 3) proposing an annotation method including building the guidelines draft, training the annotators and updating the guidelines, and corpus construction. Then, a risk factor annotated corpus based on de-identified discharge summaries and progress notes from 600 patients was developed. Built with the help of clinicians, this corpus has an inter-annotator agreement (IAA) F1-measure of 0.968, indicating a high reliability. To the best of our knowledge, this is the first annotated corpus concerning CVD risk factors in CEMRs and the guidelines for capturing CVD risk factor annotations from CEMRs were proposed. The obtained document-level annotations can be applied in future studies to monitor risk factors and CVD over the long term.

Comments:	32 pages, 3 figures, 3 tables
Subjects:	Computation and Language (cs.CL)
ACM classes:	I.2.7
Cite as:	arXiv:1611.09020 [cs.CL]
	(or arXiv:1611.09020v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1611.09020

Submission history

From: Jia Su [view email]
[v1] Mon, 28 Nov 2016 08:20:54 UTC (422 KB)
[v2] Fri, 3 Mar 2017 08:52:27 UTC (2,129 KB)

Computer Science > Computation and Language

Title:Developing a cardiovascular disease risk factor annotated corpus of Chinese electronic medical records

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Developing a cardiovascular disease risk factor annotated corpus of Chinese electronic medical records

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators