A Pattern Discovery Model for Effective Text Mining

Pipanmaekaporn, Luepol; Li, Yuefeng

doi:10.1007/978-3-642-31537-4_42

Luepol Pipanmaekaporn²⁰ &
Yuefeng Li²⁰

Part of the book series: Lecture Notes in Computer Science ((LNAI,volume 7376))

Included in the following conference series:

International Workshop on Machine Learning and Data Mining in Pattern Recognition

6153 Accesses

Abstract

The quality of extracted features is the key issue to text mining due to the large number of terms, phrases, and noise. Most existing text mining methods are based on term-based approaches which extract terms from a training set for describing relevant information. However, the quality of the extracted terms in text documents may be not high because of lot of noise in text. For many years, some researchers make use of various phrases that have more semantics than single words to improve the relevance, but many experiments do not support the effective use of phrases since they have low frequency of occurrence, and include many redundant and noise phrases. In this paper, we propose a novel pattern discovery approach for text mining. This approach first discovers closed sequential patterns in text documents for identifying the most informative contents of the documents and then utilise the identified contents to extract useful features for text mining. We develop a novel fusion method based on Dempster-Shafer’s evidential reasoning which allows to combine the pieces of document to discover the knowledge (features). To evaluate the proposed approach, we adopt the feature extraction method for information filtering (IF). The experimental results conducted on Reuters Corpus Volume 1 and TREC topics confirm that the proposed approach could achieve excellent performance.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Subscribe and save

Springer+ Basic

$34.99 /Month

Get 10 units per month
Download Article/Chapter or eBook
1 Unit = 1 Article or 1 Chapter
Cancel anytime

Buy Now

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 39.99; Price excludes VAT (USA)

Softcover Book: USD 54.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Preview

Unable to display preview. Download preview PDF.

Interpretation of text patterns

Article 22 February 2018

Extracting Various Types of Informative Web Content via Fuzzy Sequential Pattern Mining

Interpretive Psychotherapy of Text Mining Approaches

References

Bayardo Jr., R.: Efficiently mining long patterns from databases. ACM Sigmod Record 27, 85–93 (1998)
Article Google Scholar
Bi, Y., Wu, S., Wang, H., Guo, G.: Combination of evidence-based classifiers for text categorization. In: 2011 23rd IEEE International Conference on Tools with Artificial Intelligence, pp. 422–429. IEEE (2011)
Google Scholar
Buckley, C., Salton, G., Allan, J.: The effect of adding relevance information in a relevance feedback environment. In: ACM SIGIR 17th International Conf., pp. 292–300 (1994)
Google Scholar
Buckley, C., Voorhees, E.: Evaluating evaluation measure stability. In: 23th ACM SIGIR International Conf. on Research and Development in Information Retrieval, pp. 33–40 (2000)
Google Scholar
Caropreso, M., Matwin, S., Sebastiani, F.: Statistical phrases in automated text categorization. Centre National de la Recherche Scientifique, Paris, France (2000)
Google Scholar
Cheng, H., Yan, X., Han, J., Hsu, C.: Discriminative frequent pattern analysis for effective classification. In: 23rd IEEE ICDE International Conf. on Data Engineering, pp. 716–725 (2007)
Google Scholar
Jaillet, S., Laurent, A., Teisseire, M.: Sequential patterns for text categorization. Intelligent Data Analysis 10(3), 199–214 (2006)
Google Scholar
Joachims, T.: A probabilistic analysis of the rocchio algorithm with tfidf for text categorization. In: 14th ICML International Conf. on Machine Learning, pp. 143–151 (1997)
Google Scholar
Joachims, T.: Text Categorization with Support Vector Machines: Learning with many Relevant Features. In: Nédellec, C., Rouveirol, C. (eds.) ECML 1998. LNCS, vol. 1398, pp. 137–142. Springer, Heidelberg (1998)
Chapter Google Scholar
Lefevre, E., Colot, O., Vannoorenberghe, P.: Belief function combination and conflict management. Information Fusion 3(2), 149–162 (2002)
Article Google Scholar
Malik, H., Kender, J.: High quality, efficient hierarchical document clustering using closed interesting itemsets. In: 6th IEEE ICDM International Conf. on Data Mining, pp. 991–996 (2006)
Google Scholar
Nanas, N., Vavalis, M.: A “Bag” or a “Window” of Words for Information Filtering? In: Darzentas, J., Vouros, G.A., Vosinakis, S., Arnellos, A. (eds.) SETN 2008. LNCS (LNAI), vol. 5138, pp. 182–193. Springer, Heidelberg (2008)
Chapter Google Scholar
Rose, T., Stevenson, M., Whitehead, M.: The reuters corpus volume 1-from yesterday’s news to tomorrow’s language resources. In: 3th International Conf. on Language Resources and Evaluation, pp. 29–31 (2002)
Google Scholar
Sebastiani, F.: Machine learning in automated text categorization. ACM Computing Surveys 34(1), 1–47 (2002)
Article Google Scholar
Shehata, S., Karray, F., Kamel, M.: A concept-based model for enhancing text categorization. In: Proceedings of the 13th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 629–637. ACM (2007)
Google Scholar
Smets, P.: Data fusion in the transferable belief model. In: Proceedings of the Third International Conference on Information Fusion, FUSION 2000, vol. 1, pp. PS21–PS33. IEEE (2000)
Google Scholar
Soboroff, I., Robertson, S.: Building a filtering test collection for trec 2002. In: Proceedings of the 26th Annual International ACM SIGIR Conference on Research and Development in Informaion Retrieval, p. 250. ACM (2003)
Google Scholar
Stavrianou, A., Andritsos, P., Nicoloyannis, N.: Overview and semantic issues of text mining. ACM SIGMOD Record 36(3), 23–34 (2007)
Article Google Scholar
Tan, A.: Text mining: The state of the art and the challenges. In: Proceedings of the PAKDD 1999 Workshop on Knowledge Disocovery from Advanced Databases, pp. 65–70 (1999)
Google Scholar
van der Weide, T., van Bommel, P.: Measuring the incremental information value of documents. Information Sciences 176(2), 91–119 (2006)
Article MathSciNet MATH Google Scholar
Wu, S., Li, Y., Xu, Y.: Deploying approaches for pattern refinement in text mining. In: 6th IEEE ICDM International Conf. on Data Mining, pp. 1157–1161 (2006)
Google Scholar
Wu, S., Li, Y., Xu, Y., Pham, B., Chen, P.: Automatic pattern-taxonomy extraction for web mining. In: 3th IEEE/WIC/ACM WI International Conf. on Web Intelligence, pp. 242–248 (2004)
Google Scholar
Xin, D., Cheng, H., Yan, X., Han, J.: Extracting redundancy-aware top-k patterns. In: Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 444–453. ACM (2006)
Google Scholar
Xin, D., Han, J., Yan, X., Cheng, H.: Mining compressed frequent-pattern sets. In: Proceedings of the 31st International Conference on Very Large Databases, pp. 709–720. VLDB Endowment (2005)
Google Scholar
Yan, X., Cheng, H., Han, J., Xin, D.: Summarizing itemset patterns: a profile-based approach. In: 11th ACM SIGKDD International Conf. on Knowledge Discovery in Data Mining, pp. 314–323 (2005)
Google Scholar
Zaki, M.: Generating non-redundant association rules. In: 6th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 34–43 (2000)
Google Scholar
Zhang, W., Yoshida, T., Tang, X.: Text classification using multi-word features. In: IEEE International Conference on Systems, Man and Cybernetics, ISIC 2007, pp. 3519–3524. IEEE (2007)
Google Scholar
Zhong, N., Li, Y., Wu, S.: Effective pattern discovery for text mining. IEEE Transactions on Knowledge and Data Engineering, doi: http://doi.ieeecomputersociety.org/10.1109/TKDE,211

Download references

Author information

Authors and Affiliations

School of Electrical Engineering and Computer Science, Queensland University of Technology, Brisbane, Australia
Luepol Pipanmaekaporn & Yuefeng Li

Authors

Luepol Pipanmaekaporn
View author publications
You can also search for this author in PubMed Google Scholar
Yuefeng Li
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

Institute of Computer Vision and Applied Computer Sciences, IBaI, Kohlenstraße 2, 04107, Leipzig, Germany
Petra Perner

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Pipanmaekaporn, L., Li, Y. (2012). A Pattern Discovery Model for Effective Text Mining. In: Perner, P. (eds) Machine Learning and Data Mining in Pattern Recognition. MLDM 2012. Lecture Notes in Computer Science(), vol 7376. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-642-31537-4_42

Download citation

DOI: https://doi.org/10.1007/978-3-642-31537-4_42
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-642-31536-7
Online ISBN: 978-3-642-31537-4
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics

A Pattern Discovery Model for Effective Text Mining

Abstract

Access this chapter

Subscribe and save

Buy Now

Preview

Similar content being viewed by others

Interpretation of text patterns

Extracting Various Types of Informative Web Content via Fuzzy Sequential Pattern Mining

Interpretive Psychotherapy of Text Mining Approaches

References

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Publish with us

Subscribe and save

Buy Now

Navigation

A Pattern Discovery Model for Effective Text Mining

Abstract

Access this chapter

Subscribe and save

Buy Now

Preview

Similar content being viewed by others

Interpretation of text patterns

Extracting Various Types of Informative Web Content via Fuzzy Sequential Pattern Mining

Interpretive Psychotherapy of Text Mining Approaches

References

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Share this paper

Publish with us

Search

Navigation