[Back]


Talks and Poster Presentations (with Proceedings-Entry):

A. Karatzoglou, I. Feinerer, K. Hornik:
"Nonparametric distribution analysis for text mining";
Talk: 32nd Annual Conference of the Gesellschaft für Klassifikation e.V., Hamburg; 07-16-2008 - 07-18-2008; in: "Advances in Data Analysis, Data Handling and Business Intelligence", Springer, (2009), ISBN: 978-3-642-01045-3; 295 - 305.



English abstract:
A number of new algorithms for nonparametric distribution analysis based on Maximum Mean Discrepancy measures have been recently introduced. These novel algorithms operate in Hilbert space and can be used for nonparametric two-sample tests. Coupled with recent advances in string kernels, these methods extend the scope of kernel-based methods in the area of text mining. We review these kernel-based two-sample tests focusing on text mining where we will propose novel applications and present an efficient implementation in the kernlab package. We also present an efficient and integrated environment for applying modern machine learning methods to complex text mining problems through the combined use of the tm (for text mining) and the kernlab (for kernel-based learning) R packages.

Keywords:
Kernel methods, R, text mining


"Official" electronic version of the publication (accessed through its Digital Object Identifier - DOI)
http://dx.doi.org/10.1007/978-3-642-01044-6_27


Created from the Publication Database of the Vienna University of Technology.