articleTop 1% cited
Self organization of a massive document collection
IEEE Transactions on Neural Networks · 2000 · Vol. 11(3) · pp. 574–585
Abstract
This article describes the implementation of a system that is able to organize vast document collections according to textual similarities. It is based on the self-organizing map (SOM) algorithm. As the feature vectors for the documents statistical representations of their vocabularies are used. The main goal in our work has been to scale up the SOM algorithm to be able to deal with large amounts of high-dimensional data. In a practical experiment we mapped 6,840,568 patent abstracts onto a 1,002,240-node SOM. As the feature vectors we used 500-dimensional vectors of stochastic figures obtained as random projections of weighted word histograms.
Neural Networks and ApplicationsAdvanced Clustering Algorithms ResearchData Mining Algorithms and ApplicationsComputer scienceHistogramSelf-organizing mapInformation retrievalFeature (linguistics)Word (group theory)Artificial intelligenceFeature vectorData miningDocument clustering
Funding
- Academy of Finland
Citations
911
FWCI
88.50
field-weighted impact
References
61
Percentile
100%
vs. same field & year
Citations per year
Cited by
Semantic Network Analysis as a Method for Visual Text Analytics
Procedia - Social and Behavioral Sciences · 2013 · 242 citations
Clustering of the self-organizing map
IEEE Transactions on Neural Networks · 2000 · 2,628 citations
Survey of Clustering Algorithms
IEEE Transactions on Neural Networks · 2005 · 6,086 citations
Essentials of the self-organizing map
Neural Networks · 2012 · 1,559 citations
References
WEBSOM – Self-organizing maps of document collections
Neurocomputing · 1998 · 494 citations
A Nonlinear Mapping for Data Structure Analysis
IEEE Transactions on Computers · 1969 · 3,394 citations
Exploratory Data Analysis
Biometrics · 1977 · 12,886 citations
Citation Network
How this paper connects to the literature. Drag to explore, click any node to open that paper.
