Re: K MEANS clustering

Richhiey Thomas <[email protected]>
Newsgroups gmane.comp.search.xapian.devel
Message-ID <CAPHaz=DxZNVpMaL=ut2ZbBcPj_SFrdBS-vK8-L0sXB35QhF3aw@mail.gmail.com>
Hey Parth,

Thanks for the reply.
I am considering implementing a cosine distance metric too, along with
euclidian distance because of the dimensionality issue that comes in with
K-Means and euclidian distance metric.
That does help when we deal with sparse vectors for documents. The
particular problem I'm having is representing centroids in an efficient way.
For example, when we find the mean vector of a cluster, the resultant
centroid need not be a document vector of a document belonging to that
cluster. Hence representing that cluster, which will be dense as a C++ map
is inefficient because of the number of terms associated with it and
calculating distances with that doesn't work or scale too well.
Over that, my distance calculation works over two documents. So will I need
to modify that in a way to accommodate arbitrary vectors which might not
represent document vectors?
Would be great if everyone could add there inputs on this.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.