Re: K MEANS clustering

Richhiey Thomas <[email protected]>
Newsgroups gmane.comp.search.xapian.devel
Message-ID <CAPHaz=Ci2sBSLF-j60Ak5qC-2fa-oV0xb_mLCa3zXP2rpVtzFg@mail.gmail.com>
Hey Parth,

Thanks for the reply.
I am considering implementing a cosine distance metric because of the
dimensionality issue that comes in with K-Means and euclidian distance
metric.

Currently, the way I'm finding distances between documents is finding their
terms and looking up their term frequencies which I've stored in a map. So
I've not stored a unique vector for every document. Now in KMeans, when we
find the mean of a cluster, the resultant need not be a document vector. So
representing these centroids is becoming a problem since the centroids will
be dense. Should I use a map for that too? By storing all the terms and
their avg values.
Or would it be a better approach to have a document vector for every
document stored?

Thanks.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.