suspicious output from cssdiff

Thomas Michael Hagen <[email protected]>
Newsgroups gmane.mail.spam.crm114
Message-ID <[email protected]>
i'm noticing that my program is suspiciously sure of itself after a
cycle of 'learn', where it learns newspaper articles based on keywords
in the url.

a typical url will contain a keyword such as 'sport' or 'culture'.

here's the output of cssdiff between the sport and economy categories:

[st08764@gandalf emneklassifisering]$ /usr/local/crm114/bin/cssdiff
spo.cfc eco.cfc
Sparse spectra file spo.cfc has 3397338 bins total
Sparse spectra file eco.cfc has 3397338 bins total

 File 1 total features            :    570588910
 File 2 total features            :    588381755

 Similarities between files       :            0
 Differences between files        :    579485332

 File 1 dominates file 2          :    570588910
 File 2 dominates file 1          :    588381755
[st08764@gandalf emneklassifisering]$

notice that the number of features and the times they dominate each
other is identical. it also shows absolutely no similarities between
any categories, even between bus.cfc (business) and eco.cfc (economy),
which is rather implausible.

i'm thinking this can't be right, but i don't quite understand what
the error is.

does anyone recognize this pattern?

------------------------------------------------------------------------------
This SF.net email is sponsored by:
High Quality Requirements in a Collaborative Environment.
Download a free trial of Rational Requirements Composer Now!
http://p.sf.net/sfu/www-ibm-com
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.