suspicious output from cssdiff
Thomas Michael Hagen <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
i'm noticing that my program is suspiciously sure of itself after a cycle of 'learn', where it learns newspaper articles based on keywords in the url. a typical url will contain a keyword such as 'sport' or 'culture'. here's the output of cssdiff between the sport and economy categories: [st08764@gandalf emneklassifisering]$ /usr/local/crm114/bin/cssdiff spo.cfc eco.cfc Sparse spectra file spo.cfc has 3397338 bins total Sparse spectra file eco.cfc has 3397338 bins total File 1 total features : 570588910 File 2 total features : 588381755 Similarities between files : 0 Differences between files : 579485332 File 1 dominates file 2 : 570588910 File 2 dominates file 1 : 588381755 [st08764@gandalf emneklassifisering]$ notice that the number of features and the times they dominate each other is identical. it also shows absolutely no similarities between any categories, even between bus.cfc (business) and eco.cfc (economy), which is rather implausible. i'm thinking this can't be right, but i don't quite understand what the error is. does anyone recognize this pattern? ------------------------------------------------------------------------------ This SF.net email is sponsored by: High Quality Requirements in a Collaborative Environment. Download a free trial of Rational Requirements Composer Now! http://p.sf.net/sfu/www-ibm-com