Re: Mixed 64-bit system GerH binaries / BillYscripts --> two-sided training? YES!
"C. Cau" <crm114-tRhV7SAp1/[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
On Sunday 31 August 2008, Ger Hobbelt wrote: > Spot on for OSB and colleagues (had the same thought myself a while > back), but CRM114 also has 'growing' classifiers, i.e. classifiers > without constant filesize (HyperSpace IIRC, BE and other Yes... I'm a hyperspace junkie since the very beginning of it. > 'experimentals' can also 'grow' the CSS file) and for those *I* would > not know how to 'mix&merge' the two CSS files into one. Which is kind > of a showstopper for me regarding 'bundled' CSS. I think it's "just" (again, sorry) a matter of adding more info to the single token/hash, and to the .CSS/HYP file as a whole. Hard proof I don't have, to be sure, but my feeling is that the extra info would still save same serious space when compared to 2 (or more!) .HYP files which overlap - as hashes go - for a substantial percentage. > Which doesn't mean it's not doable and a great idea besides. alas, the devil is in the ******* details, as usual... > Advocate of the Devil here: the only 'drawback' I can see is when you > intend to 're-use' single CSS files in multiple A/B comparisons based > on different CSS file combos: merging all those into one will > bottleneck the training if that can be applied to a single pair A/B > in a large set of N CSS files (as they are re-used). exactly, same fear here. hypothetical as it may sound, it might affect some of my training/testing schemes. [...] > > Hey, all we'd need to modify is "just" the whole .CSS concept, > > after all, and perhaps a few hundred thousand lines of (existing > > and working) code... :-) > > Aw, heck, trivialities, that. What's the word? "Negligible"? ;-)))) *that* was the word I couldn't remember, thank you :-)) > Anyhow, good thinking IMO. (And side note: you can even do it when > the hash/word must exist on both sides, but then the CSS total size > stays the same compared to today. When you can prove hash on one side > only, you can introduce 'negative counts' to signal side (good vs > bad) and store same amount of info in about half the total file size. > ;-) (Where I 'forget' to mention a few details, but no time for long > elaborations now.) ) that was part of the "extra info" I was thinking about, in fact. Signal strength, and/or prevalent class, would get the lion's share of it. I've never investigated how many tokens (hashes) are present in both of my .HYP files (2-way classification), but I'd expect a good percentage. > <sorry for choppy text above, but now I must go back to my prayer > beads and hope the next phonecall is a green light instead of another > failure report for my release tomorrow...> good luck! Corrado ------------------------------------------------------------------------- This SF.Net email is sponsored by the Moblin Your Move Developer's challenge Build the coolest Linux based applications with Moblin SDK & win great prizes Grand prize is a trip for two to an Open Source event anywhere in the world http://moblin-contest.org/redirect.php?banner_id=100&url=/