Re: Mixed 64-bit system GerH binaries / BillYscripts --> two-sided training? YES!
"Ger Hobbelt" <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
@Fidelis: thank you very much! I do not know if it will work for me (and quick scan of subsequent conversation here leads me to believe that I must check the source code as well to see how and if the ideas are already in there -- sorry, brain is focused on an off-topic release too much to be able to 'see' this yet. This must wait a few days until things are in the green at my customer site. But I will definitely take time to look into the material you provided; I hope there's time to read up next weekend. @Paolo: also thanks for listing the objections and observations from past experiments regarding the 'decrement' idea. It saves me a bundle of time as I was thinking about it but definitely felt it was only good enough AFTER I'd seen some solid field trials first. Regarding the bang-bang controller-like behavior: yes, that was/is a 'calculated' risk. But remember also that bang-bang controllers are not necessary instable. It just that regulating them is a little different. I have notes somewhere (dang, _where_?), anyway, the next 'layer' in the idea was to supply a non-linear weighting factor (for example the squared 'learncount' as stored in the CSS files, with a little tweak to make sure that a count of +1, 0 or -1 is all VERY close to 'negligible' (lovely word, that ;-) ) so that you get a sorta 'smooth' transition. Besides, I think (but that's my data leading me on) that +N/-N learncounts are to be ignored where N is small (shouldn't say N=1 because learncount will be immediately bumped up beyond +1/-1 when using repetitive learn algorithms like THTTR: learncount=N will mirror the number of loop iterations N for message M, so you need some sort of way to recognize both 'number of _messages_ trained which contained this word' and the 'strength (loopcount) with which this message [word] was trained when training a single message'. Sorry, choppy again, but I am still not clear on how to get that into CRM114 in a /nice/ way: dual counts: BOTH message count AND learncount (as it already exists now) PER HASH. Then decide on 'insignificance' based on 'messagecount' and filter the learncount weights which produce the final pR. [There is also a lingering thought about 'assymmetric weighting', e.g. apply assymmetric, but continuous functions to the weighting like this: g(x) = scaleG * learncount for 'good' side (i.e. a linear scale function, where g(x) is the weight value for hash x, learncount the learncount for hash x and scaleG a preconfigured constant -- this is the current CRM114 mode of operation for both good and bad sides IIRC) and b(x) = scaleB * power(learncount, powB) for 'bad' side where b(x) is weight value for hash x ans scaleB and powB preconfigured constants, so g(x) = -b(x) at and near learncount==0, i.e. the connection point between good and bad --> continuous weight curve across the edge and a behaviour which ramps up 'bad' scores much faster (or slower for powB < 1) than 'good' scores for same number of trainings on hash x.] I had some scribbels somewhere about g(x) and b(x) tweaks to make them have a 'knee' at learncount=N+1, going for 'negligible' to 'significant' beyond that point, so you get a kind of a V curve with the bottom sawed off and flatlined, resulting in a curve like a morph between U and V if you see what I mean. These ideas may not work for email; I don't know if they have been tried on small messages and what one would run into then, so I got to build it and test to see if the brain and world agree. Probably not, but checking my ideas is the only way forward. On Sun, Aug 31, 2008 at 9:36 PM, Paolo <[email protected]> wrote: [...] > ok, but as long as you keep the picture in front of you we see just a nice > brownish cardboard background :) I know. I'm sorry. Two things: (1) Even the CEO of Hobbelt inc. has an NDA on this. 2) I've found that a few folks from the competition (not on this list) are very interested in what I am doing here. It's research started about 15 years ago and funded by me during that time. If it works out, it's a very probable money maker and it has some signs of being just that. And the smell of green is attracting very nice and intelligent people who I'd love to deal with... on /my/ terms. I think that's reasonable, right? That's why I keep a very tight leash on my conversations when they consider our work. If you have ever worked on patentable research, you'll know how it is. Sorry, can't share. That's also why I don't want you to bother about me and my particular CRM114 issues; got to solve them myself anyhow. Bottom line is: I [intend to] use CRM114 for very non-standard stuff, so I run into all kinds of very non-standard issues -- which makes it fun for me too. ;-) The discussions here about CRM114 fundamentals have helped me a *lot*, so thank you all for that. My product can work without it, but when I would run the tuning test sets, data or other bits into this ML to benefit from your additional expertise on matters CRM114, there is a significant risk that the competition will be able to see what I am doing and get a pretty darn good idea on the how as well. Google sometimes is NOT your friend, okay? The spin-off that I can give (and have tried to give) you is work on CRM114 that I have done and intend to keep doing in public. I am the decision maker here, so nobody but 'available time' can shut that down. My lopsided way of saying 'thank you' by trying to give something back. Right now, CRM114 is on the backburner here in the lab, because it has been decided by the team that we go to initial production with the field-tested setup, which does not include CRM114 in the pipeline. It has been a very hectic summer, but when the wrinkles for this phase have been ironed out and it is running satisfactory, then I am able to pick up CRM114 again at (I hope) close-to-full blast and see if I can push the new ideas to become as successful or even better, while 'improving' on CRM114 as it is. The full-auto-tuning/selfadjusting possibilities are a definite lure for us. ( I am both funder and 'IT geek' on this project, by the way. ;-) I picked CRM114 after comparing a multitude of packages. It still is the only one for me, particularly because it is one of the few packages that come with significant support from publishing members of the research community (most important of course: BillY). Unless lab results show me CRM114/statistical filtering is a fatal train of thought at *concept level* for my purposes I won't stop with CRM114. It *may* be that either the technology or, God forbid, my own abilities are not up to par regarding the application of statistical filters to our subject matter (other team members bring other abilities to the table here), but I think I still have some learning ability left in me and investing into a mere thought for 15 years might be a sign that I am a wee bit stubborn some times as well. ) >> into the system in full-auto. Kinda like this: you're SETI listening >> for extraterrestial signals and you just don't know what to look for > > ah, ok ... so you're looking for ET ;) Yep. And as you may have read: ET already had a latte on our porch. Forgot to get his bike on the way out, though. I'm still waiting for the Giger guys from Alien though. I hear they can throw a heck of a party. ;-) > eh not that easy, you've to choose whom you want the grace, else they mess > up ;) Nah, I just don't tell 'em I also tried to collect at the neighbours'. Anyway, looks like the deities came through collectively (thank you, my revered Almightinessirs and -nesses!) as I've got greenlighted for tomorrow just half an hour ago: 'good results!' was the text (instead of the 'doesn't look good' grumble I was getting so used to during this week and last. > PS: I'd rather move such discussions on -discuss or -devel. Uh, yeah. Oops. Sorry, should I move overthere now? -- Met vriendelijke groeten / Best regards, Ger Hobbelt -------------------------------------------------- web: http://www.hobbelt.com/ http://www.hebbut.net/ mail: [email protected] mobile: +31-6-11 120 978 -------------------------------------------------- ------------------------------------------------------------------------- This SF.Net email is sponsored by the Moblin Your Move Developer's challenge Build the coolest Linux based applications with Moblin SDK & win great prizes Grand prize is a trip for two to an Open Source event anywhere in the world http://moblin-contest.org/redirect.php?banner_id=100&url=/