Re: Mixed 64-bit system GerH binaries / BillYscripts --> two-sided training? YES!

"Ger Hobbelt" <[email protected]>
Newsgroups gmane.mail.spam.crm114
Message-ID <[email protected]>
@Fidelis: thank you very much! I do not know if it will work for me
(and quick scan of subsequent conversation here leads me to believe
that I must check the source code as well to see how and if the ideas
are already in there -- sorry, brain is focused on an off-topic
release too much to be able to 'see' this yet. This must wait a few
days until things are in the green at my customer site. But I will
definitely take time to look into the material you provided; I hope
there's time to read up next weekend.

@Paolo: also thanks for listing the objections and observations from
past experiments regarding the 'decrement' idea. It saves me a bundle
of time as I was thinking about it but definitely felt it was only
good enough AFTER I'd seen some solid field trials first.

Regarding the bang-bang controller-like behavior: yes, that was/is a
'calculated' risk. But remember also that bang-bang controllers are
not necessary instable. It just that regulating them is a little
different. I have notes somewhere (dang, _where_?), anyway, the next
'layer' in the idea was to supply a non-linear weighting factor (for
example the squared 'learncount' as stored in the CSS files, with a
little tweak to make sure that a count of +1, 0 or -1 is all VERY
close to 'negligible' (lovely word, that ;-) ) so that you get a sorta
'smooth' transition. Besides, I think (but that's my data leading me
on) that +N/-N learncounts are to be ignored where N is small
(shouldn't say N=1 because learncount will be immediately bumped up
beyond +1/-1 when using repetitive learn algorithms like THTTR:
learncount=N will mirror the number of loop iterations N for message
M, so you need some sort of way to recognize both 'number of
_messages_ trained which contained this word' and the 'strength
(loopcount) with which this message [word] was trained when training a
single message'.

Sorry, choppy again, but I am still not clear on how to get that into
CRM114 in a /nice/ way: dual counts: BOTH message count AND learncount
(as it already exists now) PER HASH. Then decide on 'insignificance'
based on 'messagecount' and filter the learncount weights which
produce the final pR.

[There is also a lingering thought about 'assymmetric weighting', e.g.
apply assymmetric, but continuous functions to the weighting like
this:
  g(x) = scaleG * learncount
for 'good' side (i.e. a linear scale function, where g(x) is the
weight value for hash x, learncount the learncount for hash x and
scaleG a preconfigured constant -- this is the current CRM114 mode of
operation for both good and bad sides IIRC) and
  b(x) = scaleB * power(learncount, powB)
for 'bad' side where b(x) is weight value for hash x ans scaleB and
powB preconfigured constants, so g(x) = -b(x) at and near
learncount==0, i.e. the connection point between good and bad -->
continuous weight curve across the edge and a behaviour which ramps up
'bad' scores much faster (or slower for powB < 1) than 'good' scores
for same number of trainings on hash x.]
I had some scribbels somewhere about g(x) and b(x) tweaks to make them
have a 'knee' at learncount=N+1, going for 'negligible' to
'significant' beyond that point, so you get a kind of a V curve with
the bottom sawed off and flatlined, resulting in a curve like a morph
between U and V if you see what I mean.

These ideas may not work for email; I don't know if they have been
tried on small messages and what one would run into then, so I got to
build it and test to see if the brain and world agree. Probably not,
but checking my ideas is the only way forward.


On Sun, Aug 31, 2008 at 9:36 PM, Paolo <[email protected]> wrote:
[...]
> ok, but as long as you keep the picture in front of you we see just a nice
> brownish cardboard background :)

I know. I'm sorry.
Two things: (1) Even the CEO of Hobbelt inc. has an NDA on this. 2)
I've found that a few folks from the competition (not on this list)
are very interested in what I am doing here. It's research started
about 15 years ago and funded by me during that time. If it works out,
it's a very probable money maker and it has some signs of being just
that. And the smell of green is attracting very nice and intelligent
people who I'd love to deal with... on /my/ terms. I think that's
reasonable, right? That's why I keep a very tight leash on my
conversations when they consider our work. If you have ever worked on
patentable research, you'll know how it is. Sorry, can't share.
That's also why I don't want you to bother about me and my particular
CRM114 issues; got to solve them myself anyhow. Bottom line is: I
[intend to] use CRM114 for very non-standard stuff, so I run into all
kinds of very non-standard issues -- which makes it fun for me too.
;-) The discussions here about CRM114 fundamentals have helped me a
*lot*, so thank you all for that. My product can work without it, but
when I would run the tuning test sets, data or other bits into this ML
to benefit from your additional expertise on matters CRM114, there is
a significant risk that the competition will be able to see what I am
doing and get a pretty darn good idea on the how as well. Google
sometimes is NOT your friend, okay?
The spin-off that I can give (and have tried to give) you is work on
CRM114 that I have done and intend to keep doing in public. I am the
decision maker here, so nobody but 'available time' can shut that
down. My lopsided way of saying 'thank you' by trying to give
something back. Right now, CRM114 is on the backburner here in the
lab, because it has been decided by the team that we go to initial
production with the field-tested setup, which does not include CRM114
in the pipeline. It has been a very hectic summer, but when the
wrinkles for this phase have been ironed out and it is running
satisfactory, then I am able to pick up CRM114 again at (I hope)
close-to-full blast and see if I can push the new ideas to become as
successful or even better, while 'improving' on CRM114 as it is. The
full-auto-tuning/selfadjusting possibilities are a definite lure for
us. ( I am both funder and 'IT geek' on this project, by the way. ;-)
I picked CRM114 after comparing a multitude of packages. It still is
the only one for me, particularly because it is one of the few
packages that come with significant support from publishing members of
the research community (most important of course: BillY). Unless lab
results show me CRM114/statistical filtering is a fatal train of
thought at *concept level* for my purposes I won't stop with CRM114.
It *may* be that either the technology or, God forbid, my own
abilities are not up to par regarding the application of statistical
filters to our subject matter (other team members bring other
abilities to the table here), but I think I still have some learning
ability left in me and investing into a mere thought for 15 years
might be a sign that I am a wee bit stubborn some times as well. )

>> into the system in full-auto. Kinda like this: you're SETI listening
>> for extraterrestial signals and you just don't know what to look for
>
> ah, ok ... so you're looking for ET ;)

Yep. And as you may have read: ET already had a latte on our porch.
Forgot to get his bike on the way out, though. I'm still waiting for
the Giger guys from Alien though. I hear they can throw a heck of a
party. ;-)

> eh not that easy, you've to choose whom you want the grace, else they mess
> up ;)

Nah, I just don't tell 'em I also tried to collect at the neighbours'.

Anyway, looks like the deities came through collectively (thank you,
my revered Almightinessirs and -nesses!) as I've got greenlighted for
tomorrow just half an hour ago: 'good results!' was the text (instead
of the 'doesn't look good' grumble I was getting so used to during
this week and last.


> PS: I'd rather move such discussions on -discuss or -devel.

Uh, yeah. Oops. Sorry, should I move overthere now?



-- 
Met vriendelijke groeten / Best regards,

Ger Hobbelt

--------------------------------------------------
web: http://www.hobbelt.com/
 http://www.hebbut.net/
mail: [email protected]
mobile: +31-6-11 120 978
--------------------------------------------------

-------------------------------------------------------------------------
This SF.Net email is sponsored by the Moblin Your Move Developer's challenge
Build the coolest Linux based applications with Moblin SDK & win great prizes
Grand prize is a trip for two to an Open Source event anywhere in the world
http://moblin-contest.org/redirect.php?banner_id=100&url=/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.