Re: dstttr

Ger Hobbelt <[email protected]>
Newsgroups gmane.mail.spam.crm114
Message-ID <[email protected]>
The best for an 'official description'[*] of DSTTTR I can do is this
(and it looks from that 3.0 pR increment check in your code you found
the very same article) but when this is really it, then at least
mealtrainer doesn't do DSTTTR like that, as it doesn't check for
continuous improvement (which doesn't always happen anyway) but rather
continues training until the threshold is finally passed, as I
described in the previous email:

[*] i.e. exclusively authored by peers listed in Burke's Peerage For
Spam Filtering

http://archive.netbsd.se/?ml=crm114-general&a=2005-11&t=1455142

on 11/6/05 4:16 PM, Fidelis Assis  wrote:
> wsy wrote:
>> Here's the newest description of TOER - note that it's renamed DSTTTR
>> because it is, after all, a thick-threshold two-sided learner.
>>
>> Here's the text as modified, now in the mainline:
>>
>> DSTTTR ­ Double Sided Thick Threshold Test and Reinforce Training
>>
>> DSTTTR is a training method that can be described as ³smart DSTTT².
>> We start with the same procedure as SSTTT ­ if an incoming unknown
>> text isn¹t scored correctly or scored correctly but not correctly
>> enough (typically, by a margin of at least 10 to 20 pR units) it
>> gets trained just as in SSTTT.  Then the text is re-classified
>> again, and the improvement measured.  If the new classification
>> doesn¹t meet the threshold requirement, or didn¹t improve in the
>> correct direction by at least some smaller margin (typically 3 pR
>> units) then the text gets trained out of the incorrect classes.  The
>> difference here is that in straight DSTTT, the train-out-of action
>> is based solely on the before-training pR values, but in DSTTTR the
>> decision to proceed with training out-of class is determined by the
>> pR value after training and retesting.  The reports are that DSTTTR
>> works even better than TOE in most circumstances.
>
> Good, this describes the kind of DSTTT implemented in
> mailfilter_light.crm, but it's not a new definition for TOER! SSTTT *is*
> a new definition for TOER, they (SSTTT and TOER) are exactly the same.
> I'd say that the method you describe above is much better than TOE, not
> just even better.
>
> Now, wrt to the name, DSTTTR, I think that the new name should reflect
> the main difference between the normal DSTTT and the new one. Both
> unlearn from the opposite class if needed, the difference is, as you
> said, what pR value is used to determine if an unlearn is needed: The
> before or the after-training value. So, what about BF-DSTTT for
> Before-Training DSTTT and AT-DSTTT for After-Training DSTTT?


-- 
Met vriendelijke groeten / Best regards,

Ger Hobbelt

--------------------------------------------------
web:    http://www.hobbelt.com/
        http://www.hebbut.net/
mail:   [email protected]
mobile: +31-6-11 120 978
--------------------------------------------------

------------------------------------------------------------------------------
Open Source Business Conference (OSBC), March 24-25, 2009, San Francisco, CA
-OSBC tackles the biggest issue in open source: Open Sourcing the Enterprise
-Strategies to boost innovation and cut costs with open source participation
-Receive a $600 discount off the registration fee with the source code: SFAD
http://p.sf.net/sfu/XcvMzF8H
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.