Re: Moving forward ...

Laird Breyer <[email protected]> Thu, 26 Feb 2004 10:59:33 +1000
Newsgroups gmane.ietf.asrg.filtering
Message-ID <20040226005933.GB30647@ender>
Here are some thoughts on some of these issues.

> * spam filtering isn't the only potential filter running.  A filter
> header specification should be extensible to other filtering types.

Filtering headers add to the information processed downstream. 
To prevent unwanted feedback (ie a procmail filter catching on to
a SpamAssassin recommendation), the spec needs to be so that added
headers can be recognized by all and skipped. This suggests a common
prefix such as "X-SPAM-", so that the programs downstream don't need
extensive logic to identify and skip headers.

Similarly, adding spam scores or "convenience" functions (e.g. URLs)
inside the body of an email should be prohibited. Anything that makes
it so the downstream filter cannot see exactly the same message as 
the upstream filter is unstable (ie if you remove a filter in the
chain, you've sabotaged training for all filters downstream).

> * The definition of spam isn't universal and filtering technologies may
> have subtle differences in their implementation of spam detection.  To
> this end, labelling a message as spam may not be consistent. (unwanted
> vs. UCE or some combination thereof)

Moreover, the labelling may not even be consistent over several
versions of the same filter...

> 
> * filtering decisions are dependant on the state of the filtering (virus
> definitions, software version) at the time of message scanning resulting
> in different results for the same message at different times.  It is
> important for filters to be able to add meta data to the label.

This is bad. I'd much prefer a system where my filter can easily
ignore anything done by other filters. Let the user decide how to 
integrate all the various recommendations.

> 
> * some filters develop a confidence as well as a label, allowing later
> processing to be tuned to reduce errors.  Some filters are unable to
> parse these confidence levels.  A common format for confidence is also
> required.

Is the confidence intended for the originating filter only, or for
every filter downstream (including the MUA). Confidence on its own
doesn't mean much without reference to the filter algorithm. Other
filters which use different algorithms would have to be intimately
familiar with the workings of the originating filter to make use of 
the stated confidence.

> 
> * Labels may be added to the message prior to gateway filtering (on the
> outgoing MTA?) and may not be trusted.  How to tell if a label was added
> by the trusted MTA?  What should an MTA do if it receives an incoming
> labelled message?
> 	- this message was virus scanned by ...

I'm not sure trusting is the way to go. I'd rather leave this to the
user. In other words, when a message arrives with a spoofed label from
filter X, then:

If independent filter Y sees the label, it should ignore it.
If filter X sees it, it should recompute the label anyway and overwrite.
If MUA sees it, it's up to the user to decide.

Again, I really don't like the idea that filter should somehow
cooperate with an upstream filtering system.

> 
> * Recipients may believe that a message was labelled in error and want
> to provide as a defect to the administrator.  This administrator (or the
> filter vendor) will need more meta data about the filtering than a
> single judgement label.  More state data will be required to diagnose
> and adjust/reject the feedback.
> 

This suggests a different type of labelling is already used by some filters:
instead of advisory labels only, the message also carries state
information. I think that means we need to discuss two different
types of headers separately:

1) headers which advise the downstream programs.
2) headers which are meant to be private to the filter.


-- 
Laird Breyer.