Re: Realisticness of header rearrangement

Laird Breyer <[email protected]> Tue, 6 Apr 2004 22:10:38 +1000
Newsgroups gmane.ietf.asrg.filtering
Message-ID <20040406121038.GA13165@ender>
On Apr 06 2004, Bill Yerazunis wrote:

>    > And the second use isn't a time stamp any more, therefore its confusing.
>    > An explicit event-ID would be clearer (not confusing as a timestamp
>    > being
>    > used for something other than marking the time something happened))
> 
>    The received time stamp is only a hack, but it's the only one that has
>    a hope of working. 
> 
> Unless each filter is trusted to re-edit timestamps of the incoming
> text, recieved time stamps are forgeable.

I don't quite see it. Care to explain what you have in mind? 

To me it seems clear that whatever header editing occurs on the
spammer side, it occurs before the SMTP server writes the Received:
line. So any header lines which can *prove* that they've seen the
Received: line are safe, assuming the spammers don't have a trojan
filter behind the SMTP server.

> 
> I'm still puzzled.  It seems to me that the simple solution is to 
> work on the assumption that the spammer _CAN_ forge any header;
> they just can't remove headers that are already present.

That's the most basic assumption. The assumption some of us have got
going in this thread is that filters which live behind the company
SMTP server are fully trustable, because we've installed them. The
natural question is, can a filter prove it's living behind the SMTP server?


> In that case, the obvious solution is to have all prefilters either
> say nothing, or say "spam" (and perhaps some quantification of that).
> 
> With that doctrine, a spammer -cannot- forge a useful header; they
> can only hurt themselves.

> Is there any reason (beyond the fact that it means a filter can't say
> anything good) that this doctrine has been ignored?

It hasn't been ignored, we just haven't discussed it much, yet ;-)
As far as I'm concerned, it's a worthwhile doctrine.
It's nowhere written that we can't have several partial solutions. 

The obvious issue with it is, as you point out, that filters must
either say "spam" or say nothing. I'd extend this a bit and say that
filters also aren't allowed to indicate degrees of spamminess.

Otherwise, for example, a spammer might try to spoof a tag saying
"spam/unsure", on the assumption that the "unsure" might shunt the 
email into a special folder which is looked at more often than the
junk folder. 

Similarly, if a well known filter writes a spamminess score between 0 and 1,
with 0 being close to ham, then the spammer might prefer to spoof
"spam score 0.1" instead of "spam score 0.9". Maybe this will again
trigger a safety system to let the mail through.

The other obvious issue is that the doctrine is not really suitable
for taggers which have several categories, eg popfile. That's going
slightly outside of spam filtering, but it's a natural direction even
for spam filtering. 

For example, think of a spam filter which distinguishes several types
of spam (pr0n, nigerian, etc). A spoofing campaign could have the goal
of blurring the distinction between the categories, ie poisoning the filter.

Just my 2 cents on this so far.

-- 
Laird Breyer.