Re: Comments on IM2000
James Craig Burley <[email protected]> 15 Apr 2005 22:49:39 -0000
| Newsgroups | gmane.mail.im2000 |
|---|---|
| Message-ID | <[email protected]> |
>On Thu, Apr 14, 2005 at 09:34:26PM -0000, James Craig Burley wrote:
>> I didn't explain that very concisely, I guess. I block email with
>> envelope senders *known* to *belong* to spammers.
>
>Oh, so you mean the tiny proportion of spam which does not carry a forged
>return-path. Clearly, the more people who implement this check, the higher
>the incentive for spammers to forge their return addresses.
Well, of course. As I said, I don't have any data on how much of the
UBE I *don't* get is blocked by this mechanism. I've used it for only
a few months, and doubt it is having all *that* much effect. But it
was fairly easy to set up, and I use it mainly out of curiousity and
because it just naturally followed from my stumbling on the site that
provides the data after receiving lots of spam from one domain that,
upon being googled by me, turned out to be owned by a known spammer
and on this data base (which was how I found it)!
I do think SPF and/or DomainKeys -- basically, any scheme designed to
reduce forgery across the board -- might make this kind of data base
more useful, and perhaps even necessary.
So this data base, which is trivially easy and cheap for me to
download every 24 hours and crunch into a qmail badmailfrom file
(largish, but no apparent need to cdb-ize it yet), makes it fairly
cheap and easy to sorta "leach" off those who are making SPF and
DomainKeys happen -- the sites that are incentivizing spammers to
*not* forge envelope-sender addresses, so they can control their own
SPF and DK records.
Once a data base like this is available, it helps not only with
envelope senders, but with HELO hostnames, and with any URLs found in
untrusted email content.
Of course, it "works" only insofar as one assumes *any* content served
via a domain owned by a "known spammer" is undesirable. The usual
caveats therefore apply ("I'm not a known spammer!", "How do I get off
this list??", "slashdot.org is a known spammer, it keeps sending me
summaries of articles I never requested!", etc.).
>Maybe we need to think of adding some sub-classification to "spam". The vast
>majority of the spam I see is from scumbags. They want to get their
>fraudulent message into my inbox, and either get eyeballs on their website,
What website? The one using the domain name they own, the one they
point to via a static IP, or the one that belongs to somebody else but
whose HTTP server they've somehow 0wned?
This part of our discussion is *tightly* related to the key "feature"
of IM2000 over SMTP.
IM2000 assumes a message (notification) recipient *will* have to
contact some entity chosen by, and acting on behalf of, the sender.
That could be a bare IP address, though I notice some IM2000
proponents suggest "outlawing" that in the design (I don't see why one
should bother, since any arbitrary domain name can be easily pointed
to any IP address, but nevermind that).
The main utility of this distinction is that a sender must therefore
have some kind of "server presence" acting on his behalf, probably in
the form of a domain name. (To me, that's just another one or two
points of failure in the email system, but I'll continue....)
Problem is, there are an infinite number of domain names out there
(versus only about 4 billion IPv4 addresses), and they are not all
under control of one genial central authority that can quickly and
reliably determine who is and isn't a spammer.
So, there'll need to be data bases of "trusted" and, certainly,
"spammer-owned" domain names available for IM2000 MUAs to cross-check
against in order to know to not bother pulling down (or unpinning) any
message content tied to message notifications from problematic domain
names.
As it happens, the data base I'm accessing is an example of exactly
that. I don't know how long it'll last, or how useful it really is,
but as someone is reliably providing it for (basically) free and it
seems to be well-maintained -- my cron job emails me diffs every 24
hours, so I can spot obvious problems -- I don't see the problem with
using it.
Put another way: if you *aren't* using such a data base to block email
sent by known spammers, in an IM2000 world, you *will* be using it to
do exactly that.
>or get me to phone or fax them for the con-trick to begin. They don't care
>if I can E-mail them back or not, and they don't care if any of their
>millions of messages bounces. It's a positive boon if the E-mail is
>difficult to track back to source (since they are almost certainly involved
>in fraud anyway)
The biggest problem with this line of thought, as illustrated by my
expanding on your mention of web sites, is that spammers for
*products* available online (versus, say, Nigerian spam) are highly
incentivized to provide easily reachable contact points in their
spams.
So, spammers obviously trying to dodge content filters, throwing tons
of arbitrary spaces and weird characters into what looks, on my screen
running my primitive MUA, almost like a pointillistic painting,
nevertheless insist on including a URL like http://www.whatever.ru, a
URL that is easily detected as the message comes in via SMTP and its
domain name looked up in the same badmailfrom file (should I bother to
patch my SMTP server accordingly).
Now, to the extent there has to be enough people to be duped into
contacting the spammers, the spam will make such contact easy.
Two obvious ways to do that: include a URL that is
machine-recognizable and -readable; and use a "real" domain name for
the envelope sender.
Both approaches yield a contact address that any 'bot can easily
compare to a data base of IP addresses and domain names known to be
*owned* -- not just forged -- by spammers.
So this takes a fairly easy shot at part of the problem. What doesn't
it handle?
- It doesn't scale up well, because of the sheer number of domain
names that can easily be obtained worldwide.
But improvements to DNS and domain-registration data base technology
can mitigate this somewhat, perhaps entirely. I.e. it *should* be
easy, given any domain name, to do non-human-assisted recursive
lookups on the owner of that domain name name and/or its various
sorts of hosts (registrant, ISP, and so on) until a sufficiently
certain threshhold of clarity ("trusted", meaning if spam is
received it can be reported and will be acted upon, or
"spammer-owned", or "untrusted") is reached.
- It doesn't detect spam sent with forged or useless envelope senders
*and* URLs in the email, but which might contain alternate
representations of contact info that machine readers can't pick up
(the classic "visit this URL [which is actually represented by a
JPEG or GIF]" case).
- It doesn't detect spam sent with forged envelope senders and/or URLs
that point to hijacked sites -- those not, strictly speaking,
"owned" by spammers, but sufficiently 0wned such that their own
*servers* are serving spam, exporting SPF/DK records, serving as
IM2000 mail stores, and so on.
- It doesn't detect spam sent with entirely legitimate envelope
senders and/or URLs, whose owners simply decide, on occasion, to
send UBE to which some of their recipients object as being spam,
despite normally being a major source of legit email. (Imagine if
AOL occasionally sent adverts for its services to all known non-AOL
users on the Internet, from [email protected]. Nothing
*technical* can stop this, if the message is designed to get past
the content filters of the day, because nobody can really block all
email from aol.com and still be said to be "using" email. Only
AI-like content analysis can help users avoid seeing email like
this, at least as high-priority email.)
>> >But your definition of spam is "sent by a program which does not correctly
>> >implement Internet RFCs".
>>
>> No no no no no! My definition of "spam" is everyone else's -- UBE.
>> It so happens that a substantial % of UBE-sending software can't
>> tolerate trivial variations in SMTP server responses, such as
>> multiline greetings.
>
>And so does another percentage of legitimate E-mail sources. Maybe your trap
>catches more spam than non-spam - today. But it will have false positives
>today, and as the spammers change, it will get worse.
I don't believe I've had any false positives of which I've ever become
aware, as a result of my running an SMTP server that requires a bit
more conformance on the part of a client, but I don't have a broad
range of sources of legitimate email these days.
But, you're right, as the spammers adapt, more spam will get past
these sorts of measures.
Given the trivial cost of implementing them, and the fact that they're
obviously still blocking a ton of spam for my domain, they've been
*totally* worth it.
I agree this sort of tactic is hardly a reason to *not* roll out
IM2000.
It does raise the issue of, how serious is the spam problem *now*,
*really*, such that we contemplate moving to a whole new system that
will still apparently require many of the same "augmentations" (RBLs,
whitelists, blacklists, etc.) that we are already using with SMTP in
order for IM2000 to actually stop spam?
>I think what you were saying before is that it's the *content* which
>distinguishes a spam from a non-spam; I won't argue too strongly against
>that. Certainly, a human being paid to delete spam from your inbox would do
>a very good job, just by examining each message.
>
>But for any sort of automated spam control to work, the only other useful
>definition of spam I can see is "stuff that is sent by spammers".
I think those two paragraphs really make the point. Until and unless
we have adequate "help" sifting through email such that *only* the
content (and its purported sender) is required as input, we *will*
need to continue coordinating and communicating information about
spammers, such as identification, tactics, etc.
So, let's say that content-only analysis is capable of solving N% of
the spam problem. I don't care what you think N is -- anywhere from 0
to 99.999 is okay.
That N% being handled by content analysis, and content analysis not
giving a rat's behind about whether the content arrived via SMTP,
IM2000, HTTP, FTP, or carrier pigeon, it ceases to be an issue in any
debate over *transport* technology.
So, we're now focusing on the remaining portion of spam -- 100% minus
N%, which we'll now treat as if it's 100% of the existing spam that is
*potentially* stopped by measures other than pure content analysis.
Now let's say that the sort of coordination and communication needed
to help everyone defined "spammers" in the sentence you use, "stuff
that is sent by spammers" -- constitutes X% of the remaining spam.
That is, we can theoretically eliminate X% of spam if we know exactly
*who* spams and *reliably* identify any incoming email (via SMTP or
IM2000) as coming from a known spammer.
Now, the pertinent question becomes, does requiring a sender to
provide a message store make *that* much difference, in terms of our
ability to reach the theoretical maximum for X, compared to what SMTP
is evolving to, in terms of putting practical requirements on senders
to inject a message from a source that is not immediately identifiable
as a known source of spam?
As long as spammers can 0wn machines that serve as message stores for
their spam, and own domain names that point to those stores, or simply
provide IP addresses pointing to them in their message notifications,
it seems obvious, to me, that the answer is "no".
>If you were able to block all communication from spammers, you'd block all
>spam (by definition). You'd also block all non-spam sent by spammers, but
>since they are scum, most people don't care. You might care if your job is
>to run an abuse-tracking service, and you need to be able to communicate
>with spammers.
Well, you are already assuming that it is necessary, or nearly so, to
block *all* email sent by known spammers (that is, from machines they
own, machines they 0wn, and domain names they own or 0wn), else
blocking based solely on content analysis would be sufficient, and
everything else would be purely a question of efficient utilization of
resources between sender and recipient running the content analysis.
I'm saying that, from your point of view, you're right, but IM2000
really doesn't seem to offer much of a bang for the huge buck$ it'll
take to roll out as a *replacement* for SMTP. (Would IM2000 have been
a better "starting point" for us, 20-whatever years ago, than SMTP? I
think not. How about SMTP with tracking instead of bounces? Perhaps.
SMTP with no assumed handing off of responsibility for a message along
with its content? Almost certainly.)
Whereas, from *my* point of view, maybe let's take another look at the
content-analysis side of things, take inspiration from the fundamental
IM2000 theory of using the bulk nature of UBE against the spammer,
combine information on the apparent persistence of a sender in
notifying/resending a given email (which might tend to decrease
"likely spam score") with information on the apparent breadth of
recipients of similar emails (which might tend to increase "likely
spam score") with whatever content analysis can be efficiently done at
that point in time, and provide the end user with a display
summarizing *all* available email -- including likely spam -- with the
latter simply identified, as a collection, as such.
How does this help "stop" spam? It doesn't. But it makes spam much
more expensive in order to penetrate the "market" to the same extent.
(Spammers have to keep issuing IM2000-style notifications or SMTP
deliveries to overcome ecrulisting. That's a lot of work for
bulk-email senders, and it'll attract lots more attention from
upstream providers, etc.)
And it makes recipients' behaviors much more opaque to the spammer,
because they rarely actually see their spam get *rejected* -- it
merely sorta "hangs around" and then maybe disappears, gets "bounced",
gets flagged as "finally read" by a user who knew to simply skim
through it quickly as part of a "lemme read my big chunk 'o' waiting
spam now" exercise, or whatever.
And, especially with naive end users, the best thing you can do for
them, to convince them that "offers" are really scams and/or spams, is
to present those emails on their screens not intermingled with
otherwise-legitimate ones, but grouped together in a way that tells
them "this appears to be junk, and I [the content-analysis bot] have
put all the 419-ish stuff in one chunk, all the
enlarge-your-weiner-ish stuff in another, and all the notifications
that you've just won a lottery in yet another".
After all, even the stereotypical grandma on the Internet for the
first time is much more likely to yawn after reading the 5th 419-type
scam in a 60-second period, starting when she saw her first-ever such
scam. Then it'll be easier for her to just click the "yeah, I get it,
this sort of thing is junk" button, which might automatically tell her
friends' MUAs, propagate the information upstream to conserve
resources (but still try to hide as much info as possible from
spammers), and so on.
How can spammers respond to this? Well, they'll have to work harder
to *constantly* distinguish their spam, as well as their sources, from
other spam, because they'll be in even more direct competition with
each other for priority on as many end users' MUA screens as possible
and for *unique* positions on those screens. 1000 different people
spamming the entire Internet with ads for Cialis? They'll have to
*each* get pretty creative in their prose-writing capabilities! (The
best will presumably move on to writing political speeches and other
more "legit" jobs.)
Generally, my point of view is, anything we can do cheaply that
requires spammers to spend, by comparison, more time and money, is
worth looking at.
I don't think IM2000 is necessarily cheap enough compared to what
it'll require of spammers.
I do think IM2000 has the kernel of one of several Big Ideas that can
make ubiquitous spam slowly dissipate over time, as it becomes less
profitable.
(And I think it's better for society, overall, to let the market teach
potential spammers not to get into the biz in the first place rather
than to find and imprison them. It's cheaper; the spammers are not
otherwise hardened criminals; and it's better if potential spammers
are nudged into more-productive activities than doing prison work. I
don't think spam is as corrosive a thing, in our society, as illegal
drugs, so I don't believe a War on Spam, in the usual sense, is
necessary or likely to be an efficient use of society's resources,
even though I'll admit I've possibly harbored personal fantasies about
punishments for spammers who are caught and captured. ;-)
>You'd block all non-spam sent by machines which have been hacked into by
>spammers. That's an unfortunate consequence, but hacked machines really
>shouldn't be on the network in the first place.
Right; hacked machines are not unlike people with split personalities
(or whatever it's called), where one personality is reasonable, even
useful, another a used-car-salesman type who won't stop trying to get
you to buy something.
Content analysis, if sufficiently sophisticated, can make a
distinction in such a case; if not, probably the best approach is to
ignore almost *anything* the source says or does.
But more severe measures, well outside the scope of any SMTP/IM2000
discussion, are needed when a machine (or person) becomes outright
"violent" in the milieu (as in participating in a DOS or dDOS attack).
>IP-based RBLs work surprisingly well. If you combine a blacklist of IP
>blocks owned by spammers (e.g. spamhaus), one of open relays and proxies,
>and a dynamic one which reacts in real time to new spam sources (e.g.
>spamcop), they work pretty well even in the current world. The problems are
>to do with identifying new sources quickly enough, and more importantly the
>collateral damage when you list an IP address which belongs to a mailserver
>shared between spammers and legitimate users.
>
>To avoid that damage, you need an indication with each message of the
>account identity used to submit it. This could be done with SMTP AUTH; but
>inertia means it won't happen. It is an inherent capability of IM2000.
I am still trying to understand how inertia won't mean *IM2000* won't
happen. There just doesn't seem to be that much of a difference, in
terms of what *really* has to happen in the real world, between an
SMTP AUTH world and an IM2000 world.
Whether it's a trusted message store or a trusted SMTP server, those
messages gotta get in their somehow and therefore have to come from
what that store or server *itself* considers a "trusted" (therefore
authenticated) source.
Sure, I understand that there will be some who will enthusiastically
jump on the IM2000 adoption bandwagon because it's a cool new tech and
spammers won't be exploiting it yet.
But, you know what? It's a whole lot easier to just run another SMTP
server, basically whitelisting anyone who contacts it, on another
port, using slightly different SMTP command verbs, and provide free
clients to your friends for them to use when they contact you.
That'd give adopters of that approach the same sweet sense of joy that
early adoption of IM2000 would give, without having to set up any mail
stores or even engage in any advocacy.
And I don't think "early adopters" describes the mind-set of the heavy
hitters in the email industry overall.
(But my background includes decades of catering to Fortran
programmers, so I might be excessively biased against rosy assumptions
concerning uptake on new tech, having been burnt by them myself on
many occasions. ;-)
>> >However, this means you'll refuse to accept mail
>> >from "good" mail sources which also happen not to implement the RFCs in the
>> >expected way.
>>
>> No I won't. *Their* clients will just have trouble transmitting email
>> to me. That's *their* problem -- I haven't "refused" it at all.
>
>I fail to see the difference between causing a problem when someone tries to
>send you a mail (such that it cannot be delivered), and refusing to accept
>the mail.
I'm not "causing a problem". I'm providing a multiline greeting,
which happens to state terms for usage of the server in legal lingo.
The user's SMTP *client* is having a problem with that; that is
*their* problem, though they are free to contact me and be whitelisted
in such a way that they get a one-line greeting, if it's that
important to them.
*Many* SMTP servers on the Internet use these sorts of measures,
because they've found that maybe 99% of more of all SMTP client
connections that "fail" because of them are known sources of spam, and
it makes more sense to just let those connections fail than to keep
doing various expensive data base lookups to figure out who they are,
where they've moved to, etc.
Similar measures include rejecting: "presumptuous" clients, also known
as "early talkers"; clients that simply refuse to allow for more than
a one-minute delay in response to a command; and clients that treat
temporary rejections as permanent or as cause for retries so immediate
that the server can infer that the source is spam.
All these measures "break" known spam software. That they might
"break" a handful of narrowly distributed, but legitimate, SMTP
clients has always been understood. Some purveyors of such clients
have responded by fixing the bugs, which is a good thing, since most
of these measures merely stress logic that needed to be in better
working order anyway. ("Early talkers" are, simply put, dangerously
broken.)
>Most people are *users* of E-mail software. They do not have the technical
>ability to *fix* the E-mail software that they use.
Such users do not typically run their own SMTP clients that connect
directly to MXes of destination hosts. Their smarthosts run them, or
their IT departments run them.
However, users of *spam-blasting* software often run their own
(broken) SMTP clients. Hey presto, foil them and you block spam,
without any heavy lifting or false positives *at all* (since "false
positive" actually means "incorrectly categorizing an incoming
connection/email as spam", and no incoming connection or email is
categorized at all -- the client merely finds it can't succeed at a
delivery, which it's supposed to handle properly, as that sort of
thing is going to happen sometimes anyway).
>> All of the "fixes" spammers need to apply to jump those hurdles
>> represent additional cost, especially writing and deploying new
>> software.
>
>They *will* happen, and soon (i.e. within months at most), as surely as the
>widespread implementation of MAIL FROM domain checks made spammers change to
>sending out mail with real (but forged) E-mail addresses as the envelope
>sender. The cost is minimal.
Right, it's an arms race, we all know that already. None of what you
have said on that topic is news to me; I doubt it's news to anyone
else here.
>> No, as is widely recognized. It's a temporary tactic, one which
>> raises the bar in terms of increasing the expense of sending UBE (and
>> the visibility that one is doing so), which is kinda the whole point
>> of IM2000, right?
>
>No.
Then what *is* the point of IM2000, when the various web pages
promoting it proclaim that it puts more of an economic burden on
*senders* than does SMTP?
>I don't support any change in the E-mail system that we use today, which can
>be bypassed just by the spammer changing tactic. Any solution has to be
>*strong*. Otherwise, it's a total waste of time.
>
>And incidentally, being on this list doesn't mean I think we *should* all
>switch to IM2000. I'm still thinking about it :-)
>
>It's clear that IM2000 still allows spam to be sent; ISTM however that *any*
>concept of a 'mailbox' which arbitary people are allowed to drop mail
>into, will also support the delivery of spam.
>
>At the moment, I think the best we can achieve is the near-instant
>"cancellation" of accounts being used to send spam; and the blacklisting of
>IP sources which allow unrestricted creation of new accounts, or unlimited
>sending of messages from a single account, or are entirely controlled by
>spammers.
>
>It has to be reactive like this, since spammers can create new on-line
>identities at will.
>
>IM2000 suits this sort of cancellation pretty well, and the level of changes
>required across the Internet to achieve the same thing with SMTP would be of
>similar order of magnitude.
Are you including the cost of rolling out IM2000 as a *wholesale*
replacement for SMTP in your calculations? By "wholesale", I mean the
costs to get to the point where pretty much *nobody* is bothering to
accept SMTP email coming from arbitrary sources anymore, because
IM2000 is working so well, versus the alternative of just continuing
to use SMTP and focus those resources on doing other things.
I find it hard to believe you are, because those costs are *huge*. It
seems, to me, that it'll be much less expensive to incrementally
augment SMTP, as much of a bodge as it has become, in the same ways
we'd have to augment an otherwise-pure, and entirely theoretical,
IM2000 world, in order to achieve basically the same results.
Again, I think the best thing for any IM2000 enthusiast to do *now* is
*not* to whip up a prototype of IM2000, rather, to whip up a quick
SMTP implementation (MTA, MUA, plus other goo) that mimicks an IM2000
"experience" fairly closely, so people can get a feel for how well
it'd work in the context of a system that *is* already experiencing
lots of spam.
Maybe we should try designing that prototype -- call it something like
IM1999, and if subsequent generations are needed, decrement to name
them IM1998, IM1997, etc. -- on another list?
--
James Craig Burley
Software Craftsperson
<http://www.jcb-sc.com>