Re: Email Web of Trust - Problem Statement
"Peter J. Holzer" <[email protected]> Tue, 9 Mar 2004 17:42:54 +0100
| Newsgroups | gmane.ietf.asrg.smtpverify |
|---|---|
| Message-ID | <[email protected]> |
On 2004-03-08 15:38:48 -0800, Mark Baugher wrote: > At 02:18 PM 3/8/2004, Alan DeKok wrote: > >Mark Baugher <[email protected]> wrote: > >> > Taken over multiple intermediary hops, these numbers give you some > >> >probability that any one message is a priori going to be spam. It's > >> >not perfect, but it's a start. > >> > >> Yes, but it begs the question of what is spam. > > > > Nope. Spam is whatever a domain decides is spam. > > I could be missing something here, but "web of trust" to me means that a > collection of principals have decided to trust each other to some degree > for some specific access or authorization to some resource. So I don't > understand your response. One principal that speaks for a particular > domain might have its own definition, but where does the web come into play? A simple example may make it clearer. I'm using persons (Alice, Bob, ...) etc. for the principals in the web. They could be machines running an MTA, or domains, or organizations. I'm using only one metric of trust. The metric is the probability that a given message from that person is interesting (i.e., not spam). This is sufficient to answer Alan's "simple question": > > - If the message was passed from person to person, rather than being > > delivered directly, what are the odds it would be marked along the > > way to be spam? Alice and Bob are friends. They share interest in most things, almost any mail they send each other is considered interesting by the receiver. So they have awarded each other a fairly high trust level, say 0.99. Bob has a friend named Charlie. He's a nice enough guy, but he always forwards (not very funny) jokes and powerpoint presentations to Bob, so Bob only considers half of Charlies messages interesting and gives him a trust level of 0.50. Now Charlie sends mail to Alice. Alice doesn't know Charlie yet, but she agrees with Bob on 99% of all messages and Bob says 50% of Charlie's messages are interesting. So she assumes that 49.5% of Charlie's messages will be interesting to her (erring on the side of caution). When Alice has received a few mails from Charlie, she can make up her own mind about his mails. Let's say she thinks his jokes are disgusting, and she never received any other mails from him, but since she knows he's a friend of Bob she doesn't quite set his trust level to 0, but to 0.05, just in case he might write something else but jokes someday. At that point Alice and Bob have different opinions about whether Charlie's mails are spam or not. Bob decides 50% of Charlie's mails are spam, Alice decides 95% of his mails are spam. Now Debbie knows both Alice and Bob, but not well. She has assigned them trust levels .6 and 0.7 respectively. So when Charlie sends her a mail, there are two paths: Through Alice or through Bob. The trust through Alice is 0.6*0.05 = 0.03. The trust through Bob is 0.7*0.5 = 0.35. On average that's a trust level of 0.19. > >> One needs a very clear definition of what is spam in order to trust > >> that some domain does or does not originate spam. And there could > >> be more than one metric such as originating UBE promotions versus > >> nefarious scams to bilk people out of their savings. > > > > More metrics make it more complicated. The simple question is: > > > > - If the message was passed from person to person, rather than being > > delivered directly, what are the odds it would be marked along the > > way to be spam? > > I think I'm missing something. Is the "web" the chain of persons > (relays?), specifically, or is it more general, such as a web of mail > operators that trust each others authorization decisions? It could be either or both. I'm leaning in favour of persons (or, more accurately, mailboxes or personae, to allow a person to have several independent identities), Alan and Yakov prefer larger entities like MTAs or domains. For the basic idea of a web of trust, it doesn't make much difference. Maybe the distinction isn't even necessary in practices, as you could publish trust records like "I believe that messages signed with PGP Key Id 0x6F2ABB2D are never spam" or "I believe that 20% of all messages sent through the MTA with the IP address 143.130.50.112 are spam". hp -- _ | Peter J. Holzer | I think we need two definitions: |_|_) | Sysadmin WSR | 1) The problem the *users* want us to solve | | | [email protected] | 2) The problem our solution addresses. __/ | http://www.hjp.at/ | -- Phillip Hallam-Baker on spam [demime 0.99d.1 removed an attachment of type application/pgp-signature]