Re: testing corpus?
Rick Wesson <[email protected]> Fri, 26 Jun 2026 20:02:39 +0000
| Newsgroups | gmane.mail.spam.spamassassin.general |
|---|---|
| Message-ID | <F_efzO3oM1_kmcrjbpDpyPZKD9BlWZsD_Dvy9dtvBQMFfadfVoWlX1yqDWEsARYmWN9k4v-ngmnf76dwWFiy9NkGBp9DSB9mj6qqsZ-yAcw=@support-intelligence.com> |
Bill, Thanks for the thoughts, I appreciate it. Is there a best way to get spam f= or research purposes? At one time I got to put mx records on a bunch of dom= ains, but I no longer have that capability. How do folks to it in 2026? -rick On Friday, June 26th, 2026 at 12:55 PM, Bill Cole <sausers-20150205@billmai= l.scconsult.com> wrote: > On 2026-06-26 at 14:48:18 UTC-0400 (Fri, 26 Jun 2026 18:48:18 +0000) > Rick Wesson <[email protected]> > is rumored to have said: >=20 > > Hello, > > > > I've been doing evaluations of various clustering algorthims. I'd like > > to do some testing to see if I can improve SA speed by adding some > > clustering capabilities. Is there a goto corpus that you use for > > testing and baseline. > > > > I've attached a draft of some of my research using malware corpus and > > I'd like to begin working on doing the same with spam. >=20 > As far as I know, there's no shareable corpus of email that is useful > for spam research. The problems are that there is a very loose consensus > on what is or is not spam, what spam and non-spam people get has huge > variability, and personally identifying information can be a significant > clue in identifying spam. >=20 > The corpora that I've used for research have all been built from the > mail flows of the people backing the research. >=20 >=20 >=20 > -- > Bill Cole > [email protected] or [email protected] > (AKA @[email protected] and many *@billmail.scconsult.com > addresses) > Please keep discussion mailing list replies *on-list* > Not Currently Available For Hire >