Re: Large scale htdig
Aaron <[email protected]>
| Newsgroups | gmane.mail.archives.mail-archive |
|---|---|
| Message-ID | <[email protected]> |
Actually, I think you guys have a great system over there. In fact, we considered you before we began our own htdig project. The problem is I am setting this up for all of Apple's public facing mailing lists, which brings up a few issues... 1. The main issue is that we have no guarantee that you would be around tomorrow. It is mentioned many times on your page. Not that this is a really bad thing, but for a large corporation, it's not ideal. 2. The ads. Again, not a horribly bad thing (and understandable in your case), but not something we wanted to have on our search engine. 3. Scaleability. Currently we have over 900,000 documents that need to be indexed and kept up to date. There are over 100 mailing lists and we average roughly 500 posts overall a day. I didn't want to break anything ;) As I said, I think you have a great system going, and I am often looking through archives for projects Im working on that you host. I think this is just something we are going to host in-house. Thanks for the tip, I created separate databases with separate configs for each list. This seems to have helped a bunch. Any other words of wisdom? =) Thanks again! -=Aaron On Sep 22, 2004, at 3:30 PM, Jeff Breidenbach wrote: >> I was wondering if anyone could offer some words of advice on setting >> up htdig on such a large scale. > > The Mail Archive uses a separate htdig index for each list, which > helps. The service has also spent quite a few CPU-years doing > incremental indexing. I'm curious - why are you rolling your own > system as opposed to using the The Mail Archive? Is there something we > should be improving? > > Cheers, > Jeff > > _______________________________________________ > Gossip mailing list > [email protected] > http://jab.org/cgi-bin/mailman/listinfo/gossip