Odds of Recurrence for Hash Function

Chris Babcock <[email protected]>
Newsgroups gmane.mail.spam.crm114
Organization ASCII King Games
Message-ID <[email protected]>
Until I am able to process mail in a single file format, I have been
using the hash command to generate file names for storing copies of the
output. The advantage to this is that if there is a problem with the
program that the mail script is feeding, users sometimes panic and
resend a message... repeatedly. By naming the file with a hash of the
content, I eliminate duplicate messages if I need to do a restore.

Since the hash is a finite size and there are an infinite number of
possible content bodies, I'm assuming that there is a risk of two
different messages creating the same hash code. How many messages would
I have to process to have ~10% chance of a duplicate hash on a 32-bit
system? Are there any recommendations for salt values so that
concatenating a pair of hash values would minimize the risk of
overwriting a message with a non-duplicate? 

Chris

-------------------------------------------------------------------------
This SF.Net email is sponsored by the Moblin Your Move Developer's challenge
Build the coolest Linux based applications with Moblin SDK & win great prizes
Grand prize is a trip for two to an Open Source event anywhere in the world
http://moblin-contest.org/redirect.php?banner_id=100&url=/
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.