Odds of Recurrence for Hash Function
Chris Babcock <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Organization | ASCII King Games |
| Message-ID | <[email protected]> |
Until I am able to process mail in a single file format, I have been using the hash command to generate file names for storing copies of the output. The advantage to this is that if there is a problem with the program that the mail script is feeding, users sometimes panic and resend a message... repeatedly. By naming the file with a hash of the content, I eliminate duplicate messages if I need to do a restore. Since the hash is a finite size and there are an infinite number of possible content bodies, I'm assuming that there is a risk of two different messages creating the same hash code. How many messages would I have to process to have ~10% chance of a duplicate hash on a 32-bit system? Are there any recommendations for salt values so that concatenating a pair of hash values would minimize the risk of overwriting a message with a non-duplicate? Chris ------------------------------------------------------------------------- This SF.Net email is sponsored by the Moblin Your Move Developer's challenge Build the coolest Linux based applications with Moblin SDK & win great prizes Grand prize is a trip for two to an Open Source event anywhere in the world http://moblin-contest.org/redirect.php?banner_id=100&url=/