Re: Hashing, and Ted's concerns about wastefulness
Graham Klyne <[email protected]>
| Newsgroups | gmane.ietf.medfree |
|---|---|
| Message-ID | <[email protected]> |
At 15:16 02/02/99 -0800, Bill Newman wrote: >Ted wrote (quoting Graham) > >> Yes, but we need to specify exactly how that is represented. Remember >> that most hashes on files care about every bit of whitespace and every >> control character; if we don't specify the representation very well, >> we will end up with conflicting hashes on equivalent feature expressions. > >How bad a problem is this? Is it is worth trying to assign the same >hash values to differently-represented-but-ultimately-equivalent >feature sets? It is a feature, but it adds complexity to the >specification and to each implementation. Good question. I don't have a clear answer, but I *suspect* that "differently-represented-but-ultimately-equivalent feature sets" won't be a big issue. I have an unsubstantiated view that automatically recognizing the same feature set from different authors will not be a big requirement. >Do you anticipate a pattern of usage where this feature would >significantly increase efficiency? Personally, no. Also, even if we did not require the hash to be computed over a fully canonicalized value, a combination of whitespace normalization and presentation conventions might still match a good number of separately authored cases. One thing I do strongly believe: this matter does not impact correct functioning of the system (to the extent that appropriate dereferencing mechanisms exist); rather, it reates to possible efficiency of operation. #g ------------ Graham Klyne ([email protected])