Re: Binary comparison of files
Greg Young <[email protected]>
| Newsgroups | gmane.comp.windows.devel.dotnet.advanced |
|---|---|
| Message-ID | <[email protected]> |
It depends on what the expected variation is ... if it were 1 random byte in the file then you would both be O(n) but you have twice as large of a C On Wed, Jan 28, 2009 at 1:26 PM, Geoff Taylor <[email protected]> wrote: >> If all you need to know is that the files are different; I would use a >> hash. CRC32 is fairly fast. For more reliability you can use a >> cryptographic hash. >> >> If you want to know what the differences are; the performance would >> depend on how you plan on processing the differences. > > Surely a hash will require processing all the file before returning its > result? So if the first byte is different, it's still going to read in all > of both files and calculate both hashes. > > Why not just go for a simpler approach using buffered IO. You read two > chunks (one chunk from each file) into byte buffers then do a byte-by-byte > comparison, breaking the loop at the first byte that doesn't match. > > Since they're small files, you could probably just do a ReadAllBytes instead > of reading into buffers. That would give you the length of both, so you > could check that first, giving you another shortcut to a fast negative. > > Here's a similar method that uses assertions instead of failure conditions > that should show you what I mean: > > public static void Compare (string referenceFilename, string > testFilename) > { > byte[] referenceFile = File.ReadAllBytes (referenceFilename); > byte[] testFile = File.ReadAllBytes (testFilename); > Assert.AreEqual (testFile.Length, referenceFile.Length, "Files > are of different lengths. Reference file is {0} bytes, test file is {1} > bytes.", referenceFile.Length, testFile.Length); > for (int counter = 0; counter < referenceFile.Length; counter++) > { > if (referenceFile [counter] != testFile [counter]) > { > Assert.Fail ("Files do not match (at position " + > counter + " - [before '" + indicatorString + "'])."); > } > } > > return; > } > > I'm pretty sure this'll be faster than using a hash (although I'd love to > see a comparison of timings). > > You might be able to make the method faster by looking at how the data is > read in. > > Good luck, > > Geoff > > =================================== > View archives and manage your subscription(s) at http://peach.ease.lsoft.com/archives > -- It is the mark of an educated mind to be able to entertain a thought without accepting it. =================================== View archives and manage your subscription(s) at http://peach.ease.lsoft.com/archives