Re: Binary comparison of files

Geoff Taylor <[email protected]>
Newsgroups gmane.comp.windows.devel.dotnet.advanced
Message-ID <[email protected]>
I think what he means is that my original code will be more  
'efficient' if the common case is that the files are the same length,  
since there's only one file-accessing system call per file. Your code  
would have two file-accessing system calls in that situation, one for  
getting the file length and one for reading the content.

Of course if the common case is that the files have different sizes,  
then the opposite applies and you code is more 'efficient'.

Me, I'm not going to make the assumption that the .NET calls map that  
cleanly on to filesystem IO operations, or how fast those IO  
operations will be in real-world use.

It'd be cool to see some figures for performance. But not tonight...

     Geoff

On 28 Jan 2009, at 20:46, Kim Major <[email protected]> wrote:

> "... I believe this is one aspect you thought your solution may be  
> more
> efficient..."
>
> Maybe I don't understand what you mean, but even without testing:
>
>    byte[] referenceFile = File.ReadAllBytes (referenceFilename);
>    byte[] testFile = File.ReadAllBytes (testFilename);
>    Assert.AreEqual (testFile.Length, referenceFile.Length, "Files  
> are of
> different lengths.  Reference file is {0} bytes, test file is {1}  
> bytes.",
> referenceFile.Length, testFile.Length);
>
> Must be much slower than:
>
>    FileInfo src = new FileInfo(referenceFilename);
>    FileInfo dst = new FileInfo(testFilename);
>    Assert.AreEqual (dst.Length, src.Length, "Files are of different
> lengths.  Reference file is {0} bytes, test file is {1} bytes.",  
> src.Length,
> dst.Length);
>
> The second approach only reads the file info to determine the size.
>
>
> Kim Major
> Renaissance Computer Systems Ltd.
> Blog: http://blogs.microsoft.co.il/blogs/kim
> http://www.renaissance.co.il
>
>
> -----Original Message-----
> From: Discussion of advanced .NET topics.
> [mailto:[email protected]] On Behalf Of Eddie Lascu
> Sent: Wednesday, January 28, 2009 10:14 PM
> To: [email protected]
> Subject: Re: [ADVANCED-DOTNET] Binary comparison of files
>
>
> Kim,
>
> I am sure you have noticed that Geoff's code also compares the  
> lengths of
> the two files and only if they match proceeds to compare them byte  
> by byte.
> It's true, it does it after reading the content of the two files and I
> believe this is one aspect you thought your solution may be more  
> efficient.
>
> Regards,
> Eddie
>
>
>
> -----Original Message-----
> From: Discussion of advanced .NET topics.
> [mailto:[email protected]] On Behalf Of Kim Major
> Sent: Wednesday, January 28, 2009 3:03 PM
> To: [email protected]
> Subject: Re: [ADVANCED-DOTNET] Binary comparison of files
>
> The code below also reads the whole file " byte[] referenceFile =
> File.ReadAllBytes (referenceFilename);"
>
> You would probably be better off by comparing 1) the size and 2)  
> then two
> file streams byte by byte.
> Like this:
> http://jopinblog.wordpress.com/2008/05/06/compare-files-method-in-c-unit-tes
> ting/
>
>
> Kim Major
> Renaissance Computer Systems Ltd.
> Blog: http://blogs.microsoft.co.il/blogs/kim
> http://www.renaissance.co.il
>
>
> -----Original Message-----
> From: Discussion of advanced .NET topics.
> [mailto:[email protected]] On Behalf Of Eddie Lascu
> Sent: Wednesday, January 28, 2009 9:53 PM
> To: [email protected]
> Subject: Re: [ADVANCED-DOTNET] Binary comparison of files
>
> Geoff,
>
> Are you working on the same project as me? It's like you are doing  
> exactly
> what I need. I agree 100% with your comments. Hashing is not the best
> approach because it will require the processing of the whole file.
>
> Thanks bunches,
> Eddie
>
>
>
> -----Original Message-----
> From: Discussion of advanced .NET topics.
> [mailto:[email protected]] On Behalf Of Geoff  
> Taylor
> Sent: Wednesday, January 28, 2009 1:27 PM
> To: [email protected]
> Subject: Re: [ADVANCED-DOTNET] Binary comparison of files
>
>> If all you need to know is that the files are different; I would  
>> use a
>> hash.  CRC32 is fairly fast.  For more reliability you can use a
>> cryptographic hash.
>>
>> If you want to know what the differences are; the performance would
>> depend on how you plan on processing the differences.
>
> Surely a hash will require processing all the file before returning  
> its
> result?  So if the first byte is different, it's still going to read  
> in all
> of both files and calculate both hashes.
>
> Why not just go for a simpler approach using buffered IO.  You read  
> two
> chunks (one chunk from each file) into byte buffers then do a byte- 
> by-byte
> comparison, breaking the loop at the first byte that doesn't match.
>
> Since they're small files, you could probably just do a ReadAllBytes  
> instead
> of reading into buffers.  That would give you the length of both, so  
> you
> could check that first, giving you another shortcut to a fast  
> negative.
>
> Here's a similar method that uses assertions instead of failure  
> conditions
> that should show you what I mean:
>
>        public static void Compare (string referenceFilename, string
> testFilename)
>        {
>            byte[] referenceFile = File.ReadAllBytes  
> (referenceFilename);
>            byte[] testFile = File.ReadAllBytes (testFilename);
>            Assert.AreEqual (testFile.Length, referenceFile.Length,  
> "Files
> are of different lengths.  Reference file is {0} bytes, test file is  
> {1}
> bytes.", referenceFile.Length, testFile.Length);
>            for (int counter = 0; counter < referenceFile.Length;  
> counter++)
>            {
>                if (referenceFile [counter] != testFile [counter])
>                {
>                    Assert.Fail ("Files do not match (at position " +
> counter + " - [before '" + indicatorString + "']).");
>                }
>            }
>
>            return;
>        }
>
> I'm pretty sure this'll be faster than using a hash (although I'd  
> love to
> see a comparison of timings).
>
> You might be able to make the method faster by looking at how the  
> data is
> read in.
>
> Good luck,
>
>            Geoff
>
> ===================================
> View archives and manage your subscription(s) at
> http://peach.ease.lsoft.com/archives
>
> ===================================
> View archives and manage your subscription(s) at
> http://peach.ease.lsoft.com/archives
>
> ===================================
> View archives and manage your subscription(s) at
> http://peach.ease.lsoft.com/archives
>
> ===================================
> View archives and manage your subscription(s) at
> http://peach.ease.lsoft.com/archives
>
> ===================================
> View archives and manage your subscription(s) at http://peach.ease.lsoft.com/archives

===================================
View archives and manage your subscription(s) at http://peach.ease.lsoft.com/archives
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.