Re: Pull request to add parallel scanning to clamscan

Michal Marek <[email protected]>
Newsgroups gmane.comp.security.virus.clamav.devel
Message-ID <[email protected]>
On 2017-06-20 11:33, Mark Allan wrote:
> From the commit message you said "build a list of files first and
> then spawn N children to scan the files in parallel."
> 
> Does this actually iterate *all* the files and directories before
> starting the first scan?

Yes.


> If you're scanning a large directory tree,
> how much overhead does this add prior to scanning the first file?

Unless you are scanning a really slow NFS or CIFS mount, it's negligible
compared to the time it takes to process the content. Initializing the
database does take noticeable time on startup.


> Alternatively, as clamscan already iterates through directories, does
> it maintain a count of the number of concurrent calls to 'scanfile()'
> and fire off another one at that point as necessary?

That would of course be an option, but it would require incrementally
passing paths to the children / threads. With the current approach, I
only need a pipe, which is the simplest synchronization primitive one
can think of :).

Michal
_______________________________________________
http://lurker.clamav.net/list/clamav-devel.html
Please submit your patches to our Bugzilla: http://bugs.clamav.net

http://www.clamav.net/contact.html#ml
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.