Re: (slightly OT?) Disk usage of find

Karl Vogel <[email protected]>
Newsgroups gmane.linux.debian.user
Message-ID <[email protected]>
>> On Sat 16 May 2026 at 14:07:30 (-0400), [email protected] wrote:

> I tend to avoid the use of find (1) because I never really
> "internalized" the syntax) and (2) on the assumption that it hits the
> disk "harder" than locate.

  It does.  "locate" reads a list of files that's been compressed using
  repeating strings on each line, so it hits the CPU a little harder but
  it has to read the entire file.

  At the very least, "find" has to read directories recursively to get
  filenames, and if you use flags like "-mtime", it has to stat() each file
  as it goes.  GNU find tries to optimize this by avoiding repeated stat()
  calls, but they add up especially on spinning rust.

  I tried finding files modified in the last hour on a spinning-rust
  filesystem with several million files, and it took over an hour to
  complete.  ZFS is *much* better at this -- I had created a snapshot of
  that filesystem at midnight, and listing the adds/deletes/etc between
  the snapshot and the current filesystem took less than 15 seconds.

  I found that SQlite competes very favorably with "locate", even for
  filesystems containing over 20 million files, but that's a longer posting.

> I mean, updatedb runs once a day (on my systems) and I assume (I know)
> that its database is mostly cached somewhere, but I don't really know
> how it finds new files.

  It doesn't until you rerun updatedb.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.