Skip to content

104.7 Find system files and place files in the correct location

Weight: 2

Candidates should be thoroughly familiar with the Filesystem Hierarchy Standard (FHS), including typical file locations and directory classifications.

Objectives

  • Understand the correct locations of files under the FHS.
  • Find files and commands on a Linux system.
  • Know the location and purpose of important files and directories as defined in the FHS.

Terms

find, locate, updatedb, whereis, which, type, /etc/updatedb.conf

FHS

Filesystem Hierarchy Standard (FHS) is a document describing the Linux/Unix file hierarchy. It is very useful to know these because it lets you easily find what you are looking for as a system admin.

directory usage
/ Primary hierarchy root and root directory of the entire file system hierarchy
/bin Essential command binaries
/boot Static files of the boot loader
/dev Device files
/etc Host-specific system configuration
/lib Essential shared libraries and kernel modules
/media Mount point for removable media
/mnt Mount point for mounting a filesystem temporarily
/opt Add-on application software packages
/sbin Essential system binaries
/srv Data for services provided by this system
/tmp Temporary files
/usr Secondary hierarchy
/var Variable data
/home User home directories (optional)
/lib Alternate format essential shared libraries (optional)
/root Home directory for the root user (optional)

/usr is the second level of the hierarchy. It contains shareable, read-only data. It can be shared between systems, although present practice does not often do this.

The /var filesystem contains variable data files, including spool directories and files, administrative and logging data, and transient and temporary files. Some portions of /var are not shareable between different systems, but others, such as /var/mail, /var/cache/man, /var/cache/fonts, and /var/spool/news, may be shared.

Four more directories, plus more detail on some:

  • /run run-time variable data, such as process ID files
  • /proc a virtual filesystem holding data about running processes, which you already met in 101.1
  • /boot holds not just the bootloader files but the kernel itself and the initial RAM disk
  • /srv holds data served by the system, so a web server's pages might live in /srv/www

Compliance with the FHS is not mandatory, but nearly every distribution follows it. The full specification is FHS 3.0, published by the Linux Foundation.

The easiest way to hold the whole layout in your head is by what changes and who owns it:

   never changes while running
   /boot   kernel and bootloader
   /bin /sbin /lib   programs and libraries needed to boot
   /usr    everything else installed by packages (read-only)

   changes constantly
   /var    logs, mail, print queues, caches, databases
   /run    process IDs and runtime state, cleared at boot
   /tmp    scratch files, usually cleared at boot

   yours, or the machine's own
   /etc    this machine's configuration
   /home   users' files
   /root   root's own home
   /opt    software installed outside the package manager
   /srv    data this machine serves to others

   not real files at all
   /dev    device nodes
   /proc   process and kernel data
   /sys    hardware and device data

The /bin versus /sbin split is the one to remember: /sbin holds system administration tools, and /lib exists to hold the shared libraries those two directories need. That is why all three must be on the root filesystem and cannot live on a separate partition, since they are needed before anything else can be mounted.

Then there are temporary files, which the objective expects you to know because the three locations behave differently:

   /tmp       may be wiped at every boot. Never keep
              anything here you care about.

   /var/tmp   also temporary, but survives a reboot.

   /run       runtime state of running processes, such as
              .pid files. Must be cleared at boot.
              On some systems /var/run is a symlink to /run.

Real world use: a long conversion job writing scratch files should use /var/tmp, because if the machine reboots halfway through, files in /tmp may be gone.

Path

A general linux install has a lot of files, 741341 files in my case. So how does the shell find and run a command? This is done by a variable called PATH:

$ echo $PATH
/home/nagato/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/usr/games:/usr/local/games;/home/nagato/bin/

And for the root user:

# echo $PATH
/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin

As you can see, this is a list of directories separated with colons. Obviously you can change your path with export PATH=$PATH:/usr/new/dir, or put this in .bashrc to make it permanent.

Compare the two lists carefully. Root's PATH puts /usr/local/sbin and /usr/sbin first and has no /home/nagato/bin at all. A normal user's PATH leads with their own bin directory. Neither contains ., the current directory, for the security reason covered in 103.1.

Locating files

which, type and whereis

The which command shows the first appearance of the command given in the path. In other words which mkfs will tell you what will be run if you issue the mkfs command.

$ which mkfs
/usr/bin/mkfs
nagato@ubuntuserver:~$ which ping
/usr/bin/ping
nagato@ubuntuserver:~$ which -a ping
/usr/bin/ping
/bin/ping

Use the -a switch to show all appearances in the path and not only the first one.

That -a output is the interesting one. There are two ping binaries on this system. Without -a you only learn about /usr/bin/ping, because it comes first in PATH and is therefore the one that actually runs.

But where is the cd command?

nagato@ubuntuserver:~$ which cd
nagato@ubuntuserver:~$ type cd
cd is a shell builtin

As you can see, which did not find anything for cd, so we tried it with type to see what it is.

$ type type
type is a shell builtin
$ type for
for is a shell keyword
$ type mkfs
mkfs is /sbin/mkfs
$ type chert
bash: type: chert: not found

The type command is more general than which and also understands and shows the bash keywords.

This is the practical difference between the two. which only searches PATH for files, so anything that is not a file on disk is invisible to it. type asks bash itself, so it can also report builtins, keywords, aliases and functions.

   which   searches PATH for a FILE
           finds:  /usr/bin/ping
           misses: cd, for, aliases

   type    asks BASH what this word means
           finds:  everything, and says which kind

Another useful command in this category is whereis. Unlike which, whereis shows man pages and source code of programs alongside their binary location.

$ whereis mkfs
mkfs: /usr/sbin/mkfs /usr/share/man/man8/mkfs.8.gz

Tip: there is also a whatis command, try it.

The type options mirror which:

$ type -a locate
locate is /usr/bin/locate
locate is /bin/locate
$ type -t locate
file
$ type -t ll
alias

-a shows every match, and -t prints just the kind: alias, keyword, function, builtin or file. The -t form is the one to use inside a script, since it gives one plain word instead of a sentence.

And the whereis filters:

$ whereis locate
locate: /usr/bin/locate /usr/share/man/man1/locate.1.gz
$ whereis -b locate
locate: /usr/bin/locate
$ whereis -m locate
locate: /usr/share/man/man1/locate.1.gz

-b limits to binaries, -m to man pages, -s to source code.

Which of the three to reach for:

   which     "what will run if I type this?"      one path
   type      "what kind of thing is this word?"   any kind
   whereis   "where is everything about this?"    binary, man, source

find

We have already seen this command in chapter 103.3, but let's see a couple of new switches.

  • The -type limits by kind, so -type f matches regular files and -type d matches directories.
  • The -user and -group specify a specific user and group
  • The -maxdepth tells find how deep it should go into the directories.
$ find /tmp/ -maxdepth 1 -user nagato | head
/tmp/1
/tmp/2

Or even find the files not belonging to any user or group with -nouser and -nogroup.

Like other tests, you can add a ! just before any phrase to negate it. So this will find files not belonging to nagato: find . ! -user nagato

It is also very common to use it for files with specific strings in their names:

$ sudo find /etc -iname "*vmware*"
/etc/vmware-tools
/etc/vmware-tools/scripts/vmware

Real world use for -nouser: after deleting a user account, every file that user owned is left with a numeric UID that matches nobody. find /home -nouser lists exactly those orphaned files so you can reassign or remove them.

A set of attribute tests pairs well with the ownership ones:

  • -readable, -writable, -executable match files the current user can read, write or run. On a directory, -executable means one you can enter.
  • -perm NNNN matches files with exactly that permission, so -perm 0664 finds rw-rw-r-- and nothing else.
  • -perm -NNNN with a leading minus matches files with at least those permissions. -perm -644 also matches 664 and 775.
  • -empty matches empty files and directories.
  • -size N with suffixes c bytes, k kibibytes, M mebibytes, G gibibytes, and + or - for bigger or smaller.

-maxdepth counting is easy to get wrong. Given this tree:

directory
├── clients.txt
├── partners.txt -> clients.txt
└── somedir
  ├── anotherdir
  └── clients.txt

-maxdepth 1 searches only the current directory. -maxdepth 2 reaches into somedir. -maxdepth 3 reaches into anotherdir. -mindepth N works the other way, skipping anything shallower than N levels.

Two more for controlling where find goes: -mount stops it descending into mounted filesystems, and -fstype limits it to one type, as in find /mnt -fstype exfat -iname "*report*".

The quoting point is worth repeating from 103.3. The pattern must be quoted:

$ find . -name '*.jpg'
./pixel_3a_seethrough_1.jpg
./Mate3.jpg
./Expert.jpg

$ find . -name '*.jpg*'
./pixel_3a_seethrough_1.jpg
./Pentaro.jpg.zip
./Mate3.jpg
./Expert.jpg

The only change was a second * on the end, and it pulled in Pentaro.jpg.zip. The first pattern means "ends in .jpg", the second means "has .jpg somewhere in it". Without the quotes, bash would expand the pattern before find ever saw it.

And the time tests, which also appeared in 103.3 but are worth restating:

  • -amin N, -cmin N, -mmin N for minutes
  • -atime N, -ctime N, -mtime N for 24 hour periods

-ctime and -cmin are the most far-reaching, because any attribute change counts, including a permission change. Almost anything you do to a file will trigger them.

A combined example:

$ find ~ -iname "*report*" -perm 0644 -atime 10 -size +1M
$ find . -mtime -1 -size +100M

The second one reads as: modified less than 24 hours ago, and bigger than 100 MiB.

locate & updatedb

You tried find and loved it, but there is an issue with it: it does a live active search. This can slow down your system, or put too much pressure on your disk, or on larger file systems take too long. To solve this, there is a faster command:

$ locate networking
/etc/cloud/cloud.cfg.d/subiquity-disable-cloudinit-networking.cfg
/snap/core20/1408/usr/lib/python3/dist-packages/cloudinit/distros/networking.py
/snap/core20/1408/usr/lib/python3/dist-packages/cloudinit/distros/__pycache__/networking.cpython-38.pyc
/snap/core20/1826/usr/lib/python3/dist-packages/cloudinit/distros/networking.py
/snap/core20/1826/usr/lib/python3/dist-packages/cloudinit/distros/__pycache__/networking.cpython-38.pyc
/usr/lib/python3/dist-packages/cloudinit/distros/networking.py
/usr/lib/python3/dist-packages/cloudinit/distros/__pycache__/networking.cpython-310.pyc
/usr/lib/python3/dist-packages/sos/report/plugins/networking.py
/usr/lib/python3/dist-packages/sos/report/plugins/__pycache__/networking.cpython-310.pyc

And it is fast:

$ time locate kernel / | wc -l
16989

real    0m0.091s
user    0m0.094s
sys 0m0.034s

Read that timing. It found 16989 matches in about nine hundredths of a second. No live search of a disk could do that, which is the whole point.

This is fast because its data comes from a database created with updatedb (stored at /var/lib/mlocate.db; its package is called plocate on debian). Usually this command runs automatically, with a cronjob, on a daily basis. Its configuration file is /etc/updatedb.conf or /etc/sysconfig/locate:

$ cat /etc/updatedb.conf
PRUNE_BIND_MOUNTS="yes"
# PRUNENAMES=".git .bzr .hg .svn"
PRUNEPATHS="/tmp /var/spool /media /home/.ecryptfs"
PRUNEFS="NFS nfs nfs4 rpc_pipefs afs binfmt_misc proc smbfs autofs iso9660 ncpfs coda devpts ftpfs devfs mfs shfs sysfs cifs lustre tmpfs usbfs udf fuse.glusterfs fuse.sshfs curlftpfs ecryptfs fusesmb devtmpfs"

You can update the db by running updatedb as root.

The trade-off between the two commands:

   find                        locate
   ---------------------------------------------------
   walks the real disk         reads a database
   slow on a big filesystem    near instant
   always current              only as fresh as the
                               last updatedb run
   can test size, time,        matches the path text
   owner, permissions          and little else

That staleness is the catch. A file created an hour ago will not be in the database yet, and a file deleted an hour ago will still be listed. There is an option for the second half of that problem:

$ locate -e jpg

-e makes locate check that each file still exists before printing it. It cannot help with files created since the last update, and for those you have to run updatedb yourself.

The other locate options:

$ locate -i .jpg

-i ignores case, so .JPG files show up too. Several patterns can be given at once, separated by spaces, and they are treated as OR by default:

$ locate -i zip jpg
/home/carol/Downloads/Expert.jpg
/home/carol/Downloads/Mate1_old.JPG
/home/carol/Downloads/OPENMSXPIHAT.zip
/home/carol/Downloads/gbs-control-master.zip

-A changes that to AND, matching only files that contain every pattern:

$ locate -A .jpg .zip
/home/carol/Downloads/Pentaro.jpg.zip

And -c counts instead of listing:

$ locate -c .jpg
1174

One thing that surprises people: locate matches anywhere in the path, not just the extension. Searching for jpg also returns jpg_specs.doc, because the letters appear in the name. It is a text match on the whole path, not a file type filter.

The four settings in /etc/updatedb.conf:

  • PRUNEFS= filesystem types to skip, space separated and case insensitive. Note the list above includes nfs, tmpfs and proc, which are network shares and memory-only filesystems, so there is no point indexing them.
  • PRUNENAMES= directory names to skip anywhere they appear. The commented-out example, .git .bzr .hg .svn, keeps version control internals out of the index.
  • PRUNEPATHS= specific paths to skip, here /tmp /var/spool /media, since those hold throwaway or removable data.
  • PRUNE_BIND_MOUNTS= yes or no. When yes, directories mounted in a second place with mount --bind are indexed only once instead of twice.

Real world use for PRUNEPATHS: adding a large backup or media directory keeps updatedb from spending its nightly run indexing files you will never search for.

Summary

I have a Linux system laid out according to the Filesystem Hierarchy Standard, which is why I can sit down at an unfamiliar distribution and still know where things are. Configuration for this machine is in /etc, users' files are in /home with root's own in /root, device nodes are in /dev, and process and kernel data appear as virtual files in /proc and /sys. Programs live in /bin and /sbin with their libraries in /lib, and those three must be on the root filesystem because they are needed before anything else can be mounted. Everything else installed by the package manager sits under /usr, which is read-only in normal use, while /opt holds software installed outside the package manager and /srv holds data this machine serves to others. Anything that changes while the system runs lives in /var, including logs, mail and print queues.

For temporary files the standard gives me three places with different rules. /tmp may be wiped at every boot, /var/tmp survives a reboot, and /run holds runtime state like process ID files and must be cleared at boot. On some systems /var/run is just a symlink to /run.

To find a command I have three tools that answer three different questions. which searches PATH and tells me the one file that would actually run, with -a showing every match. type asks bash instead, so it also knows about builtins, keywords, aliases and functions, and type -t reports just which of those a word is. whereis casts wider still, returning the binary, the man page and the source, with -b, -m and -s to narrow it down.

To find a file there are two approaches. find walks the real filesystem, which makes it slow but always accurate, and it can test far more than a name: -user and -group for ownership, -nouser and -nogroup for orphaned files, -perm for exact or minimum permissions, -size with unit suffixes, -empty, the time tests -mtime, -atime and -ctime, and -maxdepth and -mindepth to control how deep it goes, with ! in front of any test to negate it. locate takes the opposite approach, reading a database built by updatedb rather than touching the disk, which makes it almost instant but only as current as the last update. -i ignores case, -c counts, -A requires every pattern to match, and -e checks that a file still exists before listing it. What updatedb indexes is controlled by /etc/updatedb.conf, through PRUNEFS, PRUNENAMES, PRUNEPATHS and PRUNE_BIND_MOUNTS.