Skip to content

103.3 Perform basic file management

Weight: 4

Candidates should be able to use the basic Linux commands to manage files and directories.

Objectives

  • Copy, move and remove files and directories individually.
  • Copy multiple files and directories recursively.
  • Remove files and directories recursively.
  • Use simple and advanced wildcard specifications in commands.
  • Using find to locate and act on files based on type, size, or time.
  • Usage of tar, cpio, and dd.

Terms

cp, find, mkdir, mv, ls, rm, rmdir, touch, tar, cpio, dd, file, gzip, gunzip, bzip2, bunzip2, xz, unxz, file globbing

Wildcards and file globbing

File globbing is a shell capability that lets you tell things like:

  • All files
  • everything which starts with A
  • all files with 3-letter names which end in A or B or C

To do so you need to know about these characters:

  • * means any string
  • ? means any single character
  • [ABC] matches A, B, or C
  • [a-k] matches a, b, c, ..., k (both lower-case and upper-case)
  • [0-9a-z] matches all digits and numbers
  • [!x] means NOT X.

Knowing these, you can create your patterns. For example:

command meaning
rm * delete all files in this directory
ls A*B show all files starting with A and ending with B
cp ???.* /tmp Copy all files with 3 characters, then a dot then whatever (even nothing) to /tmp
rmdir [a-z]* remove all empty directories which start with a letter

The important thing to understand is that the shell expands the pattern before the command ever runs. The command never sees the *:

   you type:     rm *.txt
                    |
                    |  bash expands the pattern first
                    v
   rm actually runs as:   rm a.txt b.txt notes.txt

This is why a wildcard works with every command. It is not a feature of rm or ls, it is a feature of bash.

Each wildcard with real listings. The asterisk, matching zero, one or more characters:

$ ls lpic-*.txt

That lists files like lpic-1.txt and lpic-2.txt. It can appear anywhere in the pattern, and more than once:

$ rm *ate*

That removes any filename containing the letters ate anywhere inside it.

The question mark, matching exactly one character:

$ ls
last.txt lest.txt list.txt third.txt past.txt
$ ls l?st.txt
last.txt lest.txt list.txt
$ ls ??st.txt
last.txt lest.txt list.txt past.txt

The first pattern needed an l at the front, so past.txt was left out. The second accepted any two characters, so past.txt matched too.

Bracketed characters, matching one character from a set:

$ ls l[aef]st.txt
last.txt lest.txt
$ ls l[a-z]st.txt
last.txt lest.txt list.txt

Ranges can be combined, which is how you match a shape rather than a name:

$ ls
student-1A.txt student-2A.txt student-3.txt
$ ls student-[0-9][A-Z].txt
student-1A.text student-2A.txt

That pattern means: the text student-, then a digit, then an uppercase letter, then .txt. student-3.txt failed because it has no uppercase letter.

And the wildcards combine freely:

$ ls
last.txt lest.txt list.txt third.txt past.txt
$ ls [plf]?st*
last.txt lest.txt list.txt past.txt

$ ls
file1.txt file.txt file23.txt fom23.txt
$ ls f*[0-9].txt
file1.txt file23.txt fom23.txt

In the last one, file.txt did not match because the pattern requires at least one digit before .txt.

Real world use for [!x]: cleaning a directory but keeping one thing. rm *.log deletes every log, while a negated set lets you keep a group of them.

general commands

listing with ls

ls is used to list directories and files. You can provide an absolute or relative path. If omitted the "." will be used as a target.

nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 16K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato  29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato  24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 116 Aug 14 04:44 note_to_self

First field indicates if this is a file (-) or directory (d).

Some common switches are:

  • -l is for long (more info for each file)
  • -1 will print one file per line
  • -t sorts based on modification date
  • -r reverses the search (so -tr is reverse time, newer files at the bottom)

You can mix switches. A famous one is -ltrh (long + human readable sizes + reverse time).

Two more switches, with the plain form first:

$ ls
Desktop Downloads emp_salary file1 Music Public Videos
Documents emp_name examples.desktop file2 Pictures Templates

Long format, which is where the detail lives:

$ ls -l
total 60
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Desktop
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Documents
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Downloads
-rw-r--r-- 1 frank frank 21 Sep 7 12:59 emp_name
-rw-r--r-- 1 frank frank 20 Sep 7 13:03 emp_salary
-rw-r--r-- 1 frank frank 8980 Apr 1 2018 examples.desktop
-rw-r--r-- 1 frank frank 10 Sep 1 2018 file1
-rw-r--r-- 1 frank frank 10 Sep 1 2018 file2
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Music
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Pictures
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Public
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Templates
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Videos

Three file type characters: - for a regular file, d for a directory, and c for a special file.

Human readable sizes with -h. Compare the size column against the listing above:

$ ls -lh
total 60K
drwxr-xr-x 2 frank frank 4.0K Apr 1 2018 Desktop
-rw-r--r-- 1 frank frank 8.8K Apr 1 2018 examples.desktop
-rw-r--r-- 1 frank frank 10 Sep 1 2018 file1

8980 became 8.8K. That is all -h does, it converts the raw byte count into K, M or G.

Hidden files with -a:

$ ls -a
. .dbus file1 .profile
.. Desktop file2 Public
.bash_history .dmrc .gconf .sudo_as_admin_successful

Files whose names start with a dot are hidden by default. -a shows them, which is how you see .bash_history and the config files in a home directory.

The general syntax is ls OPTIONS FILE, and when FILE is left out, the current directory is used.

Copy (cp), Move (mv), and Delete (rm)

cp

This will copy files from one place or name to another place or name. If the target is a directory, all sources will be copied there.

cp source destination

A common switch is -r (or -R) which copies recursively (directories and their contents). So for copying a directory called A to /tmp/ you can issue cp -r A /tmp/.

Here is what happens without -r, which is the mistake people make first:

$ tree mydir
mydir
|_file1
|_newdir
  |_file2
  |_insidenew
  |_lastdir
3 directories, 2 files
$ mkdir newcopy
$ cp mydir newcopy
cp: omitting directory 'mydir'
$ cp -r mydir newcopy
$ tree newcopy
newcopy
|_mydir
  |_file1
  |_newdir
  |_file2
  |_insidenew
  |_lastdir
4 directories, 2 files

Plain cp refused and said omitting directory. Adding -r copied the whole tree.

Recursion is worth picturing, because the same idea shows up in ls -R, cp -r and rm -r:

   without -r              with -r
   students/               students/
     level1/  <- skipped     level1/
     level2/  <- skipped       (and everything inside it)
     frank                   level2/
                               (and everything inside it)
                             frank

Relative and absolute source paths:

$ cp file1 dir2
$ cp dir1/file1 dir2
$ cp /home/frank/Documents/file2 /home/frank/Documents/Backup

The first two are relative, taken from where you are standing now. The third starts with /, so it works from anywhere. A path starting with / is absolute, everything else is relative.

Copying the contents of a directory rather than the directory itself uses a wildcard:

$ cp -r animal/* forest

Note the difference. cp -r animal forest puts a folder named animal inside forest. cp -r animal/* forest puts the contents of animal directly into forest, with no wrapper folder.

mv

Will move or rename files or directories. It works like the cp command. If you are moving a file on the same file system, the inode will not change.

In general:

  • If the target is an existing directory, then all sources are copied into the target
  • If the target directory does not exist, then the source must be only one directory which will be renamed to the target directory.
  • If the target is a file, then the source must be only one file so rename will happen.

These look like "formulas" but they are common sense.

The inode detail matters. Moving inside the same filesystem only changes the name entry, so it is instant even for a huge file. Moving across filesystems has to copy every byte and then delete the original, so it is slow and the inode does change.

The two shapes:

$ mv myfile.txt /home/frank/Documents
$ mv old_file_name new_file_name

The first moves, the second renames. It is the same command, and which one happens depends only on whether the target is an existing directory.

By default mv will not ask before overwriting. -i makes it ask:

$ mv -i old_file_name new_file_name
mv: overwrite 'new_file_name'?

And -f forces the overwrite with no question at all.

rm

Removes (deletes) files. You can do this recursively using the -r switch, or even prevent it from checking for confirmations using the -f (force) switch. So a rm -rf / means delete everything from the file system.

This is worth repeating. A recursive remove on an important system directory can leave the machine unusable. There is no undo. Use -r only when you are certain about what is inside the directory.

$ rm file1
$ rm -i file1
rm: remove regular file 'file1'?
$ rm -f file1
$ rm file1 file2 file3

Trying to remove a directory without -r fails:

$ rm newcopy/
rm: cannot remove 'newcopy/': Is a directory
$ rm -r newcopy/

The difference between rm -r and rmdir is small but real. rmdir only succeeds if the directory is empty. rm -r works either way, which is why it is both more useful and more dangerous.

Real world use for rm -ri: deleting a directory you are not fully sure about. It asks about each item, so you see what is going before it goes.

notes

Normally, the cp command will copy a file over an existing copy, if the existing file is writable. On the other hand, mv will not move or rename a file if the target exists. This is highly dependent on your system's configuration. But in all cases, you can overcome this using the -f switch.

  • -f (--force) will cause cp to try overwriting the target.
  • -i (--interactive) will ask a Y/N question (deleting / overwriting).
  • -b (--backup) will make backups of overwritten files
  • -p will preserve the attributes.

Real world use for -p: copying a web directory to a new server. Without -p, every file gets today's date and the copying user's ownership. With -p, timestamps, ownership and permissions survive the copy.

Creating (mkdir) and removing (rmdir) directories

The mkdir command creates directories.

nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 16K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato  29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato  24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 116 Aug 14 04:44 note_to_self
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ mkdir new_dir
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 20K
-rw-rw-r-- 1 nagato nagato  207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato  116 Aug 14 04:44 note_to_self
drwxrwxr-x 2 nagato nagato 4.0K Aug 14 04:57 new_dir

If you want to create a tree of directories, you can use the -p switch to tell mkdir to create the parent directories if needed:

nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ mkdir -p 1/2/3
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ tree
.
├── 1
│   └── 2
│       └── 3
├── data.txt
├── info.txt
├── new_dir
├── note_to_self
└── tasks.txt

Without -p, mkdir 1/2/3 fails, because directory 1 does not exist yet and mkdir will not invent it. With -p, it creates each missing level on the way down.

If you need to delete a directory the command is rmdir and you can also use -p for nested removing:

nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ rmdir -p 1/2/3
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ rmdir new_dir
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ tree
.
├── data.txt
├── info.txt
├── note_to_self
└── tasks.txt

0 directories, 4 files

If you are using rmdir to remove a directory, it MUST BE EMPTY. That is why many people use rm -rf directory_name to delete a non-empty directory and whatever is in it.

Real world use for mkdir -p in scripts: it does not complain if the directory already exists. So mkdir -p /var/backups/daily is safe to run every night, whether or not the folder is there.

touch

touch will create an empty file if it does not exist, or update the modification date of a file if it already exists. The default time is now, but you can specify other times too.

nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 16K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato  29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato  24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 116 Aug 14 04:44 note_to_self
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ touch new_file
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 16K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato  29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato  24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 116 Aug 14 04:44 note_to_self
-rw-rw-r-- 1 nagato nagato 0 Aug 14 05:08 new_file
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ touch note_to_self
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 16K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato  29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato  24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 0 Aug 14 05:08 new_file
-rw-rw-r-- 1 nagato nagato 116 Aug 14 05:08 note_to_self

Look at the last listing carefully. note_to_self kept its size of 116 bytes, but its time changed from 04:44 to 05:08, and it moved to the bottom because the listing is sorted by time. Touching an existing file does not empty it.

Or you can specify times. It is possible to use -d and give dates, or use -t and give a timestamp in the form [[CC]YY]MMDDhhmm[.ss]

$ touch -t 200908121510.59 file1
$ touch -d 11am file2
$ touch -d "last fortnight" file3
$ touch -d "yesterday 6am" file4
$ touch -d "2 days ago 12:00" file5
$ touch -d "tomorrow 02:00" file6
$ touch -d "5 Nov" file3
$ ls -ltrh file?
-rw-rw-r-- 1 nagato nagato 0 Aug 12  2009 file1
-rw-rw-r-- 1 nagato nagato 0 Aug 12 12:00 file5
-rw-rw-r-- 1 nagato nagato 0 Aug 13 06:00 file4
-rw-rw-r-- 1 nagato nagato 0 Aug 14  2022 file2
-rw-rw-r-- 1 nagato nagato 0 Aug 15  2022 file6
-rw-rw-r-- 1 nagato nagato 0 Nov  5  2022 file3

Note that -d accepts plain English like yesterday 6am and 2 days ago 12:00, while -t wants the strict digit format. 200908121510.59 reads as year 2009, month 08, day 12, hour 15, minute 10, second 59.

It is also possible to use another file's time, with the -r switch (for --reference):

nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -l /etc/debian_version
-rw-r--r-- 1 root root 13 Aug 22  2021 /etc/debian_version
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ touch -r /etc/debian_version file1
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 20K
-rw-rw-r-- 1 nagato nagato 0 Aug 22  2021 file1

file1 took the exact date of /etc/debian_version, which is Aug 22 2021.

Two more options. -a changes only the access time, -m changes only the modification time, and using both together changes both:

$ touch -am file3

It also notes that touch can create several files at once:

$ touch file1 file2 file3

file

To determine the type of a file, you should use the file command. It looks into the file and determines its type.

nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ file file1
file1: empty
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ file note_to_self
note_to_self: ASCII text
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ file /bin/bash
/bin/bash: ELF 64-bit LSB pie executable, x86-64, version 1 (SYSV), dynamically linked, interpreter /lib64/ld-linux-x86-64.so.2, BuildID[sha1]=33a5554034feb2af38e8c75872058883b2988bc5, for GNU/Linux 3.2.0, stripped

The -i switch prints the mime format.

Breaking down that third line, since it is dense:

  • ELF 64-bit LSB pie executable = a Linux executable, 64 bit.
  • x86-64 = the CPU architecture it was built for.
  • dynamically linked, interpreter /lib64/ld-linux-x86-64.so.2 = it needs shared libraries at run time, loaded by that program. This ties back to objective 102.3.
  • stripped = the debug symbols were removed to make the file smaller.

The key point is that file reads the contents, not the name. Renaming a JPEG to notes.txt does not fool it.

dd

The dd command copies data from its input to its output (say files or devices). You may use it just like copy:

nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ dd if=note_to_self of=new_file
0+1 records in
0+1 records out
116 bytes copied, 0.00141561 s, 81.9 kB/s
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ cat new_file
I will continue learning... and if I get confused, I'll repeat the last section once more till everything is clear!
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$
  • if is the Input File
  • of is the Output File

Note the unusual syntax. dd uses option=value rather than the normal -o value or --option=value style. It is the odd one out among Unix commands.

But commonly people use it to read and write from block devices. For example, this will read all the sectors from /dev/sdb and write them to a file named backup.dd. Later you can restore this backup by swapping the if and of and writing from backup.dd to /dev/sdb.

# dd if=/dev/sda of=backup.dd bs=4096

or even:

# dd if=/dev/sda2 | gzip > backup.dd.gzip

Another common usage is creating files of specific sizes:

$ dd if=/dev/zero of=1g.bin bs=1G count=1

or even writing your iso files to a USB disk to have a live bootable USB:

$ sudo dd if=ubuntu.iso of=/dev/sdc bs=2048

Caution: here you are writing directly on a block device. If you do something wrong you will ruin your disk and need to reformat it.

Two useful options. status=progress makes dd report as it works, since by default it prints nothing until it finishes:

$ dd status=progress if=oldfile of=newfile

And conv=ucase converts the text to upper case while copying:

$ dd if=oldfile of=newfile conv=ucase

Real world use for bs=: this sets the block size, how much data is moved per read and write. A larger block size like bs=4M makes writing an ISO to USB much faster than the tiny default, because it means far fewer trips to the device.

find

The find command helps us to find files based on different criteria. Look at this:

$ find . -iname "[a-j]*"
./howcool.sort
./alldata
./mydir/howcool.sort
./mydir/newDir/insideNew
./howcool
  • The first parameter says where we should search (including subdirectories).
  • The -name switch indicates the criteria. Here iname means search files with this name and ignore the character cases (z equals Z).

Another common switch is -type to indicate the type we are searching for (f for regular files, d for directories, and l for symbolic links):

nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ find . -type d -iname "[a-j]*"
./directory
./directory/innder_one

The general shape is find STARTING_PATH OPTIONS EXPRESSION. The basic case:

$ find . -name "myfile.txt"
./myfile.txt
$ find /home/frank -name "*.png"
/home/frank/Pictures/logo.png
/home/frank/screenshot.png

Note the quotation marks around "*.png". They matter. Without them bash expands the * before find ever sees it, and you get the wrong results or an error. This connects directly to the quoting section in 103.1.

Two more criteria: -not returns results that do not match, and -maxdepth N limits how many levels deep the search goes.

If you want to search for file sizes do as below:

command meaning
-size 100c files which are exactly 100 characters/bytes (you can also use b)
-size +100k files which are more than 100 kilobytes
-size -20M files smaller than 20Megabytes
-size +2G files bigger than 2Gigabytes

So this will find all files ending in tmp with sizes between 1M and 100M in the /var/ directory:

find /var -iname '*tmp' -size +1M -size -100M

You can find all empty files with find . -size 0b or find . -empty.

Note how the range was built: two -size tests in the same command. find joins conditions with AND by default, so a file has to satisfy both.

The same idea, hunting for large files:

$ sudo find /var -size +2G
/var/lib/libvirt/images/debian10.qcow2
/var/lib/libvirt/images/rhel8.qcow2

Another useful search criterion is time. These are some of the options:

switch meaning samples
-amin Access Minutes -amin 40 means "files accessed exactly 40min ago", -amin +40 means files accessed more than 40min ago, and -amin -40 means files accessed less than 40min ago
-cmin Status Change Min -cmin +60 file status changed before last hour
-mmin Modified Minutes -mmin -60 will give us files modified in last hour
-atime access time in days -atime +1 means files accessed more than 1 day ago (which means 2 days and more)
-ctime Status Changed in Days
-mtime Modified days
-newer Newer than reference -newer file1 will give you files which are newer than file1

If you add the -daystart switch to -mtime or -atime it means that we want to consider days as calendar days, starting at midnight.

The + and - signs are the part people get wrong, so here they are as a picture, using -mmin as the example:

   now
    |
    |<---- 60 min ---->|
    |                  |
    +------------------+------------------------>  into the past
       -mmin -60            -mmin +60
    (changed inside      (changed before
     the last hour)       the last hour)

           -mmin 60 = changed exactly 60 minutes ago

A real example, searching the whole system for config files touched in the last week:

$ sudo find / -name "*.conf" -mtime 7
/etc/logrotate.conf

Acting on files

We can execute commands or do other actions on files with various switches:

switch meaning
-ls will run ls -dils on each file
-print will print the full name of the files on each line

But the best way to run commands on found files is the -exec switch. You can point to the file with '{}' or {} and finish your command with \;.

For example, this will remove all empty files in this directory and its subdirectories:

find . -empty -exec rm '{}' \;

or this will rename all htm files to html:

find . -name "*.htm" -exec mv '{}' '{}l' \;

Since deleting found files is a common task, there is a switch for it: -delete.

The -exec syntax has three parts that all matter:

   find . -name "*.conf" -exec chmod 644 '{}' \;
                                          |    |
                                          |    +-- the ; ends the command.
                                          |        It is escaped as \; so
                                          |        bash does not eat it.
                                          |
                                          +-- {} is a placeholder. find
                                              substitutes each result here.
                                              Quoted, so filenames with
                                              spaces or special characters
                                              are safe.

Another example combines find with grep to search by content instead of by name:

$ find . -type f -exec grep "lpi" '{}' \; -print
./.bash_history
Alpine/M
helping/M

And the shorthand for deleting:

$ find . -name "*.bak" -delete

Real world use: clearing out old backups automatically. find /var/backups -name "*.backup" -mtime +30 -delete removes every backup file older than thirty days. This is the sort of line that goes in a cron job.

Compression

gzip & gunzip

Straight forward, one gzips a file and one ungzips a file. In place:

nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 20K
-rw-rw-r-- 1 nagato nagato  207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato  0 Aug 14 05:08 new_file
-rw-rw-r-- 1 nagato nagato  116 Aug 14 05:08 note_to_self
drwxrwxr-x 3 nagato nagato 4.0K Aug 14 05:20 directory
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 20K
-rw-rw-r-- 1 nagato nagato  171 Aug 14 04:43 tasks.txt.gz
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato  0 Aug 14 05:08 new_file
-rw-rw-r-- 1 nagato nagato  116 Aug 14 05:08 note_to_self
drwxrwxr-x 3 nagato nagato 4.0K Aug 14 05:20 directory
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ gunzip tasks.txt.gz
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 20K
-rw-rw-r-- 1 nagato nagato  207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato  0 Aug 14 05:08 new_file
-rw-rw-r-- 1 nagato nagato  116 Aug 14 05:08 note_to_self
drwxrwxr-x 3 nagato nagato 4.0K Aug 14 05:20 directory
  • gzip preserves time
  • gzip creates the new compressed file with the same name but with a .gz ending
  • gzip removes the original files after creating the compressed file (you can keep the input file with the -k switch)

Watch the timestamps across those three listings. tasks.txt went in at 04:43 and came back out at 04:43, and the size went 207, then 171, then 207 again. That is the "preserves time" point, shown rather than stated.

bzip2 & bunzip2

bzip2 is another compressing tool. It works just like the famous gzip but with a different compression algorithm.

nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ bzip2 tasks.txt
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls  -ltrh
total 20K
-rw-rw-r-- 1 nagato nagato  172 Aug 14 04:43 tasks.txt.bz2
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato  0 Aug 14 05:08 new_file
-rw-rw-r-- 1 nagato nagato  116 Aug 14 05:08 note_to_self
drwxrwxr-x 3 nagato nagato 4.0K Aug 14 05:20 directory
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ bunzip2 tasks.txt.bz2
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls
data.txt  directory  info.txt  new_file  note_to_self  tasks.txt

Choosing between them: gzip is faster but compresses a little less, bzip2 is slower but compresses a little more. For most work either is fine.

xz & unxz

Another compression and decompression tool, just like gzip and bzip2.

nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ xz tasks.txt
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 24K
-rw-rw-r-- 1 nagato nagato  224 Aug 14 04:43 tasks.txt.xz
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato  116 Aug 14 05:08 note_to_self
drwxrwxr-x 3 nagato nagato 4.0K Aug 14 05:20 directory
-rw-rw-r-- 1 nagato nagato  116 Aug 14 07:51 new_file
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ unxz tasks.txt.xz
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 24K
-rw-rw-r-- 1 nagato nagato  207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato  116 Aug 14 05:08 note_to_self
drwxrwxr-x 3 nagato nagato 4.0K Aug 14 05:20 directory
-rw-rw-r-- 1 nagato nagato  116 Aug 14 07:51 new_file

Please note that compressing a small text file makes it larger. This is normal in small files because of all the headers and metadata.

The numbers show it. tasks.txt is 207 bytes. Compressed it becomes 224 bytes with xz, which is bigger than the original. Compare that to 171 with gzip and 172 with bzip2 on the same file. Every format adds a header, and on a tiny file that header can cost more than the compression saves.

In some cases, commands like unxz are just calls to xz --decompress.

Archiving with tar & cpio

Sometimes we need to create an archive file container of many other files. This operation is different from compressing. It combines files into one and then extracts them again. Archiving is mostly used in backups, moving files to a new location (say via email), and such. This is done with cpio and tar.

The difference between the two ideas, since they get confused:

   archiving  =  many files  -->  one file      (same total size)
                 tar, cpio

   compressing =  one file  -->  one smaller file
                  gzip, bzip2, xz

   tar -czf  does both, in that order

tar

TapeARchive, or tar, is the most common archiving tool. It automatically creates an archive file from a directory and all its subdirectories.

Common switches are:

switch meaning
-cf myarchive.tar create file named myarchive.tar
-xf myarchive.tar extract a file named myarchive.tar
-z compress the archive with gzip after creating it
-j compress the archive with bzip2 after creating it
-v verbose! print a lot of data about what is happening
-r append new files to the currently available archive

If you issue absolute paths, tar removes the starting slash (/) for safety reasons when creating an archive. If you want to override, use the -p option.

tar can work with tapes and other storage. That is why we use -f to tell it that we are working with files.

The three main operations running, with output:

$ tar -cvf archive.tar stuff
stuff/
stuff/service.conf
$ tar -xvf archive.tar
stuff/
stuff/service.conf

Extracting somewhere other than here, with -C:

$ tar -xvf archive.tar -C /tmp
$ ls /tmp
stuff

Archiving several directories at once, by listing them:

$ tar -cvf archive.tar stuff1 stuff2

And the compressed forms, which are the ones you will actually type most days:

$ tar -czvf name-of-archive.tar.gz stuff
$ tar -cjvf name-of-archive.tar.bz stuff
$ tar -xzvf archive.tar.gz

Only one operation switch is allowed at a time. -c creates, -x extracts, -t lists what is inside without unpacking it.

Real world use for -t: before extracting an archive somebody sent you, run tar -tvf archive.tar to see what is in it and where it will land. This is how you avoid an archive that unpacks fifty loose files into your current directory.

cpio

Gets a list of files and creates an archive (one file). This file can be used later to extract the original files.

$ ls | cpio -o > allfilesls.cpio
3090354 blocks
  • -o tells cpio to create an output from its input

Please note that cpio does not look into the folders. So mostly we use it with find:

find . -name "*" | cpio -o > myarchivefind.cpio

To extract the original files:

mkdir extract
mv myarchivefind.cpio extract
cd extract
cpio -id < myarchivefind.cpio
  • -d will create the folders
  • -i is for extract

The name means "copy in, copy out", which is exactly the -i and -o pair.

The big difference from tar is where the file list comes from:

   tar:   you name the directory,  tar walks it itself
          tar -cvf out.tar stuff/

   cpio:  something else makes the list, cpio just reads it
          find . -name "*" | cpio -o > out.cpio
                    |
                    +--> this is why cpio needs find or ls

That is also why cpio needs -d when extracting. It is only ever handed a flat list of names, so it has to be told to rebuild the directories.

Summary

I have a Linux system where nearly everything is a file, so these commands are the ones I use constantly. I look around with ls, and -l gives me the long listing where the first character tells me the type, - for a regular file, d for a directory, c for a special file. Adding -h turns raw byte counts into readable sizes, and -a reveals the hidden dot files. I create empty files or update timestamps with touch, and when I want to know what a file really is, file looks inside it rather than trusting the name.

For moving things around I use cp to copy, mv to move or rename, and rm to delete. All three need -r when a directory is involved, because without it they refuse or fail. The -i option makes them ask first, -f makes them stop asking, and -p on cp keeps timestamps and ownership intact. Directories come from mkdir, with -p creating any missing parents, and go away with rmdir if they are empty or rm -r if they are not. rm -rf deserves real care, since there is nothing to undo it.

Wildcards are handled by the shell, not by the commands, which is why *, ? and [abc] work everywhere. The shell expands the pattern first and the command only ever sees the finished list of names. That is also why I have to quote a pattern when I pass it to find, so find receives the pattern itself instead of whatever bash already matched. With find I search by name, by -type, by -size and by time, where + means older or larger and - means newer or smaller. Then -exec runs a command on each result, using {} as the placeholder and \; to close it, or -delete when removing is all I want.

Compression and archiving are two separate jobs that I often do together. gzip, bzip2 and xz each shrink a single file and each replace the original unless I use -k, and on a tiny file they can make it bigger because of the header. tar bundles a whole tree into one archive, and adding -z or -j compresses that bundle in the same command. cpio does the same bundling but reads its file list from find or ls on standard input, which is the real difference between the two. For raw copying at the block level there is dd, with its unusual if= and of= syntax, which is what I reach for when writing an ISO to a USB stick or imaging a whole disk.