103.3 Perform basic file management¶
Weight: 4
Candidates should be able to use the basic Linux commands to manage files and directories.
Objectives
- Copy, move and remove files and directories individually.
- Copy multiple files and directories recursively.
- Remove files and directories recursively.
- Use simple and advanced wildcard specifications in commands.
- Using find to locate and act on files based on type, size, or time.
- Usage of tar, cpio, and dd.
Terms
cp, find, mkdir, mv, ls, rm, rmdir, touch, tar, cpio, dd, file, gzip, gunzip, bzip2, bunzip2, xz, unxz, file globbing
Wildcards and file globbing¶
File globbing is a shell capability that lets you tell things like:
- All files
- everything which starts with A
- all files with 3-letter names which end in A or B or C
To do so you need to know about these characters:
*means any string?means any single character[ABC]matches A, B, or C[a-k]matches a, b, c, ..., k (both lower-case and upper-case)[0-9a-z]matches all digits and numbers[!x]means NOT X.
Knowing these, you can create your patterns. For example:
| command | meaning |
|---|---|
| rm * | delete all files in this directory |
| ls A*B | show all files starting with A and ending with B |
| cp ???.* /tmp | Copy all files with 3 characters, then a dot then whatever (even nothing) to /tmp |
| rmdir [a-z]* | remove all empty directories which start with a letter |
The important thing to understand is that the shell expands the pattern before the command ever runs. The command never sees the *:
you type: rm *.txt
|
| bash expands the pattern first
v
rm actually runs as: rm a.txt b.txt notes.txt
This is why a wildcard works with every command. It is not a feature of rm or ls, it is a feature of bash.
Each wildcard with real listings. The asterisk, matching zero, one or more characters:
That lists files like lpic-1.txt and lpic-2.txt. It can appear anywhere in the pattern, and more than once:
That removes any filename containing the letters ate anywhere inside it.
The question mark, matching exactly one character:
$ ls
last.txt lest.txt list.txt third.txt past.txt
$ ls l?st.txt
last.txt lest.txt list.txt
$ ls ??st.txt
last.txt lest.txt list.txt past.txt
The first pattern needed an l at the front, so past.txt was left out. The second accepted any two characters, so past.txt matched too.
Bracketed characters, matching one character from a set:
Ranges can be combined, which is how you match a shape rather than a name:
$ ls
student-1A.txt student-2A.txt student-3.txt
$ ls student-[0-9][A-Z].txt
student-1A.text student-2A.txt
That pattern means: the text student-, then a digit, then an uppercase letter, then .txt. student-3.txt failed because it has no uppercase letter.
And the wildcards combine freely:
$ ls
last.txt lest.txt list.txt third.txt past.txt
$ ls [plf]?st*
last.txt lest.txt list.txt past.txt
$ ls
file1.txt file.txt file23.txt fom23.txt
$ ls f*[0-9].txt
file1.txt file23.txt fom23.txt
In the last one, file.txt did not match because the pattern requires at least one digit before .txt.
Real world use for [!x]: cleaning a directory but keeping one thing. rm *.log deletes every log, while a negated set lets you keep a group of them.
general commands¶
listing with ls¶
ls is used to list directories and files. You can provide an absolute or relative path. If omitted the "." will be used as a target.
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 16K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 116 Aug 14 04:44 note_to_self
First field indicates if this is a file (-) or directory (d).
Some common switches are:
-lis for long (more info for each file)-1will print one file per line-tsorts based on modification date-rreverses the search (so-tris reverse time, newer files at the bottom)
You can mix switches. A famous one is -ltrh (long + human readable sizes + reverse time).
Two more switches, with the plain form first:
$ ls
Desktop Downloads emp_salary file1 Music Public Videos
Documents emp_name examples.desktop file2 Pictures Templates
Long format, which is where the detail lives:
$ ls -l
total 60
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Desktop
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Documents
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Downloads
-rw-r--r-- 1 frank frank 21 Sep 7 12:59 emp_name
-rw-r--r-- 1 frank frank 20 Sep 7 13:03 emp_salary
-rw-r--r-- 1 frank frank 8980 Apr 1 2018 examples.desktop
-rw-r--r-- 1 frank frank 10 Sep 1 2018 file1
-rw-r--r-- 1 frank frank 10 Sep 1 2018 file2
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Music
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Pictures
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Public
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Templates
drwxr-xr-x 2 frank frank 4096 Apr 1 2018 Videos
Three file type characters: - for a regular file, d for a directory, and c for a special file.
Human readable sizes with -h. Compare the size column against the listing above:
$ ls -lh
total 60K
drwxr-xr-x 2 frank frank 4.0K Apr 1 2018 Desktop
-rw-r--r-- 1 frank frank 8.8K Apr 1 2018 examples.desktop
-rw-r--r-- 1 frank frank 10 Sep 1 2018 file1
8980 became 8.8K. That is all -h does, it converts the raw byte count into K, M or G.
Hidden files with -a:
$ ls -a
. .dbus file1 .profile
.. Desktop file2 Public
.bash_history .dmrc .gconf .sudo_as_admin_successful
Files whose names start with a dot are hidden by default. -a shows them, which is how you see .bash_history and the config files in a home directory.
The general syntax is ls OPTIONS FILE, and when FILE is left out, the current directory is used.
Copy (cp), Move (mv), and Delete (rm)¶
cp¶
This will copy files from one place or name to another place or name. If the target is a directory, all sources will be copied there.
A common switch is -r (or -R) which copies recursively (directories and their contents). So for copying a directory called A to /tmp/ you can issue cp -r A /tmp/.
Here is what happens without -r, which is the mistake people make first:
$ tree mydir
mydir
|_file1
|_newdir
|_file2
|_insidenew
|_lastdir
3 directories, 2 files
$ mkdir newcopy
$ cp mydir newcopy
cp: omitting directory 'mydir'
$ cp -r mydir newcopy
$ tree newcopy
newcopy
|_mydir
|_file1
|_newdir
|_file2
|_insidenew
|_lastdir
4 directories, 2 files
Plain cp refused and said omitting directory. Adding -r copied the whole tree.
Recursion is worth picturing, because the same idea shows up in ls -R, cp -r and rm -r:
without -r with -r
students/ students/
level1/ <- skipped level1/
level2/ <- skipped (and everything inside it)
frank level2/
(and everything inside it)
frank
Relative and absolute source paths:
The first two are relative, taken from where you are standing now. The third starts with /, so it works from anywhere. A path starting with / is absolute, everything else is relative.
Copying the contents of a directory rather than the directory itself uses a wildcard:
Note the difference. cp -r animal forest puts a folder named animal inside forest. cp -r animal/* forest puts the contents of animal directly into forest, with no wrapper folder.
mv¶
Will move or rename files or directories. It works like the cp command. If you are moving a file on the same file system, the inode will not change.
In general:
- If the target is an existing directory, then all sources are copied into the target
- If the target directory does not exist, then the source must be only one directory which will be renamed to the target directory.
- If the target is a file, then the source must be only one file so rename will happen.
These look like "formulas" but they are common sense.
The inode detail matters. Moving inside the same filesystem only changes the name entry, so it is instant even for a huge file. Moving across filesystems has to copy every byte and then delete the original, so it is slow and the inode does change.
The two shapes:
The first moves, the second renames. It is the same command, and which one happens depends only on whether the target is an existing directory.
By default mv will not ask before overwriting. -i makes it ask:
And -f forces the overwrite with no question at all.
rm¶
Removes (deletes) files. You can do this recursively using the -r switch, or even prevent it from checking for confirmations using the -f (force) switch. So a rm -rf / means delete everything from the file system.
This is worth repeating. A recursive remove on an important system directory can leave the machine unusable. There is no undo. Use -r only when you are certain about what is inside the directory.
Trying to remove a directory without -r fails:
The difference between rm -r and rmdir is small but real. rmdir only succeeds if the directory is empty. rm -r works either way, which is why it is both more useful and more dangerous.
Real world use for rm -ri: deleting a directory you are not fully sure about. It asks about each item, so you see what is going before it goes.
notes
Normally, the cp command will copy a file over an existing copy, if the existing file is writable. On the other hand, mv will not move or rename a file if the target exists. This is highly dependent on your system's configuration. But in all cases, you can overcome this using the -f switch.
-f(--force) will cause cp to try overwriting the target.-i(--interactive) will ask a Y/N question (deleting / overwriting).-b(--backup) will make backups of overwritten files-pwill preserve the attributes.
Real world use for -p: copying a web directory to a new server. Without -p, every file gets today's date and the copying user's ownership. With -p, timestamps, ownership and permissions survive the copy.
Creating (mkdir) and removing (rmdir) directories¶
The mkdir command creates directories.
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 16K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 116 Aug 14 04:44 note_to_self
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ mkdir new_dir
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 20K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 116 Aug 14 04:44 note_to_self
drwxrwxr-x 2 nagato nagato 4.0K Aug 14 04:57 new_dir
If you want to create a tree of directories, you can use the -p switch to tell mkdir to create the parent directories if needed:
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ mkdir -p 1/2/3
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ tree
.
├── 1
│ └── 2
│ └── 3
├── data.txt
├── info.txt
├── new_dir
├── note_to_self
└── tasks.txt
Without -p, mkdir 1/2/3 fails, because directory 1 does not exist yet and mkdir will not invent it. With -p, it creates each missing level on the way down.
If you need to delete a directory the command is rmdir and you can also use -p for nested removing:
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ rmdir -p 1/2/3
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ rmdir new_dir
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ tree
.
├── data.txt
├── info.txt
├── note_to_self
└── tasks.txt
0 directories, 4 files
If you are using rmdir to remove a directory, it MUST BE EMPTY. That is why many people use rm -rf directory_name to delete a non-empty directory and whatever is in it.
Real world use for mkdir -p in scripts: it does not complain if the directory already exists. So mkdir -p /var/backups/daily is safe to run every night, whether or not the folder is there.
touch¶
touch will create an empty file if it does not exist, or update the modification date of a file if it already exists. The default time is now, but you can specify other times too.
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 16K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 116 Aug 14 04:44 note_to_self
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ touch new_file
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 16K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 116 Aug 14 04:44 note_to_self
-rw-rw-r-- 1 nagato nagato 0 Aug 14 05:08 new_file
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ touch note_to_self
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 16K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 0 Aug 14 05:08 new_file
-rw-rw-r-- 1 nagato nagato 116 Aug 14 05:08 note_to_self
Look at the last listing carefully. note_to_self kept its size of 116 bytes, but its time changed from 04:44 to 05:08, and it moved to the bottom because the listing is sorted by time. Touching an existing file does not empty it.
Or you can specify times. It is possible to use -d and give dates, or use -t and give a timestamp in the form [[CC]YY]MMDDhhmm[.ss]
$ touch -t 200908121510.59 file1
$ touch -d 11am file2
$ touch -d "last fortnight" file3
$ touch -d "yesterday 6am" file4
$ touch -d "2 days ago 12:00" file5
$ touch -d "tomorrow 02:00" file6
$ touch -d "5 Nov" file3
$ ls -ltrh file?
-rw-rw-r-- 1 nagato nagato 0 Aug 12 2009 file1
-rw-rw-r-- 1 nagato nagato 0 Aug 12 12:00 file5
-rw-rw-r-- 1 nagato nagato 0 Aug 13 06:00 file4
-rw-rw-r-- 1 nagato nagato 0 Aug 14 2022 file2
-rw-rw-r-- 1 nagato nagato 0 Aug 15 2022 file6
-rw-rw-r-- 1 nagato nagato 0 Nov 5 2022 file3
Note that -d accepts plain English like yesterday 6am and 2 days ago 12:00, while -t wants the strict digit format. 200908121510.59 reads as year 2009, month 08, day 12, hour 15, minute 10, second 59.
It is also possible to use another file's time, with the -r switch (for --reference):
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -l /etc/debian_version
-rw-r--r-- 1 root root 13 Aug 22 2021 /etc/debian_version
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ touch -r /etc/debian_version file1
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 20K
-rw-rw-r-- 1 nagato nagato 0 Aug 22 2021 file1
file1 took the exact date of /etc/debian_version, which is Aug 22 2021.
Two more options. -a changes only the access time, -m changes only the modification time, and using both together changes both:
It also notes that touch can create several files at once:
file¶
To determine the type of a file, you should use the file command. It looks into the file and determines its type.
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ file file1
file1: empty
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ file note_to_self
note_to_self: ASCII text
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ file /bin/bash
/bin/bash: ELF 64-bit LSB pie executable, x86-64, version 1 (SYSV), dynamically linked, interpreter /lib64/ld-linux-x86-64.so.2, BuildID[sha1]=33a5554034feb2af38e8c75872058883b2988bc5, for GNU/Linux 3.2.0, stripped
The -i switch prints the mime format.
Breaking down that third line, since it is dense:
ELF 64-bit LSB pie executable= a Linux executable, 64 bit.x86-64= the CPU architecture it was built for.dynamically linked, interpreter /lib64/ld-linux-x86-64.so.2= it needs shared libraries at run time, loaded by that program. This ties back to objective 102.3.stripped= the debug symbols were removed to make the file smaller.
The key point is that file reads the contents, not the name. Renaming a JPEG to notes.txt does not fool it.
dd¶
The dd command copies data from its input to its output (say files or devices). You may use it just like copy:
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ dd if=note_to_self of=new_file
0+1 records in
0+1 records out
116 bytes copied, 0.00141561 s, 81.9 kB/s
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ cat new_file
I will continue learning... and if I get confused, I'll repeat the last section once more till everything is clear!
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$
ifis the Input Fileofis the Output File
Note the unusual syntax. dd uses option=value rather than the normal -o value or --option=value style. It is the odd one out among Unix commands.
But commonly people use it to read and write from block devices. For example, this will read all the sectors from /dev/sdb and write them to a file named backup.dd. Later you can restore this backup by swapping the if and of and writing from backup.dd to /dev/sdb.
or even:
Another common usage is creating files of specific sizes:
or even writing your iso files to a USB disk to have a live bootable USB:
Caution: here you are writing directly on a block device. If you do something wrong you will ruin your disk and need to reformat it.
Two useful options. status=progress makes dd report as it works, since by default it prints nothing until it finishes:
And conv=ucase converts the text to upper case while copying:
Real world use for bs=: this sets the block size, how much data is moved per read and write. A larger block size like bs=4M makes writing an ISO to USB much faster than the tiny default, because it means far fewer trips to the device.
find¶
The find command helps us to find files based on different criteria. Look at this:
$ find . -iname "[a-j]*"
./howcool.sort
./alldata
./mydir/howcool.sort
./mydir/newDir/insideNew
./howcool
- The first parameter says where we should search (including subdirectories).
- The
-nameswitch indicates the criteria. Hereinamemeans search files with this name and ignore the character cases (z equals Z).
Another common switch is -type to indicate the type we are searching for (f for regular files, d for directories, and l for symbolic links):
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ find . -type d -iname "[a-j]*"
./directory
./directory/innder_one
The general shape is find STARTING_PATH OPTIONS EXPRESSION. The basic case:
$ find . -name "myfile.txt"
./myfile.txt
$ find /home/frank -name "*.png"
/home/frank/Pictures/logo.png
/home/frank/screenshot.png
Note the quotation marks around "*.png". They matter. Without them bash expands the * before find ever sees it, and you get the wrong results or an error. This connects directly to the quoting section in 103.1.
Two more criteria: -not returns results that do not match, and -maxdepth N limits how many levels deep the search goes.
If you want to search for file sizes do as below:
| command | meaning |
|---|---|
| -size 100c | files which are exactly 100 characters/bytes (you can also use b) |
| -size +100k | files which are more than 100 kilobytes |
| -size -20M | files smaller than 20Megabytes |
| -size +2G | files bigger than 2Gigabytes |
So this will find all files ending in tmp with sizes between 1M and 100M in the /var/ directory:
You can find all empty files with find . -size 0b or find . -empty.
Note how the range was built: two -size tests in the same command. find joins conditions with AND by default, so a file has to satisfy both.
The same idea, hunting for large files:
$ sudo find /var -size +2G
/var/lib/libvirt/images/debian10.qcow2
/var/lib/libvirt/images/rhel8.qcow2
Another useful search criterion is time. These are some of the options:
| switch | meaning | samples |
|---|---|---|
| -amin | Access Minutes | -amin 40 means "files accessed exactly 40min ago", -amin +40 means files accessed more than 40min ago, and -amin -40 means files accessed less than 40min ago |
| -cmin | Status Change Min | -cmin +60 file status changed before last hour |
| -mmin | Modified Minutes | -mmin -60 will give us files modified in last hour |
| -atime | access time in days | -atime +1 means files accessed more than 1 day ago (which means 2 days and more) |
| -ctime | Status Changed in Days | |
| -mtime | Modified days | |
| -newer | Newer than reference | -newer file1 will give you files which are newer than file1 |
If you add the -daystart switch to -mtime or -atime it means that we want to consider days as calendar days, starting at midnight.
The + and - signs are the part people get wrong, so here they are as a picture, using -mmin as the example:
now
|
|<---- 60 min ---->|
| |
+------------------+------------------------> into the past
-mmin -60 -mmin +60
(changed inside (changed before
the last hour) the last hour)
-mmin 60 = changed exactly 60 minutes ago
A real example, searching the whole system for config files touched in the last week:
Acting on files¶
We can execute commands or do other actions on files with various switches:
| switch | meaning |
|---|---|
| -ls | will run ls -dils on each file |
| will print the full name of the files on each line |
But the best way to run commands on found files is the -exec switch. You can point to the file with '{}' or {} and finish your command with \;.
For example, this will remove all empty files in this directory and its subdirectories:
or this will rename all htm files to html:
Since deleting found files is a common task, there is a switch for it: -delete.
The -exec syntax has three parts that all matter:
find . -name "*.conf" -exec chmod 644 '{}' \;
| |
| +-- the ; ends the command.
| It is escaped as \; so
| bash does not eat it.
|
+-- {} is a placeholder. find
substitutes each result here.
Quoted, so filenames with
spaces or special characters
are safe.
Another example combines find with grep to search by content instead of by name:
And the shorthand for deleting:
Real world use: clearing out old backups automatically. find /var/backups -name "*.backup" -mtime +30 -delete removes every backup file older than thirty days. This is the sort of line that goes in a cron job.
Compression¶
gzip & gunzip¶
Straight forward, one gzips a file and one ungzips a file. In place:
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 20K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 0 Aug 14 05:08 new_file
-rw-rw-r-- 1 nagato nagato 116 Aug 14 05:08 note_to_self
drwxrwxr-x 3 nagato nagato 4.0K Aug 14 05:20 directory
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 20K
-rw-rw-r-- 1 nagato nagato 171 Aug 14 04:43 tasks.txt.gz
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 0 Aug 14 05:08 new_file
-rw-rw-r-- 1 nagato nagato 116 Aug 14 05:08 note_to_self
drwxrwxr-x 3 nagato nagato 4.0K Aug 14 05:20 directory
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ gunzip tasks.txt.gz
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 20K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 0 Aug 14 05:08 new_file
-rw-rw-r-- 1 nagato nagato 116 Aug 14 05:08 note_to_self
drwxrwxr-x 3 nagato nagato 4.0K Aug 14 05:20 directory
- gzip preserves time
- gzip creates the new compressed file with the same name but with a .gz ending
- gzip removes the original files after creating the compressed file (you can keep the input file with the
-kswitch)
Watch the timestamps across those three listings. tasks.txt went in at 04:43 and came back out at 04:43, and the size went 207, then 171, then 207 again. That is the "preserves time" point, shown rather than stated.
bzip2 & bunzip2¶
bzip2 is another compressing tool. It works just like the famous gzip but with a different compression algorithm.
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ bzip2 tasks.txt
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 20K
-rw-rw-r-- 1 nagato nagato 172 Aug 14 04:43 tasks.txt.bz2
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 0 Aug 14 05:08 new_file
-rw-rw-r-- 1 nagato nagato 116 Aug 14 05:08 note_to_self
drwxrwxr-x 3 nagato nagato 4.0K Aug 14 05:20 directory
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ bunzip2 tasks.txt.bz2
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls
data.txt directory info.txt new_file note_to_self tasks.txt
Choosing between them: gzip is faster but compresses a little less, bzip2 is slower but compresses a little more. For most work either is fine.
xz & unxz¶
Another compression and decompression tool, just like gzip and bzip2.
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ xz tasks.txt
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 24K
-rw-rw-r-- 1 nagato nagato 224 Aug 14 04:43 tasks.txt.xz
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 116 Aug 14 05:08 note_to_self
drwxrwxr-x 3 nagato nagato 4.0K Aug 14 05:20 directory
-rw-rw-r-- 1 nagato nagato 116 Aug 14 07:51 new_file
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ unxz tasks.txt.xz
nagato@lpicnagato:~/lpic1-practice-iso/100/103.3$ ls -ltrh
total 24K
-rw-rw-r-- 1 nagato nagato 207 Aug 14 04:43 tasks.txt
-rw-rw-r-- 1 nagato nagato 29 Aug 14 04:43 info.txt
-rw-rw-r-- 1 nagato nagato 24 Aug 14 04:44 data.txt
-rw-rw-r-- 1 nagato nagato 116 Aug 14 05:08 note_to_self
drwxrwxr-x 3 nagato nagato 4.0K Aug 14 05:20 directory
-rw-rw-r-- 1 nagato nagato 116 Aug 14 07:51 new_file
Please note that compressing a small text file makes it larger. This is normal in small files because of all the headers and metadata.
The numbers show it. tasks.txt is 207 bytes. Compressed it becomes 224 bytes with xz, which is bigger than the original. Compare that to 171 with gzip and 172 with bzip2 on the same file. Every format adds a header, and on a tiny file that header can cost more than the compression saves.
In some cases, commands like unxz are just calls to xz --decompress.
Archiving with tar & cpio¶
Sometimes we need to create an archive file container of many other files. This operation is different from compressing. It combines files into one and then extracts them again. Archiving is mostly used in backups, moving files to a new location (say via email), and such. This is done with cpio and tar.
The difference between the two ideas, since they get confused:
archiving = many files --> one file (same total size)
tar, cpio
compressing = one file --> one smaller file
gzip, bzip2, xz
tar -czf does both, in that order
tar¶
TapeARchive, or tar, is the most common archiving tool. It automatically creates an archive file from a directory and all its subdirectories.
Common switches are:
| switch | meaning |
|---|---|
-cf myarchive.tar |
create file named myarchive.tar |
-xf myarchive.tar |
extract a file named myarchive.tar |
| -z | compress the archive with gzip after creating it |
| -j | compress the archive with bzip2 after creating it |
| -v | verbose! print a lot of data about what is happening |
| -r | append new files to the currently available archive |
If you issue absolute paths, tar removes the starting slash (/) for safety reasons when creating an archive. If you want to override, use the -p option.
tar can work with tapes and other storage. That is why we use -f to tell it that we are working with files.
The three main operations running, with output:
Extracting somewhere other than here, with -C:
Archiving several directories at once, by listing them:
And the compressed forms, which are the ones you will actually type most days:
$ tar -czvf name-of-archive.tar.gz stuff
$ tar -cjvf name-of-archive.tar.bz stuff
$ tar -xzvf archive.tar.gz
Only one operation switch is allowed at a time. -c creates, -x extracts, -t lists what is inside without unpacking it.
Real world use for -t: before extracting an archive somebody sent you, run tar -tvf archive.tar to see what is in it and where it will land. This is how you avoid an archive that unpacks fifty loose files into your current directory.
cpio¶
Gets a list of files and creates an archive (one file). This file can be used later to extract the original files.
-otells cpio to create an output from its input
Please note that cpio does not look into the folders. So mostly we use it with find:
To extract the original files:
-dwill create the folders-iis for extract
The name means "copy in, copy out", which is exactly the -i and -o pair.
The big difference from tar is where the file list comes from:
tar: you name the directory, tar walks it itself
tar -cvf out.tar stuff/
cpio: something else makes the list, cpio just reads it
find . -name "*" | cpio -o > out.cpio
|
+--> this is why cpio needs find or ls
That is also why cpio needs -d when extracting. It is only ever handed a flat list of names, so it has to be told to rebuild the directories.
Summary¶
I have a Linux system where nearly everything is a file, so these commands are the ones I use constantly. I look around with ls, and -l gives me the long listing where the first character tells me the type, - for a regular file, d for a directory, c for a special file. Adding -h turns raw byte counts into readable sizes, and -a reveals the hidden dot files. I create empty files or update timestamps with touch, and when I want to know what a file really is, file looks inside it rather than trusting the name.
For moving things around I use cp to copy, mv to move or rename, and rm to delete. All three need -r when a directory is involved, because without it they refuse or fail. The -i option makes them ask first, -f makes them stop asking, and -p on cp keeps timestamps and ownership intact. Directories come from mkdir, with -p creating any missing parents, and go away with rmdir if they are empty or rm -r if they are not. rm -rf deserves real care, since there is nothing to undo it.
Wildcards are handled by the shell, not by the commands, which is why *, ? and [abc] work everywhere. The shell expands the pattern first and the command only ever sees the finished list of names. That is also why I have to quote a pattern when I pass it to find, so find receives the pattern itself instead of whatever bash already matched. With find I search by name, by -type, by -size and by time, where + means older or larger and - means newer or smaller. Then -exec runs a command on each result, using {} as the placeholder and \; to close it, or -delete when removing is all I want.
Compression and archiving are two separate jobs that I often do together. gzip, bzip2 and xz each shrink a single file and each replace the original unless I use -k, and on a tiny file they can make it bigger because of the header. tar bundles a whole tree into one archive, and adding -z or -j compresses that bundle in the same command. cpio does the same bundling but reads its file list from find or ls on standard input, which is the real difference between the two. For raw copying at the block level there is dd, with its unusual if= and of= syntax, which is what I reach for when writing an ISO to a USB stick or imaging a whole disk.