103.2 Process text streams using filters¶
Weight: 2
Candidates should be able to apply filters to text streams.
Objectives
Send text files and output streams through text utility filters to modify the output using standard UNIX commands found in the GNU textutils package.
Terms
bzcat, cat, cut, head, less, md5sum, nl, od, paste, sed, sha256sum, sha512sum, sort, split, tail, tr, uniq, wc, xzcat, zcat
Streams¶
In UNIX, a lot of data is in TEXT form. Log files, configurations, data, and so on. Filtering this data means taking an input stream of text and performing some conversion on the text before sending it to an output stream. In this context, a stream is nothing more than "a sequence of bytes that can be read or written using library functions that hide the details of an underlying device from the application".
In simple words, a text stream is an input of text from a keyboard, a file, a network device, and so on, which can be viewed, changed and examined via text util commands.
Modern programming environments and shells, including bash, use three standard I/O streams:
- stdin is the standard input stream, which provides input to commands.
- stdout is the standard output stream, which displays output from commands.
- stderr is the standard error stream, which displays error output from commands.
Here we are talking about stdin, and viewing or manipulating it via different commands and utilities. You will see more about these streams, and how to combine commands to pipe inputs and outputs, in chapter 103.4.
+---------------------+
stdin (0) -->| a filter command |--> stdout (1)
| cat, cut, sed, ... |
+---------------------+
|
+--> stderr (2)
This is the Unix philosophy: write programs that do one thing well, and write them to work together. Piping is what makes them work together.
If you do not tell a command where to read from, it reads from your keyboard. Typing cat on its own shows this clearly, because it just repeats whatever you type:
$ cat
This is a test
This is a test
Hey!
Hey!
It is repeating everything I type!
It is repeating everything I type!
^C
A short review of the redirection operators before the filters themselves:
$ cat > mytextfile
This is a test
I hope cat is storing this to mytextfile as I redirected the output
I will hit ctrl+c now and check this
^C
$ cat mytextfile
This is a test
I hope cat is storing this to mytextfile as I redirected the output
I will hit ctrl+c now and check this
The > sign tells cat to send its output into the file instead of to the screen. Using it between two files copies one to the other:
diff printed nothing, which means the two files are the same. The >> operator appends instead of overwriting, and now diff has something to report:
$ echo 'This is my new line' >> mynewtextfile
$ diff mynewtextfile mytextfile
4d3
< This is my new line
The | pipe sends the output of one program into another:
$ cat mytextfile | grep this
I hope cat is storing this to mytextfile as I redirected the output
I will hit ctrl+c now and check this
$ cat mytextfile | grep -i this
This is a test
I hope cat is storing this to mytextfile as I redirected the output
I will hit ctrl+c now and check this
The second run added -i, which ignores case, so the line starting with a capital This was matched too.
Viewing commands¶
cat¶
This command simply outputs its input stream, or the filename you give it. As with most commands, if you do not give input to it, it will read the data from the keyboard.
nagato@funlife:~/w/lpic/101$ cat > mydata
test
this is the second line
bye
nagato@funlife:~/w/lpic/101$ cat mydata
test
this is the second line
bye
When entering the input via the keyboard, ctrl+d will end the stream.
You can also provide more than one input file name:
nagato@funlife:~/w/lpic/101$ cat mydata directory_data
test
this is the second line
bye
total 0
-rw-rw-r-- 1 nagato nagato 0 Jan 4 17:33 12
-rw-rw-r-- 1 nagato nagato 0 Jan 4 17:33 62
-rw-rw-r-- 1 nagato nagato 0 Jan 4 17:33 neda
-rw-rw-r-- 1 nagato nagato 0 Jan 4 17:33 nagato
-rw-rw-r-- 1 nagato nagato 0 Jan 4 17:33 you
-rw-rw-r-- 1 nagato nagato 0 Jan 4 17:34 amir
-rw-rw-r-- 1 nagato nagato 0 Jan 4 17:37 directory_data
The name comes from "concatenate", and this is why. The two files were joined end to end into one stream.
Some common cat switches are -n to show line numbers, -s to squeeze blanks, -T to show tabs, and -v to show non-printing characters.
Real world use for -T: a Makefile fails with a confusing error, and the cause is spaces where a tab was required. cat -T Makefile prints every tab as ^I so you can see instantly which line is wrong.
Real world use for -s: a config file with long runs of blank lines becomes readable, since -s collapses each run down to a single blank line.
bzcat, xzcat, zcat, gzcat¶
These are used to directly cat the bz, xz, and Z and gz compressed files. They let you see the contents of compressed files without uncompressing them first.
Which one to use depends on how the file was compressed:
A full example. A file named ftu.txt is created holding the list of commands for this objective, then compressed:
Note that gzip removed the original file and left only ftu.txt.gz. It prints nothing while doing so, unless you add -v for verbose output. Now zcat reads the compressed file directly:
$ zcat ftu.txt.gz
bzcat
cat
cut
head
less
md5sum
nl
od
paste
sed
sha256sum
sha512sum
sort
split
tail
tr
uniq
wc
xzcat
zcat
Real world use: rotated log files are compressed, so /var/log/syslog.2.gz cannot be read with cat. Running zcat /var/log/syslog.2.gz | grep error searches yesterday's log without unpacking it and without leaving a large uncompressed file behind.
less¶
This is a powerful tool to view larger text files. It can paginate, search and move in text files.
There is another command called more. It is more familiar for people coming from the DOS environment and not very common in the Linux world. Do not use it. Remember: less is more than more.
Some less common commands are as follows.
| Command | Usage |
|---|---|
| q | Exit |
| /foo | Searches for foo |
| n | Next (search) |
| N | Previous (search) |
| ?foo | Search backward for foo |
| G | Go to end |
| nG | Go to line n |
| PageUp, PageDown, UpArrow, DownArrow | You guess! |
A practical point: you could pipe cat into less, but there is no reason to:
$ sudo cat /var/log/syslog # scrolls past too fast to read
$ sudo less /var/log/syslog # do this instead
less opens files by itself, so the extra cat and the extra pipe do nothing useful.
od¶
This command dumps files, showing files in formats other than text. Normal behavior is OctalDump, showing in base 8:
nagato@funlife:~/w/lpic/101$ od mydata
0000000 062564 072163 072012 064550 020163 071551 072040 062550
0000020 071440 061545 067543 062156 066040 067151 005145 074542
0000040 005145
0000042
Not good enough for normal human beings. Let's use some switches:
- -t will tell what format to print:
-t afor showing only named characters,-t cfor showing escaped chars. You can summarize the two above to-aand-c - -A for choosing how to present the offset field:
-A dfor Decimal,-A ofor Octal,-A xfor hex,-A nfor None
od is very useful to find problems in your text files, say finding out if you are using tabs or correct line endings.
First, what the columns mean. The first column is the byte offset for each line. Then each of the eight columns after it holds the values of the data in that column. Hexadecimal instead of octal, with -x:
$ od -x ftu.txt
0000000 7a62 6163 0a74 6163 0a74 7563 0a74 6568
0000020 6461 6c0a 7365 0a73 646d 7335 6d75 6e0a
0000040 0a6c 646f 700a 7361 6574 730a 6564 730a
0000060 6168 3532 7336 6d75 730a 6168 3135 7332
0000100 6d75 730a 726f 0a74 7073 696c 0a74 6174
0000120 6c69 740a 0a72 6e75 7169 770a 0a63 7a78
0000140 6163 0a74 637a 7461 000a
0000151
The readable version, with -c, which shows characters instead of numbers:
$ od -c ftu.txt
0000000 b z c a t \n c a t \n c u t \n h e
0000020 a d \n l e s s \n m d 5 s u m \n n
0000040 l \n o d \n p a s t e \n s e d \n s
0000060 h a 2 5 6 s u m \n s h a 5 1 2 s
0000100 u m \n s o r t \n s p l i t \n t a
0000120 i l \n t r \n u n i q \n w c \n x z
0000140 c a t \n z c a t \n
0000151
Now every newline shows up as a visible \n. Adding -An drops the offset column, leaving just the characters:
$ od -An -c ftu.txt
b z c a t \n c a t \n c u t \n h e
a d \n l e s s \n m d 5 s u m \n n
l \n o d \n p a s t e \n s e d \n s
h a 2 5 6 s u m \n s h a 5 1 2 s
u m \n s o r t \n s p l i t \n t a
i l \n t r \n u n i q \n w c \n x z
c a t \n z c a t \n
Real world use for od -c: a shell script written on Windows fails with a strange error. Running od -c script.sh shows \r \n at the end of every line instead of just \n. Those \r characters are the Windows line endings, and they are what broke the script.
Choosing parts of files¶
split¶
Will split the files. It is very useful for transferring HUGE files on smaller media, say splitting a 3TB file into 8GB parts and moving them to another machine with a USB Disk.
nagato@funlife:~/w/lpic/101$ cat mydata
hello
this is the second line
but as you can see we are
still writing
and this is getting longer
.
.
and longer
and longer!
nagato@funlife:~/w/lpic/101$ ls
mydata
nagato@funlife:~/w/lpic/101$ split -l 2 mydata
nagato@funlife:~/w/lpic/101$ ls
mydata xaa xab xac xad xae
nagato@funlife:~/w/lpic/101$ cat xab
but as you can see we are
still writing
- By default, split uses xaa, xab, xac and so on for output file names. It can be changed with
split -l 2 mydata output, which splits mydata into outputaa, outputab and so on, 2 lines per file. - The
-l 2splits 2 lines per file. It is possible to use-b 42to split every 42 bytes, or even-n 5to force 5 output files. - If you want numeric output (x00, x01, and so on) use the
-doption.
Need to join these files? cat them with cat x* > originalfile.
mydata (9 lines)
|
| split -l 2
v
xaa xab xac xad xae
|
| cat x* > originalfile
v
originalfile (9 lines again)
Note the shape of the rejoin command. cat x* works because the default names sort alphabetically in the right order. This is exactly why split names them that way.
head and tail¶
Will show the beginning (head) or end (tail) of text files. By default, they will show 10 lines, but you can change it with -n20 or -20.
tail -f follows the new lines which are being written at the end of the file. Very useful.
To prove the ten line default, pipe head into nl, which numbers the lines it receives:
$ sudo head /var/log/syslog | nl
1 Nov 12 08:04:30 hypatia rsyslogd: [origin software="rsyslogd" swVersion="8.1910.0"] rsyslogd was HUPed
2 Nov 12 08:04:30 hypatia systemd[1]: logrotate.service: Succeeded.
3 Nov 12 08:04:30 hypatia systemd[1]: Started Rotate log files.
4 Nov 12 08:04:30 hypatia vdr: [928] video directory scanner thread started (pid=882, tid=928, prio=low)
5 Nov 12 08:04:30 hypatia vdr: [882] registered source parameters for 'A - ATSC'
6 Nov 12 08:04:30 hypatia vdr: [882] registered source parameters for 'C - DVB-C'
7 Nov 12 08:04:30 hypatia vdr: [882] registered source parameters for 'S - DVB-S'
8 Nov 12 08:04:30 hypatia vdr: [882] registered source parameters for 'T - DVB-T'
9 Nov 12 08:04:30 hypatia vdr[882]: vdr: no primary device found - using first device!
10 Nov 12 08:04:30 hypatia vdr: [929] epg data reader thread started (pid=882, tid=929, prio=high)
The same check on tail, counting lines with wc -l instead:
And changing the count with -n:
Real world use for tail -f: you restart a web server in one terminal while running tail -f /var/log/nginx/error.log in another. Errors appear on screen the moment they are written, so you see the failure as it happens instead of hunting for it afterwards.
cut¶
The cut command will cut one or more columns from a file. Good for separating fields.
Let's cut the first field of a file.
nagato@funlife:~/w/lpic/101$ cat howcool
nagato 5
sina 6
rubic 2
you 12
nagato@funlife:~/w/lpic/101$ cut -f1 howcool
nagato
sina
rubic
you
The default delimiter is TAB. Use -dx to change it to "x", or -d' ' to change it to space.
It is also possible to cut fields 1, 2 and 3 with -f1-3, or only characters with index 4, 5, 7, 8 from each line with -c4,5,7,8.
A real world case is /etc/passwd, where fields are separated by colons:
Reading it right to left: -f1 takes the first field, which is the user name, and -d: says the fields are split by colons. So this lists the names of every user whose login shell is bash.
The same trick counts users and groups. Field 3 is the UID and field 4 is the GID:
The sort -u in the middle matters. Two accounts can share a UID, so without removing duplicates the count would be wrong.
Modifying streams¶
nl¶
This command is for showing line numbers.
nagato@funlife:~/w/lpic/101$ nl mydata | head -3
1 hello
2 this is the second line
3 but as you can see we are
cat -n will also number lines.
The difference between the two is worth knowing: nl skips blank lines by default, while cat -n numbers every line including the blank ones.
sort & uniq¶
Sorts its inputs.
nagato@funlife:~/w/lpic/101$ cat uses
you fedora
nagato ubuntu
rubic windows
neda mac
nagato@funlife:~/w/lpic/101$ cat howcool
nagato 5
sina 6
rubic 2
you 12
nagato@funlife:~/w/lpic/101$ sort howcool uses
nagato 5
nagato ubuntu
neda mac
rubic 2
rubic windows
sina 6
you 12
If you want a reverse sort, use the -r switch.
If you want to sort NUMERICALLY, so that 9 is lower than 19, use -n.
Real world use for -n: sorting file sizes or process memory figures. Without -n, plain text sorting puts 100 before 2, because it compares character by character and 1 comes before 2.
And uniq removes duplicate entries from its input. Normal behavior is removing only the duplicated lines, but you can change its behavior, for example the -f1 switch forces it not to check the first field.
nagato@funlife:~/w/lpic/101$ uniq what_i_have.txt
laptop
socks
tshirt
ball
socks
glasses
nagato@funlife:~/w/lpic/101$ sort what_i_have.txt | uniq
ball
glasses
laptop
socks
tshirt
nagato@funlife:~/w/lpic/101$
As you can see, the input HAS TO BE sorted for uniq to work.
This is the single most important thing about uniq, so here it is as a picture. uniq only ever compares a line against the one directly above it:
unsorted sorted
laptop ball
socks <-- kept glasses
tshirt laptop
ball socks <-- these two are now
socks <-- kept again, socks <-- next to each other,
the two socks so one is removed
were never
neighbours
uniq has great switches:
nagato@funlife:~/w/lpic/101$ cat what_i_have.txt
laptop
socks
tshirt
ball
socks
glasses
nagato@funlife:~/w/lpic/101$ sort what_i_have.txt | uniq -c #show count of each item
1 ball
1 glasses
1 laptop
2 socks
1 tshirt
nagato@funlife:~/w/lpic/101$ sort what_i_have.txt | uniq -u #show only non-repeated items
ball
glasses
laptop
tshirt
nagato@funlife:~/w/lpic/101$ sort what_i_have.txt | uniq -d #show only repeated items
socks
Real world use for sort | uniq -c: finding which IP address hits your web server most. Cut the IP column out of the access log, sort it, then uniq -c gives you a count per address.
paste¶
The paste command pastes lines from two or more files side by side. You cannot do this in a general text editor with ease.
nagato@funlife:~/w/lpic/101$ cat howcool
nagato 5
sina 6
rubic 2
you 12
nagato@funlife:~/w/lpic/101$ cat uses
you fedora
nagato ubuntu
rubic windows
neda mac
nagato@funlife:~/w/lpic/101$ paste howcool uses
nagato 5 you fedora
sina 6 nagato ubuntu
rubic 2 rubic windows
you 12 neda mac
Note that paste joins by line position, not by matching content. Line 1 of the first file lands next to line 1 of the second, even though nagato ended up beside you fedora. It does not look anything up.
tr¶
The tr command translates characters in the stream. For example, tr 'ABC' '123' will replace A with 1, B with 2, and C with 3 in the provided stream. It is a pure filter and does not accept the input file name. If needed you can pipe cat into it, see chapter 103.4.
nagato@funlife:~/w/lpic/101$ cat mydata
hello
this is the second line
but as you can see we are
still writing
and this is getting longer
.
.
and longer
and longer!
nagato@funlife:~/w/lpic/101$ cat mydata | tr 'and' 'AND'
hello
this is the second liNe
but As you cAN see we Are
still writiNg
AND this is gettiNg loNger
.
.
AND loNger
AND loNger!
Note: all 'a's are replaced with 'A'.
This output surprises people, so it is worth spelling out. tr does not swap the word "and" for "AND". It builds a character by character map:
That is why line became liNe. The n in the middle of an unrelated word was still mapped.
Real world use for tr -s: squeezing repeated characters. Running ls -l | tr -s ' ' collapses runs of spaces into one, which makes the output usable by cut, since cut counts each space as its own separator.
sed¶
sed is stream editor. It is POWERFUL and can do things that are not far from magic. Just like most of the tools we have seen so far, sed can work as a filter or take input from a file. Sed is a great tool for replacing text using regular expressions. If you need to replace A with B only once in each line in a stream, just issue sed 's/A/B/':
nagato@funlife:~/w/lpic/101$ cat uses
you fedora
nagato ubuntu
rubic windows
neda mac
nagato@funlife:~/w/lpic/101$ sed 's/ubuntu/debian/' uses
you fedora
nagato debian
rubic windows
neda mac
nagato@funlife:~/w/lpic/101$
The pattern for changing EVERY occurrence of A to B in a line is sed 's/A/B/g'.
Remember escape characters? They also work here, and this will replace every space with a tab:
nagato@funlife:~/w/lpic/101$ cat mydata
hello
this is the second line
but as you can see we are
still writing
and this is getting longer
.
.
and longer
and longer!
nagato@funlife:~/w/lpic/101$ sed 's/ /\t/g' mydata > mydata.tab
nagato@funlife:~/w/lpic/101$ cat mydata.tab
hello
this is the second line
but as you can see we are
still writing
and this is getting longer
.
.
and longer
and longer!
Three more sed forms are worth knowing, because they cover what sed can do besides substitution.
Print only matching lines, with -n and p:
The -n tells sed to print nothing on its own, and p prints only what matched. Without -n, every line would print, and the matches would print twice.
Delete matching lines, with d:
$ sed /cat/d < ftu.txt
cut
head
less
md5sum
nl
od
paste
sed
sha256sum
sha512sum
sort
split
tail
tr
uniq
wc
Substitute, the same s form as above:
$ sed s/cat/dog/ < ftu.txt
bzdog
dog
cut
head
less
md5sum
nl
od
paste
sed
sha256sum
sha512sum
sort
split
tail
tr
uniq
wc
xzdog
zdog
And editing the file itself, rather than printing to the screen:
-i means edit in place. The text you put right after -i becomes the extension of a backup copy of the original. Writing plain -i with nothing after it overwrites your file with no backup at all, so -i.backup is the safer habit.
Real world use for sed -n '$=': this prints the number of the last line, which is the same as counting lines. It works as an alternative to wc -l:
That counts the processors on the machine, since /proc/cpuinfo has one processor line per CPU.
One more sed detail worth keeping. Selecting a field by number is not something sed does directly, so matching has to be precise. Searching /etc/passwd for group 1000 with a plain /1000/ also matches a user whose UID is 1000, which is wrong. Anchoring the colons and the character after fixes it:
$ sed -n /:1000:[A-Z]/p mypasswd | cut -d: -f5 | cut -d, -f1
Dave Edwards
Emma Jones
Frank Cassidy
Grace Kearns
Henry Adams
John Chapel
Reading that chain: sed selects the right lines, the first cut takes field 5 which is the comment field, and the second cut splits that field on commas to keep only the full name.
Getting stats¶
wc¶
The wc is word count. It counts the lines, words and bytes in the input stream.
Reading those three numbers in order: 9 lines, 25 words, 121 bytes.
It is very common to count the line numbers with the -l switch.
Real world use: grep processor /proc/cpuinfo | wc -l tells you how many CPUs the machine has, without reading the whole file yourself.
-¶
You should know that if you put - instead of a filename, the data will be replaced from the pipe (or keyboard stdin).
nagato@funlife:~/w/lpic/101$ wc -l mydata | cat mydata - mydata
hello
this is the second line
but as you can see we are
still writing
and this is getting longer
.
.
and longer
and longer!
9 mydata
hello
this is second line
but as you can see we are
still writing
and this is getting longer
.
.
and longer
and longer!
Reading that command: cat was given three inputs in a row, mydata, then -, then mydata again. The - in the middle is where the piped output of wc -l was inserted, which is why the line 9 mydata sits between the two copies of the file.
Hashing¶
A hash function is any function that can be used to map data of arbitrary size to fixed size values. There are different hashes and we use them for different purposes. For example, a site may hash your password in its database to keep it secure, then check the hash of the provided password against the hash it already has during logins. A site may also provide the hash of a file so you can be sure that you have downloaded the correct file.
The hashing algorithms covered in LPIC1 are:
- md5sum
- sha256sum
- sha512sum
You can check any file, or an input stream's hash, with something like this:
nagato@ocean:~$ md5sum /tmp/myfile.txt
8183aa57a23658efe7ba7aebe60816bc /tmp/myfile.txt
nagato@ocean:~$ sha256sum /tmp/myfile.txt
7ddcfda184b55ee06b0c81e0ad136b1aa4a86daeb1078bcaeccc246eb2c8693b /tmp/myfile.txt
nagato@ocean:~$ sha512sum /tmp/myfile.txt
79e5d789528e5e55fc1bddcb381afd56e896b1b452347a76777fb38d76c9754278700036f35df2a53c4d53d3e3623538a8b9ed155a3fd5275e667bdbf3c0b359 /tmp/myfile.txt
As you can see, sha512sum creates a longer hash which is more secure.
The other half of this is verifying a file rather than just printing a hash. Save the hash into a file first:
$ sha256sum ftu.txt
345452304fc26999a715652543c352e5fc7ee0c1b9deac6f57542ec91daf261c ftu.txt
$ sha256sum ftu.txt > sha256.txt
Then check the file against it with -c:
Now change the file by even one line, and check again:
$ echo "new entry" >> ftu.txt
$ sha256sum -c sha256.txt
ftu.txt: FAILED
sha256sum: WARNING: 1 computed checksum did NOT match
original file --> hash --> 345452304fc26999a715...
|
| saved in sha256.txt
v
file after download --> hash --> compare
|
same? -------+------- different?
| |
OK FAILED
file is intact file is corrupt
or was tampered with
Real world use: Linux distributions publish SHA256SUMS files next to their ISO downloads for exactly this. A Debian mirror, for example, lists MD5SUMS, SHA1SUMS, SHA256SUMS and SHA512SUMS alongside the .iso files. After downloading an installer image, you run sha256sum -c against that file. OK means the download is a true copy. FAILED means either the download broke, or the file is not what the project published.
Summary¶
I have a Linux system where most of the data I care about is plain text, so the tools in this objective are the ones I reach for daily. Every one of them works the same way: text comes in on stdin, comes out on stdout, and errors go to stderr. That shared shape is what lets me chain them together with a pipe, and it is why small single purpose commands are enough to do complicated work.
For looking at files I use cat for short ones and less for long ones. When a file is compressed I use zcat, bzcat or xzcat depending on how it was packed, which saves me from unpacking a rotated log just to search it. When a file misbehaves in a way I cannot see, od -c shows me the hidden characters, and that is usually how I find a stray tab or a Windows line ending.
To pull out just the part I want, head and tail give me the first or last ten lines, tail -f follows a log as it is written, cut picks out fields by delimiter, and split breaks a large file into pieces that cat can join back later. To change a stream, sort orders it, uniq collapses repeated neighbouring lines, tr maps characters one to one, and sed does substitution and much more besides. The trap I have to remember is that uniq only compares each line to the one above it, so it is nearly always sort first and uniq second.
For counting there is wc, with -l for lines being the form I use most. For checking that a file is exactly what it should be, md5sum, sha256sum and sha512sum turn any file into a fixed length fingerprint. Saving that fingerprint and later running the command with -c tells me OK or FAILED, which is how I confirm a downloaded ISO arrived intact.