103.4 Use streams, pipes and redirects¶
Weight: 4
Candidates should be able to redirect streams and connect them to efficiently process textual data. Tasks include redirecting standard input, standard output, and standard error, piping the output of one command to the input of another command, using the output of one command as arguments to another command, and sending output to both stdout and a file.
Objectives
- Redirecting standard input, standard output, and standard error.
- Pipe the output of one command to the input of another command.
- Use the output of one command as arguments to another command.
- Send output to both stdout and a file.
Terms
tee, xargs
These features help us to control the input and output of commands and do things like saving the output of a command to a file, getting the input of a command from another command, or separating the normal output from errors. We have already used them in previous sections, but let's learn more and deepen our understanding of these.
Redirecting standard IO¶
On a Linux system, most shells use streams for input and output. These streams can be from and toward various things including keyboards, block devices (hard disks, USB sticks) and files.
We have 3 different standard streams:
- STDIN is the standard input stream, which provides input to a command.
- STDOUT is the standard output stream, which includes the output of a command.
- STDERR is the standard error stream, which includes the error output of a command.
The 0, 1 and 2 numbering indicates STDIN, STDOUT and STDERR accordingly. For example, if you want to redirect the standard error, you can use 2> and STDERR will be redirected.
The underlying idea has a name: these numbers are file descriptors. A file descriptor is just an integer the system attaches to a data channel, so a process can refer to it by number. Every standard Linux process starts with three of them already open. They are also reachable as the special devices /dev/stdin, /dev/stdout and /dev/stderr.
+-------------------+
keyboard --0--> | | --1--> screen
(stdin) | a process | (stdout)
| | --2--> screen
+-------------------+ (stderr)
redirection = swapping what sits on the other end of an arrow
This is why programs can be written without caring where their data comes from. A program just reads from stdin, and the shell decides whether that is a keyboard, a file, or another program.
These are the other redirections you can use:
| Operator | Usage |
|---|---|
| > | Redirect STDOUT to a file; Overwrite if exists |
| >> | Redirect STDOUT to a file; Append if exists |
| 2> | Redirect STDERR to a file; Overwrite if exists |
| 2>> | Redirect STDERR to a file; Append if exists |
| &> | Redirect both STDOUT and STDERR; Overwrite if exists |
| &>> | Redirect both STDOUT and STDERR; Append if exists |
| < | Redirect STDIN from a file |
| <> | Redirect STDIN from the file and send the STDOUT to it |
Some examples:
$ ls
bob jack nagato linus sara who_uses_what.txt
$ ls x*
ls: x*: No such file or directory
$ ls j*
jack nagato
$ ls j* x* > output 2> errors
$ cat output
jack
nagato
$ cat errors
ls: x*: No such file or directory
$ cat who_uses_what.txt
nagato, fedora
linux, fedora
bob, ubuntu
jack, arch
sara, fedora
$ tr ' ', '' < who_uses_what.txt
tr: empty string2
$ cat who
$ tr ',', '|' < who_uses_what.txt
nagato| fedora
linux| fedora
bob| ubuntu
jack| arch
sara| fedora
Look at the fourth command carefully. ls j* x* > output 2> errors split one command's output into two different files. The successful listing went into output, and the error message went into errors. That is the whole point of having stdout and stderr as separate channels.
It is also possible to use &1, &2 and &0 to refer to the target of STDOUT, STDERR and STDIN. In this case ls > file1 2>&1 means redirect output to file1 and output stderr to the same place as stdout (file1).
Be careful. ls 2>&1 > file1 means print stderr to the current location of stdout (the terminal) and then change stdout to file1.
That ordering trap is the most common mistake in this objective, so here it is drawn out:
ls > file1 2>&1 ls 2>&1 > file1
step 1: stdout -> file1 step 1: stderr -> wherever
stdout is NOW = terminal
step 2: stderr -> wherever
stdout is NOW = file1 step 2: stdout -> file1
result: both in file1 result: stdout in file1,
stderr on screen
Bash reads redirections left to right. 2>&1 copies wherever stdout is pointing at that moment, not wherever it ends up later.
The reverse form also exists: 1>&2 redirects stdout to stderr. Here is why 2>&1 matters so much in practice. You can pipe stdout into another program, but you cannot pipe stderr directly. So error messages have to be merged into stdout first before another program can read them.
There is also noclobber, a safety option that stops > from silently destroying an existing file:
$ set -o noclobber
$ cat /proc/cpu_info 2>/tmp/error.txt
-bash: /tmp/error.txt: cannot overwrite existing file
Turn it off again with set +o noclobber or set +C. The short form to turn it on is set -C. To make it permanent it goes in your Bash profile.
Appending still works even with noclobber on, because >> adds rather than overwrites:
$ cat /proc/cpu_info 2>>/tmp/error.txt
$ cat /tmp/error.txt
cat: /proc/cpu_info: No such file or directory
cat: /proc/cpu_info: No such file or directory
The message appears twice because the second run appended to the first.
Input redirection works the same way but the data flows right to left:
The suppressed descriptor here is 0, so < is really 0<.
Real world use for 2>: running a script that prints useful output mixed with harmless warnings. ./script.sh 2>/dev/null keeps the real output on screen and throws away the noise.
sending to null¶
In Linux the /dev/null device works like an abyss. You can send anything there and it disappears without being any burden on your system. So it is normal to say:
$ ls j* x* > file1
ls: x*: No such file or directory
$ ls j* x* > file1 2>/dev/null
$ cat file1
jack
nagato
In the first command the error still printed to the screen, because only stdout was redirected. In the second, stderr was sent to /dev/null and vanished.
/dev/null is writable by every user, and nothing written there can ever be recovered. /dev/null is also a special file and is not affected by noclobber, so discarding output always works even with that option on.
here-documents¶
Many shells have here-documents (also called here-docs) as a way of input. You use << and a WORD, and then whatever you input is considered stdin until you give only the WORD on one line.
$ tr ' ' '.' << END_OF_DATA
> this is a line
> and then this
>
> we'll still type
> and,
> done!
> END_OF_DATA
this.is.a.line
and.then.this
we'll.still.type
and,
done!
Here-documents are very useful if you are writing scripts and automated tasks.
The word after << is yours to pick. END_OF_DATA has no special meaning, and EOF is just the most common choice by habit:
There is also a one line version, the here string, written with three less than symbols:
A string with spaces has to be quoted. Without quotes only the first word becomes the here string and the rest are passed as arguments to the command instead.
Real world use for a here string: sending one calculation to bc, which only reads from stdin:
Real world use for a here-document: writing a config file from inside a script, without needing a separate template file on disk.
Pipes¶
With the pipe (|), you can redirect STDOUT, STDIN and STDERR between multiple commands all on one command line. When you do command1 | command2, command1 is executed but its STDOUT is redirected as STDIN into command2.
$ cat who_uses_what.txt
nagato, fedora
linux, fedora
bob, ubuntu
jack, arch
sara, fedora
$ cut -f2 -d, who_uses_what.txt | sed -e 's/ //g' | sort | uniq -c | sort -nr
3 fedora
1 ubuntu
1 arch
If you need to start your pipeline with the contents of a file, start with cat filename | ... or use a < stdin redirect.
That five stage pipeline is worth walking through, because it is a real pattern you will reuse:
cut -f2 -d, take the second field, split on commas -> " fedora"
|
sed -e 's/ //g' delete the spaces -> "fedora"
|
sort group identical lines next to each other
|
uniq -c count each group -> "3 fedora"
|
sort -nr sort by that number, highest first
Note the sort before uniq -c. That is the rule from 103.2 showing up again, uniq only sees neighbours.
Pipes are one of the super strong and super amazing features in the UNIX world. They let you create new tools by combining tools that do atomic things.
The difference between a pipe and a redirect: a redirect sends data to a file or a file descriptor. A pipe sends it to another running process:
$ cat /proc/cpuinfo | wc
208 1184 6096
$ cat /proc/cpuinfo | grep 'model name' | uniq
model name : Intel(R) Xeon(R) CPU X5355 @ 2.66GHz
The machine in that example has many CPUs, so model name repeats many times, and uniq collapsed them into one line.
Pipes and redirects mix freely in the same command:
One more thing to know: all the commands in a pipeline start at the same time. They are not run one after another. Data flows through them as it is produced, which is why tail -f something | grep error keeps working live instead of waiting for the first command to finish.
xargs¶
The xargs utility reads space, tab, newline and end-of-file delimited strings from the standard input and executes the provided utility with the strings as their arguments.
$ ls
bob file1 nagato output who_uses_what.txt
errors jack linus sara
$ ls | xargs echo these are files:
these are files: bob errors file1 jack nagato linus output sara who_uses_what.txt
If you do not give any command to xargs, echo will be the default command.
The reason xargs exists is worth stating, because a pipe alone does not solve it. Some commands read from stdin, and some only accept arguments. rm is the second kind. A pipe hands data to stdin, so ls | rm does nothing useful. xargs bridges the gap:
ls | grep .tmp | rm does nothing, rm ignores stdin
ls | grep .tmp | xargs rm works, xargs turns the
stream into arguments
stream on stdin --> [xargs] --> command arg1 arg2 arg3
One common switch is -I. This is useful if you need to pass stdin arguments in the middle, or even at the start, of your commands. Use it like this: xargs -I SOMETHING echo here is SOMETHING end:
$ cat who_uses_what.txt
nagato, fedora
linus, fedora
bob, ubuntu
jack, arch
sara, fedora
$ cat who_uses_what.txt | xargs -I DATA echo name is DATA is the choice.
name is nagato, fedora is the choice.
name is linus, fedora is the choice.
name is bob, ubuntu is the choice.
name is jack, arch is the choice.
name is sara, fedora is the choice.
Two more useful switches:
-L 1breaks based on new lines-n 1tells xargs to invoke the provided utility after receiving 1 argument.
-I solves a real problem. mv needs the source before the destination, but xargs puts arguments at the end by default. -I gives you a placeholder to control the position:
$ find . -mindepth 2 -name '*avi' -print0 -o -name '*mp4' -print0 -o -name '*mkv' -print0 | xargs -0 -I PATH mv PATH ./
There is also the filename safety problem, which matters in real use. If a path has a space in it, xargs splits it into two arguments and everything breaks. The fix is a pair of matching options:
$ find . -name '*avi' -print0 -o -name '*mp4' -print0 -o -name '*mkv' -print0 | xargs -0 du | sort -n
find -print0 separates results with a null character instead of a newline, and xargs -0 reads them that way. Since a null can never appear inside a filename, nothing can be split by accident. Note that -print0 has to be repeated for each search criterion.
A worked example, using xargs to feed found paths into another program:
$ find /usr/share/icons -name 'debian*' | xargs identify -format "%f: %wx%h\n"
debian-swirl.svg: 48x48
debian-swirl.png: 22x22
debian-swirl.png: 32x32
debian-swirl.png: 256x256
debian-swirl.png: 48x48
debian-swirl.png: 16x16
debian-swirl.png: 24x24
debian-swirl.svg: 48x48
Note that -format there belongs to identify, not to xargs.
A practical caution: if you are already using find, its own -exec option often does the same job, so xargs -n 1 after a find may be unnecessary.
There is also command substitution, a different way to use one command's output inside another. Backquotes, or the newer $() form, replace the command with its output:
$ mkdir `date +%Y-%m-%d`
$ ls
2019-09-05
$ rmdir 2019-09-05
$ mkdir $(date +%Y-%m-%d)
$ ls
2019-09-05
The same method stores output in a variable:
The difference from xargs is the shape of the result. Command substitution pastes the output straight into the command line as text. xargs runs the command repeatedly, feeding it the input piece by piece, which handles large lists that would be too long to paste as one line.
tee¶
The problem with redirection is that you cannot see the progress of your commands in the same terminal. The tee utility solves this. If you need to see the output on screen and also save it to a file, tee is your friend. Give it one or more filenames and it will do the trick.
$ ls -1 | tee allfiles myfiles
bob
errors
file1
jack
nagato
linus
output
sara
who_uses_what.txt
$ cat allfiles myfiles
bob
errors
file1
jack
nagato
linus
output
sara
who_uses_what.txt
bob
errors
file1
jack
nagato
linus
output
sara
who_uses_what.txt
The -a switch will append to files if they exist.
If you need to save stderr too, first redirect it to stdout.
The name comes from a T shaped pipe fitting, which is exactly what it does to the data:
Here is the stderr warning proved with a real failure. Without the redirect, the error shows on screen but the file ends up empty:
With 2>&1 placed before the pipe, tee captures it:
$ make 2>&1 | tee log.txt
make: *** No targets specified and no makefile found. Stop.
$ cat log.txt
make: *** No targets specified and no makefile found. Stop.
This is the same ordering rule from earlier. The redirect has to happen before the pipe, because a pipe only ever carries stdout.
Another example puts tee at the end of a pipeline, keeping the result visible and saved at once:
$ grep 'model name' </proc/cpuinfo | uniq | tee cpu_model.txt
model name : Intel(R) Xeon(R) CPU X5355 @ 2.66GHz
$ cat cpu_model.txt
model name : Intel(R) Xeon(R) CPU X5355 @ 2.66GHz
Real world use for tee -a: watching a long install or build run while keeping a growing log of every attempt, instead of overwriting the log each time.
Summary¶
I have a Linux system where every process starts with three channels already open, and each one has a number. stdin is 0 and normally comes from my keyboard, stdout is 1 and normally goes to my screen, and stderr is 2 and also goes to my screen but carries error messages instead of results. These numbers are file descriptors, and redirecting is nothing more than telling bash to point one of them somewhere else.
For output I use > to write to a file and >> to append. Putting a 2 in front, as in 2> or 2>>, moves the error channel instead of the normal one, and &> moves both at once. For input I use < to feed a file into a command. When I want to throw output away entirely, /dev/null swallows it and nothing comes back. If I want protection against overwriting a file by accident, set -o noclobber makes > refuse, and set +C turns that back off.
The form I have to be careful with is 2>&1, because bash reads redirections left to right. Writing > file 2>&1 sends both channels to the file, but writing 2>&1 > file sends errors to the terminal and only the normal output to the file. The same rule explains why make 2>&1 | tee log.txt captures errors while make | tee log.txt does not, since a pipe only ever carries stdout.
Beyond redirects I have three ways to connect commands. A pipe with | sends one program's output straight into the next program's input, and all the stages run at the same time. Command substitution with $() or backquotes replaces a command with its output so I can use it as an argument or store it in a variable. And xargs turns a stream into arguments, which is what I need for commands like rm and mv that ignore stdin completely, with -I letting me choose where each argument lands and find -print0 paired with xargs -0 keeping filenames with spaces from breaking apart. When I want output on the screen and in a file at the same time, tee splits the stream in two, and -a makes it append rather than overwrite. For multi line input typed straight into a command I use a here-document with <<, and for a single line a here string with <<<.