Skip to content

103.7 Search text files using regular expressions

Weight: 3

Candidates should be able to manipulate files and text data using regular expressions. This objective includes creating simple regular expressions containing several notational elements as well as understanding the differences between basic and extended regular expressions. It also includes using regular expression tools to perform searches through a filesystem or file content.

Objectives

  • Create simple regular expressions containing several notational elements.
  • Understand the differences between basic and extended regular expressions.
  • Understand the concepts of special characters, character classes, quantifiers, and anchors.
  • Use regular expression tools to perform searches through a filesystem or file content.
  • Use regular expressions to delete, change and substitute text.

Terms

grep, egrep, fgrep, sed, regex(7)

Regex

Regular expression, Regex, regex is a pattern to describe what you want to match from a text. For example a and ad both match nagato. d. is a deeper example, because . means anything, so d. will match the last two characters of nagato. In this section, we will cover the grep (generalised regular expression processor) command. It has different regex dialects, in short Basic regex and Extended regex.

The building block has a name: an atom. An atom is the basic element of a regular expression, usually a single character. Most characters mean themselves, but a few have special meaning:

  • . matches any character
  • ^ matches the beginning of a line
  • $ matches the end of a line

A useful detail: ^ is a literal character everywhere except at the start of the expression, and $ is literal everywhere except at the end. So a^b really looks for a caret in the middle.

Regex basics

Simple match You can simply write down whatever you want to match and regex will search for that.

Regex Will match
a after, mina, banana, jadi
na narator, mina, nananana batman, sonar

Repeating

  • The * means repeating the previous character 0 or more times.
  • The + means repeating the previous character 1 or more times.
  • The ? means zero or one repeats.
  • {n,m} means the item needs to match at least n times, but not more than m times.
Regex Will match Note
a*b ab, aaab, aaaaab, aaabthis
a*b b, sober we should have zero or more as and then a b
a+b ab, aab, aaabenz will not match sober or b because there needs to be at least one a
a?b ab, aab, b, batman (zero a then b)

The quantifier always applies to the atom right before it, not to the whole pattern. In a*b the star belongs to the a alone, which is why b on its own matches. This is the single most common misreading of a regex.

   a*b       ->   [a repeated 0+ times][b]
                   ^^^^^^^^^^^^^^^^^^^
                   the * only reaches back one atom

Alternation (|)

If you say a|b it will match a or b.

Character Classes

The dot (.) means any character. So .. will match anything with at least two characters in it. You can also create your own classes with [abc], which will match a or b or c, and [a-z] which matches a to z.

You can also refer to digits with \d.

Regex is case-sensitive.

Ranges

There are shorthands for commonly used classes. Named classes start with [: and end with :].

Range Meaning
[:alnum:] Alphanumeric characters
[:blank:] Space and tab characters
[:digit:] The digits 0 through 9 (equivalent to 0-9)
[:upper:] and [:lower:] Upper and lower case letters, respectively
^ (negation) As the first character after [ in a character class, negates the sense of the remaining characters

A commonly used regex is .* which matches any character, zero or any length.

The full list of named classes: [:alnum:], [:alpha:], [:ascii:], [:blank:], [:cntrl:], [:digit:], [:graph:] (printable except space), [:lower:], [:print:] (printable including space), [:punct:], [:space:] (space, \f, \n, \r, \t, \v), [:upper:], and [:xdigit:] (hex digits 0 through F).

One rule that catches people: a named class can only be used inside brackets. So it is [[:digit:]] with two sets of brackets, not [:digit:] on its own. The outer brackets are the atom, the inner part is the class name.

Matching at specific locations

  • The caret ^ means beginning of the string.
  • The dollar $ means the end of the string.

Samples

  • ^a.* Matches anything that starts with a.
  • ^a.*b$ Matches anything that starts with a and ends with b.
  • ^a.*\d+.*b$ Matches anything starting with a, having some digits in the middle, and ending with b.
  • ^(l|b)oo Matches anything that starts with l or b and then has oo
  • [f-h]|[A-K]$ The last character should be f to h (small) or A to K (capital)

The difference between the two RE standards is one of the stated objectives, so it is worth laying out plainly.

   feature        extended (ERE)      basic (BRE)
   ------------------------------------------------
   zero or more   *                   *          same
   one or more    +                   \+
   zero or one    ?                   \?
   bounds         {2,4}               \{2,4\}
   alternation    |                   \|
   grouping       (abc)               \(abc\)

   in BRE, the plain characters + ? { } ( ) | are
   literal, and you add a backslash to make them special.
   in ERE it is the other way round.

grep uses basic by default. grep -E and egrep switch to extended.

Bounds come in three forms:

  • {i} exactly i times, so [[:blank:]]{2} matches exactly two blanks
  • {i,} at least i times, so [[:blank:]]{2,} matches two or more
  • {i,j} between i and j times, so xyz{2,4} matches xy followed by two to four z

And back references. A group in parentheses can be referred to later by number:

   ([[:digit:]])\1     matches a digit repeated, like 33 or 77

\1 means "the same text the first group matched". More groups become \2, \3 and so on. In basic REs the parentheses need escaping as \( and \).

One more rule about matching: when a short substring matches and a longer one starting at the same point also matches, the longer one wins.

grep

The grep command can search inside files.

$ grep funk words
Garfunkel
Garfunkel's
funk
funked
funkier
funkiest
funking
funk's
funks
funky

These are the most common switches:

switch meaning
-c just show the count
-v reverse the search
-n show line numbers
-l show only file names
-i case insensitive
-r Read all files under each directory, recursively
$ grep a *txt
friends.txt:Rosha
friends.txt:Xavier
friends.txt:Krishna
friends.txt:Mary
my_thinkgs.txt:laptop
$ grep z *txt
$ grep z *txt -i
friends.txt:Zee
$ grep x words -i -c
2264
$ grep z *txt -i -l
friends.txt
$ grep Z *txt -v
friends.txt:Rosha
friends.txt:Jim
friends.txt:Xavier
friends.txt:Krishna
friends.txt:Mary
my_thinkgs.txt:laptop
my_thinkgs.txt:pillow
my_thinkgs.txt:shorts
my_thinkgs.txt:t-shirt
$ grep Z friends.txt -v
Rosha
Jim
Xavier
Krishna
Mary
$

Two things to notice in that block. First, grep z *txt returned nothing while grep z *txt -i found Zee, which is the case sensitivity rule in action. Second, the file name prefix like friends.txt: appears when searching several files but disappears when searching one, as in the last command.

As another example, let's search all of /etc for files containing an IP address, and send the errors, mostly "you do not have permission to read this", to /dev/null:

$ egrep -r "192.168.(1|0)." /etc/ 2> /dev/null
/etc/privoxy/config:#      address 192.168.0.1 on your local private network
/etc/privoxy/config:#      (192.168.0.0) and has another outside connection with a
/etc/privoxy/config:#        listen-address  192.168.0.1:8118
/etc/avahi/hosts:# 192.168.0.1 router.local
/etc/dhcp/dhclient-exit-hooks.d/rfc3442-classless-routes:#   192.168.10.0/24 via 192.168.1.1
/etc/hosts:192.168.1.22 atiteltestbed
/etc/hosts:192.168.100.244 adpsms
/etc/ppp/options:# ms-dns 192.168.1.1
/etc/ppp/options:# ms-dns 192.168.1.2
/etc/ppp/options:# ms-wins 192.168.1.50
/etc/ssl/openssl.cnf:# proxy = # set this as far as needed, e.g., http://192.168.1.1:8080
/etc/cups/cups-browsed.conf:# BrowseAllow 192.168.1.12
/etc/cups/cups-browsed.conf:# BrowseAllow 192.168.1.0/24
/etc/proxychains4.conf:## Exclude connections to 192.168.1.0/24 with port 80
/etc/sane.d/kodakaio.conf:#net 192.168.1.2 0x4041
/etc/sane.d/epsonds.conf:# net 192.168.1.123
/etc/sane.d/airscan.conf:#ip    = 192.168.0.1    ; blacklist by address
/etc/sane.d/saned.conf:#192.168.0.1
/etc/fwupd/redfish.conf:# ex: https://192.168.0.133:443

Note the 2> /dev/null on the end, which is the redirection from 103.4. Searching all of /etc as a normal user produces many permission errors, and this throws them away so only the results remain.

Fuller names for the same options, plus several more:

  • -c or --count show the count instead of the lines
  • -i or --ignore-case
  • -f FILE or --file=FILE read the pattern from a file
  • -n or --line-number
  • -v or --invert-match
  • -H or --with-filename always print the file name
  • -z or --null-data treat the input as null separated, which pairs with find -print0

The -H option matters more than it looks. The file name prefix appears automatically with several files but not with one, so a find -exec grep loses track of which file each line came from:

$ find /usr/share/doc -type f -exec grep -i '3d modeling' "{}" \; | cut -c -100
artistic aspects of 3D modeling. Thus this might be the application you are
This major approach of 3D modeling has not been supported

Adding -H fixes it:

$ find /usr/share/doc -type f -exec grep -i -H '3d modeling' "{}" \; | cut -c -100
/usr/share/doc/openscad/README.md:artistic aspects of 3D modeling. Thus this might be the applicatio
/usr/share/doc/opencsg/doc/publications.html:This major approach of 3D modeling has not been support

There are also context lines, which is grep -1 or -C 1, printing one line either side of each match:

$ find /usr/share/doc -type f -exec grep -i -H -1 '3d modeling' "{}" \; | cut -c -100
/usr/share/doc/openscad/README.md-application Blender), OpenSCAD focuses on the CAD aspects rather t
/usr/share/doc/openscad/README.md:artistic aspects of 3D modeling. Thus this might be the applicatio
/usr/share/doc/openscad/README.md-looking for when you are planning to create 3D models of machine p

Look at the separators. A colon after the file name marks the actual match, a minus sign marks a context line. That is how you tell them apart.

Real world use for -C: reading an error out of a log. The error line alone rarely tells you enough, but grep -C 3 "error" app.log shows what happened just before and after.

Two more examples showing anchors doing useful work:

$ grep '^options' /etc/modprobe.d/alsa-base.conf
options snd-pcsp index=-2
options snd-usb-audio index=-2
options bt87x index=-2
options cx88_alsa index=-2
# fdisk -l | grep '^Disk /dev/sd[ab]'
Disk /dev/sda: 320.1 GB, 320072933376 bytes, 625142448 sectors
Disk /dev/sdb: 7998 MB, 7998537728 bytes, 15622144 sectors

Real world use for grep -v ^#: config files are mostly comments. grep -v ^# /etc/services shows only the lines that actually do something.

Extended grep

Regex is cool and grep is awesome, so many people have tried adding to them or inventing their own variants. One is GNU Extended grep. This dialect of regex does not need much escaping, and you can use it via the -E switch or by using egrep instead of the normal grep. For example, | in an extended regex means "or". So you can do egrep "a|b" words to match anything with an a or a b.

This example uses branching to search for two phrases at once:

$ find /usr/share/doc -type f -exec egrep -i -H -1 '3d (modeling|printing)' "{}" \; | cut -c -100

Either 3D modeling or 3D printing matches. With plain grep, that same pattern would look for a literal (modeling|printing) with the brackets and bar included, and find nothing.

-o prints only the matching part of the line rather than the whole line.

Real world use for -o: pulling values out of a log. grep -o '[0-9]\{1,3\}\.[0-9]\{1,3\}\.[0-9]\{1,3\}\.[0-9]\{1,3\}' access.log gives you a clean list of IP addresses with nothing else attached, ready to pipe into sort | uniq -c.

Fixed grep

If you need to search for exact strings, and not interpret them as a regex, use grep -F or fgrep, so that fgrep this$ will not go for the end of the line and will find this$that instead.

Real world use: searching for something that contains regex characters, like an IP address or a file path. In grep 192.168.1.1, every dot means "any character", so it also matches 192a168b1c1. With fgrep 192.168.1.1 the dots are just dots.

sed

In previous lessons, we saw simple sed usage, and now I have great news for you: sed understands regex. You can use the -r switch to tell sed that we are using regexes.

$ sed -r "s/(Z|R|J)/starts with ZRJ/" friends.txt
starts with ZRJee
starts with ZRJosha
starts with ZRJim
Xavier
Krishna
Mary

Common switches:

switch meaning
-r use advanced regex
-n suppress output, you can use p at the end of your regex ( /something/p ) to print the output
sed -rn "s/happy/HAPPY/p" words
HAPPY
slapHAPPY
unHAPPY

Note what -rn together did there. Without -n, every line in words would print and only some would be changed. With -n plus the p flag, only the changed lines appear.

sed's structure in more depth: an instruction is a single character, optionally preceded by an address saying which lines it applies to. An address can be a line number, a range, or a regular expression between slashes.

Using factor on the numbers 1 to 12 as sample data:

$ factor `seq 12`
1:
2: 2
3: 3
4: 2 2
5: 5
6: 2 3
7: 7
8: 2 2 2
9: 3 3
10: 2 5
11: 11
12: 2 2 3

Delete line 1, with the address 1 and the instruction d:

$ factor `seq 12` | sed 1d
2: 2
3: 3
4: 2 2
...

A range uses a comma:

$ factor `seq 12` | sed 1,7d
8: 2 2 2
9: 3 3
10: 2 5
11: 11
12: 2 2 3

Several instructions separated by semicolons, quoted so the shell does not read the semicolon itself:

$ factor `seq 12` | sed "1,7d;11d"
8: 2 2 2
9: 3 3
10: 2 5
12: 2 2 3

And an address that is a regular expression rather than a number:

$ factor `seq 12` | sed "1d;/:.*2.*/d"
3: 3
5: 5
7: 7
9: 3 3
11: 11

The pattern :.*2.* finds a 2 anywhere after the colon, so every even number and every number with 2 as a factor got deleted.

   sed  [address]  instruction

   1d              delete line 1
   1,7d            delete lines 1 to 7
   /^#/d           delete every line starting with #
   s/a/b/          substitute, no address, so every line

Other instructions, beyond d:

  • c TEXT replaces the whole matching line with TEXT
  • a TEXT adds TEXT on a new line after the match
  • r FILE inserts the contents of a file after the match
  • w FILE writes the matching line out to a file
  • s/FIND/REPLACE/ substitutes, the one used most

The c instruction in action:

$ factor `seq 12` | sed "1d;/:.*2.*/c REMOVED"
REMOVED
3: 3
REMOVED
5: 5
REMOVED
7: 7
REMOVED
9: 3 3
REMOVED
11: 11
REMOVED

For substitution, only the first match on each line is replaced unless you add the g flag, so s/hda/sda/ changes one and s/hda/sda/g changes all.

Back references can do real work too. This keeps only the last field of every line:

# lastb -d -a --time-format notime | grep -v '[0-9]$' | sed -e 's/.* \(.*\)$/\1/' | sort | uniq | head -n 10
116-25-254-113-on-nets.com
132.red-88-20-39.staticip.rima-tde.net
145-40-33-205.power-speed.at
tor.laquadrature.net
tor.momx.site
ua-83-226-233-154.bbcust.telenor.se
vmd38161.contaboserver.net
vmd60532.contaboserver.net
vmi488063.contaboserver.net
vmi515749.contaboserver.net

Reading that pipeline stage by stage:

   lastb ...              list failed SSH login attempts
        |
   grep -v '[0-9]$'       drop lines ending in a digit,
                          which are the ones with no hostname
        |
   sed 's/.* \(.*\)$/\1/' keep only the text after the last space.
                          \( \) remembers it, \1 replaces the
                          whole line with just that part
        |
   sort | uniq            collapse the repeated hostnames
        |
   head -n 10             show ten

The result is a clean list of the hosts attacking the server, ready to feed into firewall rules. This is a good picture of how these small tools combine.

Regular expressions turn up in two more places. find can match whole paths with -regex or -iregex, using -regextype posix-extended to switch dialect:

$ find /usr/share/fonts -regextype posix-extended -iregex '.*(dejavu|liberation).*sans.*(italic|oblique).*'
/usr/share/fonts/dejavu/DejaVuSans-BoldOblique.ttf
/usr/share/fonts/dejavu/DejaVuSans-Oblique.ttf
/usr/share/fonts/liberation/LiberationSans-BoldItalic.ttf
/usr/share/fonts/liberation/LiberationSans-Italic.ttf

And less takes a regex at its / search prompt. Inside a man page, typing ^ *-o jumps straight to the description of the -o option, since options sit at the start of a line after some spaces. Pressing & instead filters the display down to matching lines only.

Summary

I have a Linux system where a regular expression is a small pattern language for describing text I want to find. The building blocks are atoms: an ordinary character means itself, . means any character, ^ anchors to the start of a line and $ to the end. Square brackets make a set, so [abc] matches one of three characters and [a-z] matches a range, while a leading ^ inside the brackets flips it into "none of these". Named classes like [:digit:] and [:alnum:] are shorthands, and they only work inside brackets, which is why they always appear doubled as [[:digit:]].

Quantifiers control how many times the atom right before them may repeat. * is zero or more, + is one or more, ? is zero or one, and braces give exact counts like {2} or {2,4}. The part I have to keep straight is that a quantifier reaches back only one atom, so in a*b the star belongs to the a and a lone b still matches. Grouping with parentheses lets me apply a quantifier to more than one character and also lets me refer back to what was captured, using \1 for the first group.

The two dialects are the thing this objective really tests. In extended regular expressions the characters + ? { } ( ) | are special as they stand. In basic ones they are literal, and I have to write \+, \?, \{, \( and \| to get the special meaning. Plain grep is basic, while grep -E and egrep are extended. fgrep and grep -F turn the pattern matching off completely and search for the exact characters, which is what I want when the thing I am hunting for contains dots or dollar signs of its own.

For searching, grep filters a stream line by line, with -i for case, -v to invert, -c to count, -n for line numbers, -l for file names only, -r to walk a directory, -o to print just the matched part and -C to show surrounding context. For changing text, sed applies instructions to each line, optionally limited by an address that can be a line number, a range, or a regex between slashes. d deletes, c replaces the line, a appends after it, and s/find/replace/ substitutes, with a trailing g to catch every match on the line rather than just the first. Piping the two together, usually with sort and uniq after them, is how I turn a messy log into a short useful list.