109.3 Basic network troubleshooting¶
Weight: 4
Candidates should be able to troubleshoot networking issues on client hosts.
Objectives
- Manually configure network interfaces, including viewing and changing the configuration of network interfaces using iproute2.
- Manually configure routing, including viewing and changing routing tables and setting the default route using iproute2.
- Debug problems associated with the network configuration.
- Awareness of legacy net-tools commands.
Terms
ip, hostname, ss, ping, ping6, traceroute, traceroute6, tracepath, tracepath6, netcat, ifconfig, netstat, route
A method for troubleshooting¶
The best test of a network is to use the application itself. When "I cannot open web pages" lands on your desk, work outward, one step at a time, and each step rules out one layer:
1. interface does it have an IP address and netmask? is it UP? ip addr, ip link
|
2. gateway can I reach my own router? ping <gateway>
|
3. internet can I reach a server by IP address? ping 4.2.2.4
|
4. DNS can I reach it by name? ping google.com, dig
|
5. the path where exactly do packets get lost? traceroute, tracepath
Two families of tools do these jobs:
| Job | Modern (iproute2) | Legacy (net-tools) |
|---|---|---|
| interfaces and addresses | ip address, ip link |
ifconfig |
| routing table | ip route |
route, netstat -r |
| connections and listening ports | ss |
netstat |
| neighbour (ARP) table | ip neighbour |
arp |
ifconfig, route and netstat are considered legacy and may be missing on new systems. Installing the net-tools package brings them back. You need to recognise them for the exam, and use ip and ss in practice.
hostname prints the name of the machine you are on, a quick check that you are working on the right box:
The ip command¶
ip configures almost everything about networking. Each part has a subcommand, and each subcommand has its own man page: man ip-address, man ip-route, man ip-link. The SEE ALSO section of man ip lists them all. Add help after a subcommand for a quick summary:
$ ip address help
Usage: ip address {add|change|replace} IFADDR dev IFNAME [ LIFETIME ]
[ CONFFLAG-LIST ]
ip address del IFADDR dev IFNAME [mngtmpaddr]
ip address {save|flush} [ dev IFNAME ] [ scope SCOPE-ID ]
[ to PREFIX ] [ FLAG-LIST ] [ label LABEL ] [up]
ip address [ show [ dev IFNAME ] [ scope SCOPE-ID ] [ master DEVICE ]
(...)
ip address, ip addr and ip a are the same command. The man page uses the full name: man ip-address, not man ip-addr.
Step 1: the interface¶
Listing interfaces¶
ip link lists the network interfaces and their state. ls /sys/class/net and ifconfig -a (all interfaces, even those that are down) do the same:
$ ip link
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN mode DEFAULT group default qlen 1000
link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00
2: enp0s3: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state UP mode DEFAULT group default qlen 1000
link/ether 08:00:27:54:18:57 brd ff:ff:ff:ff:ff:ff
3: enp0s8: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state UP mode DEFAULT group default qlen 1000
link/ether 08:00:27:ab:11:3e brd ff:ff:ff:ff:ff:ff
$ ls /sys/class/net
enp0s3 enp0s8 lo
Checking the address¶
An interface needs a correct IP address and netmask to work. If an interface has no inet line, or is state DOWN, that is the problem. ip addr show displays them:
$ ip addr show
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN group default
link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00
inet 127.0.0.1/8 scope host lo
valid_lft forever preferred_lft forever
inet6 ::1/128 scope host
valid_lft forever preferred_lft forever
2: wlp3s0: <NO-CARRIER,BROADCAST,MULTICAST,UP> mtu 1500 qdisc pfifo_fast state DOWN group default qlen 1000
link/ether f0:de:f1:62:c5:73 brd ff:ff:ff:ff:ff:ff
3: enp0s25: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc mq state UP group default qlen 1000
link/ether 8c:a9:82:7b:89:06 brd ff:ff:ff:ff:ff:ff
inet 192.168.1.35/24 brd 192.168.1.255 scope global dynamic enp0s25
valid_lft 254836sec preferred_lft 254836sec
inet6 fe80::8ea9:82ff:fe7b:8906/64 scope link
valid_lft forever preferred_lft forever
What to look at: the wired card enp0s25 is UP with address 192.168.1.35/24, which looks fine. The Wi-Fi card wlp3s0 is DOWN with NO-CARRIER: nothing is connected to it. The legacy view of the same card:
$ ifconfig
enp0s25 Link encap:Ethernet HWaddr f0:de:f1:62:c5:73
inet addr:192.168.1.35 Bcast:192.168.1.255 Mask:255.255.255.0
inet6 addr: fe80::8ea9:82ff:fe7b:8906/64 Scope:Link
UP BROADCAST RUNNING MULTICAST MTU:1500 Metric:1
RX packets:231586 errors:0 dropped:0 overruns:0 frame:0
TX packets:200220 errors:0 dropped:0 overruns:0 carrier:0
collisions:0 txqueuelen:1000
RX bytes:198053888 (198.0 MB) TX bytes:51583154 (51.5 MB)
Setting an address¶
Configuring needs root. With ip, the same command works for IPv4 and IPv6:
ifconfig accepts the netmask in several forms, but needs the word add for IPv6:
# ifconfig enp1s0 192.168.50.50/24
# ifconfig eth2 192.168.50.50 netmask 255.255.255.0
# ifconfig eth2 192.168.50.50 netmask 0xffffff00
# ifconfig enp0s8 add 2001:db8::10/64
Changes made with ip or ifconfig are lost at reboot. Permanent settings belong in the configuration files (109.2).
Up, down and MTU¶
ip link set changes low-level settings, like turning an interface on or off and its MTU (the largest packet it sends). ifconfig can do both too:
# ip link set dev enp0s8 down
# ip link show dev enp0s8
3: enp0s8: <BROADCAST,MULTICAST> mtu 1500 qdisc pfifo_fast state DOWN mode DEFAULT group default qlen 1000
link/ether 08:00:27:ab:11:3e brd ff:ff:ff:ff:ff:ff
# ifconfig enp0s8 up
# ip link set enp0s8 mtu 2000
# ifconfig enp0s3 mtu 1500
| Task | ip | ifconfig |
|---|---|---|
| bring up / down | ip link set dev enp0s8 up / down |
ifconfig enp0s8 up / down |
| set the MTU | ip link set enp0s8 mtu 2000 |
ifconfig enp0s8 mtu 2000 |
| add an IPv4 address | ip addr add 192.168.5.5/24 dev enp0s8 |
ifconfig enp0s8 192.168.5.5/24 |
| add an IPv6 address | ip addr add 2001:db8::10/64 dev enp0s8 |
ifconfig enp0s8 add 2001:db8::10/64 |
Step 2 to 4: testing with ping¶
ping sends an ICMP echo request. If the target is reachable, it answers with an echo reply carrying the same data. Use ping6 for IPv6. Without -c, ping runs until you press Ctrl+C. With -c 3, it sends three packets and stops:
$ ping -c 3 192.168.50.2
PING 192.168.50.2 (192.168.50.2) 56(84) bytes of data.
64 bytes from 192.168.50.2: icmp_seq=1 ttl=64 time=0.525 ms
64 bytes from 192.168.50.2: icmp_seq=2 ttl=64 time=0.419 ms
64 bytes from 192.168.50.2: icmp_seq=3 ttl=64 time=0.449 ms
--- 192.168.50.2 ping statistics ---
3 packets transmitted, 3 received, 0% packet loss, time 2006ms
rtt min/avg/max/mdev = 0.419/0.464/0.525/0.047 ms
$ ping6 -c 3 2001:db8::10
PING 2001:db8::10(2001:db8::10) 56 data bytes
64 bytes from 2001:db8::10: icmp_seq=1 ttl=64 time=0.425 ms
(...)
No answer does not always mean the host is down. Many firewalls block ICMP on purpose, because echo packets can carry hidden data out of a network.
Now walk the steps. First find the gateway in the routing table and ping it:
$ ip route show
default via 192.168.70.1 dev enp0s1
192.168.70.0/24 dev enp0s1 proto kernel scope link src 192.168.70.2
$ ping 192.168.70.1
PING 192.168.70.1 (192.168.70.1) 56(84) bytes of data.
64 bytes from 192.168.70.1: icmp_seq=1 ttl=64 time=1.11 ms
64 bytes from 192.168.70.1: icmp_seq=2 ttl=64 time=0.855 ms
^C
Then a server on the internet by IP address, and then by name:
$ ping 4.2.2.4
PING 4.2.2.4 (4.2.2.4) 56(84) bytes of data.
64 bytes from 4.2.2.4: icmp_seq=1 ttl=50 time=108 ms
64 bytes from 4.2.2.4: icmp_seq=2 ttl=50 time=111 ms
^C
$ ping google.com
ping: unknown host google.com
The IP works but the name does not: that is a DNS problem. The machine cannot turn google.com into an address. Check the resolver configuration, and here there is no nameserver line at all:
$ cat /etc/resolv.conf
# Dynamic resolv.conf(5) file for glibc resolver(3) generated by resolvconf(8)
# DO NOT EDIT THIS FILE BY HAND -- YOUR CHANGES WILL BE OVERWRITTEN
dig then shows how a name is resolved, and by which server:
$ dig google.com
(...)
;; ANSWER SECTION:
google.com. 293 IN A 216.58.214.46
;; Query time: 120 msec
;; SERVER: 4.2.2.4#53(4.2.2.4)
Server 4.2.2.4 resolved google.com to 216.58.214.46. If ping name fails but dig name (or dig @8.8.8.8 name) returns an address, the resolver config is the fault, not the network. DNS tools are covered in detail in 109.4.
Routing¶
How routing decides¶
An IP address has two parts: the network part and the host part. The netmask (or prefix length) says how many bits are the network part:
192.168.130.5/20
192 168 130 5
11000000 10101000 10000010 00000101
11111111 11111111 11110000 00000000 <- 20 bits of netmask
network = 192.168.128.0 host = 2.5
To send a packet, the machine compares the network part of the destination with its routing table:
- A matching network in the table: send the packet there.
- No match, but a default route exists: send it to the default gateway.
- No match and no default route: the packet is dropped, with
Network is unreachable.
Viewing the routing table¶
Three commands show it: ip route, route, and netstat -r:
$ ip route
default via 10.0.2.2 dev enp0s3 proto dhcp metric 100
10.0.2.0/24 dev enp0s3 proto kernel scope link src 10.0.2.15 metric 100
192.168.150.0/24 dev enp0s8 proto kernel scope link src 192.168.150.200
$ route
Kernel IP routing table
Destination Gateway Genmask Flags Metric Ref Use Iface
default 10.0.2.2 0.0.0.0 UG 100 0 0 enp0s3
10.0.2.0 0.0.0.0 255.255.255.0 U 100 0 0 enp0s3
192.168.150.0 0.0.0.0 255.255.255.0 U 0 0 0 enp0s8
$ netstat -r
Kernel IP routing table
Destination Gateway Genmask Flags MSS Window irtt Iface
default 10.0.2.2 0.0.0.0 UG 0 0 0 enp0s3
10.0.2.0 0.0.0.0 255.255.255.0 U 0 0 0 enp0s3
192.168.150.0 0.0.0.0 255.255.255.0 U 0 0 0 enp0s8
Add -n (route -n, netstat -nr) to show numbers instead of names, so default appears as 0.0.0.0:
$ netstat -nr
Destination Gateway Genmask Flags Iface
0.0.0.0 192.168.70.1 0.0.0.0 UG enp0s1
192.168.70.0 0.0.0.0 255.255.255.0 U enp0s1
ip -6 route, route -6 or netstat -6r.
Reading one line of ip route:
default via 10.0.2.2 dev enp0s3 proto dhcp metric 100
| | | | |
destination gateway interface added by cost
The default route goes through gateway 10.0.2.2 on interface enp0s3, was added by DHCP, and has a cost (metric) of 100. When there is no scope, the scope is global. proto kernel means the kernel added the route itself, because the network is directly connected to an interface.
The Flags column of route:
| Flag | Means |
|---|---|
U |
the route is up |
G |
it goes through a gateway |
! |
a reject route, it will not be used |
n |
the route is not cached |
Metric is the cost used by routing protocols, and Ref is a reference count; the Linux kernel uses neither. Use counts how often the route was looked up.
Fixing a missing default route¶
A classic case: the gateway answers, but the internet does not:
$ ping 8.8.8.8
connect: Network is unreachable
$ ping 192.168.1.1
PING 192.168.1.1 (192.168.1.1) 56(84) bytes of data.
64 bytes from 192.168.1.1: icmp_seq=1 ttl=254 time=3.03 ms
Network is unreachable means there is no route at all. The local network works, so the default gateway is missing. Here is the same fault created on purpose, then fixed with ip route:
$ sudo ip route del default
$ ip route show
192.168.70.0/24 dev enp0s1 proto kernel scope link src 192.168.70.2
$ ping 4.2.2.4
ping: connect: Network is unreachable
$ ip route add default via 192.168.70.1
RTNETLINK answers: Operation not permitted
$ sudo ip route add default via 192.168.70.1
$ ping 4.2.2.4
PING 4.2.2.4 (4.2.2.4) 56(84) bytes of data.
64 bytes from 4.2.2.4: icmp_seq=1 ttl=55 time=316 ms
Changing routes needs root, hence the first error. Like addresses, routes added this way are lost at reboot.
Adding and removing routes¶
You can also add a route to one specific network:
| Task | ip | route |
|---|---|---|
| add a default route | ip route add default via 192.168.70.1 |
route add default gw 192.168.70.1 |
| delete the default route | ip route del default |
route del default |
| add a network route | ip route add 2001:db8:1::/64 via 2001:db8::3 |
route -6 add 2001:db8:1::/64 gw 2001:db8::3 |
| delete it | ip route del 2001:db8:1::/64 via 2001:db8::3 |
route -6 del 2001:db8:1::/64 gw 2001:db8::3 |
Note the words: ip uses via, route uses gw. And route needs -6 for IPv6 routes:
# ping6 -c 2 2001:db8:1::20
connect: Network is unreachable
# ip route add 2001:db8:1::/64 via 2001:db8::3
# ping6 -c 2 2001:db8:1::20
PING 2001:db8:1::20(2001:db8:1::20) 56 data bytes
64 bytes from 2001:db8:1::20: icmp_seq=1 ttl=64 time=0.529 ms
64 bytes from 2001:db8:1::20: icmp_seq=2 ttl=64 time=0.438 ms
# ip route del 2001:db8:1::/64 via 2001:db8::3
# ping6 -c 2 2001:db8:1::20
connect: Network is unreachable
Step 5: tracing the path¶
traceroute and traceroute6¶
traceroute shows every router between you and the destination, so you see where packets get lost. It sends packets with a growing TTL (time to live): the first packet expires at the first router, the next one at the second router, and so on. Each router sends back a "TTL exceeded" message, which reveals it:
$ traceroute 4.2.2.4
traceroute to 4.2.2.4 (4.2.2.4), 30 hops max, 60 byte packets
1 192.168.70.1 (192.168.70.1) 0.619 ms * 0.673 ms
2 10.192.0.1 (10.192.0.1) 421.898 ms 728.657 ms 728.617 ms
3 162.221.202.253 (162.221.202.253) 728.597 ms 728.564 ms 728.552 ms
4 * * *
5 207.35.48.241 (207.35.48.241) 728.289 ms 728.265 ms 728.251 ms
(...)
10 * * ae6.4.edge2.SanJose1.level3.net (4.69.220.185) 920.786 ms
11 d.resolvers.level3.net (4.2.2.4) 601.295 ms 921.003 ms 920.970 ms
How to read it:
- Each line is one router (a hop). Hop 1 is your own router, hop 2 is your ISP, and hop 11 is the destination.
- The three times are the round trip of three test packets.
- Names like
d.resolvers.level3.netcome from a reverse DNS lookup of the router's address. * * *means no reply came back. Often that router simply blocks these packets. If the stars continue to the end, the last router that answered is usually where the path breaks.
The same kind of trace, short and annotated:
$ traceroute 4.2.2.4
1 192.168.70.1 0.619 ms # my router
2 10.192.0.1 421.898 ms # my ISP
4 * * * # this hop blocks ICMP
11 d.resolvers.level3.net (4.2.2.4) 601 ms # destination
By default traceroute sends UDP packets to port 33434 and up. Firewalls often block them, so you can switch protocol (both need root):
| Option | Uses |
|---|---|
-I |
ICMP echo requests, like ping. Destinations answer these more often than UDP |
-T -p 80 |
TCP to a port you know is open, here 80. Gets through most firewalls |
Use traceroute6 for IPv6. On most systems traceroute also accepts an IPv6 address directly.
tracepath and tracepath6¶
tracepath works like traceroute. It also discovers the MTU along the path: it sends a very large packet, and the router with the smallest MTU answers with its size. The last line shows the smallest MTU of the whole path (pmtu), which helps with connections that break when packets are split into fragments:
$ tracepath 192.168.1.20
1?: [LOCALHOST] pmtu 1500
1: 10.0.2.2 0.321ms
1: 10.0.2.2 0.110ms
2: 192.168.1.20 2.714ms reached
Resume: pmtu 1500 hops 2 back 64
Unlike traceroute, tracepath does not handle IPv6 by itself. You must use tracepath6:
$ tracepath 2001:db8::11
tracepath: 2001:db8::11: Address family for hostname not supported
$ tracepath6 2001:db8::11
1?: [LOCALHOST] 0.027ms pmtu 1500
1: net2.example.net 0.917ms reached
Resume: pmtu 1500 hops 1 back 1
Connections and listening ports: ss and netstat¶
ss (modern) and netstat (legacy) show which ports are open and which connections exist. They take the same main options:
| Option | Shows |
|---|---|
-a |
all sockets |
-l |
listening sockets only |
-t |
TCP |
-u |
UDP |
-n |
numbers, no name lookups for addresses or ports |
-p |
the process that owns each socket |
-r |
(netstat only) the routing table |
The combination you will type most is -tulpn (or -tulnp): TCP and UDP, listening, with process, as numbers. It answers "is my service actually listening?":
# netstat -tulnp
Active Internet connections (only servers)
Proto Recv-Q Send-Q Local Address Foreign Address State PID/Program name
tcp 0 0 0.0.0.0:22 0.0.0.0:* LISTEN 892/sshd
tcp 0 0 127.0.0.1:25 0.0.0.0:* LISTEN 1141/master
tcp6 0 0 :::22 :::* LISTEN 892/sshd
tcp6 0 0 ::1:25 :::* LISTEN 1141/master
udp 0 0 0.0.0.0:68 0.0.0.0:* 692/dhclient
# ss -tulnp
Netid State Recv-Q Send-Q Local Address:Port Peer Address:Port
udp UNCONN 0 0 *:68 *:* users:(("dhclient",pid=693,fd=6))
tcp LISTEN 0 128 *:22 *:* users:(("sshd",pid=892,fd=3))
tcp LISTEN 0 100 127.0.0.1:25 *:* users:(("master",pid=1099,fd=13))
tcp LISTEN 0 128 [::]:22 [::]:* users:(("sshd",pid=892,fd=4))
tcp LISTEN 0 100 [::1]:25 [::]:* users:(("master",pid=1099,fd=14))
A shorter view of the same idea:
$ ss -tulpn
Netid State Local Address:Port Process
tcp LISTEN 127.0.0.1:631 (cupsd)
tcp LISTEN 0.0.0.0:22 (sshd)
tcp LISTEN 0.0.0.0:25 (master)
The SSH server listens on port 22 on every address (0.0.0.0, * or [::]). The mail server (master, port 25) listens only on 127.0.0.1, so other machines cannot reach it. That detail alone explains many "I cannot connect" reports. Recv-Q counts data received but not yet read by the program, and Send-Q data sent but not yet acknowledged.
netcat: test any port¶
nc (netcat) is a swiss-army knife: it reads and writes raw data over TCP, UDP and Unix sockets, with IPv4 or IPv6. Unlike telnet it scripts cleanly. It is the simplest way to test whether a port is reachable, without the real application. Start a listener with -l on one machine:
and connect from another, then type:
The text typed on one side appears on the other. Press Ctrl+C to stop. -u uses UDP instead of TCP.
Another quick test, with both ends on one machine:
$ nc -l 1337 # listen on port 1337
$ nc localhost 1337 # connect and type; text appears on the listener
Are you enjoying the LPIC?
Some versions of netcat have -e, which runs a program and connects it to the network. This makes a crude remote shell, so be careful with it:
$ hostname
net2
$ nc -u -e /bin/bash -l 1234
$ hostname
net1
$ nc -u net2.example.net 1234
hostname
net2
pwd
/home/emma
The commands typed on net1 ran on net2. Not every nc supports -e; check its man page.
Summary¶
I troubleshoot outward, one layer at a time. First hostname reminds me which machine I am on. Then the interface: ip addr show shows whether it has the right address and netmask and ip link whether it is UP; I fix them with ip addr add and ip link set dev ... up|down|mtu, and I recognise the legacy ifconfig forms too. Then I ping the gateway, then a server by IP like 4.2.2.4, then the same server by name; when the IP works but the name fails, the problem is DNS and I check /etc/resolv.conf and use dig. I remember that a host which ignores ping may only be behind a firewall.
The routing table decides where packets go: a matching network first, then the default route, otherwise Network is unreachable. I read it with ip route show, route -n or netstat -nr, adding -6 for IPv6, and I know the U and G flags. If the gateway answers but the internet does not, the default route is usually missing, and ip route add default via <gateway> fixes it (route add default gw in the old syntax). All of these changes are temporary until I put them in the configuration files.
To find where packets get lost I use traceroute (with -I for ICMP or -T -p for TCP when UDP is blocked), where * * * marks a hop that does not answer, or tracepath, which also reports the path MTU; the IPv6 variants are traceroute6 and tracepath6. ss -tulpn, or the older netstat -tulpn, tells me whether a service is listening, on which address and which process owns it, and nc lets me open or test any TCP or UDP port by hand to see whether it passes traffic. Throughout, I know the iproute2 tools ip and ss are the modern replacements for the net-tools ifconfig, route and netstat.