107.3 Localisation and internationalisation¶
Weight: 3
Candidates should be able to localise a system in a different language than English. As well, an understanding of why LANG=C is useful when scripting.
Objectives
- Configure locale settings and environment variables.
- Configure timezone settings and environment variables.
Terms
/etc/timezone, /etc/localtime, /usr/share/zoneinfo/, LC_*, LC_ALL, LANG, TZ, /usr/bin/locale, tzselect, timedatectl, date, iconv, UTF-8, ISO-8859, ASCII, Unicode
The earth is round and large, so people live in different timezones and speak different languages. That gives Linux a few jobs:
- show programs and messages in the user's own language,
- know the local timezone, so scheduled jobs run at the right moment and log times are correct (very important when you investigate a security problem),
- tell programs about the user's preferences: language, number and date format, currency,
- store and show text correctly, which needs an agreed character encoding.
All distributions use the same basic files, variables and commands for this. This page covers time first, then language and encoding.
Timezones and UTC¶
A timezone is a region of the world that has the same clock time. Timezones are measured from the prime meridian (longitude 0). The time there is called UTC (Coordinated Universal Time). Every timezone is described by its offset from UTC. GMT (Greenwich Mean Time) is used as another name for UTC in offset names.
Real timezones follow country and city borders, so they are usually named after a large city or region, like America/Sao_Paulo or Europe/Paris. The offset name works too, but be careful with the sign:
GMT-5means the region is 5 hours behind UTC (UTC is 5 hours ahead),GMT+3means the region is 3 hours ahead of UTC.
UTC lets people agree on a moment wherever they are: "start the change at 02:30 UTC" means the same instant for everybody, and each person converts it to local time.
A good habit for servers: keep the hardware clock in UTC and let each case choose its own timezone. Cloud servers usually run in UTC to avoid confusion between machines, while a user who logs in remotely can still see their local time.
Checking the date and time¶
date shows the current date, time and timezone. A + lets you choose the format:
$ date
Fri Jun 2 02:07:35 PM EDT 2023
$ date +'%Y%m%d-%H%M'
20230602-1408
$ date
Mon Oct 21 10:45:21 -03 2019 # -03: three hours behind UTC, so GMT-3
cal prints a month calendar. On systemd systems, timedatectl gives more detail:
$ timedatectl
Local time: Fri 2023-06-02 14:23:34 EDT
Universal time: Fri 2023-06-02 18:23:34 UTC
RTC time: Fri 2023-06-02 18:23:34
Time zone: America/New_York (EDT, -0400)
System clock synchronized: no
NTP service: active
RTC in local TZ: no
RTC is the hardware clock, and RTC in local TZ: no means it is kept in UTC, as recommended.
Choosing a timezone with tzselect¶
Timezone names are not always easy to guess, and one zone can have several names. tzselect asks you a few questions (continent, country, region) and prints the right name. It comes with the GNU C library tools, so every distribution has it. Older systems had tzconfig.
$ tzselect
Please identify a location so that time zone rules can be set correctly.
Please select a continent, ocean, "coord", or "TZ".
1) Africa 7) Europe
2) Americas 8) Indian Ocean
3) Antarctica 9) Pacific Ocean
4) Asia 10) coord - I want to use geographical coordinates.
5) Atlantic Ocean 11) TZ - I want to specify the timezone using the Posix TZ format.
6) Australia
#? 2
Please select a country whose clocks agree with yours.
...
9) Brazil
...
#? 9
Please select one of the following time zone regions.
...
8) Brazil (southeast: GO, DF, MG, ES, RJ, SP, PR, SC, RS)
...
#? 8
The following information has been given:
Brazil
Brazil (southeast: GO, DF, MG, ES, RJ, SP, PR, SC, RS)
Therefore TZ='America/Sao_Paulo' will be used.
Is the above information OK?
1) Yes
2) No
#? 1
You can make this change permanent for yourself by appending the line
TZ='America/Sao_Paulo'; export TZ
to the file '.profile' in your home directory; then log out and log in again.
Here is that TZ value again, this time on standard output so that you
can use the /usr/bin/tzselect command in shell scripts:
America/Sao_Paulo
Note that tzselect does not change anything. It only tells you the name and suggests a TZ line, like Atlantic/Bermuda if you choose Atlantic Ocean and then Bermuda.
The TZ variable¶
The TZ environment variable sets the timezone for your shell session only, whatever the system default is. Add this line to ~/.profile to keep it for your future logins:
Or use it for one command, to see the time somewhere else. env runs the command with the same environment, except for the variable you change:
$ env TZ='Africa/Cairo' date
Mon Oct 21 15:45:21 EET 2019
$ env TZ='Asia/Tokyo' date
Sat Jun 3 03:25:40 AM JST 2023
The system timezone: /etc/timezone and /etc/localtime¶
The data for every timezone lives in /usr/share/zoneinfo/, as binary files organized by name. The file for America/Sao_Paulo is /usr/share/zoneinfo/America/Sao_Paulo. These files also hold the daylight saving time rules (when clocks move by an hour). Those rules change from time to time, so keep the files up to date: the normal package upgrade of your distribution does that.
The system timezone is set by two files:
| File | Content |
|---|---|
/etc/localtime |
the timezone data the system uses. It should be a symbolic link to a file in /usr/share/zoneinfo/ |
/etc/timezone |
the timezone name as plain text, for example America/Sao_Paulo |
$ ls -l /etc/localtime
lrwxrwxrwx 1 root root 38 Apr 26 09:41 /etc/localtime -> ../usr/share/zoneinfo/America/New_York
$ cat /etc/timezone
America/Sao_Paulo
To change the system timezone, point /etc/localtime to the right file. A symbolic link is better than a copy, because a copy can cause problems during later upgrades. On Debian-based systems the zone name is also written in plain text in /etc/timezone (older Red Hat used /etc/sysconfig/clock). With systemd, timedatectl set-timezone Asia/Tokyo updates the link for you.
In /etc/timezone, an offset-based name must start with Etc/. So for GMT+3 you write:
Locales and the LANG variable¶
A locale is the set of language and regional settings. The most basic one is the LANG variable, which most programs read to choose their language. Its format is:
pt_BR.UTF-8
| | |
| | character encoding
| region code (ISO-3166): Brazil
language code (ISO-639): Portuguese
So en_US.UTF-8 means English, US variant, UTF-8 encoding.
The locale command (/usr/bin/locale) shows all the locale variables currently in use:
$ locale
LANG=en_US.UTF-8
LC_CTYPE="en_US.UTF-8"
LC_NUMERIC="en_US.UTF-8"
LC_TIME="en_US.UTF-8"
LC_COLLATE="en_US.UTF-8"
LC_MONETARY="en_US.UTF-8"
LC_MESSAGES="en_US.UTF-8"
LC_PAPER="en_US.UTF-8"
LC_NAME="en_US.UTF-8"
LC_ADDRESS="en_US.UTF-8"
LC_TELEPHONE="en_US.UTF-8"
LC_MEASUREMENT="en_US.UTF-8"
LC_IDENTIFICATION="en_US.UTF-8"
LC_ALL=
$ locale -a # list the installed locales
locale -a lists every locale installed on the system.
The LC_* variables¶
Each LC_* variable controls one part of the locale. If it is not set, it takes its value from LANG. The ones to know:
| Variable | Controls |
|---|---|
LC_COLLATE |
alphabetical order, for example the order in which files are listed |
LC_CTYPE |
how characters are treated, for example which are uppercase or lowercase |
LC_MESSAGES |
the language of program messages (mostly GNU programs) |
LC_MONETARY |
the currency symbol and money format |
LC_NUMERIC |
the format of non-money numbers: thousand and decimal separators |
LC_TIME |
the date and time format |
LC_PAPER |
the standard paper size |
LC_ALL |
overrides all the others, including LANG |
You do not have to use one locale for everything. You can keep the language in Brazilian Portuguese and set only LC_NUMERIC to the American format, or set LC_TIME="en_GB.UTF-8" to get only British style dates.
LC_ALL is the strongest. It is normally empty, and you set it to override everything for a while. Look at date on a system set to pt_BR.UTF-8:
export LC_ALL=fa_IR.UTF-8 changes all the settings for the session to that locale, with no exception, and unset LC_ALL gives control back to the other variables.
So the order of priority is: LC_ALL first, then the individual LC_*, then LANG.
LANG=C in scripts¶
Locale settings change how programs sort text and format numbers. A script that sorts a list can give different results on machines with different locales. To get the same, predictable result everywhere, set LANG=C in scripts. The C locale does a simple byte-by-byte comparison (so sorting is in plain binary order) and uses default English. Because it is so simple, it is also faster than other locales.
Setting the locale for the system and for users¶
The system-wide locale is set in /etc/locale.conf, written like normal shell variables:
On systemd systems, localectl reads and changes it:
localectl also manages the console keyboard layout.
A user can choose a different locale by setting LANG for the current session, or for future sessions by adding it to ~/.bash_profile or ~/.profile. Until the user logs in, though, programs that do not belong to a user (like the login screen of the display manager) still use the system locale.
On Debian based systems, dpkg-reconfigure locales adds locales and sets the default, and the system-wide default is stored in /etc/default/locale (on systemd and Red Hat it is /etc/locale.conf). It is not an exam topic, but it is useful.
Character encoding¶
Computers only store numbers. A character is just a number linked to a symbol. If two systems link different numbers to the same character, text from one looks broken on the other. So they must agree on an encoding, or know how to convert between them.
| Encoding | What it is |
|---|---|
| ASCII | American Standard Code for Information Interchange, the first widely used standard. 7 bits, so only 128 characters: English letters, digits and punctuation. No characters for other languages |
| ISO-8859 | a family of sets that keep ASCII and add characters for other languages (Thai, Arabic and more). For example ISO-8859-1 for Western European. Old, and should be replaced by Unicode |
| Unicode | gives a unique number to every character of every language, plus symbols like ¾, ♠ and π |
| UTF-8 | the most common way to store Unicode, variable-width. Backward compatible with ASCII. It uses 8-bit code units, and one character can take one or more of them |
Use UTF-8 unless you have a reason not to. All modern operating systems use Unicode by default, and with UTF-8 your files work practically everywhere.
Converting files with iconv¶
When a file shows strange characters, it was probably written in another encoding. iconv converts it. -f is from, -t is to:
$ iconv -f ISO-8859-1 -t UTF-8 original.txt > converted.txt
$ iconv -f ISO-8859-1 -t UTF-8 -o converted.txt original.txt # -o instead of >
$ iconv -l # list all encodings
$ iconv -f WINDOWS-1258 -t UTF-8 /tmp/myfile.txt > /tmp/utf8.txt
$ iconv -f UTF-8 -t ASCII//TRANSLIT in.txt # //TRANSLIT: approximate (é -> e), stay ASCII
The long forms are --from-code, --to-code, --output and --list. You will rarely need it today, but you must know it for the exam, and it is a lifesaver when you get an old file from another country.
Summary¶
I split this objective into time and language. A timezone is an offset from UTC (careful: GMT-5 is five hours behind), and servers are best kept with the hardware clock in UTC. date and timedatectl show the time and zone, tzselect only helps me find the right zone name, and the TZ variable sets the zone for my session (in ~/.profile, or env TZ=... date for a one-off command). The system zone comes from /etc/localtime, a symbolic link to a file in /usr/share/zoneinfo/, with the plain-text name also written in /etc/timezone on Debian (offset names start with Etc/). For language, LANG in the form language_REGION.encoding is the overall default, each LC_* variable tunes one part like LC_TIME or LC_NUMERIC, and LC_ALL overrides them all; the locale command lists them, and the system default lives in /etc/locale.conf or comes from localectl. In scripts I use LANG=C, because it forces predictable byte-order sorting and default English output on any machine. Encodings went from 7-bit ASCII to ISO-8859 to Unicode, stored as UTF-8, which is the sensible default today, and iconv -f FROM -t TO converts a file between them.