Understanding ASCII
The Encoding Patterns That Make ASCII Useful
ASCII is too common to be ignored, even though it's one of the most important fundamental places of software infra. It's also one of the best examples of encoding design.
In this article, we demistify the design behind ASCII's encoding.
ASCII table is divided into seven sections as shown below:
There's a notable pattern in the layout of ASCII: the digits, uppercase letters, and lowercase letters are arranged in consecutive ranges, but separated by punctuation intentionally. It looks like this:
0x00 - 0x1F control characters 0x20 SPACE 0x21 - 0x2F punctuation (!!!) 0x30 - 0x39 '0' - '9' 0x3A - 0x40 punctuation (!!!) 0x41 - 0x5A 'A' - 'Z' 0x5B - 0x60 punctuation (!!!) 0x61 - 0x7A 'a' - 'z' 0x7B - 0x7E punctuation (!!!) 0x7F DEL
The punctuation section between uppercase and lowercase letters spans 0x20, which implies that uppercase letters can be transformed into lowercase letters with a simple arithmetic operation, and vice versa.
The encodings of uppercase and lowercase letters are 0x41–0x5A and 0x61–0x7A, respectively, which means their difference is 0x20. The pattern of 0x?1–0x?A is useful for encoding the position of a character in the alphabet. More precisely, the lower 5 bits indicate the character's position in the alphabet:
'A' & 0x1F = 1 'B' & 0x1F = 2 ... 'Z' & 0x1F = 26
To achieve this encoding, punctuation ranges are inserted between digits, uppercase letters, and lowercase letters. At first glance, the layout seems seemingly arbitrary. However, it is deliberately tailored to support the 0x?1–0x?A pattern. This also makes it easy to transform the digits '0'–'9' into the values 0–9 with a simple arithmetic operation, such as subtracting 0x30.
The first 32 code points are control characters, which means the printable characters start at 0x20. So it's easy to check whether it's >= 0x20 to confirm it's printable. However there's an exception, DEL is 0x7F which is not printable and standing at the end of ASCII like a gatekeeper.