diff options
Diffstat (limited to 'src')
| -rw-r--r-- | src/blog/2023-06-16-regex/regex.md | 201 | ||||
| -rw-r--r-- | src/blog/blog.md | 1 | ||||
| -rw-r--r-- | src/blog/feed.xml | 7 |
3 files changed, 209 insertions, 0 deletions
diff --git a/src/blog/2023-06-16-regex/regex.md b/src/blog/2023-06-16-regex/regex.md new file mode 100644 index 0000000..e6c46f1 --- /dev/null +++ b/src/blog/2023-06-16-regex/regex.md | |||
| @@ -0,0 +1,201 @@ | |||
| 1 | # UNIX text filters, part 0 of 3: regular expressions | ||
| 2 | |||
| 3 | One of the most important features of UNIX and its descendants, if | ||
| 4 | not *the* most important feature, is input / output redirection: | ||
| 5 | the output of a command can be displayed to the user, written to a | ||
| 6 | file or used as the input for another command seamlessly, without | ||
| 7 | the the program knowing which of these things is happening. This | ||
| 8 | is possible because most UNIX programs use *plain text* as their | ||
| 9 | input/output language, which is understood equally well by the three | ||
| 10 | types of users - humans, files and other running programs. | ||
| 11 | |||
| 12 | Since this is such a fundamental feature of UNIX, I thought it would | ||
| 13 | be nice to go through some of the standard tools that help the user | ||
| 14 | take advantage of it. At first I thought of doing this as part of | ||
| 15 | my *man page reading club* series, but in the end I decided to give | ||
| 16 | them their own space. My other series has also been going on for | ||
| 17 | more than a year now, so it is a good time to end it and start a | ||
| 18 | new one. | ||
| 19 | |||
| 20 | Let me then introduce you to: **UNIX text filters**. | ||
| 21 | |||
| 22 | ## Text filters | ||
| 23 | |||
| 24 | For the purpose of this blog series, a *text filter* is a program | ||
| 25 | that reads plain text from standard input and writes a modified, | ||
| 26 | or *filtered* version of the same text to standard output. According | ||
| 27 | to the introductory paragraph, this definition includes most UNIX | ||
| 28 | programs; but we are going to focus on the following three, in | ||
| 29 | increasing order of complexity: | ||
| 30 | |||
| 31 | * grep | ||
| 32 | * sed | ||
| 33 | * awk | ||
| 34 | |||
| 35 | In order to unleash the true power of these tools, we first need | ||
| 36 | to grasp the basics of | ||
| 37 | [regular expressions](https://en.wikipedia.org/wiki/Regular_expression). | ||
| 38 | And what better way to do it than following the dedicated | ||
| 39 | [OpenBSD manual page](https://man.openbsd.org/OpenBSD-7.3/re_format)? | ||
| 40 | |||
| 41 | ## (Extended) regular expressions | ||
| 42 | |||
| 43 | Regular expressions, or regexes for short, are a convenient way to | ||
| 44 | describe text patterns. They are commonly used to solve genering | ||
| 45 | string matching problems, such as determining if a given piece | ||
| 46 | of text is a valid URL. Many standard UNIX tools, including the three | ||
| 47 | we are going to cover in this series, support regexes. | ||
| 48 | |||
| 49 | Let's deal with the nasty part first: even within POSIX, there is | ||
| 50 | not one single standard for regular expressions; there are at least | ||
| 51 | two of them: Basic Regular Expressions (BREs) and Extended Regular | ||
| 52 | Expressions (ERE). As it always happens when there is more than one | ||
| 53 | standard for the same thing, other people decided to come up with | ||
| 54 | another version to replace all previous "standards", so we have also | ||
| 55 | [PCREs](https://en.wikipedia.org/wiki/Perl_Compatible_Regular_Expressions), | ||
| 56 | and probably more. [Things got out of hand quickly](https://xkcd.com/927). | ||
| 57 | |||
| 58 | In this post I am going to follow the structure of | ||
| 59 | [re_format(7)](https://man.openbsd.org/OpenBSD-7.3/re_format) and | ||
| 60 | present *extended* regular expresssions first. After that I'll point | ||
| 61 | out the differences with *basic* regular expressions. | ||
| 62 | |||
| 63 | The goal is not to provide a complete guide to regexes, but rather | ||
| 64 | an introduction to the most important features, glossing over the | ||
| 65 | nasty edge-cases. Keep also in mind that I am in no way an expert | ||
| 66 | on the subject: we are learning together, here! | ||
| 67 | |||
| 68 | ### The basics | ||
| 69 | |||
| 70 | You can think of a regular expression as a *pattern*, or a *rule*, | ||
| 71 | that describes which strings are "valid" (they are *matched* by the | ||
| 72 | regular expression) and which are not. As a trivial example, the | ||
| 73 | regular expression `hello` matches only the string "hello". A less | ||
| 74 | trivial example is the regex `.*` that matches *any* string. I'll | ||
| 75 | explain why in a second. | ||
| 76 | |||
| 77 | Beware not to confuse regular expressions with *shell globs*, i.e. | ||
| 78 | the rules for shell command expansion. Although they use similar | ||
| 79 | symbols to achieve a similar goal, they are not the same thing. See | ||
| 80 | [my post on sh(1)](../2022-09-13-sh-1) or | ||
| 81 | [glob(7)](https://man.openbsd.org/OpenBSD-7.3/glob.7) for an | ||
| 82 | explanation on shell globs. | ||
| 83 | |||
| 84 | ### General structure and terminology | ||
| 85 | |||
| 86 | A general regex looks something like this: | ||
| 87 | |||
| 88 | ``` | ||
| 89 | piece piece piece ... | piece piece piece ... | ... | ||
| 90 | ``` | ||
| 91 | |||
| 92 | A sequence of *pieces* is called a *branch*, and a regex is a | ||
| 93 | sequence of branches separated by pipes `|`. Pieces are not separated | ||
| 94 | by spaces, they are simply concatenated. | ||
| 95 | |||
| 96 | The pipes `|` are read "or": a regex matches a given string if any | ||
| 97 | of its branches does. A branch matches a given string if the latter | ||
| 98 | can be written as a sequence of strings, each matching one of the | ||
| 99 | pieces, in the given order. | ||
| 100 | |||
| 101 | Before going into what pieces are exactly, consider the following | ||
| 102 | example: | ||
| 103 | |||
| 104 | ``` | ||
| 105 | hello|world | ||
| 106 | ``` | ||
| 107 | |||
| 108 | This regex matches both the string "hello" and the string "world", | ||
| 109 | and nothing else. The pieces are the single letters composing the | ||
| 110 | two words, and as you can see they are juxtaposed without spaces. | ||
| 111 | |||
| 112 | But what else is a valid piece? In general, a piece is made up of | ||
| 113 | an *atom*, optionally followed by a *multiplier*. | ||
| 114 | |||
| 115 | ### Atoms | ||
| 116 | |||
| 117 | As we have already seen, the most simple kind of atom is a single | ||
| 118 | character. The most *general* kind of atom, on the other hand, is | ||
| 119 | a whole regular expression enclosed in parentheses `()`. Yes, regexes | ||
| 120 | are recursive. | ||
| 121 | |||
| 122 | There are some special characters: for example, a single dot `.` | ||
| 123 | matches *any* single character. The characters `^` and `$` match | ||
| 124 | an empty string at the beginning and at the end of a line, respectively. | ||
| 125 | If you want to match a special character as if it was regular, say | ||
| 126 | because you want to match strings that represent values in the | ||
| 127 | dollar currency, you can *escape* them with a backslash. For example | ||
| 128 | `\$` matches the string "$". | ||
| 129 | |||
| 130 | The last kind of atoms are *bracket expressions*, which consist of | ||
| 131 | lists of characters enclosed in brackets `[]`. A simple list of | ||
| 132 | characters in brackets, like `[xyz]`, matches any character in the | ||
| 133 | list, unless the first character is a `^`, in which case it matches | ||
| 134 | every character *not* in the list. Two characters separated by a | ||
| 135 | dash `-` denote a range: for example `[a-z]` matches every lowercase | ||
| 136 | letter and `[1-7]` matches all digits from 1 to 7. | ||
| 137 | |||
| 138 | You can also use cetain special sets of characters, like `[:lower:]` | ||
| 139 | to match every lowercase letter (same as `[a-z]`), `[:alnum:]` to | ||
| 140 | match every alphanumeric character or `[:digit:]` to match every | ||
| 141 | decimal digit. Check the | ||
| 142 | [man page](https://man.openbsd.org/OpenBSD-7.3/re_format) | ||
| 143 | for the full list. | ||
| 144 | |||
| 145 | ### Multipliers | ||
| 146 | |||
| 147 | The term "multiplier" does not appear anywhere in the manual page, I | ||
| 148 | made it up. But I think it fits, so I'll keep using it. | ||
| 149 | |||
| 150 | Multipliers allow you to match an atom repeating a specified or | ||
| 151 | unspecified amount of times. The most general one is the *bound* | ||
| 152 | multiplier, which consists of one or two comma-separated numbers | ||
| 153 | enclosed in braces `{}`. | ||
| 154 | |||
| 155 | In its most simple form, the multiplier `{n}` repeats the multiplied | ||
| 156 | atom `n` times. For example, the regex `a{7}` is equivalent to the | ||
| 157 | regex `aaaaaaa` (and it matches the string "aaaaaaa"). | ||
| 158 | |||
| 159 | The form `{n,m}` matches *any number* between `n` and `m` of copies | ||
| 160 | of the preceeding atom. For example `a{2,4}` is equivalent to | ||
| 161 | `aa|aaa|aaaa`. If the integer `m` is not specified, the multiplied | ||
| 162 | atom matches any string that consists of *at least* `n` copies of | ||
| 163 | the atom. | ||
| 164 | |||
| 165 | Now we can explain very quickly the more common multipliers `+`, | ||
| 166 | `*` and `?`: they are equivalent to `{1,}`, `{0,}` and `{0,1}` | ||
| 167 | respectively. That is to say, `+` matches at least one copy of the | ||
| 168 | atom, `*` matches any number of copies (including none) and `?` | ||
| 169 | matches either one copy or none. | ||
| 170 | |||
| 171 | ## Basic regular expressions | ||
| 172 | |||
| 173 | Basic regular expressions are less powerful than their extended | ||
| 174 | counterpart (with one exception, see below) and require more | ||
| 175 | backslashes, but it is worth knowing them, because they are used | ||
| 176 | by default in some programs (for example [ed(1)](../2022-12-24-ed)). | ||
| 177 | The main differences between EREs and BREs are: | ||
| 178 | |||
| 179 | * BREs consist of one single branch, i.e. there is no `|`. | ||
| 180 | * Multipliers `+` and `?` do not exist. | ||
| 181 | * You need to escape parentheses `\(\)` and braces `\{\}` to | ||
| 182 | use them with their special meaning. | ||
| 183 | |||
| 184 | There is one feature of BREs, called *back-reference*, that is | ||
| 185 | absent in EREs. Apparently it makes the implementation much more | ||
| 186 | complex, and it makes BREs more powerful. I noticed the author of | ||
| 187 | the manual page despises back-references, so I am not going to learn | ||
| 188 | them out of respect for them. | ||
| 189 | |||
| 190 | ## Conclusion | ||
| 191 | |||
| 192 | Regexes are a powerful tool, and they are more than worth knowing. | ||
| 193 | But, quoting from the manual page: | ||
| 194 | |||
| 195 | ``` | ||
| 196 | Having two kinds of REs is a botch. | ||
| 197 | ``` | ||
| 198 | |||
| 199 | I hope you enjoyed this post, despite the lack of practical examples. | ||
| 200 | If you want to see more applications of regular expressions, stay | ||
| 201 | tuned for the next entries on grep, sed and awk! | ||
diff --git a/src/blog/blog.md b/src/blog/blog.md index be9015d..6024e0c 100644 --- a/src/blog/blog.md +++ b/src/blog/blog.md | |||
| @@ -5,6 +5,7 @@ | |||
| 5 | 5 | ||
| 6 | ## 2023 | 6 | ## 2023 |
| 7 | 7 | ||
| 8 | * 2023-06-16 [UNIX text filters, part 0 of 3: regular expressions](2023-06-16-regex) | ||
| 8 | * 2023-05-05 [I had to debug C code on a smartphone](2023-05-05-debug-smartphone) | 9 | * 2023-05-05 [I had to debug C code on a smartphone](2023-05-05-debug-smartphone) |
| 9 | * 2023-04-10 [The big rewrite](2023-04-10-the-big-rewrite) | 10 | * 2023-04-10 [The big rewrite](2023-04-10-the-big-rewrite) |
| 10 | * 2023-03-30 [The man page reading club: dc(1)](2023-03-30-dc) | 11 | * 2023-03-30 [The man page reading club: dc(1)](2023-03-30-dc) |
diff --git a/src/blog/feed.xml b/src/blog/feed.xml index 58df461..f7582dd 100644 --- a/src/blog/feed.xml +++ b/src/blog/feed.xml | |||
| @@ -9,6 +9,13 @@ Thoughts about software, computers and whatever I feel like sharing | |||
| 9 | </description> | 9 | </description> |
| 10 | 10 | ||
| 11 | <item> | 11 | <item> |
| 12 | <title>UNIX text filters, part 0 of 3: regular expressions</title> | ||
| 13 | <link>https://sebastiano.tronto.net/blog/2023-06-16-regex</link> | ||
| 14 | <description>UNIX text filters, part 0 of 3: regular expressions</description> | ||
| 15 | <pubDate>2023-06-16</pubDate> | ||
| 16 | </item> | ||
| 17 | |||
| 18 | <item> | ||
| 12 | <title>I had to debug C code on a smartphone</title> | 19 | <title>I had to debug C code on a smartphone</title> |
| 13 | <link>https://sebastiano.tronto.net/blog/2023-05-05-debug-smartphone</link> | 20 | <link>https://sebastiano.tronto.net/blog/2023-05-05-debug-smartphone</link> |
| 14 | <description>I had to debug C code on a smartphone</description> | 21 | <description>I had to debug C code on a smartphone</description> |
