The Basics: Literals and Metacharacters
A regular expression is a pattern that describes a set of strings. At its simplest, a regex is just a literal string: the pattern hello matches the text "hello". Things get powerful when you introduce metacharacters — characters with special meaning.
There are 12 metacharacters in regex. Everything else is a literal match.
. ^ $ * + ? { } [ ] \ | ( )
To match a metacharacter literally, escape it with a backslash. To match a literal dot, use \.. To match a literal backslash, use \\.
| Metacharacter | Meaning | Example |
|---|---|---|
. | Any character except newline | h.t matches "hat", "hit", "hot" |
^ | Start of string/line | ^Hello matches "Hello world" |
$ | End of string/line | world$ matches "Hello world" |
| | Alternation (OR) | cat|dog matches "cat" or "dog" |
\ | Escape character | \. matches a literal dot |
Character Classes
Character classes match one character from a defined set. Square brackets define a custom class. Shorthand notations cover common cases.
| Pattern | Matches | Equivalent |
|---|---|---|
[abc] | a, b, or c | — |
[^abc] | Any character except a, b, c | — |
[a-z] | Any lowercase letter | — |
[A-Za-z0-9] | Any alphanumeric character | — |
\d | Any digit | [0-9] |
\D | Any non-digit | [^0-9] |
\w | Word character | [A-Za-z0-9_] |
\W | Non-word character | [^A-Za-z0-9_] |
\s | Whitespace | [ \t\n\r\f\v] |
\S | Non-whitespace | [^ \t\n\r\f\v] |
Try these patterns in the QTool Regex Tester to see matches highlighted in real time.
Quantifiers
Quantifiers control how many times a pattern must match. By default, quantifiers are greedy — they match as much as possible. Add ? after a quantifier to make it lazy (match as little as possible).
| Quantifier | Meaning | Greedy Example |
|---|---|---|
* | 0 or more | ab*c matches "ac", "abc", "abbc" |
+ | 1 or more | ab+c matches "abc", "abbc" (not "ac") |
? | 0 or 1 | colou?r matches "color", "colour" |
{n} | Exactly n times | \d{4} matches "2026" |
{n,} | n or more times | \d{2,} matches "42", "123", "9999" |
{n,m} | Between n and m times | \d{2,4} matches "42", "123", "2026" |
The pattern <.+> applied to <b>bold</b> matches the entire string <b>bold</b> because .+ is greedy. Use <.+?> to match only <b> (the first tag). This is one of the most common regex mistakes.
Anchors and Boundaries
Anchors do not match characters — they match positions in the string.
| Anchor | Position | Example |
|---|---|---|
^ | Start of string | ^Start matches only at beginning |
$ | End of string | end$ matches only at end |
\b | Word boundary | \bword\b matches "word" not "sword" |
\B | Non-word boundary | \Bword matches "sword" not "word" |
Word boundaries are essential for matching whole words. Without \b, the pattern cat matches inside "caterpillar", "concatenate", and "scatter".
Groups and Capturing
Parentheses () serve two purposes: grouping (to apply quantifiers to a sequence) and capturing (to extract parts of a match).
# Extract date components
Pattern: (\d{4})-(\d{2})-(\d{2})
Input: 2026-02-21
Match: 2026-02-21
Group 1: 2026 (year)
Group 2: 02 (month)
Group 3: 21 (day)
# Named groups (clearer in code)
Pattern: (?<year>\d{4})-(?<month>\d{2})-(?<day>\d{2})
# Non-capturing group (grouping without capture)
Pattern: (?:https?|ftp):// # Groups http/https/ftp but does not capture
Backreferences
Backreferences match the same text that was captured by a previous group. \1 refers to the first group, \2 to the second, and so on.
# Match repeated words
Pattern: \b(\w+)\s+\1\b
Matches: "the the", "is is", "very very"
# Match HTML tags (opening = closing)
Pattern: <(\w+)>.*?</\1>
Matches: <b>text</b>, <div>content</div>
Lookaheads and Lookbehinds
Lookaround assertions check what comes before or after a position without including it in the match. They are "zero-width" — they assert a condition but consume no characters.
| Syntax | Name | Meaning |
|---|---|---|
(?=pattern) | Positive lookahead | What follows matches pattern |
(?!pattern) | Negative lookahead | What follows does not match |
(?<=pattern) | Positive lookbehind | What precedes matches pattern |
(?<!pattern) | Negative lookbehind | What precedes does not match |
# Match a number only if followed by "px"
\d+(?=px)
"12px 3em 24px" -> matches 12, 24
# Match a word NOT followed by a comma
\w+(?!,)
"apple, banana cherry" -> matches "banana", "cherry"
# Match a number preceded by $
(?<=\$)\d+
"$100 and 200" -> matches 100
# Password validation: at least one digit and one uppercase
^(?=.*\d)(?=.*[A-Z]).{8,}$
Lookarounds are easier to understand when you can see them work. Paste these patterns into the Regex Tester with sample text and watch which parts match and which parts are asserted but not consumed.
10 Patterns You Will Actually Use
These are production-ready patterns for common tasks. Test them in the QTool Regex Tester before dropping them into your code.
# 1. Email (practical, not RFC-complete)
^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$
# 2. URL (http and https)
https?://[^\s/$.?#].[^\s]*
# 3. IPv4 address
^(?:(?:25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)\.){3}(?:25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)$
# 4. Date (YYYY-MM-DD)
^\d{4}-(0[1-9]|1[0-2])-(0[1-9]|[12]\d|3[01])$
# 5. Phone number (US format)
^(\+1)?[-.\s]?\(?\d{3}\)?[-.\s]?\d{3}[-.\s]?\d{4}$
# 6. Hex color code
^#?([a-fA-F0-9]{6}|[a-fA-F0-9]{3})$
# 7. Strong password (8+ chars, uppercase, lowercase, digit, special)
^(?=.*[a-z])(?=.*[A-Z])(?=.*\d)(?=.*[@$!%*?&])[A-Za-z\d@$!%*?&]{8,}$
# 8. HTML tags (simple extraction)
<(\w+)(?:\s[^>]*)?>(.*?)</\1>
# 9. CSS hex/rgb color values
(?:#[a-fA-F0-9]{3,8}|rgba?\([^)]+\)|hsla?\([^)]+\))
# 10. Import statements (JS/TS)
^import\s+(?:{[^}]+}|\w+)(?:\s*,\s*(?:{[^}]+}|\w+))?\s+from\s+['"][^'"]+['"];?$
The Regex Builder on QTool lets you construct patterns with interactive visual blocks. Select character classes, quantifiers, and groups from a menu instead of writing syntax from memory.
Regex Tools
Frequently Asked Questions
Greedy quantifiers (*, +, {n,m}) match as many characters as possible, then backtrack if needed. Lazy quantifiers (*?, +?, {n,m}?) match as few characters as possible, then expand if needed. For example, given the string <b>bold</b> and <b>more</b>, the greedy pattern <b>.*</b> matches <b>bold</b> and <b>more</b> (everything between the first <b> and the last </b>), while the lazy pattern <b>.*?</b> matches <b>bold</b> (stopping at the first </b>). Use lazy quantifiers when you want the shortest possible match.
A practical email regex is: ^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$. This matches most real-world email addresses. It checks for one or more valid characters before the @, a domain name with dots, and a top-level domain of at least 2 characters. Note that the full RFC 5322 email specification allows many edge cases that no simple regex can cover. For production use, this pattern catches 99% of valid emails. Always combine regex validation with an actual confirmation email for critical applications.
Lookaheads and lookbehinds are zero-width assertions that check for a pattern without including it in the match. A positive lookahead (?=pattern) asserts that what follows matches the pattern. A negative lookahead (?!pattern) asserts that what follows does not match. A positive lookbehind (?<=pattern) asserts that what precedes matches the pattern. A negative lookbehind (?<!pattern) asserts that what precedes does not match. For example, \d+(?= dollars) matches "100" in "100 dollars" but not in "100 euros". The word "dollars" is checked but not included in the match result.
The most effective way to learn regex is through interactive practice with immediate feedback. Start with a regex testing tool like QTool's Regex Tester where you can see matches highlighted in real-time as you type. Begin with literal matches, then learn character classes, quantifiers, and anchors. Practice with real tasks from your own codebase: extracting URLs from text, validating form inputs, or parsing log files. Avoid trying to memorize every feature upfront. Learn the basics and look up advanced features like lookaheads and backreferences when you actually need them.
Different programming languages implement different regex engines with varying feature support. JavaScript uses the ECMAScript regex engine, while Python uses its own re module. Key differences include: Python supports lookbehind assertions with variable length while JavaScript requires fixed length. Python uses re.DOTALL for dot to match newlines while JavaScript uses the s flag. Named groups use (?P<name>) in Python but (?<name>) in JavaScript. Python has re.VERBOSE for readable patterns while JavaScript does not. Always test your regex in the target language's environment and consult language-specific documentation for edge cases.