What Is Regex (And Why Learn It)
A regular expression (regex) is a pattern that describes a set of strings. Instead of searching for a specific piece of text, you search for anything that matches a pattern.
Consider this: you have a log file with 10,000 lines and you need to find every IP address. You could read each line manually. Or you could use the regex pattern \d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3} and find them all in under a second.
Regex is used everywhere in software development:
- Form validation — Check if an email, phone number, or URL is formatted correctly.
- Search and replace — Find all dates in
MM/DD/YYYYformat and convert them toYYYY-MM-DD. - Log parsing — Extract timestamps, error codes, or IP addresses from server logs.
- Data cleaning — Strip unwanted characters, normalize whitespace, or extract structured data from unstructured text.
- Code refactoring — Rename variables, update function signatures, or migrate API calls across an entire codebase.
Every major programming language supports regex: JavaScript, Python, Java, Go, Ruby, PHP, C#, Rust, and more. The core syntax is the same across all of them. Learn it once, use it everywhere.
Open QTool's Regex Tester in another tab. As you read each section, paste the examples in and experiment. Interactive practice is the fastest way to learn regex.
Literal Matching: The Basics
The simplest regex is a literal string. The pattern cat matches the text "cat" wherever it appears.
Pattern: cat
Text: "The cat sat on the caterpillar"
Matches: "The [cat] sat on the [cat]erpillar"
^^^ ^^^
Pattern: 404
Text: "Error 404: Not Found (code 404)"
Matches: "Error [404]: Not Found (code [404])"
Regex is case-sensitive by default. The pattern Cat does not match "cat". You can change this with the i flag (covered in the Flags section).
Regex finds all occurrences of the pattern within the text (when using the global flag). It does not just find the first one. This is what makes it powerful for search-and-replace operations.
Metacharacters and Special Characters
Certain characters have special meaning in regex. These are called metacharacters. To match them literally, you need to escape them with a backslash (\).
| Character | Meaning | To Match Literally |
|---|---|---|
. |
Any single character (except newline) | \. |
* |
Zero or more of the preceding element | \* |
+ |
One or more of the preceding element | \+ |
? |
Zero or one of the preceding element | \? |
^ |
Start of string (or line) | \^ |
$ |
End of string (or line) | \$ |
[ ] |
Character class | \[ and \] |
( ) |
Grouping / capturing | \( and \) |
{ } |
Quantifier range | \{ and \} |
| |
Alternation (OR) | \| |
\ |
Escape character | \\ |
Pattern: c.t
Text: "cat cot cut c3t c_t c t"
Matches: [cat] [cot] [cut] [c3t] [c_t] [c t]
Pattern: 192\.168\.1\.1
Text: "Server at 192.168.1.1 responded"
Matches: "Server at [192.168.1.1] responded"
(Without escaping the dots: 192.168.1.1 would also match "192x168y1z1")
Character Classes
A character class matches one character from a set. You define the set inside square brackets.
| Pattern | Matches | Example |
|---|---|---|
[abc] |
Any one of a, b, or c | [abc] matches "a" in "apple" |
[a-z] |
Any lowercase letter | [a-z] matches "h" in "Hello" |
[A-Z] |
Any uppercase letter | [A-Z] matches "H" in "Hello" |
[0-9] |
Any digit | [0-9] matches "4" in "Room 4B" |
[a-zA-Z0-9] |
Any alphanumeric character | Combine ranges in one class |
[^abc] |
Any character EXCEPT a, b, or c | [^0-9] matches any non-digit |
Shorthand Character Classes
Regex provides shorthand notation for common character classes. These save typing and improve readability.
| Shorthand | Equivalent | Matches |
|---|---|---|
\d |
[0-9] |
Any digit |
\D |
[^0-9] |
Any non-digit |
\w |
[a-zA-Z0-9_] |
Any "word" character |
\W |
[^a-zA-Z0-9_] |
Any non-word character |
\s |
[ \t\n\r\f] |
Any whitespace |
\S |
[^ \t\n\r\f] |
Any non-whitespace |
Pattern: \d\d\d-\d\d\d\d
Text: "Call 555-1234 or 800-5678"
Matches: "Call [555-1234] or [800-5678]"
Pattern: [aeiou]
Text: "Hello World"
Matches: "H[e]ll[o] W[o]rld"
Pattern: [^aeiou\s]
Text: "Hello World"
Matches: "[H]e[l][l]o [W]o[r][l][d]" (consonants only)
Quantifiers: How Many
Quantifiers specify how many times the preceding element should be matched.
| Quantifier | Meaning | Example |
|---|---|---|
* |
Zero or more | ab*c matches "ac", "abc", "abbc" |
+ |
One or more | ab+c matches "abc", "abbc" (not "ac") |
? |
Zero or one (optional) | colou?r matches "color" and "colour" |
{3} |
Exactly 3 | \d{3} matches "123" (not "12" or "1234") |
{2,4} |
Between 2 and 4 | \d{2,4} matches "12", "123", "1234" |
{3,} |
3 or more | \w{3,} matches words with 3+ characters |
Greedy vs Lazy Quantifiers
By default, quantifiers are greedy: they match as much text as possible. Adding ? after a quantifier makes it lazy: it matches as little as possible.
Text: "<b>bold</b> and <b>more</b>"
Greedy: <b>.*</b>
Match: [<b>bold</b> and <b>more</b>] (one big match)
Lazy: <b>.*?</b>
Matches: [<b>bold</b>] and [<b>more</b>] (two separate matches)
Using .* (greedy dot-star) is the most common source of unexpected regex behavior. When in doubt, use .*? (lazy) or be more specific about what characters you expect.
Anchors: Where to Match
Anchors do not match characters. They match positions within the string.
| Anchor | Matches | Example |
|---|---|---|
^ |
Start of string (or line with m flag) |
^Hello matches "Hello" only at the start |
$ |
End of string (or line with m flag) |
world$ matches "world" only at the end |
\b |
Word boundary | \bcat\b matches "cat" but not "caterpillar" |
\B |
Not a word boundary | \Bcat matches "cat" in "scat" but not "cat" alone |
Pattern: ^Error
Text: "Error: file not found\nWarning: low disk\nError: timeout"
Matches: [Error]: file not found (only the first "Error" at line start)
Pattern: \bcat\b
Text: "The cat chased the caterpillar into the bobcat's cave"
Matches: "The [cat] chased the caterpillar into the bobcat's cave"
(Only the standalone word "cat", not "caterpillar" or "bobcat")
Groups and Capturing
Parentheses ( ) serve two purposes: they group parts of a pattern together, and they capture the matched text for later use.
Pattern: (\d{4})-(\d{2})-(\d{2})
Text: "Date: 2026-02-14"
Full match: "2026-02-14"
Group 1: "2026" (year)
Group 2: "02" (month)
Group 3: "14" (day)
Named Groups
Named groups make your regex self-documenting and easier to reference in code.
const dateRegex = /(?<year>\d{4})-(?<month>\d{2})-(?<day>\d{2})/;
const match = '2026-02-14'.match(dateRegex);
console.log(match.groups.year); // "2026"
console.log(match.groups.month); // "02"
console.log(match.groups.day); // "14"
import re
pattern = r'(?P<year>\d{4})-(?P<month>\d{2})-(?P<day>\d{2})'
match = re.search(pattern, '2026-02-14')
print(match.group('year')) # "2026"
print(match.group('month')) # "02"
print(match.group('day')) # "14"
Non-Capturing Groups
When you need grouping but not capturing (to avoid polluting your match groups), use (?:...).
Pattern: (?:https?|ftp)://(\S+)
Text: "Visit https://example.com"
Full match: "https://example.com"
Group 1: "example.com" (only the domain is captured)
(The protocol group is NOT captured because of ?:)
Backreferences
You can reference a previously captured group within the same pattern using \1, \2, etc.
Pattern: \b(\w+)\s+\1\b
Text: "The the quick brown fox fox jumped"
Matches: "[The the]" and "[fox fox]"
(Finds repeated words)
Try these group patterns in QTool's Regex Playground, which visualizes capture groups and highlights each group in a different color.
Alternation (OR)
The pipe character | acts as a logical OR. It matches the pattern on either side.
Pattern: cat|dog|bird
Text: "I have a cat and a dog"
Matches: "I have a [cat] and a [dog]"
Pattern: (Mon|Tue|Wed|Thu|Fri|Sat|Sun)day
Text: "Monday and Friday are my favorites"
Matches: "[Monday] and [Friday] are my favorites"
Pattern: \.(jpg|jpeg|png|gif|webp)$
Text: "photo.jpg"
Matches: "photo[.jpg]" (matches image file extensions)
Lookaheads and Lookbehinds
Lookarounds are "zero-width assertions." They check if a pattern exists ahead or behind the current position without including it in the match. Think of them as conditions that must be true, but the matched text stays out of the result.
| Syntax | Name | What It Does |
|---|---|---|
(?=...) |
Positive lookahead | What follows MUST match |
(?!...) |
Negative lookahead | What follows must NOT match |
(?<=...) |
Positive lookbehind | What precedes MUST match |
(?<!...) |
Negative lookbehind | What precedes must NOT match |
// Positive lookahead: match a number followed by "px"
Pattern: \d+(?=px)
Text: "width: 100px; height: 200em;"
Matches: "[100]px" (matches 100, not 200)
// Negative lookahead: match "http" NOT followed by "s"
Pattern: http(?!s)
Text: "http://old.com and https://new.com"
Matches: "[http]://old.com" (matches insecure only)
// Positive lookbehind: match digits preceded by "$"
Pattern: (?<=\$)\d+
Text: "Price: $49 and EUR 30"
Matches: "Price: $[49] and EUR 30" (matches 49, not 30)
// Negative lookbehind: match digits NOT preceded by "#"
Pattern: (?<!#)\b\d+\b
Text: "Item #42 costs 99"
Matches: "Item #42 costs [99]" (matches 99, not 42)
Lookbehinds ((?<=...) and (?<!...)) were added to JavaScript in ES2018. They work in all modern browsers (Chrome 62+, Firefox 78+, Safari 16.4+, Edge 79+). In older environments, use a capturing group as a workaround.
Flags and Modifiers
Flags change how the regex engine interprets your pattern. They are added after the closing delimiter in most languages.
| Flag | Name | Effect |
|---|---|---|
g |
Global | Find all matches, not just the first. |
i |
Case-insensitive | /hello/i matches "Hello", "HELLO", "hElLo". |
m |
Multiline | ^ and $ match line starts/ends, not just string starts/ends. |
s |
Dot-all | . matches newline characters too (normally it does not). |
u |
Unicode | Enable Unicode matching. Required for emoji and non-Latin characters. |
// Find all email-like patterns, case insensitive
const emails = text.match(/[\w.+-]+@[\w-]+\.[\w.]+/gi);
// Match at the start of each line (multiline)
const headings = text.match(/^#+\s.+/gm);
// Match across newlines (dotall)
const blocks = text.match(/\{.*?\}/gs);
Practical Examples
Here are real-world regex patterns you will use regularly. Each one is tested and ready to copy.
Email Validation (Basic)
^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$
Matches: user@example.com, first.last+tag@sub.domain.co
No match: @example.com, user@, user@.com
URL Matching
https?://[^\s/$.?#].[^\s]*
Matches: https://example.com/path?q=1
http://sub.domain.co.uk/page#section
Phone Number (US Format)
^(\+1[-.\s]?)?(\(?\d{3}\)?[-.\s]?)?\d{3}[-.\s]?\d{4}$
Matches: 555-123-4567, (555) 123-4567, +1 555.123.4567
5551234567, +1-555-123-4567
Extract HTML Tags
<(\w+)([^>]*)>(.*?)</\1>
Group 1: tag name
Group 2: attributes
Group 3: content
Text: "<h1 class='title'>Hello World</h1>"
Match: tag="h1", attrs=" class='title'", content="Hello World"
Password Strength Check
^(?=.*[a-z])(?=.*[A-Z])(?=.*\d).{8,}$
Matches: "Password1", "myS3cure", "Ab12345678"
No match: "password", "PASSWORD1", "Pass1", "12345678"
Breakdown:
(?=.*[a-z]) - at least one lowercase letter (lookahead)
(?=.*[A-Z]) - at least one uppercase letter (lookahead)
(?=.*\d) - at least one digit (lookahead)
.{8,} - at least 8 characters total
Find and Replace: Date Format Conversion
const text = "Start: 02/14/2026, End: 03/15/2026";
const result = text.replace(
/(\d{2})\/(\d{2})\/(\d{4})/g,
'$3-$1-$2'
);
console.log(result);
// "Start: 2026-02-14, End: 2026-03-15"
Slug Generator
function slugify(text) {
return text
.toLowerCase()
.replace(/[^\w\s-]/g, '') // Remove non-word chars (except spaces and hyphens)
.replace(/\s+/g, '-') // Replace spaces with hyphens
.replace(/-+/g, '-') // Collapse multiple hyphens
.replace(/^-+|-+$/g, ''); // Trim hyphens from ends
}
slugify("Hello World! This is a Test");
// "hello-world-this-is-a-test"
For quick slug generation without writing code, use QTool's Slug Generator tool. And you can test all the patterns above in the Regex Tester.
Test Your Regex Patterns Live
QTool's free Regex Tester highlights matches in real time, shows capture groups, and explains your pattern. No signup required.
Open Regex Tester Regex BuilderRegex Tools
Frequently Asked Questions
A regular expression (regex or regexp) is a sequence of characters that defines a search pattern. It is used to match, search, extract, or replace text based on patterns rather than exact strings. For example, the regex \d{3}-\d{4} matches any phone number in the format 555-1234 (three digits, a hyphen, four digits). Regular expressions are supported in virtually every programming language including JavaScript, Python, Java, Go, Ruby, PHP, and C#, as well as in command-line tools like grep, sed, and awk.
Greedy quantifiers (*, +, {n,m}) match as much text as possible while still allowing the overall pattern to succeed. Lazy quantifiers (*?, +?, {n,m}?) match as little text as possible. For example, given the string "<b>bold</b> and <b>more</b>", the greedy pattern <b>.*</b> matches everything between the first <b> and the last </b>. The lazy pattern <b>.*?</b> matches only "<b>bold</b>" (stops at the first closing tag). Use lazy quantifiers when you want the shortest possible match, which is common when parsing HTML, extracting quoted strings, or matching delimited content.
A practical email validation regex is: ^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$ which matches most valid email addresses by checking for one or more alphanumeric characters (plus common special characters) before the @ symbol, followed by a domain name with at least one dot and a top-level domain of 2 or more letters. Note that the full RFC 5322 email specification is extremely complex and nearly impossible to match perfectly with regex alone. For production applications, use a simple regex for basic format checking, then verify the email actually exists by sending a confirmation message.
Lookaheads and lookbehinds (collectively called lookarounds) are zero-width assertions that check whether a pattern exists before or after the current position without including it in the match. A positive lookahead (?=pattern) asserts that what follows matches the pattern. A negative lookahead (?!pattern) asserts that what follows does NOT match. A positive lookbehind (?<=pattern) asserts that what precedes matches the pattern. A negative lookbehind (?<!pattern) asserts that what precedes does NOT match. Example: \d+(?= dollars) matches "100" in "100 dollars" but not "100" in "100 euros", because it requires "dollars" to follow without including it in the match.
Virtually every modern programming language supports regular expressions. JavaScript uses /pattern/flags syntax or the RegExp constructor. Python provides the re module with functions like re.match(), re.search(), re.findall(), and re.sub(). Java has the java.util.regex package with Pattern and Matcher classes. Go has the regexp package. Ruby has built-in regex with the =~ operator. PHP offers preg_match() and preg_replace() functions. C# uses the System.Text.RegularExpressions namespace. Command-line tools like grep, sed, and awk also use regex extensively. While the core syntax is similar across languages, there are differences in supported features and escaping rules.
Use an interactive regex tester that highlights matches in real time as you type. QTool's free Regex Tester lets you enter a pattern and test string, see all matches highlighted, view capture groups, and get an explanation of what each part of your pattern does. When debugging, start simple: build your pattern incrementally, testing each piece before adding complexity. Use the verbose/extended flag (x in Python) to add comments to complex patterns. Break long patterns into named groups for readability. The Regex Builder tool can also help you construct patterns visually without memorizing syntax.