What Are Regular Expressions?
A regular expression — usually shortened to regex or regexp — is a sequence of characters that defines a search pattern. Think of it like a metal detector for text. You describe the shape of what you are looking for, and the regex engine scans through your string to find every match.
Imagine you have a 10,000-line log file and you need to find every line that contains an IP address. You could scroll through manually, or you could write a short regex pattern that describes the structure of an IP address and let the computer find all of them in milliseconds.
Regex is built into virtually every programming language. JavaScript has RegExp, Python has the re module, Java has java.util.regex, and even command-line tools like grep and sed are built on regular expressions. Once you learn the syntax, you can use it everywhere.
Here is the simplest possible regex: the literal string hello. It matches the word "hello" inside any text. That is all regex is at its core — a pattern that describes text. Everything else is just making that description more flexible and powerful.
Basic Regex Syntax: Literal Characters and Metacharacters
Every regex pattern is made of two kinds of characters: literal characters that match themselves, and metacharacters that have special meaning.
Literal characters are straightforward. The pattern cat matches the letters c, a, t in sequence. No surprises.
Metacharacters are where regex gets its power. There are 12 characters that have special meaning in regex:
. ^ $ * + ? | \ [ ] { } ( )
Here is what each one does:
.— Matches any single character except a newline. The patternc.tmatches "cat", "cut", "c9t", and "c!t".^— Matches the start of a string (or line in multiline mode).$— Matches the end of a string (or line in multiline mode).*— Matches the preceding element zero or more times.+— Matches the preceding element one or more times.?— Matches the preceding element zero or one time (makes it optional).|— Acts as OR. The patterncat|dogmatches "cat" or "dog".\— Escapes a metacharacter so it matches literally.\.matches an actual dot.[ ]— Defines a character class (a set of characters to match).{ }— Specifies an exact repetition count for the preceding element.( )— Creates a capturing group (groups parts of a pattern together).
If you want to match a metacharacter literally — for example, an actual dot in a filename — you escape it with a backslash: \. matches a period, \* matches an asterisk, and so on.
# Match "index.html" literally (dot is escaped)
index\.html
# Match any character followed by "html"
index.html # also matches "indexXhtml", "index9html"
Key insight: Most regex confusion comes from not knowing which characters are metacharacters. Memorize the 12 listed above and you will immediately understand why patterns that look random actually make perfect sense.
Character Classes: Defining Sets of Characters
Character classes let you match any one character from a defined set. They are written inside square brackets.
# Match any vowel
[aeiou]
# Match any lowercase letter
[a-z]
# Match any digit
[0-9]
# Match any letter (upper or lower) or digit
[a-zA-Z0-9]
The hyphen inside brackets creates a range. [a-z] matches any lowercase letter from a to z. [0-9] matches any digit. You can combine multiple ranges: [a-zA-Z0-9] matches any alphanumeric character.
Negation is done with a caret inside the brackets. [^abc] matches any character that is not a, b, or c. [^0-9] matches any non-digit.
Regex also provides shorthand character classes for the most common sets:
\d— Any digit. Equivalent to[0-9].\D— Any non-digit. Equivalent to[^0-9].\w— Any "word" character: letters, digits, and underscore. Equivalent to[a-zA-Z0-9_].\W— Any non-word character. Equivalent to[^a-zA-Z0-9_].\s— Any whitespace character: space, tab, newline, carriage return.\S— Any non-whitespace character.
# Match a US zip code (5 digits)
\d\d\d\d\d
# Match a word followed by a space followed by a word
\w+\s\w+
These shorthands make patterns dramatically more readable. Instead of writing [0-9][0-9][0-9], you write \d\d\d. And once you learn quantifiers in the next section, you will write \d{3}.
Quantifiers: How Many Times to Match
Quantifiers control how many times the preceding element should be matched. Without quantifiers, each element matches exactly once.
*— Zero or more.ab*cmatches "ac", "abc", "abbc", "abbbc".+— One or more.ab+cmatches "abc", "abbc", but NOT "ac".?— Zero or one (optional).colou?rmatches both "color" and "colour".{n}— Exactly n times.\d{4}matches exactly four digits.{n,}— n or more times.\d{2,}matches two or more digits.{n,m}— Between n and m times.\d{2,4}matches two, three, or four digits.
# Match a US phone number: 3 digits, dash, 3 digits, dash, 4 digits
\d{3}-\d{3}-\d{4}
# Matches: 555-123-4567
# Match "http" or "https"
https?
# The ? makes the "s" optional
# Match one or more whitespace characters
\s+
# Useful for splitting text on whitespace
By default, quantifiers are greedy — they match as much text as possible. Adding a ? after any quantifier makes it lazy (non-greedy), matching as little as possible:
# Greedy: matches from first quote to LAST quote
".*"
# Input: "hello" and "world" → matches: "hello" and "world"
# Lazy: matches from first quote to NEAREST quote
".*?"
# Input: "hello" and "world" → matches: "hello" then "world"
The difference between greedy and lazy becomes critical when your text has multiple potential endpoints. Use lazy quantifiers when you want the shortest possible match.
▶ Build Regex Patterns Visually with Regex BuilderAnchors: Controlling Where Matches Occur
Anchors do not match characters — they match positions in the string. They answer the question: "Where should this pattern appear?"
^— Start of string.^Hellomatches "Hello" only if it appears at the very beginning.$— End of string.world$matches "world" only if it appears at the very end.\b— Word boundary. The position between a word character and a non-word character.
# Match lines that start with a number
^\d
# Match lines that end with a period
\.$
# Match the word "cat" but not "category" or "concatenate"
\bcat\b
# Validate that a string is ONLY digits (nothing else)
^\d+$
The word boundary anchor \b is especially useful. Without it, searching for cat would match inside "category", "scatter", and "concatenate". With \bcat\b, you only match the standalone word "cat".
Anchors are essential for validation. If you want to check that an entire input is a valid email, you need both ^ and $. Without them, the pattern will match a valid email substring inside a longer invalid string.
Groups and Capturing
Parentheses create groups in regex. Groups serve two purposes: they bundle parts of a pattern together, and they capture the matched text so you can reference it later.
Basic capturing groups
# Capture the area code from a phone number
(\d{3})-\d{3}-\d{4}
# Input: 555-123-4567
# Group 1 captures: "555"
# Capture both parts of a date
(\d{4})-(\d{2})-(\d{2})
# Input: 2026-02-10
# Group 1: "2026", Group 2: "02", Group 3: "10"
In JavaScript, captured groups are accessible via the match result array. In Python, use .group(1), .group(2), etc. Captured groups are numbered left to right by their opening parenthesis.
Non-capturing groups
Sometimes you need grouping for structure but do not care about capturing the content. Use (?:pattern) for a non-capturing group:
# Group without capturing (the "?:" means don't capture)
(?:https?://)?(www\.)?example\.com
# Groups "https://" for the ? quantifier but does not capture it
Non-capturing groups are slightly faster and keep your match results cleaner when you only need certain parts of the pattern captured.
Named groups
For complex patterns with many groups, numbered references get confusing fast. Named groups solve this:
# Named groups with (?<name>pattern)
(?<year>\d{4})-(?<month>\d{2})-(?<day>\d{2})
# Access in JS: match.groups.year, match.groups.month, match.groups.day
# Access in Python: match.group("year")
Named groups make your code self-documenting. Anyone reading your regex replacement or extraction logic can immediately understand what each group represents.
Lookahead and Lookbehind
Lookaheads and lookbehinds are zero-width assertions. They check whether a pattern exists ahead of or behind the current position without consuming any characters. Think of them as conditions that must be true for the match to succeed, but they do not become part of the match itself.
(?=pattern)— Positive lookahead. Succeeds ifpatternmatches ahead.(?!pattern)— Negative lookahead. Succeeds ifpatterndoes NOT match ahead.(?<=pattern)— Positive lookbehind. Succeeds ifpatternmatches behind.(?<!pattern)— Negative lookbehind. Succeeds ifpatterndoes NOT match behind.
# Match a number only if followed by "px"
\d+(?=px)
# "12px 5em 8px" → matches: "12" and "8" (not "5")
# Match a word NOT followed by a comma
\w+(?!,)
# Useful for parsing comma-separated values
# Match a number preceded by a dollar sign
(?<=\$)\d+
# "$100 and 200" → matches: "100" (not "200")
# Match a word NOT preceded by "un"
(?<!un)\w+able
# "readable unbreakable" → matches: "readable"
The classic real-world use case for lookaheads is password validation. You can require multiple conditions (uppercase, lowercase, digit, special character) without dictating the order they appear in:
# Password: at least 8 chars, one upper, one lower, one digit
^(?=.*[a-z])(?=.*[A-Z])(?=.*\d).{8,}$
# Each lookahead checks for one requirement independently
Tip: Lookbehinds have length restrictions in some regex engines. JavaScript historically did not support lookbehinds at all, though modern engines (ES2018+) now do. If you need broad compatibility, stick with lookaheads and restructure your pattern.
10 Practical Regex Examples
Theory only takes you so far. Here are ten patterns you will actually use in production, each with a clear explanation. Paste them into the Regex Tester to see them in action.
1. Email address
^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$
Matches standard email addresses. The local part allows letters, digits, dots, underscores, percent, plus, and hyphens. The domain requires at least one dot and a TLD of two or more letters.
2. URL (HTTP and HTTPS)
^https?:\/\/[^\s/$.?#].[^\s]*$
A practical URL pattern that matches HTTP and HTTPS URLs. It avoids being overly strict about domain format, making it more forgiving for real-world URLs with complex paths and query strings.
3. Phone number (flexible)
^\+?[\d\s\-\(\)]{7,15}$
Matches international phone numbers with optional plus sign, digits, spaces, hyphens, and parentheses. The length constraint of 7-15 characters prevents matching random digit strings.
4. IPv4 address
^((25[0-5]|2[0-4]\d|[01]?\d\d?)\.){3}(25[0-5]|2[0-4]\d|[01]?\d\d?)$
Validates each octet is between 0-255. This is more robust than a naive \d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3} which would accept invalid values like 999.
5. Date (YYYY-MM-DD)
^\d{4}-(0[1-9]|1[0-2])-(0[1-9]|[12]\d|3[01])$
Matches ISO 8601 date format. Validates month range (01-12) and day range (01-31). Does not catch invalid dates like February 31st — use a date library for calendar validation.
6. Hex color code
^#([a-fA-F0-9]{6}|[a-fA-F0-9]{3})$
Matches 3-character and 6-character hex colors with a leading hash. Covers standard CSS color codes like #FF5733 and shorthand like #F00.
7. HTML tag (simple)
<([a-z][a-z0-9]*)\b[^>]*>(.*?)<\/\1>
Matches a simple opening and closing HTML tag pair. The \1 backreference ensures the closing tag name matches the opening tag. Not suitable for nested HTML — use a DOM parser for that.
8. File extension filter
^.+\.(jpg|jpeg|png|gif|svg|webp|pdf)$
Matches filenames ending with common image or document extensions. Add the i flag for case-insensitive matching so it also catches .JPG and .PNG.
9. Password strength
^(?=.*[a-z])(?=.*[A-Z])(?=.*\d)(?=.*[@$!%*?&])[A-Za-z\d@$!%*?&]{8,}$
Requires at least 8 characters with a mix of uppercase, lowercase, digit, and special character. Each requirement is checked by a separate lookahead so order does not matter.
10. CSV field (handles quoted fields)
(?:^|,)("(?:[^"]|"")*"|[^,]*)
Matches individual fields in a CSV row, handling both unquoted values and quoted fields that may contain commas or escaped quotes (doubled double-quotes). Captured in group 1.
▶ Test All 10 Patterns Live in the Regex TesterRegex Flags
Flags (also called modifiers) change how the regex engine processes a pattern. They are placed after the closing delimiter in most languages:
// JavaScript
/pattern/gi
# Python
re.compile(r"pattern", re.IGNORECASE | re.MULTILINE)
g— Global. Find all matches, not just the first one. Without this flag, the engine stops after the first match.i— Case-insensitive./hello/imatches "Hello", "HELLO", and "hElLo".m— Multiline. Makes^and$match the start and end of each line, not just the entire string.s— Dotall / single-line. Makes the dot (.) match newline characters too. Without this,.stops at line breaks.u— Unicode. Enables full Unicode matching. In JavaScript, this makes\wand\dwork correctly with characters beyond the ASCII range.
# Find all emails in a document (global + case-insensitive)
/[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}/gi
# Match start of each line in multiline text
/^\d+\./gm
# Finds numbered list items like "1.", "2.", "3." at the start of any line
The g and i flags are by far the most commonly used. The m flag becomes important when working with multiline text like log files or configuration files where you need anchors to work on a per-line basis.
Best Tools for Testing Regex
Writing regex without a testing tool is like coding without a debugger — technically possible, but unnecessarily painful. Here are the tools that make regex development fast and visual.
The QTool Regex Tester lets you paste a pattern and test string, then see matches highlighted in real time. It runs entirely in your browser, requires no sign-up, and supports all JavaScript regex features including flags and named groups.
If you prefer building patterns visually rather than writing raw syntax, the QTool Regex Builder provides a point-and-click interface. Select character classes, quantifiers, and groups from a menu, and the tool assembles the pattern for you. It is ideal for beginners who are still learning the syntax.
For working with text data more broadly — counting words, analyzing character frequency, or checking readability — the Text Analyzer complements your regex workflow by helping you understand the data before you write patterns against it.
Practice tip: The fastest way to internalize regex is to keep a tester open in one tab while you work. Every time you encounter a text manipulation task, try solving it with regex first, even if you end up using a different approach in production.
Common Mistakes Beginners Make
Knowing these pitfalls upfront will save you hours of debugging:
- Forgetting to escape the dot. An unescaped
.matches any character, not a literal period.file.txtalso matches "fileXtxt". Usefile\.txtinstead. - Missing anchors on validation patterns. Without
^and$, a pattern like\d{5}matches "12345" inside "abc12345xyz". Add anchors:^\d{5}$. - Using greedy quantifiers when lazy is needed. The pattern
".*"on input"a" and "b"matches the entire string. Use".*?"for the shortest match. - Not testing edge cases. Always test against empty strings, strings with only whitespace, and strings with special characters. These are where regex patterns break most often.
- Over-complicating patterns. If you need to check whether a string contains "error", use
string.includes("error")in JavaScript. Regex is overkill for simple substring checks.
Wrapping Up: Your Regex Learning Path
Regular expressions are a skill that compounds. Every pattern you write makes the next one easier. Here is the progression that works best:
- Week 1: Literal matches, character classes (
[a-z],\d,\w), and basic quantifiers (*,+,?). - Week 2: Anchors (
^,$,\b), groups (()), and alternation (|). - Week 3: Specific quantifiers (
{n,m}), non-capturing groups ((?:)), and flags. - Week 4: Lookaheads, lookbehinds, named groups, and backreferences.
Keep the Regex Tester open as your permanent companion. Every time you encounter a text pattern in your work — log entries, user input validation, data extraction — try writing the regex. In a month, you will be writing patterns from memory.
For a ready-made collection of battle-tested patterns you can copy and use immediately, check out our companion article: 20 Regex Patterns Every Developer Needs.
Regex is not magic. It is a skill, and like every skill, it becomes second nature with practice. Start simple, build gradually, and let the tools do the heavy lifting while you learn.