Basic Syntax
A list comprehension creates a new list by applying an expression to each item in an iterable. The general form is:
[expression for item in iterable]
This is equivalent to:
result = []
for item in iterable:
result.append(expression)
Here are the simplest practical examples:
# Square every number
squares = [x**2 for x in range(10)]
# [0, 1, 4, 9, 16, 25, 36, 49, 64, 81]
# Convert strings to uppercase
names = ["alice", "bob", "charlie"]
upper = [name.upper() for name in names]
# ["ALICE", "BOB", "CHARLIE"]
# Get lengths of strings
lengths = [len(name) for name in names]
# [5, 3, 7]
# Strip whitespace from each string
raw = [" hello ", " world ", " python "]
clean = [s.strip() for s in raw]
# ["hello", "world", "python"]
# Convert a list of strings to integers
str_nums = ["1", "2", "3", "4", "5"]
nums = [int(s) for s in str_nums]
# [1, 2, 3, 4, 5]
When moving examples between languages, the Code Converter provides a useful starting point. Review generated code before using it in production.
Filtering with Conditionals
Add an if clause at the end to include only items that meet a condition:
# Even numbers only
evens = [x for x in range(20) if x % 2 == 0]
# [0, 2, 4, 6, 8, 10, 12, 14, 16, 18]
# Words longer than 3 characters
words = ["the", "quick", "brown", "fox", "jumps"]
long_words = [w for w in words if len(w) > 3]
# ["quick", "brown", "jumps"]
# Filter out None values
data = [1, None, 3, None, 5, None]
clean = [x for x in data if x is not None]
# [1, 3, 5]
# Filter by type
mixed = [1, "hello", 2.5, "world", 3]
strings_only = [x for x in mixed if isinstance(x, str)]
# ["hello", "world"]
# Files with a specific extension
files = ["app.py", "style.css", "index.html", "utils.py"]
python_files = [f for f in files if f.endswith(".py")]
# ["app.py", "utils.py"]
# Multiple conditions
nums = range(100)
special = [x for x in nums if x % 3 == 0 if x % 5 == 0]
# [0, 15, 30, 45, 60, 75, 90] (divisible by both 3 and 5)
Conditional Expressions (if-else)
When you want to transform every item but with different logic depending on a condition, put the conditional before the for:
# Replace negatives with zero
nums = [4, -1, 7, -3, 2, -8]
clamped = [x if x > 0 else 0 for x in nums]
# [4, 0, 7, 0, 2, 0]
# Label numbers as even or odd
labels = ["even" if x % 2 == 0 else "odd" for x in range(6)]
# ["even", "odd", "even", "odd", "even", "odd"]
# Grade categories
scores = [92, 67, 85, 43, 78]
grades = ["pass" if s >= 60 else "fail" for s in scores]
# ["pass", "pass", "pass", "fail", "pass"]
# Combine filter AND transform
data = [5, -2, None, 8, -1, None, 3]
result = [x * 2 if x > 0 else 0 for x in data if x is not None]
# [10, 0, 16, 0, 6]
[x for x in items if condition] — filters the list (fewer items in output). [a if condition else b for x in items] — transforms every item (same number of items in output). The if position is what matters.
Nested Comprehensions
You can nest for clauses to iterate over multiple iterables. The order follows the same order as nested for loops — outer loop first, inner loop second.
# Flatten a 2D list
matrix = [[1, 2, 3], [4, 5, 6], [7, 8, 9]]
flat = [num for row in matrix for num in row]
# [1, 2, 3, 4, 5, 6, 7, 8, 9]
# All combinations of two lists
colors = ["red", "blue"]
sizes = ["S", "M", "L"]
combos = [(c, s) for c in colors for s in sizes]
# [("red", "S"), ("red", "M"), ("red", "L"),
# ("blue", "S"), ("blue", "M"), ("blue", "L")]
# Transpose a matrix
matrix = [[1, 2, 3], [4, 5, 6], [7, 8, 9]]
transposed = [[row[i] for row in matrix] for i in range(3)]
# [[1, 4, 7], [2, 5, 8], [3, 6, 9]]
# Flatten with condition
matrix = [[1, -2, 3], [-4, 5, -6], [7, -8, 9]]
positives = [n for row in matrix for n in row if n > 0]
# [1, 3, 5, 7, 9]
If your comprehension needs more than two levels of nesting, rewrite it as a regular for loop. The goal of comprehensions is clarity, not code golf. A comprehension that takes 30 seconds to parse is worse than a 5-line loop.
Dictionary Comprehensions
Dictionary comprehensions use curly braces with a key: value pair:
# Map names to lengths
names = ["Alice", "Bob", "Charlie"]
name_lengths = {name: len(name) for name in names}
# {"Alice": 5, "Bob": 3, "Charlie": 7}
# Invert a dictionary (swap keys and values)
original = {"a": 1, "b": 2, "c": 3}
inverted = {v: k for k, v in original.items()}
# {1: "a", 2: "b", 3: "c"}
# Filter dictionary entries
scores = {"Alice": 92, "Bob": 45, "Charlie": 78, "Diana": 38}
passing = {k: v for k, v in scores.items() if v >= 60}
# {"Alice": 92, "Charlie": 78}
# Transform values
prices = {"apple": 1.20, "banana": 0.50, "cherry": 2.00}
doubled = {k: round(v * 2, 2) for k, v in prices.items()}
# {"apple": 2.40, "banana": 1.00, "cherry": 4.00}
# Create dict from two lists
keys = ["name", "age", "city"]
values = ["Alice", 30, "NYC"]
person = {k: v for k, v in zip(keys, values)}
# {"name": "Alice", "age": 30, "city": "NYC"}
# Word frequency counter
text = "the cat sat on the mat the cat"
words = text.split()
freq = {w: words.count(w) for w in set(words)}
# {"the": 3, "cat": 2, "sat": 1, "on": 1, "mat": 1}
Set Comprehensions
Set comprehensions use curly braces with a single expression. The result is a set with no duplicates:
# Unique word lengths
words = ["hello", "world", "python", "code", "is", "great"]
lengths = {len(w) for w in words}
# {2, 4, 5, 6}
# Unique first characters
names = ["Alice", "Bob", "Anna", "Charlie", "Adam"]
first_chars = {name[0] for name in names}
# {"A", "B", "C"}
# Unique file extensions
files = ["app.py", "style.css", "main.py", "index.html", "util.py"]
extensions = {f.split(".")[-1] for f in files}
# {"py", "css", "html"}
Generator Expressions
Replace square brackets with parentheses to create a generator that produces items on demand instead of building the full list in memory:
# Sum of squares without building a list
total = sum(x**2 for x in range(1000000))
# Join strings efficiently
names = ["Alice", "Bob", "Charlie"]
csv = ", ".join(name.upper() for name in names)
# "ALICE, BOB, CHARLIE"
# Check if any item matches
numbers = [2, 4, 6, 7, 8]
has_odd = any(x % 2 != 0 for x in numbers)
# True
# Find first match
data = [10, 20, 35, 40, 55]
first_odd = next(x for x in data if x % 2 != 0)
# 35
# Memory comparison
import sys
list_comp = [x**2 for x in range(10000)]
gen_expr = (x**2 for x in range(10000))
print(sys.getsizeof(list_comp)) # ~87,616 bytes
print(sys.getsizeof(gen_expr)) # ~200 bytes
Performance: Comprehensions vs. Loops
List comprehensions are not just syntactic sugar. They generate optimized bytecode that avoids the overhead of repeated list.append() method lookups.
import timeit
# For loop approach
def loop_squares():
result = []
for x in range(1000):
result.append(x ** 2)
return result
# List comprehension approach
def comp_squares():
return [x ** 2 for x in range(1000)]
# map() approach
def map_squares():
return list(map(lambda x: x ** 2, range(1000)))
# Typical results (Python 3.12):
# loop_squares: ~85 microseconds
# comp_squares: ~62 microseconds (27% faster)
# map_squares: ~70 microseconds (18% faster)
For lists under 100 items, the performance difference is invisible. Choose whichever form is most readable. For large lists (10,000+ items), comprehensions provide a meaningful speed boost. For very large datasets where you only need to iterate once, use a generator expression.
30+ Real-World Examples
Data Cleaning
# Remove empty strings
data = ["hello", "", "world", "", "python"]
clean = [s for s in data if s]
# ["hello", "world", "python"]
# Parse CSV-like data
raw = "Alice,30,NYC;Bob,25,LA;Charlie,35,Chicago"
records = [line.split(",") for line in raw.split(";")]
# [["Alice", "30", "NYC"], ["Bob", "25", "LA"], ...]
# Extract emails from text
import re
text = "Contact alice@mail.com or bob@work.org"
emails = [w for w in text.split() if "@" in w]
# ["alice@mail.com", "bob@work.org"]
# Normalize phone numbers (digits only)
phones = ["(555) 123-4567", "555.987.6543", "555-111-2222"]
normalized = ["".join(c for c in p if c.isdigit()) for p in phones]
# ["5551234567", "5559876543", "5551112222"]
File and Path Operations
from pathlib import Path
# List all Python files in a directory
py_files = [f.name for f in Path(".").glob("*.py")]
# Read lines from a file, stripped
with open("data.txt") as f:
lines = [line.strip() for line in f if line.strip()]
# Get file sizes
files = list(Path(".").glob("*"))
sizes = {f.name: f.stat().st_size for f in files if f.is_file()}
API and JSON Processing
# Extract specific fields from API response
users = [
{"id": 1, "name": "Alice", "active": True},
{"id": 2, "name": "Bob", "active": False},
{"id": 3, "name": "Charlie", "active": True},
]
# Get names of active users
active_names = [u["name"] for u in users if u["active"]]
# ["Alice", "Charlie"]
# Create lookup dictionary
user_lookup = {u["id"]: u["name"] for u in users}
# {1: "Alice", 2: "Bob", 3: "Charlie"}
# Flatten nested JSON
orders = [
{"id": 1, "items": ["apple", "banana"]},
{"id": 2, "items": ["cherry", "date"]},
]
all_items = [item for order in orders for item in order["items"]]
# ["apple", "banana", "cherry", "date"]
Math and Number Processing
# Fibonacci-like: pairwise sums
nums = [1, 1, 2, 3, 5, 8, 13]
pairwise = [nums[i] + nums[i+1] for i in range(len(nums)-1)]
# [2, 3, 5, 8, 13, 21]
# Running average
data = [10, 20, 30, 40, 50]
running_avg = [sum(data[:i+1])/(i+1) for i in range(len(data))]
# [10.0, 15.0, 20.0, 25.0, 30.0]
# Prime numbers (simple sieve)
primes = [n for n in range(2, 100)
if all(n % d != 0 for d in range(2, int(n**0.5)+1))]
# [2, 3, 5, 7, 11, 13, ...]
# Matrix multiplication (dot product of row and column)
A = [[1, 2], [3, 4]]
B = [[5, 6], [7, 8]]
result = [[sum(a*b for a, b in zip(row, col))
for col in zip(*B)] for row in A]
# [[19, 22], [43, 50]]
Common Pitfalls and How to Avoid Them
1. Side Effects in Comprehensions
# BAD: Using comprehension for side effects
[print(x) for x in range(5)] # Creates an unnecessary list of None values
# GOOD: Use a regular loop
for x in range(5):
print(x)
2. Variable Leaking (Python 2 only)
In Python 3, comprehension variables are scoped to the comprehension. This was a bug in Python 2 where [x for x in range(5)] would leave x = 4 in the outer scope. No action needed in modern Python.
3. Overly Complex Comprehensions
# BAD: Too complex to read
result = [f(x) for x in [g(y) for y in data if h(y)] if p(x) and q(x)]
# GOOD: Break it into steps
filtered = [g(y) for y in data if h(y)]
result = [f(x) for x in filtered if p(x) and q(x)]
4. Walrus Operator for Reuse (:=)
Python 3.8+ lets you assign and use a value in the same expression:
# Compute once, filter and use
results = [y for x in data if (y := expensive_function(x)) > threshold]
Related Free Tools
Frequently Asked Questions
Yes, list comprehensions are generally 10-30% faster than equivalent for loops for creating lists. The speed advantage comes from the optimized bytecode Python generates for comprehensions, which avoids the overhead of repeated list.append() method lookups and calls. However, the difference is negligible for small lists. For very large datasets where memory is a concern, consider using a generator expression instead, which processes items one at a time without building the entire list in memory.
Avoid list comprehensions when the logic requires more than two levels of nesting or multiple conditions that make the expression hard to read. If a comprehension exceeds roughly 80 characters or requires a mental pause to understand, rewrite it as a regular for loop. Also avoid comprehensions when you do not need the resulting list, such as when calling functions purely for side effects like printing or writing to a file. In those cases, a regular for loop is clearer. List comprehensions should make code more readable, not less.
A list comprehension uses square brackets [x for x in items] and builds the entire list in memory at once. A generator expression uses parentheses (x for x in items) and produces items one at a time on demand, using almost no memory regardless of the input size. Use a list comprehension when you need to access items multiple times, slice the result, or check its length. Use a generator expression when you only iterate through the items once, especially with large datasets or when passing directly to functions like sum(), min(), max(), or join(). For example, sum(x*x for x in range(1000000)) uses constant memory while sum([x*x for x in range(1000000)]) builds a million-element list first.
Yes, but the syntax depends on whether you are filtering or transforming. For filtering (only including items that meet a condition), the if goes at the end: [x for x in items if x > 0]. For transforming (choosing between two values for every item), use a ternary expression at the beginning: [x if x > 0 else 0 for x in items]. You can combine both: [x*2 if x > 0 else 0 for x in items if x != None] first filters out None values, then doubles positive numbers and replaces non-positive with zero.
Dict comprehensions use curly braces with a key: value pair: {k: v for k, v in items}. Set comprehensions use curly braces with a single expression: {x for x in items}. Both support the same filtering and nesting as list comprehensions. For example, {name: len(name) for name in names} creates a dictionary mapping names to their lengths. {word.lower() for word in words} creates a set of unique lowercase words. Dict comprehensions are commonly used to invert dictionaries ({v: k for k, v in d.items()}), filter dictionary entries ({k: v for k, v in d.items() if v > 0}), or transform values ({k: v.strip() for k, v in d.items()}).