1. Getting Started: requests + Beautiful Soup
The simplest Python web scraping stack combines requests for fetching pages and Beautiful Soup for parsing HTML. This works for any site where the data is present in the initial HTML response -- no JavaScript rendering required.
Installation
pip install requests beautifulsoup4 lxml
The lxml parser is optional but recommended. It is significantly faster than Python's built-in html.parser for large documents.
Your First Scraper
import requests
from bs4 import BeautifulSoup
# Fetch the page
url = 'https://example.com/articles'
headers = {
'User-Agent': 'Mozilla/5.0 (compatible; MyBot/1.0)'
}
response = requests.get(url, headers=headers, timeout=10)
response.raise_for_status() # Raise exception for HTTP errors
# Parse the HTML
soup = BeautifulSoup(response.text, 'lxml')
# Extract data
articles = soup.find_all('article', class_='post-card')
for article in articles:
title = article.find('h2').get_text(strip=True)
link = article.find('a')['href']
summary = article.find('p', class_='summary')
summary_text = summary.get_text(strip=True) if summary else 'No summary'
print(f'Title: {title}')
print(f'Link: {link}')
print(f'Summary: {summary_text}')
print('---')
Many websites block requests with Python's default User-Agent string (python-requests/2.x). Setting a descriptive User-Agent that identifies your bot is both polite and practical.
Using Sessions for Multiple Requests
import requests
from bs4 import BeautifulSoup
import time
session = requests.Session()
session.headers.update({
'User-Agent': 'Mozilla/5.0 (compatible; DataCollector/1.0)',
'Accept-Language': 'en-US,en;q=0.9',
})
base_url = 'https://example.com/products?page='
all_products = []
for page in range(1, 11):
response = session.get(f'{base_url}{page}', timeout=10)
soup = BeautifulSoup(response.text, 'lxml')
products = soup.select('.product-card')
for product in products:
name = product.select_one('.product-name').text.strip()
price = product.select_one('.product-price').text.strip()
all_products.append({'name': name, 'price': price})
# Be polite: wait between requests
time.sleep(2)
print(f'Collected {len(all_products)} products')
2. Parsing HTML with CSS Selectors
Beautiful Soup supports CSS selector syntax through the .select() and .select_one() methods. If you already know CSS, these will feel immediately familiar. You can test and refine your selectors with the QTool CSS Selector Tester before writing code.
Common CSS Selector Patterns
# Select by tag
soup.select('p') # All paragraphs
# Select by class
soup.select('.article-title') # Elements with class "article-title"
# Select by ID
soup.select_one('#main-content') # Element with id "main-content"
# Select by attribute
soup.select('a[href^="https"]') # Links starting with https
soup.select('img[alt]') # Images that have an alt attribute
soup.select('input[type="email"]') # Email input fields
# Descendant selectors
soup.select('div.content p') # Paragraphs inside div.content
soup.select('ul.nav > li') # Direct child li of ul.nav
# Pseudo-selectors
soup.select('tr:nth-child(even)') # Even table rows
soup.select('li:first-child') # First li in each list
# Multiple selectors
soup.select('h1, h2, h3') # All heading levels 1-3
# Combining selectors
soup.select('table.data-table tbody tr td:nth-child(2)')
# Second cell in each row of the table body
Extracting Attributes and Text
element = soup.select_one('.article-card')
# Get text content (strips nested tags)
text = element.get_text(strip=True)
# Get text with a separator between nested elements
text = element.get_text(separator=' | ', strip=True)
# Get an attribute
link = element.select_one('a')['href']
image_src = element.select_one('img')['src']
data_id = element['data-id']
# Get attribute with default (avoids KeyError)
alt_text = element.select_one('img').get('alt', 'No alt text')
# Check if element exists before accessing
price_el = element.select_one('.price')
price = price_el.text.strip() if price_el else 'N/A'
3. XPath: When CSS Selectors Are Not Enough
XPath is more verbose than CSS selectors, but it can do things CSS cannot: select elements by their text content, traverse upward to parent elements, and use complex conditional logic.
Beautiful Soup does not support XPath natively. You need lxml directly:
from lxml import html
import requests
response = requests.get('https://example.com/products')
tree = html.fromstring(response.content)
# Select by text content (CSS cannot do this)
buy_buttons = tree.xpath('//button[contains(text(), "Add to Cart")]')
# Select parent of a specific element
price_parents = tree.xpath('//span[@class="price"]/..')
# Select based on position
first_row = tree.xpath('//table/tbody/tr[1]/td/text()')
# Select with multiple conditions
links = tree.xpath(
'//a[starts-with(@href, "/products") and @class="active"]/@href'
)
# Select following sibling
descriptions = tree.xpath(
'//h3[text()="Description"]/following-sibling::p[1]/text()'
)
| Task | CSS Selector | XPath |
|---|---|---|
| Select by class | .classname |
//*[@class="classname"] |
| Select by ID | #myid |
//*[@id="myid"] |
| Select by text content | Not possible | //button[text()="Submit"] |
| Select parent | Not possible | //span[@class="price"]/.. |
| Nth child | li:nth-child(3) |
//li[3] |
4. Handling JavaScript with Playwright
Many modern websites load content via JavaScript after the initial HTML is delivered. If you fetch the page with requests and the data is missing, you need a tool that runs JavaScript. Playwright is the best choice in 2026: it is faster than Selenium, has better API design, and supports Chromium, Firefox, and WebKit.
Installation
pip install playwright
playwright install chromium
Basic Playwright Scraping
from playwright.sync_api import sync_playwright
from bs4 import BeautifulSoup
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
# Navigate and wait for content to load
page.goto('https://example.com/dynamic-content')
page.wait_for_selector('.product-list', timeout=10000)
# Get the rendered HTML
html_content = page.content()
soup = BeautifulSoup(html_content, 'lxml')
# Now parse as usual
products = soup.select('.product-item')
for product in products:
name = product.select_one('.name').text.strip()
price = product.select_one('.price').text.strip()
print(f'{name}: {price}')
browser.close()
Handling Infinite Scroll
from playwright.sync_api import sync_playwright
import time
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto('https://example.com/feed')
previous_height = 0
max_scrolls = 10
for i in range(max_scrolls):
# Scroll to bottom
page.evaluate('window.scrollTo(0, document.body.scrollHeight)')
time.sleep(2) # Wait for content to load
# Check if page grew
current_height = page.evaluate('document.body.scrollHeight')
if current_height == previous_height:
break # No new content loaded
previous_height = current_height
print(f'Scroll {i + 1}: page height = {current_height}px')
# Now extract all loaded content
items = page.query_selector_all('.feed-item')
print(f'Found {len(items)} items after scrolling')
browser.close()
Playwright launches an actual browser, which uses significantly more memory and CPU than requests. Only use it when JavaScript rendering is required. For static HTML pages, requests + Beautiful Soup is 10-50x faster.
5. Scraping Tables, Lists, and Structured Data
Tables are one of the most common scraping targets. Beautiful Soup handles them well, but pandas can be even faster for tabular data.
With Beautiful Soup
table = soup.select_one('table.data-table')
headers = [th.text.strip() for th in table.select('thead th')]
rows = []
for tr in table.select('tbody tr'):
cells = [td.text.strip() for td in tr.select('td')]
rows.append(dict(zip(headers, cells)))
# rows is now a list of dictionaries
for row in rows:
print(row)
With pandas (One-Liner)
import pandas as pd
# Reads ALL tables on the page into a list of DataFrames
tables = pd.read_html('https://example.com/statistics')
df = tables[0] # First table
# Export to CSV
df.to_csv('output.csv', index=False)
print(df.head())
When working with HTML structures, the HTML Live Preview tool lets you paste HTML and inspect its structure visually, which helps you identify the right selectors before writing code.
6. Avoiding Blocks and Detection
Websites use various techniques to detect and block scrapers. Here are the most common defenses and how to work within them responsibly.
Essential Anti-Detection Techniques
import requests
import random
import time
# Rotate User-Agent strings
USER_AGENTS = [
'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36',
'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36',
'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36',
]
session = requests.Session()
def polite_get(url):
"""Fetch a URL with randomized delays and headers."""
session.headers.update({
'User-Agent': random.choice(USER_AGENTS),
'Accept': 'text/html,application/xhtml+xml',
'Accept-Language': 'en-US,en;q=0.9',
'Accept-Encoding': 'gzip, deflate, br',
'Connection': 'keep-alive',
})
# Random delay between 1 and 3 seconds
time.sleep(random.uniform(1.0, 3.0))
response = session.get(url, timeout=15)
response.raise_for_status()
return response
- Add delays: Minimum 1-2 seconds between requests. Randomize the interval.
- Rotate User-Agents: Do not send the same User-Agent string every time.
- Use sessions: Maintain cookies like a real browser would.
- Set realistic headers: Include Accept, Accept-Language, and Referer headers.
- Respect
robots.txt: Always check and follow the site's crawling rules. - Handle errors gracefully: Back off when you receive 429 or 503 status codes.
7. robots.txt, Legal Considerations, and Ethics
Before scraping any website, you have both legal and ethical responsibilities. Ignoring these can result in IP bans, legal action, or worse.
Checking robots.txt
from urllib.robotparser import RobotFileParser
rp = RobotFileParser()
rp.set_url('https://example.com/robots.txt')
rp.read()
# Check if scraping a specific path is allowed
url = 'https://example.com/products/all'
user_agent = 'MyBot/1.0'
if rp.can_fetch(user_agent, url):
print('Allowed to scrape this URL')
else:
print('Blocked by robots.txt -- do not scrape')
# Check crawl delay
delay = rp.crawl_delay(user_agent)
if delay:
print(f'Requested crawl delay: {delay} seconds')
Legal Guidelines
- Public data: Scraping publicly accessible data is generally permissible, but always check the site's Terms of Service.
- Personal data: Scraping personal information (names, emails, profiles) requires careful consideration under GDPR, CCPA, and similar regulations.
- Copyrighted content: Do not scrape and republish copyrighted text, images, or media.
- Rate of access: Even if scraping is allowed, overwhelming a server can be considered a denial-of-service attack.
- Authentication boundaries: Scraping behind a login (especially by circumventing access controls) raises serious legal risks.
Before scraping, always check if the website offers a public API. APIs are more reliable, structured, and legal. Use the JSON Editor to inspect API responses when you find one.
Ethical Scraping Checklist
- Read and respect
robots.txt - Read the website's Terms of Service
- Add reasonable delays between requests
- Identify your bot with a descriptive User-Agent
- Do not scrape personal data without a legitimate purpose
- Cache responses to avoid redundant requests
- Contact the site owner if you need large-scale access
8. Choosing the Right Tool
| Tool | Best For | JS Support | Speed | Learning Curve |
|---|---|---|---|---|
| requests + BS4 | Static HTML pages | No | Very Fast | Easy |
| Playwright | JS-rendered pages, SPAs | Full | Moderate | Medium |
| Selenium | Legacy projects, older browsers | Full | Slow | Medium |
| Scrapy | Large-scale crawling | Via plugins | Fast (async) | Steep |
| pandas.read_html | HTML tables only | No | Very Fast | Easy |
Start with requests + Beautiful Soup. Move to Playwright only when you need JavaScript rendering. Move to Scrapy only when you need to crawl thousands of pages efficiently.
Test your CSS selectors and regex patterns before coding with the Regex Builder -- it helps you validate extraction patterns without running your full scraper.
9. Frequently Asked Questions
Is web scraping legal in 2026?
Web scraping of publicly available data is generally legal in most jurisdictions, but there are important caveats. You must respect robots.txt directives, comply with the website's terms of service, avoid scraping personal or copyrighted data without permission, and follow data protection regulations like GDPR. The 2022 US Ninth Circuit ruling in hiQ v. LinkedIn confirmed that scraping public data does not violate the Computer Fraud and Abuse Act, but this does not give blanket permission for all scraping activities.
When should I use Playwright instead of Beautiful Soup for scraping?
Use Playwright (or Selenium) when the website relies on JavaScript to render content. If you view the page source and the data you need is not present in the raw HTML but only appears after JavaScript executes, you need a browser automation tool like Playwright. Beautiful Soup with requests is faster and lighter, so use it when the data is available in the initial HTML response. A quick test: use curl or requests to fetch the page and check if the data is in the response.
What is the difference between CSS selectors and XPath for scraping?
CSS selectors use the same syntax as CSS stylesheets (e.g., div.class, #id, [attribute]) and are generally simpler and faster. XPath uses path expressions to navigate XML/HTML documents and is more powerful for complex selections like selecting by text content, navigating to parent elements, or using conditional logic. For most scraping tasks CSS selectors are sufficient and easier to read. Use XPath when you need to select elements based on their text content or traverse upward in the document tree.
How do I avoid getting blocked while web scraping?
To avoid being blocked: add delays between requests (1-3 seconds minimum), rotate User-Agent strings, respect robots.txt directives, use session objects to maintain cookies, limit concurrent requests, avoid scraping during peak hours, and consider using residential proxies for large-scale projects. Most importantly, check if the site offers an official API before scraping, as APIs are more reliable and less likely to break.
How do I parse HTML tables with Beautiful Soup?
To parse HTML tables with Beautiful Soup, find the table element using soup.find('table') or soup.select('table.classname'), then iterate over rows with table.find_all('tr'). For each row, extract cells with row.find_all(['td', 'th']). For a quicker approach with tabular data, use pandas.read_html(url) which automatically finds and parses all tables on a page into DataFrames. Beautiful Soup gives you more control over which table and which cells to extract.