Web Scraping with Python in 2026: Beautiful Soup, Playwright, and Ethics

A practical guide to extracting data from websites using Python. From basic HTML parsing to handling JavaScript-heavy pages, with a serious look at legal and ethical boundaries.

Table of Contents
  1. Getting Started: requests + Beautiful Soup
  2. Parsing HTML with CSS Selectors
  3. XPath: When CSS Selectors Are Not Enough
  4. Handling JavaScript with Playwright
  5. Scraping Tables, Lists, and Structured Data
  6. Avoiding Blocks and Detection
  7. robots.txt, Legal Considerations, and Ethics
  8. Choosing the Right Tool
  9. Frequently Asked Questions

1. Getting Started: requests + Beautiful Soup

The simplest Python web scraping stack combines requests for fetching pages and Beautiful Soup for parsing HTML. This works for any site where the data is present in the initial HTML response -- no JavaScript rendering required.

Installation

Terminal
pip install requests beautifulsoup4 lxml

The lxml parser is optional but recommended. It is significantly faster than Python's built-in html.parser for large documents.

Your First Scraper

Python
import requests
from bs4 import BeautifulSoup

# Fetch the page
url = 'https://example.com/articles'
headers = {
    'User-Agent': 'Mozilla/5.0 (compatible; MyBot/1.0)'
}
response = requests.get(url, headers=headers, timeout=10)
response.raise_for_status()  # Raise exception for HTTP errors

# Parse the HTML
soup = BeautifulSoup(response.text, 'lxml')

# Extract data
articles = soup.find_all('article', class_='post-card')

for article in articles:
    title = article.find('h2').get_text(strip=True)
    link = article.find('a')['href']
    summary = article.find('p', class_='summary')
    summary_text = summary.get_text(strip=True) if summary else 'No summary'

    print(f'Title: {title}')
    print(f'Link: {link}')
    print(f'Summary: {summary_text}')
    print('---')
Always Set a User-Agent

Many websites block requests with Python's default User-Agent string (python-requests/2.x). Setting a descriptive User-Agent that identifies your bot is both polite and practical.

Using Sessions for Multiple Requests

Python
import requests
from bs4 import BeautifulSoup
import time

session = requests.Session()
session.headers.update({
    'User-Agent': 'Mozilla/5.0 (compatible; DataCollector/1.0)',
    'Accept-Language': 'en-US,en;q=0.9',
})

base_url = 'https://example.com/products?page='
all_products = []

for page in range(1, 11):
    response = session.get(f'{base_url}{page}', timeout=10)
    soup = BeautifulSoup(response.text, 'lxml')

    products = soup.select('.product-card')
    for product in products:
        name = product.select_one('.product-name').text.strip()
        price = product.select_one('.product-price').text.strip()
        all_products.append({'name': name, 'price': price})

    # Be polite: wait between requests
    time.sleep(2)

print(f'Collected {len(all_products)} products')

2. Parsing HTML with CSS Selectors

Beautiful Soup supports CSS selector syntax through the .select() and .select_one() methods. If you already know CSS, these will feel immediately familiar. You can test and refine your selectors with the QTool CSS Selector Tester before writing code.

Common CSS Selector Patterns

Python
# Select by tag
soup.select('p')                    # All paragraphs

# Select by class
soup.select('.article-title')       # Elements with class "article-title"

# Select by ID
soup.select_one('#main-content')    # Element with id "main-content"

# Select by attribute
soup.select('a[href^="https"]')     # Links starting with https
soup.select('img[alt]')             # Images that have an alt attribute
soup.select('input[type="email"]')  # Email input fields

# Descendant selectors
soup.select('div.content p')        # Paragraphs inside div.content
soup.select('ul.nav > li')          # Direct child li of ul.nav

# Pseudo-selectors
soup.select('tr:nth-child(even)')   # Even table rows
soup.select('li:first-child')       # First li in each list

# Multiple selectors
soup.select('h1, h2, h3')          # All heading levels 1-3

# Combining selectors
soup.select('table.data-table tbody tr td:nth-child(2)')
# Second cell in each row of the table body

Extracting Attributes and Text

Python
element = soup.select_one('.article-card')

# Get text content (strips nested tags)
text = element.get_text(strip=True)

# Get text with a separator between nested elements
text = element.get_text(separator=' | ', strip=True)

# Get an attribute
link = element.select_one('a')['href']
image_src = element.select_one('img')['src']
data_id = element['data-id']

# Get attribute with default (avoids KeyError)
alt_text = element.select_one('img').get('alt', 'No alt text')

# Check if element exists before accessing
price_el = element.select_one('.price')
price = price_el.text.strip() if price_el else 'N/A'

3. XPath: When CSS Selectors Are Not Enough

XPath is more verbose than CSS selectors, but it can do things CSS cannot: select elements by their text content, traverse upward to parent elements, and use complex conditional logic.

Beautiful Soup does not support XPath natively. You need lxml directly:

Python
from lxml import html
import requests

response = requests.get('https://example.com/products')
tree = html.fromstring(response.content)

# Select by text content (CSS cannot do this)
buy_buttons = tree.xpath('//button[contains(text(), "Add to Cart")]')

# Select parent of a specific element
price_parents = tree.xpath('//span[@class="price"]/..')

# Select based on position
first_row = tree.xpath('//table/tbody/tr[1]/td/text()')

# Select with multiple conditions
links = tree.xpath(
    '//a[starts-with(@href, "/products") and @class="active"]/@href'
)

# Select following sibling
descriptions = tree.xpath(
    '//h3[text()="Description"]/following-sibling::p[1]/text()'
)
Task CSS Selector XPath
Select by class .classname //*[@class="classname"]
Select by ID #myid //*[@id="myid"]
Select by text content Not possible //button[text()="Submit"]
Select parent Not possible //span[@class="price"]/..
Nth child li:nth-child(3) //li[3]

4. Handling JavaScript with Playwright

Many modern websites load content via JavaScript after the initial HTML is delivered. If you fetch the page with requests and the data is missing, you need a tool that runs JavaScript. Playwright is the best choice in 2026: it is faster than Selenium, has better API design, and supports Chromium, Firefox, and WebKit.

Installation

Terminal
pip install playwright
playwright install chromium

Basic Playwright Scraping

Python
from playwright.sync_api import sync_playwright
from bs4 import BeautifulSoup

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()

    # Navigate and wait for content to load
    page.goto('https://example.com/dynamic-content')
    page.wait_for_selector('.product-list', timeout=10000)

    # Get the rendered HTML
    html_content = page.content()
    soup = BeautifulSoup(html_content, 'lxml')

    # Now parse as usual
    products = soup.select('.product-item')
    for product in products:
        name = product.select_one('.name').text.strip()
        price = product.select_one('.price').text.strip()
        print(f'{name}: {price}')

    browser.close()

Handling Infinite Scroll

Python
from playwright.sync_api import sync_playwright
import time

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto('https://example.com/feed')

    previous_height = 0
    max_scrolls = 10

    for i in range(max_scrolls):
        # Scroll to bottom
        page.evaluate('window.scrollTo(0, document.body.scrollHeight)')
        time.sleep(2)  # Wait for content to load

        # Check if page grew
        current_height = page.evaluate('document.body.scrollHeight')
        if current_height == previous_height:
            break  # No new content loaded
        previous_height = current_height
        print(f'Scroll {i + 1}: page height = {current_height}px')

    # Now extract all loaded content
    items = page.query_selector_all('.feed-item')
    print(f'Found {len(items)} items after scrolling')

    browser.close()
Performance Note

Playwright launches an actual browser, which uses significantly more memory and CPU than requests. Only use it when JavaScript rendering is required. For static HTML pages, requests + Beautiful Soup is 10-50x faster.

5. Scraping Tables, Lists, and Structured Data

Tables are one of the most common scraping targets. Beautiful Soup handles them well, but pandas can be even faster for tabular data.

With Beautiful Soup

Python
table = soup.select_one('table.data-table')
headers = [th.text.strip() for th in table.select('thead th')]
rows = []

for tr in table.select('tbody tr'):
    cells = [td.text.strip() for td in tr.select('td')]
    rows.append(dict(zip(headers, cells)))

# rows is now a list of dictionaries
for row in rows:
    print(row)

With pandas (One-Liner)

Python
import pandas as pd

# Reads ALL tables on the page into a list of DataFrames
tables = pd.read_html('https://example.com/statistics')
df = tables[0]  # First table

# Export to CSV
df.to_csv('output.csv', index=False)
print(df.head())

When working with HTML structures, the HTML Live Preview tool lets you paste HTML and inspect its structure visually, which helps you identify the right selectors before writing code.

6. Avoiding Blocks and Detection

Websites use various techniques to detect and block scrapers. Here are the most common defenses and how to work within them responsibly.

Essential Anti-Detection Techniques

Python
import requests
import random
import time

# Rotate User-Agent strings
USER_AGENTS = [
    'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36',
    'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36',
    'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36',
]

session = requests.Session()

def polite_get(url):
    """Fetch a URL with randomized delays and headers."""
    session.headers.update({
        'User-Agent': random.choice(USER_AGENTS),
        'Accept': 'text/html,application/xhtml+xml',
        'Accept-Language': 'en-US,en;q=0.9',
        'Accept-Encoding': 'gzip, deflate, br',
        'Connection': 'keep-alive',
    })

    # Random delay between 1 and 3 seconds
    time.sleep(random.uniform(1.0, 3.0))

    response = session.get(url, timeout=15)
    response.raise_for_status()
    return response

7. robots.txt, Legal Considerations, and Ethics

Before scraping any website, you have both legal and ethical responsibilities. Ignoring these can result in IP bans, legal action, or worse.

Checking robots.txt

Python
from urllib.robotparser import RobotFileParser

rp = RobotFileParser()
rp.set_url('https://example.com/robots.txt')
rp.read()

# Check if scraping a specific path is allowed
url = 'https://example.com/products/all'
user_agent = 'MyBot/1.0'

if rp.can_fetch(user_agent, url):
    print('Allowed to scrape this URL')
else:
    print('Blocked by robots.txt -- do not scrape')

# Check crawl delay
delay = rp.crawl_delay(user_agent)
if delay:
    print(f'Requested crawl delay: {delay} seconds')

Legal Guidelines

Check for APIs First

Before scraping, always check if the website offers a public API. APIs are more reliable, structured, and legal. Use the JSON Editor to inspect API responses when you find one.

Ethical Scraping Checklist

  1. Read and respect robots.txt
  2. Read the website's Terms of Service
  3. Add reasonable delays between requests
  4. Identify your bot with a descriptive User-Agent
  5. Do not scrape personal data without a legitimate purpose
  6. Cache responses to avoid redundant requests
  7. Contact the site owner if you need large-scale access

8. Choosing the Right Tool

Tool Best For JS Support Speed Learning Curve
requests + BS4 Static HTML pages No Very Fast Easy
Playwright JS-rendered pages, SPAs Full Moderate Medium
Selenium Legacy projects, older browsers Full Slow Medium
Scrapy Large-scale crawling Via plugins Fast (async) Steep
pandas.read_html HTML tables only No Very Fast Easy

Start with requests + Beautiful Soup. Move to Playwright only when you need JavaScript rendering. Move to Scrapy only when you need to crawl thousands of pages efficiently.

Test your CSS selectors and regex patterns before coding with the Regex Builder -- it helps you validate extraction patterns without running your full scraper.

9. Frequently Asked Questions

Is web scraping legal in 2026?

Web scraping of publicly available data is generally legal in most jurisdictions, but there are important caveats. You must respect robots.txt directives, comply with the website's terms of service, avoid scraping personal or copyrighted data without permission, and follow data protection regulations like GDPR. The 2022 US Ninth Circuit ruling in hiQ v. LinkedIn confirmed that scraping public data does not violate the Computer Fraud and Abuse Act, but this does not give blanket permission for all scraping activities.

When should I use Playwright instead of Beautiful Soup for scraping?

Use Playwright (or Selenium) when the website relies on JavaScript to render content. If you view the page source and the data you need is not present in the raw HTML but only appears after JavaScript executes, you need a browser automation tool like Playwright. Beautiful Soup with requests is faster and lighter, so use it when the data is available in the initial HTML response. A quick test: use curl or requests to fetch the page and check if the data is in the response.

What is the difference between CSS selectors and XPath for scraping?

CSS selectors use the same syntax as CSS stylesheets (e.g., div.class, #id, [attribute]) and are generally simpler and faster. XPath uses path expressions to navigate XML/HTML documents and is more powerful for complex selections like selecting by text content, navigating to parent elements, or using conditional logic. For most scraping tasks CSS selectors are sufficient and easier to read. Use XPath when you need to select elements based on their text content or traverse upward in the document tree.

How do I avoid getting blocked while web scraping?

To avoid being blocked: add delays between requests (1-3 seconds minimum), rotate User-Agent strings, respect robots.txt directives, use session objects to maintain cookies, limit concurrent requests, avoid scraping during peak hours, and consider using residential proxies for large-scale projects. Most importantly, check if the site offers an official API before scraping, as APIs are more reliable and less likely to break.

How do I parse HTML tables with Beautiful Soup?

To parse HTML tables with Beautiful Soup, find the table element using soup.find('table') or soup.select('table.classname'), then iterate over rows with table.find_all('tr'). For each row, extract cells with row.find_all(['td', 'th']). For a quicker approach with tabular data, use pandas.read_html(url) which automatically finds and parses all tables on a page into DataFrames. Beautiful Soup gives you more control over which table and which cells to extract.

269 Free Developer Tools

HTML previews, regex builders, JSON editors -- all browser-based, no signup.

Browse All Tools

Related Articles

Built by Miguel

Need a custom tool or website?

From . Delivered in 24-48h. You own the code.

View Services →