别再用requests了!用Python 3.11+的httpx和BeautifulSoup4爬取豆瓣电影Top250(附完整代码)
用Python 3.11+的httpx和BeautifulSoup4高效爬取豆瓣电影Top250
在Python爬虫领域,技术栈的迭代速度令人目不暇接。十年前流行的urllib2如今已被更现代、更高效的库所取代。本文将带你使用Python 3.11+的最新特性,结合httpx和BeautifulSoup4这两个强力工具,打造一个高效、稳定的豆瓣电影Top250爬虫。
1. 为什么选择httpx和BeautifulSoup4组合
1.1 httpx vs requests:现代HTTP客户端的优势
httpx是Python生态中新兴的HTTP客户端库,相比传统的requests,它带来了几项关键改进:
- 原生支持HTTP/2:显著提升连接效率,特别是在需要大量请求的场景下
- 完整的异步支持:与
asyncio无缝集成,轻松实现高性能并发爬取 - 更智能的连接池:自动管理连接复用,减少TCP握手开销
- 更严格的类型提示:充分利用Python 3.11+的类型系统,提高代码健壮性
import httpx
# 同步请求示例
with httpx.Client() as client:
response = client.get("https://movie.douban.com/top250")
# 异步请求示例
async with httpx.AsyncClient() as client:
response = await client.get("https://movie.douban.com/top250")
1.2 BeautifulSoup4的最新解析器选择
BeautifulSoup4作为HTML解析的标杆库,其性能很大程度上取决于底层解析器。2023年推荐使用以下组合:
| 解析器 | 速度 | 内存占用 | 容错性 | 适用场景 |
|---|---|---|---|---|
| lxml | ★★★★ | ★★★ | ★★★★ | 大多数情况首选 |
| html.parser | ★★ | ★★★★ | ★★ | 无外部依赖时使用 |
| html5lib | ★ | ★ | ★★★★★ | 处理极不规范的HTML |
from bs4 import BeautifulSoup
# 推荐使用lxml作为解析器
soup = BeautifulSoup(html_content, "lxml")
2. 豆瓣电影Top250页面结构分析
2.1 URL规律与分页处理
豆瓣电影Top250采用经典的分页模式,每页显示25部电影。通过分析,我们发现URL模式非常规律:
https://movie.douban.com/top250?start={offset}&filter=
其中offset参数遵循以下规律:
- 第1页:start=0(显示1-25名)
- 第2页:start=25(显示26-50名)
- ...
- 第10页:start=225(显示226-250名)
我们可以利用Python 3.11的math.ceil和生成器表达式高效生成所有页面URL:
import math
total_movies = 250
movies_per_page = 25
total_pages = math.ceil(total_movies / movies_per_page)
urls = [
f"https://movie.douban.com/top250?start={page * movies_per_page}&filter="
for page in range(total_pages)
]
2.2 关键数据定位与提取
豆瓣页面的电影信息主要包含在<div class="item">元素中。每个电影条目包含:
- 电影名称(中英文)
- 评分
- 评价人数
- 经典台词(可能没有)
- 导演和主演信息
- 上映年份和国家
- 电影类型
使用BeautifulSoup4提取这些信息的核心思路:
items = soup.select("div.item")
for item in items:
title = item.select_one("span.title").text
rating = item.select_one("span.rating_num").text
# 其他字段类似处理
3. 完整爬虫实现与反爬策略
3.1 基础爬虫框架搭建
我们构建一个面向对象的爬虫框架,便于扩展和维护:
class DoubanTop250Spider:
BASE_URL = "https://movie.douban.com/top250"
HEADERS = {
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64)...",
"Accept-Language": "zh-CN,zh;q=0.9",
}
def __init__(self):
self.client = httpx.Client(headers=self.HEADERS, timeout=10.0)
def fetch_page(self, page: int) -> str:
params = {"start": (page - 1) * 25, "filter": ""}
response = self.client.get(self.BASE_URL, params=params)
response.raise_for_status()
return response.text
def parse_page(self, html: str) -> list[dict]:
# 解析逻辑实现
pass
def run(self):
for page in range(1, 11):
html = self.fetch_page(page)
yield from self.parse_page(html)
3.2 应对反爬机制的关键技巧
豆瓣对爬虫有一定防护,以下是几个实用对策:
-
请求头伪装:
- 设置合理的User-Agent
- 添加Referer和Accept-Language头
-
请求频率控制:
import random import time # 在请求间添加随机延迟 time.sleep(random.uniform(1.0, 3.0)) -
IP轮换策略:
- 使用httpx的代理支持
- 考虑付费代理服务(如需要大规模爬取)
-
Cookie处理:
self.client = httpx.Client( headers=self.HEADERS, cookies={"key": "value"}, # 从浏览器获取有效cookie follow_redirects=True )
4. 数据清洗与存储方案
4.1 数据清洗与规范化
从网页抓取的数据通常需要清洗:
-
去除空白字符:
clean_text = " ".join(text.split()) # 合并多个空白字符 -
处理中英文标题:
def process_title(title_element): chinese = title_element.text english = title_element.find_next_sibling("span", class_="other") return { "chinese": chinese, "english": english.text.strip() if english else None } -
评分数据转换:
rating = float(rating_str) if rating_str else None
4.2 存储方案比较与实现
根据需求不同,可以选择多种存储方式:
方案一:CSV文件存储
import csv
def save_to_csv(movies: list[dict], filename: str):
with open(filename, "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=movies[0].keys())
writer.writeheader()
writer.writerows(movies)
方案二:SQLite数据库存储
import sqlite3
def init_db(filename: str):
conn = sqlite3.connect(filename)
conn.execute("""
CREATE TABLE IF NOT EXISTS movies (
id INTEGER PRIMARY KEY,
chinese_title TEXT NOT NULL,
english_title TEXT,
rating REAL,
votes INTEGER,
year INTEGER,
directors TEXT,
actors TEXT
)
""")
return conn
方案三:JSON格式存储
import json
def save_to_json(movies: list[dict], filename: str):
with open(filename, "w", encoding="utf-8") as f:
json.dump(movies, f, ensure_ascii=False, indent=2)
5. 性能优化与高级技巧
5.1 异步并发爬取实现
利用httpx的异步特性大幅提升爬取速度:
async def fetch_all_pages():
async with httpx.AsyncClient(headers=HEADERS) as client:
tasks = [fetch_page(client, page) for page in range(1, 11)]
return await asyncio.gather(*tasks)
async def fetch_page(client: httpx.AsyncClient, page: int):
params = {"start": (page - 1) * 25, "filter": ""}
response = await client.get(BASE_URL, params=params)
response.raise_for_status()
return response.text
5.2 利用Python 3.11新特性优化代码
-
模式匹配(Pattern Matching):
match movie.get("rating"): case float(r) if r >= 9.0: print("经典电影") case float(r) if r >= 8.0: print("推荐观看") case _: print("一般电影") -
TOML配置支持:
import tomllib with open("config.toml", "rb") as f: config = tomllib.load(f) -
更快的错误处理:
try: response = client.get(url) response.raise_for_status() except httpx.HTTPStatusError as e: print(f"请求失败: {e.response.status_code}")
5.3 异常处理与日志记录
健壮的爬虫需要完善的错误处理:
import logging
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s - %(name)s - %(levelname)s - %(message)s"
)
logger = logging.getLogger(__name__)
def safe_parse(html: str):
try:
return parse_page(html)
except Exception as e:
logger.error(f"解析失败: {str(e)}", exc_info=True)
return []
6. 项目扩展与实用建议
6.1 扩展功能思路
-
定时自动更新:
- 使用
schedule或APScheduler库设置定时任务 - 比较新旧数据,只更新变化的条目
- 使用
-
数据可视化:
import matplotlib.pyplot as plt ratings = [m["rating"] for m in movies] plt.hist(ratings, bins=10) plt.title("豆瓣Top250评分分布") plt.show() -
构建简单的Web界面:
- 使用
FastAPI或Flask展示数据 - 添加搜索和筛选功能
- 使用
6.2 生产环境部署建议
-
容器化部署:
FROM python:3.11-slim WORKDIR /app COPY requirements.txt . RUN pip install -r requirements.txt COPY . . CMD ["python", "main.py"] -
监控与告警:
- 使用
Prometheus监控爬虫运行状态 - 设置异常告警通知
- 使用
-
分布式扩展:
- 考虑使用
Celery或Dask分发爬取任务 - 使用Redis作为任务队列和结果存储
- 考虑使用
在实际项目中,我发现豆瓣对频繁请求比较敏感,建议将爬取间隔设置为3-5秒,并尽量在非高峰时段运行爬虫。对于关键业务数据,可以考虑使用官方API(如果有的话)替代网页爬取。
更多推荐


所有评论(0)