☰
Python网络爬虫之获取网络数据
2026/10/12 6:17:28 网站建设 项目流程

前言


「获取网络数据」听起来就是把网页下下来,但真实场景里数据来源远不止 HTML:有的是返回 JSON 的接口,有的是 XML 的订阅源,有的是需要 POST 提交参数才给的报表。方法选错了,要么拿不到数据,要么拿到的是一堆需要自己拼的碎片。


一个典型误解是「只要urlopen一下就能拿到内容」。实际上有三个变量决定了你要怎么写:请求方法是 GET 还是 POST、参数放在查询串还是请求体、响应是 HTML 还是 JSON。搞清这三件事,取数代码就八九不离十。


本文讲清楚这几种取数方式,以及拿到数据后怎么正确还原和分流。全部基于标准库,示例适用于 Python 3.8 及以上。Python 2 已于 2020 年 1 月 1 日停止维护,本文只给 Python 3 写法——urllib2在 3 里已并入urllib.request。


一、GET 请求与查询参数


GET 把参数放在 URL 的查询串里。拼查询串不要手写字符串,要用urlencode做百分号编码,否则中文、空格、&都会出问题。


# 适用于 Python 3.8+
import urllib.parse
import urllib.request

base = "https://www.example.com/api/items"
query = urllib.parse.urlencode({"page": 2, "size": 20, "kw": "python 教程"})
url = base + "?" + query
print("请求地址:", url)

req = urllib.request.Request(url, headers={
"User-Agent": "DataFetchBot/0.1 (contact: me@example.com)",
"Accept": "application/json",
})
with urllib.request.urlopen(req, timeout=10) as resp:
print("状态:", resp.status)
print("类型:", resp.headers.get("Content-Type"))
text = resp.read().decode("utf-8", errors="replace")
print(text[:80])

urlencode会把中文和空格编码成%xx形式。Request的headers传字典即可,需要逐条追加时用add_header(key, val)。


二、POST 与表单数据


POST 把参数放在请求体里。标准做法是用urlencode编码后.encode("utf-8")成 bytes,作为Request的data参数。给了data,请求方法默认就是 POST(也可以用method显式指定)。


# 适用于 Python 3.8+
import urllib.parse
import urllib.request

data = urllib.parse.urlencode({"q": "python", "page": 1}).encode("utf-8")
req = urllib.request.Request(
"https://www.example.com/search",
data=data,
method="POST",
headers={
"User-Agent": "DataFetchBot/0.1 (contact: me@example.com)",
"Content-Type": "application/x-www-form-urlencoded",
},
)
with urllib.request.urlopen(req, timeout=10) as resp:
print("状态:", resp.status)

三个必须记住的点:data必须是 bytes(字符串不行);表单提交要带Content-Type;data和查询串可以同时存在,但语义上别混淆。


三、请求头与身份


请求头决定了服务端怎么看你。爬虫相关的几个:User-Agent(身份标识)、Accept(期望的数据类型)、Accept-Encoding(声明支持的压缩)、Referer(来源页)、Cookie(会话态)。


# 适用于 Python 3.8+
import urllib.request

req = urllib.request.Request("https://www.example.com/api/items")
req.add_header("User-Agent", "DataFetchBot/0.1 (contact: me@example.com)")
req.add_header("Accept", "application/json")
req.add_header("Accept-Encoding", "gzip, deflate")
with urllib.request.urlopen(req, timeout=10) as resp:
print(resp.headers.get("Content-Encoding"))

add_header(key, val)一份名字只能存一个值,后调用的会覆盖前面的。注意:一旦你声明支持 gzip,就要自己负责解压——urlopen不会替你解。


四、响应数据的还原与分流


拿到 bytes 之后,第一步永远是把编码问清楚,第二步才按数据类型选解析器。




响应类型典型 Content-Type还原与解析



HTMLtext/htmlcharset 解码后交给html.parser

JSONapplication/json解码后用json.loads

XMLapplication/xml、text/xml解码后用xml.etree.ElementTree

纯文本text/plain解码即可



# 适用于 Python 3.8+
import json
import urllib.error
import urllib.request

req = urllib.request.Request(
"https://www.example.com/api/items",
headers={"User-Agent": "DataFetchBot/0.1 (contact: me@example.com)"},
)
try:
with urllib.request.urlopen(req, timeout=10) as resp:
charset = resp.headers.get_content_charset() or "utf-8"
raw = resp.read().decode(charset, errors="replace")
except urllib.error.HTTPError as e:
print("HTTP 错误:", e.code)
except urllib.error.URLError as e:
print("网络错误:", e.reason)
else:
payload = json.loads(raw)
print("顶层类型:", type(payload).__name__)

json.loads接收 str、bytes 或 bytearray(Python 3.6 起支持 bytes)。JSON 里的null、true、false会被映射成None、True、False,数字按 JSON 规则映射成int或float。


五、要更细的控制:http.client


如果需要手工控制连接、复用连接或看清响应头,可以下沉到http.client。它比urlopen更底层:不自动跟随重定向、不处理 Cookie。


# 适用于 Python 3.8+
import http.client

conn = http.client.HTTPSConnection("www.example.com", timeout=10)
conn.request("GET", "/api/items?page=1", headers={
"User-Agent": "DataFetchBot/0.1 (contact: me@example.com)",
})
resp = conn.getresponse()
print("状态:", resp.status, "原因:", resp.reason)
body = resp.read()
conn.close()
print("字节数:", len(body))

conn.request(method, url, body=None, headers={}, ...)的第二个参数是路径(含查询串),不是完整 URL。resp.getheader(name)取响应头。


实战:一个能分流 JSON 与 HTML 的取数函数


# 适用于 Python 3.8+
import gzip
import json
import urllib.error
import urllib.request

USER_AGENT = "DataFetchBot/0.1 (contact: me@example.com)"


def fetch(url, data=None, timeout=10):
"""返回 (content_type, payload)。text 类型给 str,json 类型给解析后的对象。"""
headers = {"User-Agent": USER_AGENT}
if data is not None:
headers["Content-Type"] = "application/x-www-form-urlencoded"
req = urllib.request.Request(url, data=data, headers=headers)
with urllib.request.urlopen(req, timeout=timeout) as resp:
raw = resp.read()
if resp.headers.get("Content-Encoding") == "gzip":
raw = gzip.decompress(raw)
ctype = resp.headers.get_content_type()
charset = resp.headers.get_content_charset() or "utf-8"
text = raw.decode(charset, errors="replace")
if ctype == "application/json":
return ctype, json.loads(text)
return ctype, text


if __name__ == "__main__":
try:
ctype, payload = fetch("https://www.example.com/api/items")
except urllib.error.HTTPError as e:
print("HTTP 错误:", e.code)
except urllib.error.URLError as e:
print("网络错误:", e.reason)
else:
print("类型:", ctype, "| 顶层:", type(payload).__name__)

resp.headers.get_content_type()返回小写的 MIME 类型(如application/json),用它分流比判断字符串是否含 "json" 更可靠。


常见坑点



  1. 手拼查询串


❌url + "?kw=" + keyword—— 中文和空格没转义,服务端解析错。 ✅ 用urllib.parse.urlencode({...})编码后再拼。



  1. POST 时data传字符串


❌Request(url, data="a=1")——data必须是 bytes。 ✅urllib.parse.urlencode({"a": 1}).encode("utf-8")。



  1. POST 忘了 Content-Type


❌ 只给data不给Content-Type—— 服务端可能按未知类型处理。 ✅ 表单提交补上application/x-www-form-urlencoded。



  1. 声明支持 gzip 却不解压


❌ 请求头写了Accept-Encoding: gzip,拿到字节直接 decode。 ✅ 先看Content-Encoding,用gzip.decompress处理。



  1. 硬编码 UTF-8 解析 JSON


❌json.loads(resp.read())忽略编码差异,个别响应会失败。 ✅ 先按get_content_charset()解码,再json.loads。



  1. 用getheader判断类型做字符串匹配


❌if "json" in resp.headers.get("Content-Type")—— 大小写与参数会干扰。 ✅ 用resp.headers.get_content_type()拿到规范化的 MIME 类型。



  1. http.client里传了完整 URL


❌conn.request("GET", "https://www.example.com/a")—— 第二个参数只填路径。 ✅ 主机名交给HTTPSConnection,request里写/a?x=1。



  1. 把urllib2/print语句当还能用


❌import urllib2、print text—— Python 2 已停止维护,3 里均不成立。 ✅ 用urllib.request,输出用print(...)。


总结




场景参数位置关键写法



GET 取数URL 查询串urlencode+Request(headers=...)

POST 提交请求体urlencode(...).encode("utf-8")+Content-Type

JSON 响应—先解码,再json.loads

HTML 响应—解码后交给html.parser

精细控制—http.client,自己处理重定向与 Cookie



取数的套路固定:先确定方法、再确定参数放哪、最后按 Content-Type 分流处理。把这三步做对,再加上超时、异常处理和限速,一个稳健的取数模块就成型了。

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询