前言
「获取网络数据」听起来就是把网页下下来,但真实场景里数据来源远不止 HTML:有的是返回 JSON 的接口,有的是 XML 的订阅源,有的是需要 POST 提交参数才给的报表。方法选错了,要么拿不到数据,要么拿到的是一堆需要自己拼的碎片。
一个典型误解是「只要urlopen一下就能拿到内容」。实际上有三个变量决定了你要怎么写:请求方法是 GET 还是 POST、参数放在查询串还是请求体、响应是 HTML 还是 JSON。搞清这三件事,取数代码就八九不离十。
本文讲清楚这几种取数方式,以及拿到数据后怎么正确还原和分流。全部基于标准库,示例适用于 Python 3.8 及以上。Python 2 已于 2020 年 1 月 1 日停止维护,本文只给 Python 3 写法——urllib2在 3 里已并入urllib.request。
一、GET 请求与查询参数
GET 把参数放在 URL 的查询串里。拼查询串不要手写字符串,要用urlencode做百分号编码,否则中文、空格、&都会出问题。
# 适用于 Python 3.8+
import urllib.parse
import urllib.request
base = "https://www.example.com/api/items"
query = urllib.parse.urlencode({"page": 2, "size": 20, "kw": "python 教程"})
url = base + "?" + query
print("请求地址:", url)
req = urllib.request.Request(url, headers={
"User-Agent": "DataFetchBot/0.1 (contact: me@example.com)",
"Accept": "application/json",
})
with urllib.request.urlopen(req, timeout=10) as resp:
print("状态:", resp.status)
print("类型:", resp.headers.get("Content-Type"))
text = resp.read().decode("utf-8", errors="replace")
print(text[:80])urlencode会把中文和空格编码成%xx形式。Request的headers传字典即可,需要逐条追加时用add_header(key, val)。
二、POST 与表单数据
POST 把参数放在请求体里。标准做法是用urlencode编码后.encode("utf-8")成 bytes,作为Request的data参数。给了data,请求方法默认就是 POST(也可以用method显式指定)。
# 适用于 Python 3.8+
import urllib.parse
import urllib.request
data = urllib.parse.urlencode({"q": "python", "page": 1}).encode("utf-8")
req = urllib.request.Request(
"https://www.example.com/search",
data=data,
method="POST",
headers={
"User-Agent": "DataFetchBot/0.1 (contact: me@example.com)",
"Content-Type": "application/x-www-form-urlencoded",
},
)
with urllib.request.urlopen(req, timeout=10) as resp:
print("状态:", resp.status)三个必须记住的点:data必须是 bytes(字符串不行);表单提交要带Content-Type;data和查询串可以同时存在,但语义上别混淆。
三、请求头与身份
请求头决定了服务端怎么看你。爬虫相关的几个:User-Agent(身份标识)、Accept(期望的数据类型)、Accept-Encoding(声明支持的压缩)、Referer(来源页)、Cookie(会话态)。
# 适用于 Python 3.8+
import urllib.request
req = urllib.request.Request("https://www.example.com/api/items")
req.add_header("User-Agent", "DataFetchBot/0.1 (contact: me@example.com)")
req.add_header("Accept", "application/json")
req.add_header("Accept-Encoding", "gzip, deflate")
with urllib.request.urlopen(req, timeout=10) as resp:
print(resp.headers.get("Content-Encoding"))add_header(key, val)一份名字只能存一个值,后调用的会覆盖前面的。注意:一旦你声明支持 gzip,就要自己负责解压——urlopen不会替你解。
四、响应数据的还原与分流
拿到 bytes 之后,第一步永远是把编码问清楚,第二步才按数据类型选解析器。
| 响应类型 | 典型 Content-Type | 还原与解析 |
|---|
| HTML | text/html | charset 解码后交给html.parser |
| JSON | application/json | 解码后用json.loads |
| XML | application/xml、text/xml | 解码后用xml.etree.ElementTree |
| 纯文本 | text/plain | 解码即可 |
# 适用于 Python 3.8+
import json
import urllib.error
import urllib.request
req = urllib.request.Request(
"https://www.example.com/api/items",
headers={"User-Agent": "DataFetchBot/0.1 (contact: me@example.com)"},
)
try:
with urllib.request.urlopen(req, timeout=10) as resp:
charset = resp.headers.get_content_charset() or "utf-8"
raw = resp.read().decode(charset, errors="replace")
except urllib.error.HTTPError as e:
print("HTTP 错误:", e.code)
except urllib.error.URLError as e:
print("网络错误:", e.reason)
else:
payload = json.loads(raw)
print("顶层类型:", type(payload).__name__)json.loads接收 str、bytes 或 bytearray(Python 3.6 起支持 bytes)。JSON 里的null、true、false会被映射成None、True、False,数字按 JSON 规则映射成int或float。
五、要更细的控制:http.client
如果需要手工控制连接、复用连接或看清响应头,可以下沉到http.client。它比urlopen更底层:不自动跟随重定向、不处理 Cookie。
# 适用于 Python 3.8+
import http.client
conn = http.client.HTTPSConnection("www.example.com", timeout=10)
conn.request("GET", "/api/items?page=1", headers={
"User-Agent": "DataFetchBot/0.1 (contact: me@example.com)",
})
resp = conn.getresponse()
print("状态:", resp.status, "原因:", resp.reason)
body = resp.read()
conn.close()
print("字节数:", len(body))conn.request(method, url, body=None, headers={}, ...)的第二个参数是路径(含查询串),不是完整 URL。resp.getheader(name)取响应头。
实战:一个能分流 JSON 与 HTML 的取数函数
# 适用于 Python 3.8+
import gzip
import json
import urllib.error
import urllib.request
USER_AGENT = "DataFetchBot/0.1 (contact: me@example.com)"
def fetch(url, data=None, timeout=10):
"""返回 (content_type, payload)。text 类型给 str,json 类型给解析后的对象。"""
headers = {"User-Agent": USER_AGENT}
if data is not None:
headers["Content-Type"] = "application/x-www-form-urlencoded"
req = urllib.request.Request(url, data=data, headers=headers)
with urllib.request.urlopen(req, timeout=timeout) as resp:
raw = resp.read()
if resp.headers.get("Content-Encoding") == "gzip":
raw = gzip.decompress(raw)
ctype = resp.headers.get_content_type()
charset = resp.headers.get_content_charset() or "utf-8"
text = raw.decode(charset, errors="replace")
if ctype == "application/json":
return ctype, json.loads(text)
return ctype, text
if __name__ == "__main__":
try:
ctype, payload = fetch("https://www.example.com/api/items")
except urllib.error.HTTPError as e:
print("HTTP 错误:", e.code)
except urllib.error.URLError as e:
print("网络错误:", e.reason)
else:
print("类型:", ctype, "| 顶层:", type(payload).__name__)resp.headers.get_content_type()返回小写的 MIME 类型(如application/json),用它分流比判断字符串是否含 "json" 更可靠。
常见坑点
- 手拼查询串
❌url + "?kw=" + keyword—— 中文和空格没转义,服务端解析错。 ✅ 用urllib.parse.urlencode({...})编码后再拼。
- POST 时
data传字符串
❌Request(url, data="a=1")——data必须是 bytes。 ✅urllib.parse.urlencode({"a": 1}).encode("utf-8")。
- POST 忘了 Content-Type
❌ 只给data不给Content-Type—— 服务端可能按未知类型处理。 ✅ 表单提交补上application/x-www-form-urlencoded。
- 声明支持 gzip 却不解压
❌ 请求头写了Accept-Encoding: gzip,拿到字节直接 decode。 ✅ 先看Content-Encoding,用gzip.decompress处理。
- 硬编码 UTF-8 解析 JSON
❌json.loads(resp.read())忽略编码差异,个别响应会失败。 ✅ 先按get_content_charset()解码,再json.loads。
- 用
getheader判断类型做字符串匹配
❌if "json" in resp.headers.get("Content-Type")—— 大小写与参数会干扰。 ✅ 用resp.headers.get_content_type()拿到规范化的 MIME 类型。
http.client里传了完整 URL
❌conn.request("GET", "https://www.example.com/a")—— 第二个参数只填路径。 ✅ 主机名交给HTTPSConnection,request里写/a?x=1。
- 把
urllib2/print语句当还能用
❌import urllib2、print text—— Python 2 已停止维护,3 里均不成立。 ✅ 用urllib.request,输出用print(...)。
总结
| 场景 | 参数位置 | 关键写法 |
|---|
| GET 取数 | URL 查询串 | urlencode+Request(headers=...) |
| POST 提交 | 请求体 | urlencode(...).encode("utf-8")+Content-Type |
| JSON 响应 | — | 先解码,再json.loads |
| HTML 响应 | — | 解码后交给html.parser |
| 精细控制 | — | http.client,自己处理重定向与 Cookie |
取数的套路固定:先确定方法、再确定参数放哪、最后按 Content-Type 分流处理。把这三步做对,再加上超时、异常处理和限速,一个稳健的取数模块就成型了。