- 网络安全
- 网络
- IDS
【免费下载链接】zeek
Zeek is a powerful network analysis framework that is much different from the typical IDS you may know.
导读
本文以 Zeek 的 Summary Statistics(简称 SumStats)框架为核心,系统讲解如何将海量、无界的网络事件数据流压缩为简单的统计度量,并在此基础上实现周期汇总与阈值告警。读完本文,你将掌握 SumStats 的三大核心构件(Observation、Reducer、SumStat)及其完整类型接口、全部计算插件(SUM、AVERAGE、TOPK 等)的用法、阈值与多阈值序列的配置方法,以及该框架在单机与集群部署下的内置透明性——这些能力足以支撑你独立编写诸如"统计每分钟连接数""检测扫描行为"之类的 Zeek 脚本。
框架定位:为什么需要 SumStats
测量网络流量是 Zeek 脚本中最常见的任务之一。对于有限的小样本(如指定行数的 trace 文件),直接用事件处理器累加变量即可;但在真实部署中会遇到两个难题:集群化(多个 worker 进程同时嗅探流量,数据分散)与无界数据集(流量永不停止,内存无法无限增长)。SumStats 框架正是为了在这两种条件下"消费无界数据集并使其可被度量"而设计的(见 doc/frameworks/sumstats.rst)。
框架给出的核心承诺是:同一份脚本在单进程 Zeek 与多 worker 集群上都能正确运行,集群中流量负载均衡带来的复杂性由框架内置的集群透明机制自动处理,脚本开发者可以忽略。
整体处理流程:Observation → Reducer → SumStat
Sumstats 的处理流程被拆分为三个环节,官方文档与源码(scripts/base/frameworks/sumstats/main.zeek)中的定义完全对应:
- Observation(观察):事件的某个方面被观察到后,作为一条数据点喂入框架。一条观察属于一个任意命名的观察流(observation stream),并带有一个 Key(该观察是关于谁的)以及观察值本身。
- Reducer(归约器):对观察流施加各种计算(如求和、均值、方差),把无界观察集合压缩成更小的表示。结果在每个 Reducer 内部按 Key 分别收集,因此需要控制被跟踪的 Key 总数,避免内存失控。
- SumStat(汇总统计):最终定义,把一个或多个 Reducer 在一个时间区间(称为epoch,即纪元/周期)内汇总。在 SumStat 层可以配置阈值(threshold)与越过阈值时的回调,也可以为每个 epoch 结束时提供按 Key 访问结果的回调。
核心类型接口全解
SumStats 全部定义在SumStats命名空间中(module SumStats;,见 main.zeek)。以下类型、字段与默认值均以当前仓库源码为准。
SumStats::Key —— 度量针对谁
Key是一个 record,表示"正在被收集汇总结果的对象",只有两个可选字段:
| 字段 | 类型 | 说明 |
|---|---|---|
str | string &optional | 非地址型度量的键,或地址型度量的子键。例如"按客户端 IP 统计成功 SSH 连接"时,IP 是 host 键;而"按 Host 头统计 HTTP 请求数"则是非主机型度量(同一 Host 值可能对应多个 IP),此时用str作键。 |
host | addr &optional | 该度量所适用的主机地址。 |
源码中key2str()(main.zeek#L285-L293)把 Key 格式化为sumstats_key(host=..., str=...)这样的可读字符串,方便打印或作为日志字段。
SumStats::Observation —— 单条数据点
Observation是"单次观察加入的数据",字段如下,一次只能提供单个字段:
| 字段 | 类型 | 说明 |
|---|---|---|
num | count &optional | 计数类值。 |
dbl | double &optional | 浮点值。 |
str | string &optional | 字符串值。 |
注意:在observe()的实现中(main.zeek#L492-L498),如果观察只带字符串、没有num/dbl,框架会把数值val回退为1.0——这正是SUM计算"对字符串值统计数量"的底层原因。
SumStats::Reducer —— 计算挂在哪里
Reducer是连接的枢纽:它声明自己挂接在哪个观察流上、对该流施加哪些计算:
| 字段 | 类型 | 说明 |
|---|---|---|
stream | string | Reducer 挂接的观察流标识符(必填)。 |
apply | set[SumStats::Calculation] | 要对数据点执行的计算集合(必填)。 |
pred | function(key, obs): bool &optional | 谓词函数,按 Key 决定是否接受该数据点;返回F则跳过。 |
normalize_key | function(key): Key &optional | 键归一化函数,可用来聚合或归一化整个 Key。 |
源码中create()会把 Reducer 登记到reducer_store(按流 ID 索引的 Reducer 集合,main.zeek#L246),observe()正是通过reducer_store[id]找到所有订阅该流的 Reducer 逐一处理(main.zeek#L439-L445)。
SumStats::Result / ResultTable / ResultVal —— 结果怎么存
Result是table[string] of ResultVal:按观察流 ID 索引多个 Reducer 的结果(main.zeek#L78)。ResultTable是table[Key] of Result:按 Key 索引的 SumStats 结果总表(main.zeek#L81)。ResultVal是单个流的计算结果记录,基类自带三个字段(main.zeek#L63-L74):
| 字段 | 类型/默认值 | 说明 |
|---|---|---|
begin | time | 第一条观察被加入该结果的时间。 |
end | time | 最后一条观察被加入的时间(observe()每次都会更新,见 main.zeek#L490)。 |
num | count &default=0 | 收到的观察总数。 |
其余字段(average、max、min、sum、variance、std_dev、hll_unique、unique、last_elements、samples、topk等)全部由计算插件通过redef record ResultVal +=追加,详见下文插件一节。
SumStats::SumStat —— 汇总的最终定义
SumStat记录把 Reducer 集合、epoch 周期、阈值机制与各回调组装在一起(main.zeek#L91-L144):
| 字段 | 类型 | 说明 |
|---|---|---|
name | string | SumStat 的任意名称,后续可据此引用。 |
epoch | interval | 周期区间。每个 epoch 结束时触发epoch_result回调,同时重置结果——因此基于阈值的检测值应设为本 epoch 内预期出现的量级。设为0 secs即切换到手动 epoch,需自行调用SumStats::next_epoch结束周期。 |
reducers | set[Reducer] | 该 SumStat 使用的 Reducer 集合。 |
threshold_val | function(key, result): double &optional | 对每条观察调用、从Result中提取用于阈值比较的值。只要设置了threshold或threshold_series就必须提供(create()中会校验并报错,见 main.zeek#L394-L397)。 |
threshold | double &optional | 触发threshold_crossed回调的阈值。需要多个阈值时改用threshold_series。 |
threshold_series | vector of double &optional | 阈值序列,必须按升序排列,因为某个阈值只有在前一个被越过之后才会被检查。 |
threshold_crossed | function(key, result) &optional | 阈值被越过时调用的回调。越过条件:threshold_val的返回值大于等于阈值,且一个 epoch 内每个 Key 只触发第一次。 |
epoch_result | function(ts, key, result) &optional | 每个分析周期结束时接收各 Key 结果的回调,对每个 Key 各调用一次。 |
epoch_finished | function(ts) &optional | 一个收集周期整体完成时调用,ts为该周期开始时间。 |
一个重要的集群提示(main.zeek#L86-L90):不要在回调中访问传入参数之外的任何全局状态,因为集群中无法保证回调在哪个节点上执行。
五个全局函数
| 函数 | 签名 | 作用 |
|---|---|---|
SumStats::create | function(ss: SumStat) | 创建一个汇总统计(main.zeek#L392-L437)。内部完成:校验阈值配置、登记到stats_store、初始化threshold_tracker、把每个 Reducer 注册进reducer_store、解析计算依赖、调用reset()并(非手动 epoch 时)调度finish_epoch事件。 |
SumStats::observe | function(id: string, key: Key, obs: Observation) | 向观察流添加数据点,应在脚本测得某个度量值时调用(main.zeek#L439-L503)。 |
SumStats::request_key | function(ss_name: string, key: Key): Result | 动态请求某个 SumStat Key 的当前结果。文档与源码均强调应谨慎使用,不能替代SumStat的回调机制;且只能在when语句中作为异步函数使用(见 non-cluster.zeek#L90-L100)。 |
SumStats::key2str | function(key: Key): string | 把 Key 转成简单字符串的辅助函数。 |
SumStats::next_epoch | function(ss_name: string): bool | 手动结束某 SumStat 的当前 epoch(仅当该 SumStat 以 0 周期创建为手动 epoch 时可用)。结束不是即时的——集群中需要节点间交换多条消息;集群中必须在 manager 上调用,worker 上调用无效。失败场景:SumStat 不存在,或未按手动 epoch 创建(main.zeek#L272-L283)。 |
计算插件体系:Calculation 枚举与 ResultVal 字段
所有计算类型都以插件形式实现:Calculation枚举基类只有一个PLACEHOLDER成员(main.zeek#L10-L12),每个插件通过redef enum Calculation +=扩展枚举、redef record ResultVal +=扩展结果字段,并注册到register_observe_plugins钩子。所有插件由 plugins/load.zeek 默认加载。
以SUM为例(plugins/sum.zeek):插件注册了观察函数rv$sum += val(sum.zeek#L39-L45),通过init_resultval_hook初始化$sum,并通过compose_resultvals_hook支持集群中两个结果的合并。
全部插件及其语义、相关字段如下表:
| 计算 | 插件文件 | 语义 | 涉及的 ResultVal 字段/默认值 |
|---|---|---|---|
AVERAGE | plugins/average.zeek | 数值的平均值 | average: double &optional |
HLL_UNIQUE | plugins/hll_unique.zeek | 用 HyperLogLog 估计唯一值数量 | hll_unique: count &default=0、card: opaque of cardinality、hll_error_margin、hll_confidence |
LAST | plugins/last.zeek | 在队列中保留最近 X 条观察 | last_elements: Queue::Queue;不要直接访问该字段,应使用SumStats::get_last取回元素向量;Reducer 侧配置num_last_elements: count &default=0 |
MAX | plugins/max.zeek | 最大值 | max: double &optional |
MIN | plugins/min.zeek | 最小值 | min: double &optional |
SAMPLE | plugins/sample.zeek | 从观察流中均匀随机采样 | samples: vector of Observation &default=[]、sample_elements: count &default=0、num_samples: count &default=0;Reducer 侧配置num_samples |
STD_DEV | plugins/std-dev.zeek | 标准差 | std_dev: double &default=0.0 |
SUM | plugins/sum.zeek | 数值求和;字符串值时统计字符串个数 | sum: double &default=0.0 |
TOPK | plugins/topk.zeek | 保留 top-k 列表 | topk: opaque of topk(可传给内置函数取结果);Reducer 侧配置topk_size: count &default=500 |
UNIQUE | plugins/unique.zeek | 精确统计唯一值数量 | unique: count &default=0、unique_vals: set[Observation];Reducer 侧配置unique_max: count &optional(最大存储的唯一值个数) |
VARIANCE | plugins/variance.zeek | 数值方差 | variance: double &optional、prev_avg: double &optional、var_s: double &default=0.0 |
另有 Reducer 上由插件追加的配置字段:hll_error_margin(HLL 误差率,默认0.01)、hll_confidence(HLL 置信度,默认0.95)、num_last_elements(LAST 保留条数,默认0)、num_samples(SAMPLE 采样条数,默认0)、topk_size(TOPK 列表长度,默认500)、unique_max(UNIQUE 上限)。
插件还支持依赖解析:add_observe_plugin_dependency()(main.zeek#L300-L305)声明某个计算依赖另一个计算(例如 VARIANCE 可能依赖 SUM),create()时add_calc_deps()(main.zeek#L370-L390)递归展开依赖并去重,最终按依赖顺序执行calc_store中的观察函数(main.zeek#L499-L500)。
实战一:统计周期内连接数
以下完整示例来自 doc/frameworks/sumstats-countconns.zeek(运行时需@load base/frameworks/sumstats):
@load base/frameworks/sumstats event connection_established(c: connection) { # Make an observation! # This observation is global so the key is empty. # Each established connection counts as one so the observation is always 1. SumStats::observe("conn established", SumStats::Key(), SumStats::Observation($num=1)); } event zeek_init() { # Create the reducer. # The reducer attaches to the "conn established" observation stream # and uses the summing calculation on the observations. local r1 = SumStats::Reducer($stream="conn established", $apply=set(SumStats::SUM)); # Create the final sumstat. # We give it an arbitrary name and make it collect data every minute. # The reducer is then attached and a $epoch_result callback is given # to finally do something with the data collected. SumStats::create([$name = "counting connections", $epoch = 1min, $reducers = set(r1), $epoch_result(ts: time, key: SumStats::Key, result: SumStats::Result) = { # This is the body of the callback that is called when a single # result has been collected. We are just printing the total number # of connections that were seen. The $sum field is provided as a # double type value so we need to use %f as the format specifier. print fmt("Number of connections established: %.0f", result["conn established"]$sum); }]); }要点拆解:
- 观察:每个
connection_established事件调用一次observe(),流 ID 为"conn established";因为是全局度量,Key 为空SumStats::Key();每次计数 1,所以Observation($num=1)。 - 归约:Reducer 订阅该流并
apply=set(SumStats::SUM),对数值做累加。 - 汇总:
epoch = 1min,每分钟触发一次epoch_result回调;通过result["conn established"]$sum取该流结果。由于$sum是double,打印用%.0f格式符。
官方文档给出的运行效果(对 Zeek 测试集 PCAP 执行):
$ zeek -r workshop_2011_browse.trace sumstats-countconns.zeek Number of connections established: 6实战二:用阈值检测扫描主机
下面的"玩具级"扫描检测演示了阈值机制(完整源码见 doc/frameworks/sumstats-toy-scan.zeek):
@load base/frameworks/sumstats # We use the connection_attempt event to limit our observations to those # which were attempted and not successful. event connection_attempt(c: connection) { # Make an observation! # This observation is about the host attempting the connection. # Each established connection counts as one so the observation is always 1. SumStats::observe("conn attempted", SumStats::Key($host=c$id$orig_h), SumStats::Observation($num=1)); } event zeek_init() { # Create the reducer. # The reducer attaches to the "conn attempted" observation stream # and uses the summing calculation on the observations. Keep # in mind that there will be one result per key (connection originator). local r1 = SumStats::Reducer($stream="conn attempted", $apply=set(SumStats::SUM)); # Create the final sumstat. # This is slightly different from the last example since we're providing # a callback to calculate a value to check against the threshold with # $threshold_val. The actual threshold itself is provided with $threshold. # Another callback is provided for when a key crosses the threshold. SumStats::create([$name = "finding scanners", $epoch = 5min, $reducers = set(r1), # Provide a threshold. $threshold = 5.0, # Provide a callback to calculate a value from the result # to check against the threshold field. $threshold_val(key: SumStats::Key, result: SumStats::Result) = { return result["conn attempted"]$sum; }, # Provide a callback for when a key crosses the threshold. $threshold_crossed(key: SumStats::Key, result: SumStats::Result) = { print fmt("%s attempted %.0f or more connections", key$host, result["conn attempted"]$sum); }]); }本示例与上一例的三个关键差异:
- Key 携带主机:
SumStats::Key($host=c$id$orig_h),每个发起连接的主机成为一个独立 Key,结果按 Key 分开统计。 - 阈值三元组:
$threshold = 5.0给出阈值;$threshold_val从Result中提取比较值(这里就是求和值);$threshold_crossed在 Key 越过阈值时被回调。 - 触发语义:阈值越过判断发生在每次观察插入时,
check_thresholds()(main.zeek#L507-L549)会先检查threshold_tracker确保同一 epoch 内每个 Key 只触发一次(源码注释见 main.zeek#L129-L132);threshold_crossed()负责递增跟踪计数并调用回调(main.zeek#L551-L570)。
官方文档中的运行效果(对含 nmap 主机的 PCAP 执行):
$ zeek -r nmap-vsn.trace sumstats-toy-scan.zeek 192.168.1.71 attempted 5 or more connections多阈值序列
若需多个告警档位,改用threshold_series。注意其约束(main.zeek#L123-L127):阈值必须按升序排列;实现上get_threshold_index()(main.zeek#L226-L234)记录每个 Key 已越过的阈值偏移,只有前一个阈值被越过后才检查下一个(main.zeek#L539-L546),因此阈值序列天然形成阶梯式告警。
手动 epoch
把$epoch设为0 secs即进入手动模式:create()不再调度finish_epoch事件(main.zeek#L434-L436),脚本可在任意时刻(例如某个业务事件发生时)调用SumStats::next_epoch("sumstat 名称")触发周期收尾;集群中该调用必须在 manager 上执行。
底层机制:create 与 observe 的源码级流程
create() 做了什么(main.zeek#L392-L437)
- 校验:设置了
threshold/threshold_series但未提供threshold_val时报错。 - 将 SumStat 存入
stats_store(按 name 索引)。 - 若配置了阈值,初始化该 SumStat 的
threshold_tracker。 - 遍历 Reducer:记录其所属 SumStat 名(内部字段
ssname),解析计算依赖得到有序的calc_funcs,并把 Reducer 加入reducer_store(按 stream 索引)。 - 调用
reset()清空result_store与threshold_tracker。 - 若 epoch 非 0,
schedule ss$epoch { SumStats::finish_epoch(ss) }周期性触发。
observe() 做了什么(main.zeek#L439-L503)
- 流 ID 不在
reducer_store中则直接返回(无订阅者,零开销)。 - 对每个订阅该流的 Reducer:先用
normalize_key(若配置)归一化 Key,再跑pred谓词,返回F则跳过。 - 命中阈值且无
epoch_result回调时跳过后续计数——源码注释说明这是为了避免在度量唯一性时产生状态管理问题而做的优化(main.zeek#L456-L474)。 - 初始化该 Key/流的结果
ResultVal(init_resultval()会调用init_resultval_hook,让各插件初始化自己的字段,见 main.zeek#L313-L318),递增$num、更新$end。 - 将数值(
num或dbl,缺省回退 1.0)依次喂给calc_funcs对应的插件观察函数。 - 调用
data_added(),在非集群实现中触发check_thresholds()与threshold_crossed()(non-cluster.zeek#L84-L88)。
epoch 收尾与集群合并
- 单机(non-cluster.zeek):
finish_epoch事件驱动do_finish_epoch(),按每批 50 个 Key 调度process_epoch_result事件逐批调用epoch_result回调,全部处理完后调用epoch_finished,随后reset()并调度下一周期;zeek_done()时以&priority=10处理遗留的自动 epoch。 - 集群(cluster.zeek):框架通过事件与消息在 worker 与 manager 之间同步数据,
compose_resultvals()(main.zeek#L320-L331)与compose_results()(main.zeek#L333-L351)把多个 worker 的ResultVal/Result合并(取最早的begin、最晚的end、累加num,并逐个调用插件的compose_resultvals_hook合并插件字段)。加载逻辑在load.zeek 中按Cluster::is_enabled()分支选择 cluster 或 non-cluster 实现,这也是"脚本无需关心集群差异"的机制来源。
使用注意事项总结
- Observation 一次只填一个字段:
num、dbl、str三选一;只给字符串时数值按 1.0 处理。 - 控制 Key 数量:Reducer 按 Key 分开收集结果,Key 过多会耗尽内存;必要时用
normalize_key聚合或pred过滤。 - 阈值必须配
threshold_val:否则create()直接报错。 threshold_series必须升序:后一档阈值依赖前一档被越过。- 回调内不要访问外部全局状态:集群环境下回调执行节点不确定。
- 手动 epoch:
next_epoch仅对 epoch 为 0 的 SumStat 有效,且集群中只能在 manager 调用。 request_key慎用:只应作为when语句中的异步函数使用,不要用它替代回调机制。
参考资源
- 框架 API 文档(本文主体来源):doc/scripts/base/frameworks/sumstats/main.zeek.rst
- 框架用户指南与术语、示例说明:doc/frameworks/sumstats.rst
- 核心实现:scripts/base/frameworks/sumstats/main.zeek
- 集群/单机支持:scripts/base/frameworks/sumstats/cluster.zeek、scripts/base/frameworks/sumstats/non-cluster.zeek
- 插件目录:scripts/base/frameworks/sumstats/plugins/load.zeek
- 完整可运行示例:doc/frameworks/sumstats-countconns.zeek、doc/frameworks/sumstats-toy-scan.zeek
- 网络安全
- 网络
- IDS
【免费下载链接】zeek
Zeek is a powerful network analysis framework that is much different from the typical IDS you may know.
相关推荐
Zeek SumStats 框架实战指南:面向集群与无界流量数据的流式统计与阈值检测
Zeek SumStats 框架实战指南:面向集群与无界流量数据的流式统计与阈值检测 本篇技术指南围绕 Zeek 内置的 SumStats(Summary St
网络安全网络IDSbigdata_analyse 数据处理性能优化:10个提升效率的技巧
bigdata_analyse 数据处理性能优化:10个提升效率的技巧 在大数据分析项目中,数据处理性能优化是提升整体效率的关键环节。无论你是在处理百万级的用户
网络安全网络IDSZeek SumStats 框架 MIN 计算插件深入解析:用 `SumStats::MIN` 追踪数值流的最小值
Zeek SumStats 框架 MIN 计算插件深入解析:用 SumStats::MIN 追踪数值流的最小值 SumStats::MIN 是 Zeek 汇总统
网络安全网络IDS
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考