☰
Pixie Stirling 在 Kubernetes 集群上的独立部署指南:深入解析 run_stirling_wrapper_on_k8s.sh 与镜像拉取密钥配置
2026/10/8 19:13:42 网站建设 项目流程
  • 可观测性
  • 云原生

【免费下载链接】pixie

Instant Kubernetes-Native Application Observability

项目地址:https://gitcode.com/gh_mirrors/pixie/pixie
点击查看免费下载

本文围绕 Pixie 仓库中 src/stirling/k8s/README.md 所描述的 Stirling 独立部署方案展开,讲解如何将stirling_wrapper以 Pod / DaemonSet 形式部署到 Kubernetes 集群、如何通过sops解密并安装镜像拉取密钥(image pull secret)、以及如何收集日志并完成清理。读完本文,你将能够理解run_stirling_wrapper_on_k8s.sh的完整工作流程、掌握Makefile中各个部署/删除目标的用法,并能在自己的 GKE 或自建集群上复现这一套 Stirling 观测数据采集流程。

Stirling 与 stirling_wrapper:为什么要独立部署到集群

Stirling 是 Pixie 在节点上的数据采集器(data collector),它使用 Linux 内核 API(包括 eBPF)从内核、系统库或应用本身采集应用指标与事件,主要数据包括应用 CPU、内存、网络利用率,以及 HTTP、MySQL、Postgres 等协议的网络消息事件,详见 src/stirling/README.md。

stirling_wrapper是 Stirling 的独立可执行版本(stirling_wrapper.cc),它不依赖 PEM 或 Kubernetes,可以直接在本地运行并把采集到的数据打印到 STDOUT。而 src/stirling/k8s 目录提供了一套脚本与 YAML 模板,用于把stirling_wrapper以独立容器的方式部署到 Kubernetes 集群上——这在开发、调试 Stirling 数据源(source connector)或需要在整个集群范围快速部署观测代理时非常有用。

该目录结构如下:

src/stirling/k8s/ ├── Makefile # 提供部署/删除的 make 目标 ├── README.md # 本文所依据的说明文档 ├── namespace.yaml.tmpl # Namespace 模板 ├── run_stirling_wrapper_on_k8s.sh # 一键部署主脚本 ├── stirling-daemonset.yaml.tmpl # 全节点 DaemonSet 模板 ├── stirling-pod.yaml.tmpl # 单节点 Pod 模板 ├── stirling-profiler-pod.yaml.tmpl # Profiler Pod 模板 └── strace-daemonset.yaml # strace 调试用 DaemonSet

一键脚本:run_stirling_wrapper_on_k8s.sh

run_stirling_wrapper_on_k8s.sh是这一整套流程的入口脚本。README 明确指出:在 GKE 上可以直接运行该脚本;在其他集群上,则可能需要先按下文说明配置镜像拉取密钥。

用法与参数

脚本的 usage 信息(run_stirling_wrapper_on_k8s.sh)完整如下:

Usage: run_stirling_wrapper_on_k8s.sh [-c <fastbuild|dbg|opt>] [-l] [<duration>] -l 自动打开日志 -c 选择编译模式

各参数的含义与默认值(见脚本中parse_args的默认值设置):

参数说明默认值
-c <mode>Bazel 编译模式,可选fastbuild、dbg、optdbg
-l日志收集完成后用less自动打开关闭
<duration>在集群上运行 Stirling 的秒数,脚本会等待该时长后收集日志30

一个特殊行为:当<duration>取值为0时,脚本会部署 Stirling 后立即退出,不做日志收集与清理,并提示用户稍后手动执行make delete_stirling_daemonset或kubectl delete -n ${NAMESPACE}来收尾。

完整执行流程

脚本的执行流程(结合 run_stirling_wrapper_on_k8s.sh 源码)可以分为六个阶段:

  1. 构建并推送镜像:执行

    bazel build -c "$COMPILATION_MODE" //src/stirling/binaries:stirling_wrapper_image bazel run -c "$COMPILATION_MODE" --config=stamp //src/stirling/binaries:push_stirling_wrapper_image

    其中--config=stamp用于以用户名(username)作为镜像 tag 进行推送。

  2. 清理旧实例:检查目标 Namespace 中是否存在stirling-wrapper前缀的 Pod:

    stirling_wrapper_pod_count=$(kubectl get pods -n "${NAMESPACE}" 2> /dev/null | grep -c ^stirling-wrapper || true) if [ "$stirling_wrapper_pod_count" -ne 0 ]; then make delete_stirling_daemonset sleep 5 fi

    注意脚本以bash -e运行,而grep在计数为 0 时会返回退出码 1,因此这里用|| true防止脚本意外中止。

  3. 部署 DaemonSet:执行make deploy_stirling_daemonset,把 Stirling 部署到集群的所有节点。

  4. 等待采集:sleep "$T"等待指定秒数,让 Stirling 在集群上采集数据。

  5. 收集日志:列出stirling-wrapperPod,对每个 Running 状态的 Pod 获取其所在节点名,并把kubectl logs输出保存到本地:

    LOGDIR=logs node_name="$(kubectl get pod "${pod}" -n "${NAMESPACE}" -o=custom-columns=:.spec.nodeName | xargs)" filename="${LOGDIR}/log$timestamp.${pod}.${node_name}" kubectl logs -n "${NAMESPACE}" "${pod}" > "${filename}"

    日志文件按logs/log<时间戳>.<pod名>.<节点名>命名;若指定了-l,则用less逐个打开。

  6. 清理:执行make delete_stirling_daemonset删除 DaemonSet,并kubectl delete namespace "${NAMESPACE}"删除整个 Namespace。

Namespace 与模板渲染

脚本始终在自身所在目录执行(cd "$scriptdir"),且部署目标 Namespace 固定为:

NAMESPACE=pl-${USER}

即每个开发者使用以自己系统用户名命名的独立 Namespace(如pl-oazizi)。Makefile 通过sed对模板进行变量替换:

EVAL_TMPL:=sed -e "s/{{USER}}/${USER}/g" -e "s/{{NAMESPACE}}/${NAMESPACE}/g"

模板中的{{USER}}和{{NAMESPACE}}占位符会在部署时被替换为实际值,见 Makefile。

镜像构建与推送的 Bazel 实现

脚本中的镜像构建与推送对应 src/stirling/binaries/BUILD.bazel 中的 Bazel 目标:

cc_image( name = "stirling_wrapper_image", base = ":stirling_binary_base_image", binary = ":stirling_wrapper", ) container_push( name = "push_stirling_wrapper_image", format = "Docker", image = ":stirling_wrapper_image", registry = "gcr.io", repository = "pl-dev-infra/stirling_wrapper", tag = "{BUILD_USER}", tags = ["manual"], )

可见镜像被推送到gcr.io/pl-dev-infra/stirling_wrapper:<用户名>,这与 DaemonSet 模板中image: gcr.io/pl-dev-infra/stirling_wrapper:{{USER}}的引用完全对应;--config=stamp正是为了让{BUILD_USER}被替换为当前用户。类似的,stirling_profiler的镜像目标为stirling_profiler_image/push_stirling_profiler_image,推送到gcr.io/pl-dev-infra/stirling_profiler。

关于stirling_wrapper本身:它是一个把 Stirling 采集结果直接输出到 STDOUT 的独立程序(stirling_wrapper.cc),支持若干 gflags 参数,例如:

  • --trace:动态追踪规格,可以是 PxL/IR 规格文件的路径,或<对象文件路径>:<符号名>形式的函数级追踪;
  • --print_record_batches:控制输出哪些数据表(默认http_events,mysql_events,pgsql_events,redis_events,cql_events,dns_events,tls_events);
  • --init_only、--timeout_secs、--enable_heap_profiler等。

它还注册了 SIGINT/SIGQUIT/SIGTERM/SIGHUP 的信号处理器,在退出时调用Stirling::Stop()以释放 BPF 资源,避免泄漏。

镜像拉取密钥:sops 解密与 gcloud 认证

这是 README 的核心内容。集群中的 Kubernetes 部署清单(Pod / DaemonSet / Profiler Pod)都通过imagePullSecrets引用名为image-pull-secret的密钥来拉取 Stirling 容器镜像。该密钥文件位于仓库中,但是加密的,需要使用sops解密后再安装到 Kubernetes。

脚本与 Makefile 会自动处理这一过程,但前提是gcloud auth已正确配置,sops才能成功解密 JSON 文件。如果在未运行过gcloud auth的机器上执行,sops会报“无法解密 json 文件”之类的错误。此时按以下步骤修复:

  1. 执行gcloud auth --no-launch-browser application-default login;
  2. 将命令输出的链接复制粘贴到本地浏览器;
  3. 将浏览器中显示的验证码复制回gcloud auth的提示符。

Makefile 中密钥安装的实现细节

从 Makefile 可以看到create_image_pull_secret目标的真实行为:

create_image_pull_secret: @echo "Installing image-pull-secrets (command echo suppressed)" @kubectl -n ${NAMESPACE} delete secret image-pull-secret 2> /dev/null || true @kubectl -n ${NAMESPACE} create secret docker-registry image-pull-secret \ --docker-server=https://gcr.io \ --docker-username=_json_key \ --docker-email=${USER}@pixielabs.ai \ --docker-password='$(shell sops -d ../../../../credentials/k8s/dev/image-pull-secrets.encrypted.json)'

它首先删除可能已存在的同名密钥(|| true容错),然后使用sops -d解密仓库中加密的 JSON(路径相对src/stirling/k8s解析到仓库根目录外的credentials/k8s/dev/image-pull-secrets.encrypted.json),并以_json_key用户名、服务账号 JSON 内容作为密码,创建docker-registry类型的密钥。需要注意:该加密凭据文件并不在公开镜像仓库中(credentials 目录被排除在仓库之外),因此这一流程主要面向 Pixie 内部开发环境;自建集群用户可以仿照该模式,用kubectl create secret docker-registry创建自己的拉取密钥,并在 YAML 模板中替换imagePullSecrets.name。

部署清单解析:Pod、DaemonSet 与 Profiler

目录中的三个 YAML 模板分别对应三种部署形态,均由 Makefile 的对应目标渲染后kubectl apply -f -应用:

模板Makefile 目标用途
stirling-pod.yaml.tmpldeploy_stirling_pod在单个节点上部署一个 Stirling Pod
stirling-daemonset.yaml.tmpldeploy_stirling_daemonset以 DaemonSet 在所有节点上部署
stirling-profiler-pod.yaml.tmpldeploy_stirling_profiler_pod部署 Stirling profiler Pod(带 nodeSelector 示例)

运行 Stirling 所需的容器环境

由于 Stirling 依赖 Linux 内核 API(包括 eBPF),容器化运行时必须满足以下条件(这也是 src/stirling/README.md 中“Stirling docker container environment”一节的要求):

配置项作用
privileged: true且增加SYS_PTRACE、SYS_ADMINcapabilitiesBPF 需要 root 权限
hostNetwork: true、dnsPolicy: ClusterFirstWithHostNet使用宿主机网络
hostPID: true使用宿主机 PID 命名空间,否则getpid()解析会出错
挂载/到/host(只读)、/sys到/sys(只读)访问宿主机数据文件、系统头文件(BCC 编译 C 代码需要)及调试信息
环境变量PL_HOST_PATH=/host与宿主机根文件系统挂载配套

资源限制方面,模板统一配置limits.memory: 2048Mi、requests.cpu: 100m、requests.memory: 512Mi。

DaemonSet 模板中的动态追踪示例

stirling-daemonset.yaml.tmpl 还在容器环境变量中内置了一个STIRLING_AUTO_TRACE的 PxL 动态追踪示例(占位示例,需要替换 UPID 与字段):

import pxtrace import px # This is a placeholder example. To use, update UPID and other fields. upid = "00000000-0000-0000-0000-000000000000" table_name = 'placeholder_table' tp_name = table_name @pxtrace.probe("SomeFunction") def probe_func(): return [{ 'latency': pxtrace.FunctionLatency(), 'id': pxtrace.ArgExpr('id') }] pxtrace.UpsertTracepoint(table_name, table_name, probe_func, px.uint128(upid), "30m")

这与stirling_wrapper的源码实现相呼应:GetTraceProgram()(stirling_wrapper.cc)会优先读取--traceflag,其次读取环境变量STIRLING_AUTO_TRACE;当通过环境变量传入时按PxL 格式处理,通过CompileTracepoint编译并转换为 Stirling 的 tracepoint 部署。也就是说,在集群上通过环境变量即可把一段动态追踪脚本注入stirling_wrapper,而无需修改镜像。

Profiler Pod 与 strace 调试

  • stirling-profiler-pod.yaml.tmpl部署gcr.io/pl-dev-infra/stirling_profiler:{{USER}}镜像,并带有一个nodeSelector示例(kubernetes.io/hostname: gke-dev-cluster-...-q1cg),用于把 profiler 固定调度到指定节点,args: ["625612"]为示例 PID 参数。
  • strace-daemonset.yaml 是独立的调试辅助 DaemonSet,在每个节点上运行jess/strace容器(hostPID: true、privileged: true),通过args: ["-p", "<PID>", "-f", "-v", "-s", "1024"]追踪指定进程的系统调用,用于排查 Stirling 在集群上运行时的行为。

Makefile 手动控制:部署与删除目标

虽然 README 推荐直接使用run_stirling_wrapper_on_k8s.sh,但 Makefile 也暴露了全部底层目标,方便手动分步控制。完整目标清单(Makefile):

目标依赖行为
create_namespacenamespace.yaml.tmpl创建目标 Namespace
create_image_pull_secret—安装镜像拉取密钥(命令输出被抑制)
deploy_stirling_pod模板 + 上两者在单节点部署 Stirling Pod
deploy_stirling_daemonset模板 + 上两者在所有节点部署 Stirling DaemonSet
deploy_stirling_profiler_pod模板 + 密钥 +delete_stirling_profiler_pod部署 profiler Pod(先删后建)
delete_stirling_podstirling-pod.yaml.tmpl删除 Stirling Pod
delete_stirling_daemonsetstirling-daemonset.yaml.tmpl删除 Stirling DaemonSet
delete_stirling_profiler_podstirling-profiler-pod.yaml.tmpl删除 profiler Pod

其中NAMESPACE可通过make NAMESPACE=my-ns deploy_stirling_daemonset覆盖默认的pl-${USER}。所有部署目标都依赖create_namespace与create_image_pull_secret,因此会先自动创建 Namespace 并安装密钥。

部署前后注意事项

  • GKE 与自建集群的差异:GKE 环境下凭据与网络已就绪,可直接运行主脚本;自建集群需要自行保证镜像仓库凭据可用(参照上文密钥安装一节)。
  • 旧实例清理:脚本会检测并删除旧的stirling-wrapperDaemonSet,避免与新建实例冲突;--duration 0模式下脚本不清理,需要手动收尾。
  • 日志位置:日志统一保存到脚本所在目录下的logs/,文件名携带 Pod 名与节点名,便于多节点部署时区分来源。
  • 删除即清理:脚本结束时会删除 DaemonSet 与整个pl-${USER}Namespace,因此不会在集群中残留采集器。
  • eBPF 环境:部署清单中的特权、hostPID、宿主机挂载等配置都是 Stirling 正常运行 BPF 探针的硬性前提;若自定义 YAML,务必保留这些字段,相关容器化要求可参考 src/stirling/README.md 的“Stirling docker container environment”一节。

通过以上脚本、Makefile 目标与 YAML 模板的组合,你可以在 Kubernetes 集群上快速拉起 Stirling 独立采集器、按需注入 PxL 动态追踪脚本、收集日志并自动完成清理,为 Stirling 数据源调试与集群级观测验证提供了一条可复现的完整路径。

  • 可观测性
  • 云原生

【免费下载链接】pixie

Instant Kubernetes-Native Application Observability

项目地址:https://gitcode.com/gh_mirrors/pixie/pixie
点击查看免费下载
上一篇:解决Postman集合导入失败:Bruno工具故障排除指南
下一篇:零代码构建企业级智能客服:AI-For-Beginners实战指南

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询