☰
GSD Skill 编写指南:如何写出清晰、直接、让 Agent 零歧义执行的指令
2026/10/8 7:46:01 网站建设 项目流程
  • 人工智能
  • AI Agent
  • 代码智能体
  • Agent 编排
  • CLI
  • AI 应用

【免费下载链接】gsd-2

A powerful meta-prompting, context engineering and spec-driven development system that enables agents to work for long periods of time autonomously without losing track of the big picture

项目地址:https://gitcode.com/gh_mirrors/gs/gsd-2
点击查看免费下载

导读:在 GSD 的技能(Skill)体系中,SKILL.md与其配套的references/、workflows/、templates/、scripts/文件本质上是写给 Claude 等大模型的"程序化提示词"。本篇技术指南以仓库中 be-clear-and-direct.md 为骨架,系统讲解 skill 指令编写的清晰度(Clarity)与直接性(Directness)原则:如何提供上下文、如何把模糊需求改写成具体指令、如何消除歧义、如何定义边界与成功标准,并结合仓库内create-skill技能及其源码级生态(发现机制、遥测、健康监控)给出可落地的实战方案。读完后,你将掌握一套可复用的"指令工程"方法,写出能被模型稳定解析、减少返工、节省 token 的高质量 skill。


一、黄金法则:先让"零上下文"的人读懂你的指令

be-clear-and-direct.md 开篇给出了整个 skill 编写体系的第一性法则:

Show your skill to someone with minimal context and ask them to follow the instructions. If they're confused, Claude will likely be too.(把你的 skill 展示给一个几乎没有任何背景的人,请他们照着指令执行。如果他们感到困惑,Claude 大概率也会困惑。)

这条"黄金法则"是衡量一切指令质量的终极测试:你自己拥有的项目知识、潜在假设,对模型来说并不存在。skill 的全部价值在于——把你脑中的隐式上下文,显式地写进指令里。

这与create-skill技能中"Skills Are Prompts"的原则一脉相承。在 SKILL.md 中明确指出:所有提示词最佳实践都适用于 skill 编写——要清晰、直接、使用 XML 结构,并且"假设 Claude 很聪明,只补充 Claude 不知道的上下文"。清晰指令能减少错误、提升执行质量、最小化 token 浪费,这三者正是 overview 对 skill 编写核心目标的定义。

在 GSD 生态中,这条法则还会被量化验证:skill-telemetry会记录每个 skill 的读取次数、最近使用时间与通过率(详见gsd-skill-ecosystem.md),而清晰度直接决定了这些指标的上限。


二、提供上下文信息:为任务"定位"

模型无法读取你的内心。<contextual_information>一节要求为任务提供四类框架性信息:

  • 任务结果将被用于什么(What the task results will be used for)
  • 输出面向什么受众(What audience the output is meant for)
  • 任务属于哪条工作流(What workflow the task is part of)
  • 最终目标 / 成功的完成形态(The end goal or what successful completion looks like)

官方给出的最小示例:

<context> This analysis will be presented to investors who value transparency and actionable insights. Focus on financial metrics and clear recommendations. </context>

这段上下文让模型知道:受众是重视透明度的投资者 → 因此要聚焦财务指标和清晰建议;输出是"演示/报告"场景 → 因此要避免冗长的背景铺陈。上下文帮助 Claude 做出更好的决策,产出更合适的输出。

在 GSD 的 XML 标签体系中,<context>正是use-xml-tags.md(见 use-xml-tags.md)定义的条件标签之一:当模型在开始前需要背景或情境信息时使用。而<objective>则是每个 skill必须拥有的三个标签之一,负责"说明 skill 做什么以及为什么重要,设定上下文与范围"。上下文信息与 objective 协同工作:objective 声明任务范围,context 补充任务的"社会背景"。


三、具体化(Specificity):明确说出"只要代码,别的什么都不要"

模糊指令会把决策权全部甩给模型。<specificity>一节的对比是教科书级的示范:

模糊(Vague)具体(Specific)
"Help with the report""Generate a markdown report with three sections: Executive Summary, Key Findings, Recommendations"
"Process the data""Extract customer names and email addresses from the CSV file, removing duplicates, and save to JSON format"

注意具体版指令的三个特征:输出格式明确(markdown、三个固定章节)、操作对象明确(CSV 中的字段)、去重与保存目标明确(JSON 格式)。文档总结道:"Specificity eliminates ambiguity and reduces iteration cycles"——具体化消除歧义,减少迭代轮次。

这一原则在 core-principles.md 中被进一步理论化为"自由度原则(Degrees of Freedom)":把具体程度与任务的脆弱性(fragility)匹配。数据库迁移、支付处理、安全操作属于低自由度场景,必须给出精确到"不要修改命令、不要添加额外 flag"的指令;而代码审查、内容生成等创造性任务属于高自由度场景,应以原则和启发式引导而非逐条死命令。自由度过高会让脆弱任务出错,自由度过低会僵化创造性输出。


四、顺序步骤:把流程写成可验证的操作序列

<sequential_steps>一节要求用编号列表或项目符号把指令写成顺序步骤:

<workflow> 1. Extract data from source file 2. Transform to target format 3. Validate transformation 4. Save to output file 5. Verify output correctness </workflow>

顺序步骤的价值在于建立明确预期,降低模型跳过关键操作的概率。注意第 3 步与第 5 步都包含"验证"——这在 GSD 的 skill 结构中不是偶然。从仓库源码结构看,create-skill/workflows/create-new-skill.md(见 create-new-skill.md)本身就是这种顺序哲学的产物:Step 1 确定范围 → Step 2 需求收集 → Step 3 API 调研 → Step 4 结构决策 → Step 5 建目录 → ... → Step 9 验证结构 → Step 10 激活 → Step 11 测试,每一步都带明确的产出物与检查清单。

在 GSD 的 XML 标签体系中,多步骤流程对应<workflow>或<process>标签(可互换),是"流程型"skill 的骨架标签;而 GSD 的工作流(workflow)文件同样遵循<required_reading> → <process> → <success_criteria>三段式结构(见 recommended-structure.md)。


五、模糊与清晰对照:同一任务,两种命运

原文用"匿名化客户反馈"这一案例,完整演示了从模糊到清晰的改写过程,这是全文最值得精读的对比实验。

❌ 模糊版本:

<quick_start> Please remove all personally identifiable information from these customer feedback messages: {{FEEDBACK_DATA}} </quick_start>

留给模型的未决问题:

  • 什么算 PII?(邮箱、姓名、还是公司名?)
  • 用什么替换 PII?(空字符串、掩码、还是占位符?)
  • 输出是什么格式?
  • 如果没有发现 PII 怎么办?
  • 产品名是否也要打码?

✅ 清晰版本:

<objective> Anonymize customer feedback for quarterly review presentation. </objective> <quick_start> <instructions> 1. Replace all customer names with "CUSTOMER_[ID]" (e.g., "Jane Doe" → "CUSTOMER_001") 2. Replace email addresses with "EMAIL_[ID]@example.com" 3. Redact phone numbers as "PHONE_[ID]" 4. If a message mentions a specific product (e.g., "AcmeCloud"), leave it intact 5. If no PII is found, copy the message verbatim 6. Output only the processed messages, separated by "---" </instructions> Data to process: {{FEEDBACK_DATA}} </quick_start> <success_criteria> - All customer names replaced with IDs - All emails and phones redacted - Product names preserved - Output format matches specification </success_criteria>

清晰版本为何更好(原文结论,逐条继承):

  1. 陈述了目的(季度评审演示)
  2. 提供了显式的逐步规则
  3. 清晰定义了输出格式
  4. 规定了边界情况(产品名保留、无 PII 时原样复制)
  5. 定义了成功标准

两者对比的结论非常关键:模糊版本把所有这些决策都留给 Claude,增加了与预期不一致的可能性。清晰版本的信息量几乎是模糊版本的十倍,但每条信息都直接转化为可执行约束,这正是"上下文信息密度"的胜利——GSD 的 token 优化文档 同样强调"信息密度优先于篇幅长度"。


六、展示而非仅仅描述(Show, Don't Just Tell)

<show_dont_just_tell>一节的核心理念是:当格式重要时,用示例展示,而不是用语言描述。

❌ 描述式(telling):

<commit_messages> Generate commit messages in conventional format with type, scope, and description. </commit_messages>

✅ 展示式(showing):

<commit_message_format> Generate commit messages following these examples: <example number="1"> <input>Added user authentication with JWT tokens</input> <output> feat(auth): implement JWT-based authentication Add login endpoint and token validation middleware </output> </example> <example number="2"> <input>Fixed bug where dates displayed incorrectly in reports</input> <output> fix(reports): correct date formatting in timezone conversion Use UTC timestamps consistently across report generation </output> </example> Follow this style: type(scope): brief description, then detailed explanation. </commit_message_format>

"telling"只说"用 conventional 格式",但 conventional 格式中 type 的具体取值(feat/fix)、scope 的写法、subject 与 body 的分隔、body 的详略程度,模型全部要靠猜;"showing"则通过两个完整示例把这些细微之处(精确的格式:空格、大小写、标点;语气与风格;细节层级;跨多个案例的模式)一次性固定下来。文档的结论是:Claude 从示例中学习模式,比从描述中学习更可靠。

这正是use-xml-tags.md中<examples>标签的设计动机——多示例学习(multi-shot learning)、输入/输出对、模式演示,都是<examples>标签的适用场景。在 GSD 的模板体系里,templates/目录("Claude 复制后填充的输出结构")同样贯彻"展示优先"哲学:给出结构骨架,让模型按骨架填充。


七、消除歧义:从措辞层面切断错误分支

<avoid_ambiguity>一节给出了一份可直接对照使用的"危险措辞表":

❌ 歧义表达隐含问题✅ 清晰表达
"Try to..."暗示可选,不做也无妨"Always..." / "Never..."
"Should probably..."义务不明确"Must..." / "May optionally..."
"Generally..."例外何时允许?"Always... except when..."
"Consider..."每次都做还是偶尔做?"If X, then Y" / "Always..."

原文的实操示例完整展示了两版对比:

❌ 歧义版:

<validation> You should probably validate the output and try to fix any errors. </validation>

✅ 清晰版:

<validation> Always validate output before proceeding: ```bash python scripts/validate.py output_dir/

If validation fails, fix errors and re-validate. Only proceed when validation passes with zero errors.

清晰版做了三件事:把义务提升为硬性约束(Always)、把验证动作绑定到具体命令、把失败后的处置写成确定性分支("修复后重新验证,直到零错误才继续")。这种"If X, then Y"的条件式写作,在 GSD 的验证型 skill 中被广泛复用——例如 `verify-before-complete`、`web-quality-audit` 等技能都内置了脚本验证步骤,而 [workflows-and-validation.md](https://link.gitcode.com/i/a1cb726d8b0429fd2f626244a04655cd) 是这一模式的系统化参考。 --- ## 八、定义边界情况:别让模型猜"空集"怎么处理 `<define_edge_cases>` 一节指出:**预测边界情况并定义处理方式,不要留白给模型猜测**。 **❌ 未定义边界**: ```xml <quick_start> Extract email addresses from the text file and save to a JSON array. </quick_start>

被遗留在空中的问题:没有邮箱怎么办?重复邮箱怎么办?格式非法的邮箱怎么办?JSON 的确切格式?

✅ 定义边界:

<quick_start> Extract email addresses from the text file and save to a JSON array. <edge_cases> - **No emails found**: Save empty array `[]` - **Duplicate emails**: Keep only unique emails - **Malformed emails**: Skip invalid formats, log to stderr - **Output format**: Array of strings, one email per element </edge_cases> <example_output> ```json [ "user1@example.com", "user2@example.com" ]

</example_output> </quick_start>

注意这个版本的三个细节:`<edge_cases>` 明确"没有/重复/非法"三种情况的处置;`<example_output>` 用真实 JSON 示例钉死输出格式;最后一条"Output format"既是对正常路径的规定,也是对边界情况的兜底。**边界情况定义 + 示例输出**的组合,可以让模型在"空输入"时也输出结构完全一致的 `[]`,而不是自作主张返回错误或空文本。 --- ## 九、输出格式规范:把"报告"变成逐条可校验的规格 `<output_format_specification>` 一节展示了对"生成报告"类任务的具体化: **❌ 模糊格式**: ```xml <output> Generate a report with the analysis results. </output>

✅ 精确格式:

<output_format> Generate a markdown report with this exact structure: ```markdown # Analysis Report: [Title] ## Executive Summary [1-2 paragraphs summarizing key findings] ## Key Findings - Finding 1 with supporting data - Finding 2 with supporting data - Finding 3 with supporting data ## Recommendations 1. Specific actionable recommendation 2. Specific actionable recommendation ## Appendix [Raw data and detailed calculations]

Requirements:

  • Use exactly these section headings
  • Executive summary must be 1-2 paragraphs
  • List 3-5 key findings
  • Provide 2-4 recommendations
  • Include appendix with source data </output_format>
这条规格同时具备三层约束:**结构约束**(固定的章节标题层级)、**数量约束**(摘要 1-2 段、发现 3-5 条、建议 2-4 条)、**内容约束**(附录必须含原始数据)。模型拿到它,产出就从一个"像报告的东西"收敛为"一份符合验收标准的报告"。 在 GSD 的 skill 生态里,这种"输出结构规格化"正是 `templates/` 目录的定位——为 plans、specs、configs、reports 提供"复制 + 填充"的统一输出骨架。以仓库中的 [create-workflow](https://link.gitcode.com/i/d8ec7b072728f2597612abc5dc07810f) 技能为例,其 `templates/` 目录存放了可直接套用的 YAML 工作流定义模板,实践上与"格式规格化"完全同构。 --- ## 十、决策标准:当模型必须选择时,给它客观判据 `<decision_criteria>` 一节处理的是"模型需要做决策"的场景——单纯说"你决定用哪种可视化"等于没给约束: **❌ 无判据**: ```xml <workflow> Analyze the data and decide which visualization to use. </workflow>

✅ 有判据:

<workflow> Analyze the data and select appropriate visualization: <decision_criteria> **Use bar chart when**: - Comparing quantities across categories - Fewer than 10 categories - Exact values matter **Use line chart when**: - Showing trends over time - Continuous data - Pattern recognition matters more than exact values **Use scatter plot when**: - Showing relationship between two variables - Looking for correlations - Individual data points matter </decision_criteria> </workflow>

这套判据把"哪种图表更好"的隐性审美,翻译成了 9 条可判定的显性条件。当决策基于客观条件而非主观偏好时,Claude 的选择就变得可预期、可复现——这正是core-principles.md中"高自由度任务用启发式引导"原则的具体落地。


十一、约束分级:必须做 / 锦上添花 / 禁止做

<constraints_and_requirements>一节的洞见是:清晰地分离"必须做"(must do)、"锦上添花"(nice to have)与"禁止做"(must not do)。

❌ 模糊需求:

<requirements> The report should include financial data, customer metrics, and market analysis. It would be good to have visualizations. Don't make it too long. </requirements>

遗留问题:三类内容都是必需的吗?可视化是可选项还是必选项?"too long"到底多长?

✅ 分级需求:

<requirements> <must_have> - Financial data (revenue, costs, profit margins) - Customer metrics (acquisition, retention, lifetime value) - Market analysis (competition, trends, opportunities) - Maximum 5 pages </must_have> <nice_to_have> - Charts and visualizations - Industry benchmarks - Future projections </nice_to_have> <must_not> - Include confidential customer names - Exceed 5 pages - Use technical jargon without definitions </must_not> </requirements>

分级后的收益显而易见:优先级与约束的显式化避免了期望错位。模型不会再为"要不要加图表"纠结——nice_to_have明确告诉它"有则更好、无则不影响验收";must_have中的 "Maximum 5 pages" 把模糊的 "Don't make it too long" 变成可测量的硬指标。这种"三栏式"写法在 GSD 中同样有直接对应:工作流文件中的<success_criteria>、milestone 验收清单等都沿用了"硬性标准 + 可选增强"的分级思路。


十二、成功标准:让模型知道"何时算完成"

<success_criteria>一节指出:定义成功的样子——模型如何知道自己做对了?

❌ 无成功标准:

<objective> Process the CSV file and generate a report. </objective>

✅ 有成功标准:

<objective> Process the CSV file and generate a summary report. </objective> <success_criteria> - All rows in CSV successfully parsed - No data validation errors - Report generated with all required sections - Report saved to output/report.md - Output file is valid markdown - Process completes without errors </success_criteria>

成功标准的本质是完成判据的可操作化:解析所有行、无校验错误、报告含所有必需章节、保存到指定路径、输出是合法 markdown、全程无报错——每一条都可独立验证。文档明确指出其价值:"Clear completion criteria eliminate ambiguity about when the task is done."

在 GSD 中,<success_criteria>(别名<when_successful>)是每个 skill 的三个必需标签之一(见 use-xml-tags.md)。SKILL.md 自身的验收标准就是活例——SKILL.md 的<success_criteria>规定了一个合格 skill 必须"有合法 YAML frontmatter、正文为纯 XML 结构、essential principles 内联、SKILL.md 不超过 500 行、已经过真实使用测试"。而 GSD 的skill-health机制(见gsd-skill-ecosystem.md)甚至会把"成功率低于 70%"的 skill 标记为待审查——成功标准是否清晰,会直接反映到遥测数据上。


十三、测试清晰度:把指令交给"初级工程师"

<testing_clarity>一节给出了一个自检心智模型:

Could I hand these instructions to a junior developer and expect correct results?(我能把这些指令交给一个初级工程师,并期待得到正确结果吗?)

测试流程(原文六步,完整继承):

  1. 通读你的 skill 指令
  2. 移除只有你才掌握的上下文(项目知识、未言明的假设)
  3. 找出歧义术语或模糊需求
  4. 在需要处补充具体性
  5. 让一个不掌握你上下文的人测试
  6. 根据他们的困惑持续迭代

"如果零上下文的人类都读不懂,Claude 也会读不懂。"这条与开篇的黄金法则首尾呼应,构成完整的质量闭环。在 GSD 的create-skill技能中,这步对应workflows/create-new-skill.md的 Step 11 "Test"——真实调用 skill,观察它是否问对了 intake 问题、是否加载了正确的工作流、输出是否符合预期,并"基于真实使用而非假设"迭代。

iteration-and-testing.md 将这一环节升级为系统的"评估驱动开发"(Evaluation-Driven Development):先在无 skill 的情况下让模型处理代表性任务、记录具体失败,再创建评估用例(含query、files、expected_behavior三要素)、建立基线、写最小指令、逐轮迭代对比。清晰度不是一次性写出来的,而是"测试 → 观察 → 修正"迭代出来的。


十四、完整实战示例:数据清理与函数生成

原文在<practical_examples>中给出两个跨域完整示例,全部保留如下。

示例一:数据清理

❌ 模糊:

<quick_start> Clean the data and remove bad entries. </quick_start>

✅ 清晰:

<quick_start> <data_cleaning> 1. Remove rows where required fields (name, email, date) are empty 2. Standardize date format to YYYY-MM-DD 3. Remove duplicate entries based on email address 4. Validate email format (must contain @ and domain) 5. Save cleaned data to output/cleaned_data.csv </data_cleaning> <success_criteria> - No empty required fields - All dates in YYYY-MM-DD format - No duplicate emails - All emails valid format - Output file created successfully </success_criteria> </quick_start>

注意模糊版连"bad entries"都没有定义,而清晰版把"bad"翻译成了四条可判定规则(空必填字段、日期格式、重复邮箱、非法邮箱),再用五条成功标准闭环。

示例二:代码生成

❌ 模糊:

<quick_start> Write a function to process user input. </quick_start>

✅ 清晰:

<quick_start> <function_specification> Write a Python function with this signature: ```python def process_user_input(raw_input: str) -> dict: """ Validate and parse user input. Args: raw_input: Raw string from user (format: "name:email:age") Returns: dict with keys: name (str), email (str), age (int) Raises: ValueError: If input format is invalid """

Requirements:

  • Split input on colon delimiter
  • Validate email contains @ and domain
  • Convert age to integer, raise ValueError if not numeric
  • Return dictionary with specified keys
  • Include docstring and type hints </function_specification>

<success_criteria>

  • Function signature matches specification
  • All validation checks implemented
  • Proper error handling for invalid input
  • Type hints included
  • Docstring included </success_criteria> </quick_start>
函数生成示例展示了指令工程的最高形态:**函数签名本身就是规格**(docstring 中的 Args/Returns/Raises 三区、类型注解),Requirement 列表是行为规格,success_criteria 是验收清单。模型拿到这份规格,产出的函数将"一次通过"而不是来回试探——这正是"Specificity eliminates ambiguity and reduces iteration cycles"的终极体现。 --- ## 十五、与 GSD Skill 生态的衔接:清晰指令如何被系统消费 以上所有原则,最终都落在 GSD 的实际运行机制上。理解"系统如何消费你的 skill",才能把清晰度原则用得有的放矢。根据 [gsd-skill-ecosystem.md](https://link.gitcode.com/i/5bf40ef9ae127ac364860e413be6199b) 的说明: **目录与发现**:GSD 支持两级 skill 目录——全局 `~/.agents/skills/` 与项目级 `.agents/skills/`。会话启动时全部 skill 会被枚举,其**名称 + 描述注入系统提示词的 `<available_skills>` 块**;auto-mode 下每次单元边界都会对比快照,新 skill 通过 `<newly_discovered_skills>` 自动注入。这意味着 **description 是模型发现你的 skill 的第一入口**——模糊的 description 会让 skill 在正确场景下也"隐身",这正是"清晰直接"原则延伸到元数据层面的体现。 **元数据校验**:name 必须全小写字母/数字/连字符、与目录名完全一致;description 非空、≤1024 字符、不含 XML 标签、用第三人称、必须同时说明"做什么"与"何时用"。这些约束在 [SKILL.md](https://link.gitcode.com/i/d6d056de0acca86b461d5e0eecf64475) 的 `<yaml_requirements>` 中有完整规范。 **遥测与健康**:`skill-telemetry` 记录每个 skill 的读取次数、最后使用时间,60+ 天未使用标记为 stale;`skill-health` 聚合出成功率与 token 趋势,**成功率低于 70%、token 趋势上升、60+ 天未用**都会被 `/doctor` 命令标记审查。清晰指令带来的低返工、低 token 浪费,会直接体现为这些健康指标的改善。 **渐进式披露**:所有 `references/`、`workflows/` 文件默认不会被一次性加载,而是由 SKILL.md 的 `<routing>` 按需指向(见 [core-principles.md](https://link.gitcode.com/i/68e456c73e649dd4748992452cb0a606))。这就要求把"必须无条件执行的原则"内联进 SKILL.md(它总是被加载),把"按场景取用的细节"放进 references/——这条规则与本篇的"上下文供给"原则互为表里:**只在正确的时机、给正确的信息量,本身就是一种清晰**。 --- ## 总结:清晰与直接是 skill 编写的底层协议 回顾全文,`be-clear-and-direct.md` 的九大准则构成了一个完整的指令质量协议: 1. **上下文供给**——让模型知道任务的目的、受众、流程与终态; 2. **具体化**——格式、对象、操作、保存目标全部写明; 3. **顺序步骤**——用编号清单建立执行预期; 4. **示例优于描述**——用输入/输出对展示格式与风格; 5. **消除歧义措辞**——用 Always/Never/Must/If-then 替代 Try/Consider/Generally; 6. **边界情况定义**——空集、重复、非法输入各有明确处置; 7. **输出格式规格化**——章节、数量、内容逐条校验; 8. **决策判据显式化**——把审美翻译成客观条件; 9. **成功标准可验证**——让完成时刻没有争议。 再加上"交给零上下文的人测试"这一贯穿始终的黄金法则,这套方法论的每一步都能在 GSD 仓库中找到落地证据:`create-skill` 技能的 SKILL.md 与 14 份 references 文档本身就是按照这些准则撰写的样本,而 `skill-health` 的成功率与 token 趋势遥测,则为"清晰度可以度量"提供了系统级背书。**当你开始编写下一个 skill 时,请把这句话贴在案头:如果你的指令需要解释,那它还不够清晰。**
  • 人工智能
  • AI Agent
  • 代码智能体
  • Agent 编排
  • CLI
  • AI 应用

【免费下载链接】gsd-2

A powerful meta-prompting, context engineering and spec-driven development system that enables agents to work for long periods of time autonomously without losing track of the big picture

项目地址:https://gitcode.com/gh_mirrors/gs/gsd-2
点击查看免费下载

相关推荐

上一篇:终极指南:LaTeX的起源与核心价值解析
下一篇:Kubernetes安全:命名空间隔离与网络策略全攻略

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询