「音频技术深度实战 第10章」FunASR 工具包深度实战
FunASR 1.4.16 工业级 ASR 工具包深度解析:51 个模型注册表/CIF 非自回归预测器/SANM 流式架构/OpenAI 兼容 API/vLLM 加速/车规实战对比。
FunASR 1.4.16 工业级 ASR 工具包深度解析:51 个模型注册表/CIF 非自回归预测器/SANM 流式架构/OpenAI 兼容 API/vLLM 加速/车规实战对比。
深度剖析 Deuz-AI/Deuz-SDK(⭐687)的核心架构:纯 Web API 零运行时依赖的 TypeScript AI SDK,把 Durable Execution(SessionStore + AgentCheckpoint)、Evolutionary Search(islands + MAP-Elites + UCB1 ensemble)、Budget Ledger(admission + reservation + settlement)与 Execution Policy(tool/model 强约束 + 继承收紧)四件套塞进同一个 monorepo,并附带原生 MCP / 混合 RAG / 黑板 Swarm / 手动 Compaction 等子模块。文章重点解释其"Free-function Factory + Frozen Plain Object + One-rule Top-level Spread"反 class / 反继承哲学,以及与 LangChain(class + graph)、Vercel AI SDK(thin wrapper)、AutoGen(对话驱动)、CrewAI(角色管线)等主流方案的设计哲学分歧。
从 AMAP-ML/LongHorizon-Harness(1701⭐,MIT,arXiv 2608.01964,Hugging Face Daily Papers
深度剖析 nolabs-ai/nono(⭐4.4k,Apache-2.0,Rust)如何用 Linux Landlock LSM + macOS Seatbelt 内核级能力沙箱 + Sigstore DSSE 供应链验证 + L7 网络代理 + 命令级 sub-sandbox,在零延迟、零 daemon、零容器的前提下给 Coding Agent 提供零信任执行环境。
从 Human-Agent-Society/reef(7489⭐,Apache-2.0)出发,深度解析 Harness 6 件套之外的"持续自进化层"工程化实现:四步循环(Serve/Observe/Grow/Commit)+ 5 种 dispatch mode 事件总线 + 扁平 composition tree + 不可变 artifact 发布链 + Reefine 提案者设计 first 协议。
深度剖析 mnemox-ai/tradememory-protocol 的核心架构:20 个 MCP 工具、Outcome-Weighted Memory 认知记忆模型、SHA-256 链式审计、Mnemox Control 制动器代理与 PolicyBundle 策略包。把 AI 交易 Agent 从「凭 LLM 直觉下单」变成「带历史记忆 + 风险闸门 + 可审计决策」的工程化金融基础设施。
SenseVoice-Small 深度解析:架构/25055 token 多语种词表/SANM 流式注意力/ASR+SER+AED 三合一/车规落地调优。
深度剖析 virattt/dexter 的核心架构:Claude Code 风格 Agent Harness + 21 个金融专用工具 + SQLite 混合检索 Memory + SKILL.md 工作流引擎 + WhatsApp Gateway + Cron/Heartbeat 周期任务。一种 Coding Agent × 垂直领域的新甜区范式。
几乎所有跑过 Coding Agent 的人都撞过这两面墙:
第一面:让 Agent 修一个 50k 行的 Django 迁移。开了 50 次工具调用之后,它开始忘——忘了原始目标、忘了上一个阶段的产出、忘了为什么要先迁移用户表。结果是 50 页里写满了似是而非的”重构”,最后 pytest 变红。
第二面:做完所有 phase 之后,它不肯停。因为它已经忘了”什么时候算完成”——任务计划躺在 200k token 之外的上下文里,再也回不到 attention 窗口。
为什么会这样? 因为所有 harness 默认把任务计划当成 prompt 文本:放在 system 提示的最远处,跟着一次 evaluate(model_input) 一起被遗忘。每一次 /clear、每一次 compaction、每一次上下文压缩,plan 跟着一起死。
OthmanAdi 的 planning-with-files(下文简称 PwF,27300⭐,TrendShift #1,2026-01-06 全语言榜首)用了一个最朴素的解法:把它写进磁盘,再用 hook 强制塞回去。三个 markdown 文件(task_plan.md / findings.md / progress.md)+ 5 个生命周期 hook + 一台 SHA-256 attestation 锁,让 Agent 哪怕在三小时后再 /clear,plan 也活着。
这篇文章拆它的 6 大原语,给一份可跑的 MVP,并和 3 个同类项目对比为什么只有它把”计划持久化”做成了一件机制而非建议的事。
| 失败模式 | 根因 | PwF 的解 |
|---|---|---|
| /clear 后失忆 | plan 住在 context 里 | plan 住磁盘,hook 重新注入 |
| compaction 后失忆 | context 被压缩,plan 被一起压缩掉 | PreCompact hook 把 plan 摘出来再压 |
| 50 calls 后 drift | 原 goal 滑出 attention 窗口 | PreToolUse hook 每工具调用前重读 plan |
| 不肯停 | 没有”完成”的判定门 | gated mode + check-complete.sh 数 phase 状态 |
| 执行中被 prompt injection 改 plan | plan 是纯文本,谁都能改 | SHA-256 attestation + 每次 hook 重新 hash |
| 多 worker 撞计划 | 共享一个 task_plan.md | .planning/<id>/ slug 隔离 + PLAN_ID 绑定 |
PwF 同时覆盖 4 个组件,是少见的多原语 Harness:
1 | ┌────────────────────────────────────────────────────────┐ |
把”long-running task 的计划”从模型的记忆变成磁盘上的工件,让 harness 用机械手段(hook + gate + hash)代替模型自律(”记得看 plan.md”)。
读 PwF 源码(738 个文件、inject-plan.py 60k+ 字符、19 个脚本 × 5 种语言 i18n)后,最值得拆解的 6 大原语如下。
设计哲学:不要相信”prompt 里说要做的事”。把它变成 shell 命令,让宿主机械地执行。
SKILL.md 的 frontmatter 直接声明 5 个 hook(host 是 Claude Code 时):
1 | hooks: |
双层 dispatcher(skill-hook.sh → inject-plan.py):
graph LR
subgraph Host["Claude Code 宿主"]
UE[UserPromptSubmit] --> H[skill-hook.sh]
PE[PreToolUse] --> H
POE[PostToolUse] --> H
PC[PreCompact] --> H
SE[Stop] --> H
end
H -->|"event=*"| INJ["inject-plan.py<br/>(60k 字符<br/>单进程 130 fork 优化)"
]
INJ -->|"stdout"| CTX["📥 additionalContext<br/>进入模型上下文"]
style H fill:#FFDAB9,stroke:#333,color:#333
style INJ fill:#E8D5F5,stroke:#333,color:#333
style CTX fill:#B5EAD7,stroke:#333,color:#333
style UE fill:#C7CEEA,stroke:#333,color:#333
style PE fill:#C7CEEA,stroke:#333,color:#333
style POE fill:#C7CEEA,stroke:#333,color:#333
style PC fill:#C7CEEA,stroke:#333,color:#333
style SE fill:#C7CEEA,stroke:#333,color:#333关键设计 1:shell dispatcher + Python twin。inject-plan.py 文档头注释直接写——
Why this file exists (v3.17.0). hooks/claude-hook.sh answered every lifecycle event by running resolve-plan-dir.sh and inject-plan.sh, and those scripts answer by forking: realpath, stat, sha256sum, awk, tr, mktemp, head, tail, sed, wc, four separate Python starts… One UserPromptSubmit fire forks about 130 times, one PreToolUse fire about 60. Under Git Bash on Windows a fork costs about 90 ms, so the same fire took seven to twelve seconds against the 10 s hook timeout.
解法:把 130 个 fork 压缩到一个 Python 进程(inject-plan.py),在 Windows Git Bash 上把 hook 响应时间从 7-12s 砍到 60ms。同时保留 inject-plan.sh 作为”参考实现”,让没有 Python 3 的 host 仍能跑(测试用 test_inject_plan_python_parity.py 断言两者 stdout 字节相同)。
关键设计 2:5 事件不是对称的,每个有不同职责:
| 事件 | 职责 | 增量 |
|---|---|---|
| UserPromptSubmit | 重新装配本 session 的 nudge | 重新武装 turn-marker |
| PreToolUse | 把 plan 摘要注入 additionalContext | 模型在工具调用前看到 plan |
| PostToolUse | 验证当前 plan 状态 + 一次性 nudge | 防 plan 被改坏 |
| PreCompact | 在 compaction 前转交 reminder | 计划 summary 优先进入压缩 |
| Stop | 跑 gated mode 的完成门 + 转 stdio 给 gate-stop | 强制 model 不能”假装完成” |
设计哲学:Manus 团队(Lance Martin 拆解)的 “Manipulate Attention Through Recitation”——
Recites and updates todo.md throughout tasks to push global plan into model’s recent attention span.
PwF 把 Manus 的”递归念 plan”做成了 task_plan.md / findings.md / progress.md 三件套(注意:不是 todo.md,是三个语义不同的文件):
graph TB
subgraph Disk["📂 .planning/<plan-id>/"]
TP["📋 task_plan.md<br/>目标 + phase 列表 + status"]
F["🔍 findings.md<br/>调研结论 + 引用源"]
P["📊 progress.md<br/>会话日志 + 错误记录"]
end
subgraph Hook["Hook 注入链"]
PR[PreToolUse]
PO[PostToolUse]
PC[PreCompact]
end
PR -->|"机械 re-read"| TP
PO -->|"verify"| TP
PC -->|"flush before compaction"| Disk
style TP fill:#FFDAB9,stroke:#333,color:#333
style F fill:#E8D5F5,stroke:#333,color:#333
style P fill:#FFF9C4,stroke:#333,color:#333
style PR fill:#B5EAD7,stroke:#333,color:#333
style PO fill:#B5EAD7,stroke:#333,color:#333
style PC fill:#B5EAD7,stroke:#333,color:#333三个文件的语义:
| 文件 | 内容 | 谁写 | 谁读 |
|---|---|---|---|
task_plan.md | Goal + Phase 列表 + 每个 phase 的 Status | LLM(Edit) | 但TaskPlan(每次 PreToolUse) |
findings.md | 调研结论、引用、对比表 | LLM | LLM(按需) |
progress.md | 会话日志 + 错误 + 决策 | LLM | LLM(按需,tail 注入) |
关键设计:每次 PreToolUse hook 触发时,inject-plan.py 会重新读 task_plan.md,把当前 phase 状态、goal、最近的 progress tail 拼成 <additionalContext> 注入。这正是 Manus “Manipulate Attention Through Recitation” 的字面落地:让 plan 始终出现在 attention 窗口的”最近”位置,而不是”最远”位置。
设计哲学:多 worker 跑同一份 plan 文件 = 一定撞车。PwF 用 .planning/<plan-id>/ slug 隔离每个 worker 的工件,用 PLAN_ID 环境变量当显式绑定。
resolve-plan-dir.sh 的解析顺序:
1 | 1. $1 (显式路径参数) |
关键设计 1:PLAN_ID 是绑定不是提示。如果 PLAN_ID 显式声明但解析失败,不会 fallback 到 cwd 下的其他 plan——返回错误。check-complete.sh 源码注释直接说:
Explicit selectors are bindings, not hints (issue #237). The shared resolver rejected one, so the legacy cwd fallback below must not run: answering a mistyped pin with the ROOT plan’s completion state is the same wrong-plan harm the binding removes.
关键设计 2:slug 命名合法性严格。slug_is_valid() 用 case 模式匹配拒绝路径穿越:
1 | slug_is_valid() { |
关键设计 3:resolve-plan-dir.sh 拒绝 symlink(避免 task_plan.md -> /etc/passwd 类的逃逸)。attest-plan.sh 同款拒绝:
1 | [ ! -L "${target_dir}/task_plan.md" ] || return 1 |
设计哲学:plan 文件是 LLM 的”宪法”,不能被任何工具调用悄悄改。PwF 把 plan 当成只读工件 + 写时 hash,每次 hook 注入前再 hash 一次确认。
attest-plan.sh 流程:
sequenceDiagram
participant LLM as 🤖 LLM
participant FS as 📂 task_plan.md
participant ATT as 🔒 .attestation
participant HK as 🪝 PreToolUse Hook
Note over LLM,FS: LLM 想"最终化"或"故意编辑"plan
LLM->>FS: Write/Edit task_plan.md
LLM->>ATT: sh scripts/attest-plan.sh
ATT->>FS: sha256sum(plan) → 写入 .attestation
Note over HK,FS: 下一次 hook 触发
HK->>FS: read task_plan.md
HK->>ATT: read .attestation
HK->>HK: hash(today) == hash(stored)?
alt 匹配
HK-->>LLM: 注入 plan 进 additionalContext
else 不匹配
HK-->>LLM: 输出 [PLAN TAMPERED] 警告<br/>不再注入
end
style LLM fill:#C7CEEA,stroke:#333,color:#333
style FS fill:#FFDAB9,stroke:#333,color:#333
style ATT fill:#FFB3C6,stroke:#333,color:#333
style HK fill:#B5EAD7,stroke:#333,color:#333安全事件:2026-03 的”proactive security audit”发现 prompt injection 放大向量——PreToolUse hook 重读 task_plan.md 是核心机制,但 allowed-tools 声明 WebFetch/WebSearch 创造了一条路径让不可信 web 内容进入 plan 文件、然后被每次工具调用都重新注入。v2.21.0 修:删 WebFetch/WebSearch from allowed-tools,加 Security Boundary 段。这是 PwF 把”安全”当成”机制”而不是”建议”的关键证据。
关键设计:v3 模式(autonomous / gated)默认开启 attestation,且拒绝注入未经 attest 的 plan。在长循环里这意味着:哪怕 prompt injection 改了 plan,hook 也会拒绝注入被污染的内容。
设计哲学:完成不是模型说了算。是 plan 文件里的 phase status 说了算。
check-complete.sh 在 v3 模式下是 Stop hook 的”硬关卡”。默认 advisory echo,opt-in --gate 后变成 deliberation gate,只在 ALL 5 条件同时成立时才阻止 Stop:
stateDiagram-v2
[*] --> CheckGate
CheckGate --> AllowStop: 任何 1 个不满足
CheckGate --> BlockStop: 5 个条件全部满足
state BlockStop <<choice>>
BlockStop --> AllowStop: stop_blocks >= PWF_GATE_CAP(默认 20)
BlockStop --> AllowStop: ledger 未推进(stall)
BlockStop --> AllowStop: stop_hook_active=true(已在强制续)
BlockStop --> BlockStop: emit block-decision JSON
CheckGate: 决策表<br/>1. .mode == "gate"<br/>2. 存在 in_progress phase<br/>3. stop_hook_active=false<br/>4. block 计数 < cap<br/>5. ledger 推进过
style CheckGate fill:#FFF9C4,stroke:#333,color:#333
style AllowStop fill:#B5EAD7,stroke:#333,color:#333
style BlockStop fill:#FFB3C6,stroke:#333,color:#333配套 ledger:ledger-summary.sh 把 .planning/<id>/ledger-<agent>.jsonl 转成定长 summary(tick count、phases complete/total、in_progress heading、last event type)。没有 free text、没有 timestamp——所以是 KV-cache stable(这正是 Manus 原则 1 “Design Around KV-Cache” 的字面落地)。
设计哲学:不假设 host。5 种语言 i18n × 双语言实现 × 强制力梯度。
6 层 i18n + 双语实现矩阵:
| 语言 | Shell | Python | Skill |
|---|---|---|---|
| en (default) | ✅ 19 scripts | ✅ 2 py | ✅ |
| ar (Arabic) | ✅ | ✅ | ✅ |
| de (German) | ✅ | ✅ | ✅ |
| es (Spanish) | ✅ | ✅ | ✅ |
| zh (简体) | ✅ | ✅ | ✅ |
| zht (繁体) | ✅ | ✅ | ✅ |
强制力梯度(针对 60+ 不同 host):
| 强制力 | Host |
|---|---|
| 硬阻塞(block Stop) | Claude Code, Codex, Continue |
| follow-up 注入(注入但允许停) | Cursor, Pi, Kiro |
| 仅通知(仅日志不阻塞) | 其他 |
这样设计是因为不同 host 的 hook 协议不一样:Claude Code 用 decision: "block" JSON 输出,Cursor 用 followupMessage,其他可能只有 stdout log。PwF 把”硬约束”这个动作的”严格度”参数化,让同一个 skill 在 60+ 工具里都能用。
把 PwF 的核心机制抽到一个可跑的 Python 文件里:3 文件工件 + Hook 模拟 + Attestation + Gate。
1 | """ |
运行结果:
1 | === planning-with-files MVP === |
关键点:5 个原语在 200 行里全部能跑,attestation 真做 hash 校验,gate 真按 5 条件 AND。
| 维度 | PwF(27300⭐) | Anthropic skill-creator(官方) | Aider --auto-yes | Cline Task |
|---|---|---|---|---|
| 计划持久化层 | ✅ 3 markdown 文件住磁盘 | ❌ 纯 prompt | ❌ 内存 to-do list | ⚠️ 单文件 todo.md |
| Hook 强约束 | ✅ 5 生命周期 hook + 强制注入 | ❌ prompt 引导 | ❌ 无 hook | ⚠️ PostMessage UI 提示 |
| 完成判定机制 | ✅ phase status + 5 条件 gate | ❌ 模型自己说 done | ⚠️ 仅 auto-yes 跳过确认 | ⚠️ LLM 自检 |
| Attestation 锁 | ✅ SHA-256 + tamper 报警 | ❌ 无 | ❌ 无 | ❌ 无 |
| /clear 后恢复 | ✅ PreCompact hook + 磁盘 | ❌ plan 死了 | ❌ to-do 死了 | ❌ todo 死了 |
| 跨 host 适配 | ✅ 60+ host × 双语实现 | ⚠️ Claude Code only | ❌ Aider only | ❌ VSCode only |
| 多 Agent 隔离 | ✅ slug + PLAN_ID 绑定 | ❌ 无 | ❌ 无 | ❌ 无 |
| Prompt injection 防御 | ✅ attestation + 安全 audit | ❌ 无 | ❌ 无 | ❌ 无 |
Anthropic skill-creator:是”教学框架”——告诉模型怎么写一个 skill。PwF 是”运行时机制”——让宿主机械地执行一个 skill。两者的边界是 “prompt 引导” vs “shell hook”。skill-creator 的 plan 文件存在但没有强制 hook,模型可以选择忽略;PwF 的 plan 必须被 hook 重新注入,模型没机会忽略。
Aider --auto-yes:假设 Agent 不会失败,所以关掉所有确认弹窗。PwF 假设 Agent 会失败(context 死、prompt injection、drift),所以强制让 plan 住磁盘。两个极端:信任 vs 不信任。
Cline Task:UI 层的 todo.md + 用户可见 check box。PwF 的 progress.md 不进 UI,只进 context——Cline 是”用户视角的进度板”,PwF 是”模型视角的 attention 锚”。
最本质差异:PwF 把”长任务”当成”分布式系统”问题——plan 文件是 source of truth、hook 是消息总线、gate 是 transaction commit、attestation 是 CAS lock。其他三个把”长任务”当成”prompt engineering”问题。
TrendShift #1(2026-01-06,全语言榜首)不是偶然。两个实证数字:
| 数字 | 来源 | 含义 |
|---|---|---|
| 96.7% pass (29/30) | Anthropic skill-creator formal eval | 工作流保真度极高 |
| 3/3 blind A/B wins | 盲测胜率 | 模型真的”更听话” |
| 5.0 vs 13.3 turns | 作者内部恢复基准(v1) | /clear 后恢复速度 2.6× |
| 68% token overhead | 同 eval | 性价比可接受 |
| 330 tokens/turn | 长任务稳态开销 | 极低(比一次 tool call 的 1k+ 还少) |
| 优点 | 实测 |
|---|---|
| 机制与策略分离 | plan 文件 + 模板(task_plan.md / analytics_task_plan.md / task_plan_autonomous.md)让 LLM 只管内容,harness 管流程 |
| 跨 host 解耦 | 60+ host 一份 skill,不锁死单一 harness 厂商 |
| 渐进披露 | v2 → v3 加入 autonomous / gated 模式都是 opt-in,不破坏老用户 |
| 安全优先 | 2026-03 主动做 audit,发现 WebFetch 注入放大向量并修复 |
| i18n 完整 | 5 种语言 i18n × 双语实现 |
| 测试充分 | 89 个测试文件,含 byte-identical parity test |
| 缺点 | 影响 |
|---|---|
| hook 开销 | PreToolUse 每次 60ms(Git Bash Windows 130 fork → 60ms);稳态 330 token/turn + 90 token/tool call |
| 68% token overhead | 比无结构 run 多 68% token(19,926 vs 11,899),多 17% 时间 |
| slug 解析复杂度 | resolve-plan-dir.sh 6 层 fallback 顺序,初次读源码需要 30+ 分钟 |
| 双语言 twin 维护负担 | sh + py 必须 stdout 字节相同(test_inject_plan_python_parity.py) |
| gated 模式 5 条件全部要写对 | 任何一条件失败 = 静默 allow stop = 完成门形同虚设 |
| CLI 工具链繁重 | 19 个脚本需要规划/实验 |
如果我自己复刻一个 80% 版本,最低成本是什么?
MVP 必装(1-2 人天):
/clear-recovery 命令,重新从磁盘读 planMVP 应装(1 周):
4. SHA-256 attestation(防 prompt injection)
5. Stop hook 完成门(phase status 计数)
MVP 可选(迭代时加):
6. 多 Agent slug 隔离(只有跑并行 worker 才需要)
7. i18n(只有发布给全球用户才需要)
8. 双语 twin(只有 host 在 Windows Git Bash 上才需要)
踩坑预警:
/clear 必死。PwF 用 disk + hook 是唯一 robust 解。PwF 不是另一个”Claude Code 插件”。它把 long-running task 的”计划持久化”从模型的记忆变成磁盘上的工件,让 harness 用机械手段(hook + gate + hash)代替模型自律。
对 Harness Engineering 的启示:
/clear 和 compaction 都不杀死工作。下一个值得挖的角度(选题空白):
推荐谁看:做 HCI 引擎、跑 5+ tool call 的 Agent、多 Agent 并行编排、研究 long-running task 的工程师。
项目链接:github.com/OthmanAdi/planning-with-files
参考资料:
端侧 KWS 唤醒到底怎么落地?用 k2-fsa 官方 WenetSpeech-3.3M 模型,INT8 量化到 4.8MB,单线程流式推理 80ms。