跳到主要内容
知仓学习社ZHICANG

eval-regression

Use when the user asks to check a plugin or skill with behavioral evals, compare it with HEAD, or investigate a behavioral regression.

不碰外部(只输出文字)无严重或高危命中hashgraph-online/awesome-codex-plugins

它会碰到什么

扫了多少1 个文本文件,2 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

Eval regression

Use deterministic repository tests first. Stop when the target has no behavioral diff.

Resolve the plugin from the argument or cwd. Its catalog is evals/evals.json; the shared runner is ../ai-agent-bench/scripts/run_evals.py. Require the user to choose agent, model, and effort. Never select a costly model or high effort silently.

Normal check

  1. Select cases by changed paths with --changed-from <base>, or name them with --case. Do not select the full catalog implicitly.
  2. Run the command without --run. The runner prints the cases, modes, repeat count, session count, per-session timeout, and maximum duration without starting an agent.
  3. Present that plan and stop for explicit cost approval.
  4. After approval, repeat the same command with --run. The default is candidate-only, one run per case, 180 seconds per session, and at most four sessions.
  5. Report every failed assertion, timeout, non-zero exit, duration, and token count. Missing evidence is inconclusive.

Escalation

  • Compare base and candidate only when the user asks, or when a failed candidate check needs to distinguish a regression from an existing failure. Extract the base with git archive; run the same selected cases, agent, model, effort, repeat, and timeout on both; compare reports with --compare.
  • Repeat only a failed or observably unstable case. Three repeats are a stability benchmark, not a default.
  • --all, --repeat > 1, a larger --max-sessions, or a longer timeout needs a new run plan and explicit approval.
  • Routing cases stop at the first Skill selection. A tool assertion is appropriate there because routing is the contract; they must not execute the selected skill.

Remove temporary base copies after the comparison. Do not edit or commit the target.

Deterministic assertions are tool, tool_not, clean_worktree, changed_files_exact, file_contains, file_not_contains, transcript_contains, and tool_sequence. Use a semantic judge only when no filesystem, command, tool, ordering, or assistant-output observation can express the contract.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 1,027
本站分层T1
该仓技能数1910
原文件路径plugins/reidemeister94/development-skills/skills/eval-regression/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 1910 个技能