跳到主要内容
知仓学习社ZHICANG

agent-harness

Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools,…

读凭据改身份文件执行命令写文件读文件严重 1 · 高危 8alirezarezvani/claude-skills

它会碰到什么

扫了多少26 个文本文件,653 KB
它会碰到什么读凭据改身份文件执行命令写文件读文件
命中总数15 处
命中统计严重 1 · 高 8 · 中 6 · 低 0
逐条看命中(9 条严重或高危)
  • 严重 assets/harnesses/engineering.json:1975cred-paths
    "description": "Manage environment-variable hygiene and secrets safety across local development and production. Practical auditing, drift awareness, rotation re
  • assets/harnesses/engineering-team.json:334identity-write
    "description": "Graduate a proven pattern from auto-memory (MEMORY.md) to CLAUDE.md or .claude/rules/ for permanent enforcement. Use when the user runs /si:prom
  • assets/harnesses/engineering-team.json:334identity-write
    "description": "Graduate a proven pattern from auto-memory (MEMORY.md) to CLAUDE.md or .claude/rules/ for permanent enforcement. Use when the user runs /si:prom
  • assets/harnesses/engineering-team.json:362identity-write
    "description": "Curate Claude Code's auto-memory into durable project knowledge. Analyze MEMORY.md for patterns, promote proven learnings to CLAUDE.md and .clau
  • assets/harnesses/engineering-team.json:362identity-write
    "description": "Curate Claude Code's auto-memory into durable project knowledge. Analyze MEMORY.md for patterns, promote proven learnings to CLAUDE.md and .clau
  • assets/harnesses/engineering.json:1415identity-write
    "description": "Use when the user wants their Claude agent to self-improve from past usage, asks about a nightly/offline 'sleep' or 'dream' cycle, memory/skill 
  • assets/harnesses/engineering.json:2672exec-spawn
    "description": "> Security audit and vulnerability scanner for AI agent skills before installation. Use when: (1) evaluating a skill from an untrusted source, (
  • scripts/loop_controller.py:223exec-spawn
    proc = subprocess.run(
  • scripts/loop_controller.py:224exec-shell-true
    chk["cmd"], shell=True, cwd=args.cwd,

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

Agent Harness

You are a harness operator, not a hero. The loop — not your optimism — decides when work

is done. Your job: compile the goal into tasks with checks, execute one task at a time,

let the controller adjudicate verification, and stop when the state machine says stop.

The contract

GOAL → goal_compiler → PLAN → loop_controller: [execute → verify]* → CLOSE
                                     ↑______retry (≤ max_attempts, changed approach)
                                     └── ESCALATE on exhausted budgets — never fake success

Three layers, all JSON: a committed per-domain manifest (what skills/tools/checks

exist), a per-goal plan (which tasks, which verifications, what "done" means), and a

per-run state file (the single source of truth; a fresh session resumes from it alone).

Quick start

# 0. Pick the domain manifest (18 committed under assets/harnesses/, e.g. engineering-team.json)
ls assets/harnesses/

# 1. Compile the goal (refuses vague goals with exit 3 + forcing questions)
python3 scripts/goal_compiler.py \
  --goal "audit the payments service and design an SLO with an error budget" \
  --manifest assets/harnesses/engineering.json --out plan.json

# 2. Initialize the loop state
python3 scripts/loop_controller.py init --plan plan.json --state .agent-harness/state.json

# 3. Drive the loop — repeat until directive is "close" or "escalate"
python3 scripts/loop_controller.py next --state .agent-harness/state.json
#    → {"action": "execute", "task": "T1", ...}: open the task's skill (SKILL.md at
#      skill_path), do the work with its tools, then:
python3 scripts/loop_controller.py record --state .agent-harness/state.json \
  --task T1 --phase execute --exit-code 0
#    → the controller runs the task's checks ITSELF (subprocess, timeout, evidence log):
python3 scripts/loop_controller.py verify --state .agent-harness/state.json --task T1 --cwd <repo-root>

# 4. Close — refused (exit 4) while any task is unverified and unwaived
python3 scripts/loop_controller.py close --state .agent-harness/state.json

Regenerate a manifest after skills change (diff-stable, CI-checkable):

python3 scripts/harness_manifest_builder.py --domain engineering-team \
  --repo-root <repo-root> --out-dir assets/harnesses --no-timestamp

Hard rules

  1. Never adjudicate your own verification. verify runs the checks via subprocess;

a passing record --phase verify without --evidence is rejected (exit 6). You do not

get to declare a task verified.

  1. Never modify a gate you are judged by. Check commands come from the manifest/plan.

Editing a check to make it pass is the reward-hacking failure mode

(see [references/verification_discipline.md](references/verification_discipline.md)) — same

invariant as autoresearch-agent's locked evaluator.

  1. One task at a time, writes serialized. Parallelize reading and judging, never two

tasks writing the same artifact ([references/agentic_loop_canon.md](references/agentic_loop_canon.md)).

  1. Retry means a changed approach. Same command + same input = same failure. The retry

directive says so; honor it.

  1. Budgets are terminal states, not suggestions. max_attempts_per_task → escalated

(exit 2); max_loop_iterations → escalate (exit 5). Exhausted budgets are never

reported as success — a human waives (close --waive T3 --reason "..."), you don't.

  1. Fresh context beats long context. Every next directive is executable by a new

session reading only the plan + state files. Long-running goals: run each iteration as

its own session against the durable state.

  1. State lives in .agent-harness/ — never in .agenthub/, .autoresearch/, or

docs/TC/ (those belong to sibling skills).

  1. Plan and state files are a trust boundary. verify shell-executes each task's

check command; only run the harness on plan/state files you or goal_compiler.py

produced, never on files from untrusted input (see

[references/verification_discipline.md](references/verification_discipline.md)).

Forcing questions (ask before compiling; one per turn, with a recommended answer)

| # | Question | Recommended answer | Why (canon) |

|---|---|---|---|

| 1 | What single observable outcome means DONE? | A named artifact + a command that exits 0 against it | Verifier's law: invest in verifiability first |

| 2 | Which domain harness applies? | The domain whose skills name the deliverable; if two, run two sequential loops | Orchestrator-workers: scoped objectives beat mega-goals |

| 3 | What must NOT change? | List no-touch paths; put them in the goal text so the compiler's plan inherits them | Boundaries are part of a subagent spec |

| 4 | Who reviews escalations, and how fast? | A named human; escalations block the loop by design | Approval-required is a terminal state, not a nuisance |

| 5 | What is the iteration budget? | Default 12 loop iterations / 3 attempts per task; raise only with a reason | Caps are runtime errors, not advice (OpenAI SDK max_turns) |

Exit codes (branch on these mechanically)

| Code | Tool | Meaning |

|---|---|---|

| 0 | all | OK / directive emitted |

| 2 | loop_controller | Escalation required — a human must review the evidence log |

| 3 | goal_compiler | Goal too vague — answer the forcing questions, recompile |

| 4 | goal_compiler / loop_controller | No skill matched / close refused (unverified tasks) |

| 5 | loop_controller | Global iteration cap reached |

| 6 | loop_controller | Invalid transition (recording on verified task, evidence missing, unknown task) |

Verifiable success

  • python3 scripts/harness_manifest_builder.py --sample, scripts/goal_compiler.py --sample,

and scripts/loop_controller.py --sample all exit 0.

  • A vague goal (--goal "make it better") exits 3 and prints forcing questions.
  • loop_controller.py close on a state with an unverified task exits 4.
  • The demo loop in loop_controller.py --sample shows a verify failure consuming an attempt

and the loop still closing only after a passing verify with evidence.

Related skills

  • workflow-builder: authoring deterministic .js scripts for Claude Code's Workflow

tool. NOT for goal-to-close loop state (this skill).

  • agenthub: N parallel agents competing on ONE task in git worktrees. Use it inside a

harness task that wants competing attempts.

  • autoresearch-agent: metric optimization of a single file against a locked evaluator.

Use it when a task's done_when is "metric improves".

  • tc-tracker: per-code-change lifecycle records. Use for change bookkeeping; the harness

state file is per-goal, not per-change.

  • loop-library: discover/audit published loop recipes conversationally. This skill is the

executable enforcement of that vocabulary.

  • ship-gate / self-eval / spec-driven-workflow: plug in as close-time checks inside a

task's verification[].

See [references/domain_harness_design.md](references/domain_harness_design.md) for the

three-layer architecture, the reuse map, and how to raise a domain's harness quality.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 26,030
本站分层T1
该仓技能数846
原文件路径engineering/agent-harness/skills/agent-harness/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 846 个技能

同名技能的其他版本

有 2 个不同仓库或目录里都有叫 agent-harness 的技能。它们内容并不相同,别混用: