跳到主要内容
知仓学习社ZHICANG

skillgrade-graders

Authors deterministic and LLM rubric graders for skillgrade evaluations. Use when creating scoring scripts, writing evaluation rubrics, or combining…

不碰外部(只输出文字)无严重或高危命中mgechev/skillgrade

它会碰到什么

扫了多少2 个文本文件,5 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

Skillgrade Grader Authoring

Procedures

Step 1: Identify the Grading Strategy

  1. Determine whether the task requires objective verification (deterministic) or qualitative assessment (LLM rubric).
  2. For most tasks, combine both: deterministic graders verify outcomes (weight 0.7), LLM rubrics assess approach quality (weight 0.3).

Step 2: Write a Deterministic Grader

  1. Create a script in the skill's graders/ directory (bash or TypeScript).
  2. The script must output a JSON object to stdout with the following structure:
   {"score": 0.67, "details": "2/3 checks passed", "checks": [{"name": "check-name", "passed": true, "message": "Description"}]}
  1. score (0.0–1.0) and details are required. checks is optional but recommended.
  2. Read references/grader-output-schema.md for the full output specification.
  3. Use awk for arithmetic in bash scripts — bc is not available in node:20-slim.
  4. Reference the grader in eval.yaml:
   - type: deterministic
     run: bash graders/check.sh
     weight: 0.7

Step 3: Write an LLM Rubric Grader

  1. Draft a rubric with explicit scoring criteria and point allocations.
  2. Structure the rubric into weighted sections that sum to 1.0:
   Workflow Compliance (0-0.5):
   - Did the agent follow the mandatory workflow steps?
   Efficiency (0-0.5):
   - Completed in ≤5 commands without trial-and-error?
  1. Reference the rubric in eval.yaml:
   - type: llm_rubric
     rubric: |
       [rubric text or file path]
     weight: 0.3
     provider: gemini               # optional: gemini (default) | anthropic | openai
     model: gemini-3.5-flash        # optional model override (defaults to the latest dynamically resolved flash model)
  1. For long rubrics, store in a separate file and reference by path: rubric: rubrics/quality.md.

Step 4: Combine Multiple Graders

  1. Assign weights to each grader based on importance. Weights are normalized automatically.
  2. Final reward is calculated as: Σ (grader_score × weight) / Σ weight.
  3. Example configuration:
   graders:
     - type: deterministic
       run: bash graders/check.sh
       weight: 0.7
     - type: llm_rubric
       rubric: rubrics/quality.md
       weight: 0.3

Step 5: Validate Graders

  1. Create a reference solution script that produces the expected output.
  2. Run skillgrade --validate to verify graders score the reference solution correctly.
  3. Test only deterministic graders: skillgrade --grader=deterministic (skips LLM calls, faster iteration).
  4. Test only LLM rubric graders: skillgrade --grader=llm_rubric.
  5. Run a specific eval with a specific grader type: skillgrade --eval=my-eval --grader=deterministic.
  6. If a grader returns unexpected scores, inspect the script output and adjust scoring logic.

Error Handling

  • If a deterministic grader outputs non-JSON, ensure all echo/console.log statements except the final JSON result are redirected to stderr.
  • If an LLM rubric grader returns 0.00 with a missing API key message, set the appropriate key for your provider: GEMINI_API_KEY (provider: gemini), ANTHROPIC_API_KEY (provider: anthropic), or OPENAI_API_KEY (provider: openai).
  • To use a custom/self-hosted LLM endpoint, set ANTHROPIC_BASE_URL (for provider: anthropic) or OPENAI_BASE_URL (for provider: openai) — e.g. for Ollama or vLLM.
  • If scores are inconsistent across trials, reduce rubric ambiguity by adding concrete examples of passing and failing behavior.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 710
本站分层T2
该仓技能数2
原文件路径skills/skillgrade-graders/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 2 个技能