跳到主要内容
知仓学习社ZHICANG

qa-quarto

Adversarial Quarto-vs-Beamer parity QA. A critic agent compares the Quarto HTML render to the Beamer PDF benchmark for content/visual parity; a fixe…

不碰外部(只输出文字)无严重或高危命中pedrohcgs/claude-code-my-workflow

它会碰到什么

扫了多少1 个文本文件,4 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

Adversarial Quarto vs Beamer QA Workflow

Compare Quarto HTML slides against their Beamer PDF benchmark using an iterative critic/fixer loop.

Philosophy: The Beamer PDF is the gold standard. The Quarto translation must be at least as good in every dimension.


Workflow

Phase 0: Pre-flight → Phase 1: Critic audit → Phase 2: Fixer → Phase 3: Re-audit → Loop until APPROVED (max 5 rounds)

Hard Gates (Non-Negotiable)

| Gate | Condition |

|------|-----------|

| Overflow | NO content cut off |

| Plot Quality | Interactive charts >= static plots |

| Content Parity | No missing slides/equations/text |

| Visual Regression | Quarto >= Beamer in all dimensions |

| Slide Centering | Content centered, no jumping |

| Notation Fidelity | All math verbatim from Beamer |

Phase 0: Pre-flight

  1. Locate Beamer (.tex/.pdf) and Quarto (.qmd/.html) files
  2. Check freshness (re-render if QMD newer than HTML)
  3. Verify TikZ SVGs if applicable

Phase 1: Initial Audit

Launch the quarto-critic agent to compare Beamer vs Quarto comprehensively. Report saved to quality_reports/[Lecture]_qa_critic_round1.md.

Phase 2: Fix Cycle

If not APPROVED, launch quarto-fixer agent to apply fixes (Critical → Major → Minor), re-render, and verify.

Phase 3: Re-Audit

Re-launch critic to verify fixes. Loop back to Phase 2 if needed.

Iteration Limits — loop-until-dry

This is the loop-until-dry primitive from [orchestrator-protocol.md](../../rules/orchestrator-protocol.md): the critic returns FINDINGs (the hard-gate table is the CRITICAL roll-up, per [orchestration-schemas.md](../../references/orchestration-schemas.md)); the loop converges when a round adds 0 new CRITICAL/MAJOR findings (deduped on id = sha1(file:line:locus)), not at a fixed round count.

  • Fallback cap: 5 rounds bounds a non-converging loop, then escalate to the user with remaining issues.
  • Two-strikes: the same gate failing in rounds N and N+2 is flagged for the user, not patched again ([summary-parity.md](../../rules/summary-parity.md)).
  • APPROVED iff every hard gate passes (zero CRITICAL).

Final Report

Save to quality_reports/[Lecture]_qa_final.md with hard gate status, iteration summary, and remaining issues.

Findings are validated, not just written (v2.5)

This skill's reviewers emit findings under the machine-checked contract in

[finding-schema.json](../../references/finding-schema.json). Reports are JSON arrays.

Smoke-test the harness before spending review effort — a run that fans out reviewers and

then cannot write a valid report has wasted the whole pass:

echo '[]' | python3 scripts/validate-findings.py

Then, before presenting any summary:

python3 scripts/validate-findings.py <report>.json   # exit 0 required

What the contract forces, and why:

  • rule — the documented rule or standard violated. A finding citing no rule is an

opinion, and opinions do not gate a commit.

  • failing_case — a concrete configuration under which the claim breaks, or the exact

missing hypothesis. "This could be clearer" does not validate.

  • id = sha1("<file>:<line>:<locus>") — deterministic, so dedup across rounds is

exact and the two-strikes rule is checkable rather than eyeballed.

  • mechanicaltrue only for fixes that cannot change a result (typo, cross-reference,

formatting, label). Never for an estimand, assumption, specification, inference

procedure, sample definition, or reporting language: those return to the researcher.

Apply the per-lens evidence burdens and the "does NOT count" filters in

[orchestration-schemas.md §7](../../references/orchestration-schemas.md) before

verification, so known false alarms never reach the judge. The verifier pass is

refute-biased: only verdict: "confirmed" findings ship; anything it cannot ground is

dropped, not downgraded to a warning.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。