跳到主要内容
知仓学习社ZHICANG

aaai-experiments

Use when designing or auditing AAAI experiments for the broad-AI program committee, including baselines, ablations, statistical significance, robust…

不碰外部(只输出文字)无严重或高危命中brycewang-stanford/Awesome-Journal-Skills

它会碰到什么

扫了多少1 个文本文件,5 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

AAAI Experiments

Use this before submission to ensure empirical evidence supports the AI contribution. AAAI

reviewers may come from adjacent AI subfields, so experiments must be interpretable beyond one

benchmark community.

Experiment audit

  • Map every experimental block to a claim in the introduction.
  • Compare against strong, recent, and fairly tuned baselines.
  • Include ablations that isolate mechanisms rather than removing multiple components at once.
  • Report uncertainty, variance, and statistical tests when small differences matter.
  • Test robustness to data split, prompt, seed, environment, user population, or distribution shift

when relevant.

  • For human evaluation, document task, instructions, annotator pool, quality control, aggregation,

and ethics/IRB status.

  • Report compute, hardware, data access, model size, and training/inference cost.

Claim-to-evidence ledger

Build this table before adding new experiments. It keeps the AAAI evidence package aligned with

the main text and with the reproducibility checklist.

| Manuscript claim | Required evidence | Phase-1 risk if missing | Checklist hook |

| --- | --- | --- | --- |

| New AI capability | benchmark + qualitative failure cases | broad reviewer sees only engineering | datasets, metrics, baselines |

| Better mechanism | single-factor ablations | gain looks like tuning luck | ablation and hyperparameter answers |

| Robust deployment | shift / seed / subgroup stress test | result seems brittle | variance, compute, environment |

| Social-impact or safety claim | stakeholder, harm, and misuse analysis | ethical claim looks asserted | ethics, limitations, data access |

For each row, mark ready / weak / missing and name the fastest fix that can be run before the

supplementary-material deadline. Do not leave a claim in the abstract if its evidence row is weak.

AAAI-specific review pressure

  • Phase 1 reviewers need a fast reason to trust the evidence.
  • The reproducibility checklist must match the experiment descriptions.
  • AI for Social Impact and AI Alignment claims require stronger treatment of stakeholders, harms,

risk mitigation, and scope.

  • New results usually cannot rescue the paper in rebuttal, so submit complete evidence upfront.
  • The AI-assisted review pilot is non-decisional, but it may surface checklist mismatches; make

result provenance, seeds, data splits, and limits machine-readable enough that a human SPC/AC can

quickly audit them.

Pre-rebuttal freeze rule

Before submission, decide which experiments would be impossible to add later under AAAI's rebuttal

constraints: missing baselines, missing seeds, missing supplement files, or missing reproducibility

checklist answers. Treat those as pre-submission blockers, not rebuttal TODOs. The author

response can explain and clarify submitted evidence; it should not depend on new results, URLs, or

repaired supplementary files.

Evidence triage table

Because an AAAI reviewer from an adjacent subfield must trust your numbers quickly, classify each

experimental block by how much weight it can bear and what would strengthen it.

| Block | Carries the claim when | Reviewer doubt | Cheap reinforcement |

| --- | --- | --- | --- |

| Headline benchmark | beats tuned recent baselines | "lucky seed" | seeds, variance bars |

| Ablation | isolates one mechanism | "joint removal" | single-factor toggles |

| Robustness | holds across split/shift | "one setting" | extra split or perturbation |

| Human eval | protocol is documented | "rater bias" | IRB note, inter-rater agreement |

Common AAAI experiment rejects

  • Benchmark bump with no mechanism analysis, which a broad committee reads as engineering, not AI

insight.

  • Baselines weaker than current open-source systems, so the comparison looks unfair.
  • A Social-Impact or alignment claim with no stakeholder, harm, or risk-mitigation evidence.
  • Results that rely on a closed API with no reproducible substitute for the checklist.

Worked vignette

A planning paper reports a single-seed win on one domain. Audit: the headline block "needs

robustness" and "needs variance", so the fix before the deadline is five seeds with confidence

intervals plus one extra IPC-style domain. Because new results cannot rescue this in rebuttal, the

team runs both before submission and aligns the checklist's seed answer to the supplement.

Output format

[Claim] <paper claim>
[Evidence status] sufficient / needs baseline / needs ablation / needs robustness / unclear
[Fairness issue] <compute, tuning, data, prompt, metric, human eval>
[Checklist dependency] <what checklist answer this supports>
[Pre-rebuttal blockers] <missing evidence that must be run before submission>
[Fast fix] <experiment or analysis feasible before deadline>

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。