trustworthy-experiment-insights
Assess whether experiment results are credible enough to influence product decisions. Use when checking false positive or false negative risk, under…
它会碰到什么
这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。
技能内容
Trustworthy Experiment Insights
Use this skill to decide whether an experiment result is believable enough to
shape a product or engineering decision. It focuses on false positives, false
negatives, power, replication, meta-analysis, stratified sampling, covariate
adjustment, and suspicious result review.
Source Traceability
Primary source: Next-Level A/B Testing by Leemay Nassery. Guidance is
transformed and paraphrased from Chapter 6 on false positives and negatives,
meta-analysis, metric sensitivity, stratified random sampling, covariate
adjustments, replication, longer runs, and statistical power.
Related skills:
ab-test-results-readoutfor standard experiment reporting.experiment-sensitivity-optimizationfor improving precision before or
during experiment design.
experiment-verification-monitoringfor operational validity checks.
Reference Routing
| Need | Read |
|------|------|
| Insight-quality concepts | references/core/knowledge.md |
| Credibility and follow-up rules | references/core/rules.md |
| Result-review scenarios | references/core/examples.md |
| Step-by-step credibility review | workflows/review-experiment-credibility.md |
Workflow
- Confirm the experiment was operationally valid enough to interpret.
- Check power, practical significance, and whether metrics were underpowered.
- Look for false positive risk: suspicious lift, many comparisons, early stop,
weak prior, or contradiction with prior experiments.
- Look for false negative risk: noisy metrics, small sample, low sensitivity,
or over-broad metric choice.
- Compare with similar experiments or run meta-analysis when available.
- Recommend launch, replicate, extend, investigate, or reject the result.
Output Format
# Experiment Insight Credibility Review
## Result Under Review
[Experiment, metric, observed result, and proposed decision.]
## Credibility Assessment
[Trust | Trust with caveats | Replicate | Extend | Investigate | Do not trust]
## Evidence
| Check | Finding | Risk |
|-------|---------|------|
## Follow-Up
- Replication needed:
- Longer run needed:
- Meta-analysis/comparison:
- Variance reduction opportunity:
## Decision Guidance
[What decision can be made now, and what should wait.]
Quality Bar
- Do not celebrate a result before checking whether it could be a false positive.
- Do not dismiss a flat result before checking power and sensitivity.
- Do not compare against prior experiments without noting differences in
population, metric, design, and timing.
- Do not use statistical checks to hide operational failures; verify experiment
health first.
想直接用这个技能?
本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。
它属于哪个仓库
plugins/LVTD-LLC/skills/skills/trustworthy-experiment-insights/SKILL.md