跳到主要内容
知仓学习社ZHICANG

trustworthy-experiment-insights

Assess whether experiment results are credible enough to influence product decisions. Use when checking false positive or false negative risk, under…

不碰外部(只输出文字)无严重或高危命中hashgraph-online/awesome-codex-plugins

它会碰到什么

扫了多少6 个文本文件,15 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

Trustworthy Experiment Insights

Use this skill to decide whether an experiment result is believable enough to

shape a product or engineering decision. It focuses on false positives, false

negatives, power, replication, meta-analysis, stratified sampling, covariate

adjustment, and suspicious result review.

Source Traceability

Primary source: Next-Level A/B Testing by Leemay Nassery. Guidance is

transformed and paraphrased from Chapter 6 on false positives and negatives,

meta-analysis, metric sensitivity, stratified random sampling, covariate

adjustments, replication, longer runs, and statistical power.

Related skills:

  • ab-test-results-readout for standard experiment reporting.
  • experiment-sensitivity-optimization for improving precision before or

during experiment design.

  • experiment-verification-monitoring for operational validity checks.

Reference Routing

| Need | Read |

|------|------|

| Insight-quality concepts | references/core/knowledge.md |

| Credibility and follow-up rules | references/core/rules.md |

| Result-review scenarios | references/core/examples.md |

| Step-by-step credibility review | workflows/review-experiment-credibility.md |

Workflow

  1. Confirm the experiment was operationally valid enough to interpret.
  2. Check power, practical significance, and whether metrics were underpowered.
  3. Look for false positive risk: suspicious lift, many comparisons, early stop,

weak prior, or contradiction with prior experiments.

  1. Look for false negative risk: noisy metrics, small sample, low sensitivity,

or over-broad metric choice.

  1. Compare with similar experiments or run meta-analysis when available.
  2. Recommend launch, replicate, extend, investigate, or reject the result.

Output Format

# Experiment Insight Credibility Review

## Result Under Review
[Experiment, metric, observed result, and proposed decision.]

## Credibility Assessment
[Trust | Trust with caveats | Replicate | Extend | Investigate | Do not trust]

## Evidence
| Check | Finding | Risk |
|-------|---------|------|

## Follow-Up
- Replication needed:
- Longer run needed:
- Meta-analysis/comparison:
- Variance reduction opportunity:

## Decision Guidance
[What decision can be made now, and what should wait.]

Quality Bar

  • Do not celebrate a result before checking whether it could be a false positive.
  • Do not dismiss a flat result before checking power and sensitivity.
  • Do not compare against prior experiments without noting differences in

population, metric, design, and timing.

  • Do not use statistical checks to hide operational failures; verify experiment

health first.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 1,027
本站分层T1
该仓技能数1910
原文件路径plugins/LVTD-LLC/skills/skills/trustworthy-experiment-insights/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 1910 个技能