跳到主要内容
知仓学习社ZHICANG

atc-experiments

Use when designing or auditing the evaluation of an ATC (ACM SIGOPS Annual Technical Conference, formerly USENIX ATC) systems paper — matching evide…

不碰外部(只输出文字)无严重或高危命中brycewang-stanford/Awesome-Journal-Skills

它会碰到什么

扫了多少1 个文本文件,5 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

ATC Experiments

Match the evidence to the claim. ATC is the systems community's implementation-and-measurement

venue: reviewers read for measured behavior on a real system, not asymptotics or accuracy on a

dataset. In round two, 3-4 reviewers close to your subarea will open the artifact and probe whether

the numbers are end-to-end, fair, and honest about cost. Design the evaluation so their first three

objections are already answered.

Match evidence to claim shape

| If your claim is... | The evidence ATC expects |

|---|---|

| "Faster / lower latency" | End-to-end latency including tails (p99/p999) and throughput at a matched operating point, on a described testbed |

| "Lower overhead / cheaper" | The resource cost (CPU, memory, writes, energy) measured, at matched function — not just the headline win |

| "Scales" | Measurements across a real range of load/nodes/cores with the scaling curve and where it bends |

| "More reliable / correct" | Fault-injection or crash/recovery experiments, not just steady-state runs |

| "Useful in practice" (experience) | Production-derived workloads and lessons; what broke and what generalizes |

Real testbeds and workloads

  • Describe the testbed so results are reproducible: CPU/NIC/SSD models, core counts, memory,

kernel/OS versions, network topology, and any co-location. A result without its testbed is not a

systems result.

  • Use realistic workloads. Production-derived traces, standard benchmarks, or documented

generators beat hand-picked inputs. State how the workload was obtained and why it is

representative; if it is synthetic, justify the parameters.

  • Warm-up and steady state. Say how you handled cold start, warm-up windows, and measurement

duration — systems reviewers know where transient effects hide.

Fair baselines

  • Compare against the strongest reasonable alternative, configured well (a strawman baseline is

caught immediately). If you tuned your system, tune the baseline.

  • Compare at a matched cost or operating point: same memory budget, same flash-write budget,

same load. An unmatched comparison is the classic systems-reviewer objection.

  • If no baseline exists, say so and use the unmodified system or an ablation of your own design

as the reference.

End-to-end plus microbenchmarks

ATC reviewers want both:

  • End-to-end results show the contribution matters for the whole system under a real workload.
  • Microbenchmarks isolate the mechanism, attributing the win (or cost) to your design rather

than to unrelated system effects. A paper with only end-to-end numbers cannot explain why; one

with only microbenchmarks cannot show it matters.

Tails, variance, and honest reporting

  • Report tail latency (p99, often p999), not just means — the tail is where systems pain lives.
  • Report variance across repeated runs (multiple trials, min/max or CIs). A single run is a data

point, not a result.

  • Report the cost beside the gain, at the matched operating point (see atc-writing-style). A

win with an unstated cost reads as a hidden weakness.

  • State negative or neutral regions honestly — "where the working set fits, our policy neither helps

nor hurts" builds more trust than a uniformly rosy curve.

Provenance you cannot reconstruct later

Pin these at collection time — they cannot be recovered at the deadline (see atc-reproducibility):

[Hardware]   CPU/NIC/SSD models, core/memory counts, firmware where it matters
[Software]   kernel/OS versions, library and compiler versions, config flags
[Workload]   trace source and date, generator version and seeds, request mix
[Method]     warm-up window, measurement duration, number of runs, aggregation
[Code]       commit SHAs for the system and every baseline

Special cases

  • Concurrency/nondeterminism: report the distribution and the scheduling/affinity settings, not

a lucky run.

  • Energy/power claims: name the measurement instrument and boundary (wall vs. component).
  • Security/isolation claims: state the threat model and what the measurement does and does not

cover.

  • Experience papers: the "evaluation" is the deployment itself — scale, duration, incidents, and

transferable lessons; ATC's Deployed Systems lane values this even without a new mechanism.

Output format

[Claim -> evidence] each claim mapped to the experiment that supports it; gaps flagged
[Testbed] hardware/software/workload described enough to reproduce? yes/no
[Baselines] strongest alternative, well-configured, at a matched operating point? yes/no
[Depth] end-to-end AND microbenchmarks present? tails + variance reported?
[Honesty] costs reported beside gains? neutral/negative regions stated?
[Provenance] hardware/software/workload/method/code pinned at collection time? yes/no

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。