跳到主要内容
知仓学习社ZHICANG

audit

Run a corpus-scale, STATS-ONLY PII audit over a folder of session transcripts LOCALLY and produce an aggregate report — counts by type and by layer,…

写文件无严重或高危命中glebis/claude-skills

它会碰到什么

扫了多少2 个文本文件,21 KB
它会碰到什么写文件
命中总数3 处
命中统计严重 0 · 高 0 · 中 3 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

confide:audit — corpus-scale, stats-only PII audit

Measure how much PII lives across a whole folder of sessions, without ever exposing any

of it. The audit runs the layered LOCAL detector stack from shared/confide_core.py

(regex → Natasha → local LLM) over each file and emits only aggregates. This mirrors

the real_session_eval privacy contract: read text only in-process, emit counts.

Privacy invariants (do not violate)

  • Local-only. No cloud APIs. Raw transcript text never leaves the machine.
  • Stats-only output. The report (markdown + json + optional HTML) contains ONLY

counts and rates — never a transcript substring, never a detected PII value.

  • No filenames. Per-file rows are keyed by anonymized ids own-00, own-01, …

The original path/name is never written or printed. On an unreadable file, only the

index + exception class name is recorded.

  • Safe to surface. Because it is counts-only, the aggregate report can be shared

with a cloud agent or pasted into a chat. The PII stays on the machine.

What it reports

  • n_files, total / mean / min / max document chars
  • spans_by_type (PERSON, EMAIL, PHONE, DATE, …) and spans_by_layer (regex / natasha / llm)
  • overall_redaction_rate plus the per-session redaction-rate distribution

(min / median / mean / max)

  • a coarse residual proxy: spans still detectable after redaction — ~0 on a clean RED

corpus, a leakage signal on a GREEN corpus.

Run it

Point it at a folder (recurses, processes every .md/.txt; skips confide's own

.green.md / .stats.json outputs):

python3 skills/audit/scripts/audit.py FOLDER

Options:

  • --list paths.txt — also/instead audit absolute paths listed one per line.
  • --layers regex,natasha,llm — choose detection layers (default from config).

Use --layers regex for a fully offline, deterministic pass (no models/network).

  • --out report.md — report path; a report.json sibling is written alongside.
  • --html — also write a Tufte-ish dashboard (report.html, counts only).

Writes the markdown + json report (and optional HTML) and prints the aggregate summary —

all counts only.

RED vs GREEN

  • RED (raw) corpus: sizes the PII problem before any redaction.
  • GREEN (redacted) corpus: the residual proxy and remaining spans_by_type tell you

whether redaction is holding at scale.

After running

  1. Report the aggregate summary (file count, span totals by type/layer, redaction-rate

distribution, residual proxy) — never paste PII.

  1. If residual is non-trivial on a GREEN corpus, point the user at confide:anon to

re-redact and confide:red to probe re-identification risk.

Setup

Layer availability (Natasha, local LLM via Ollama) comes from config — run

confide:setup if they aren't installed. --layers regex always works offline.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。