audit
Run a corpus-scale, STATS-ONLY PII audit over a folder of session transcripts LOCALLY and produce an aggregate report — counts by type and by layer,…
它会碰到什么
这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。
技能内容
confide:audit — corpus-scale, stats-only PII audit
Measure how much PII lives across a whole folder of sessions, without ever exposing any
of it. The audit runs the layered LOCAL detector stack from shared/confide_core.py
(regex → Natasha → local LLM) over each file and emits only aggregates. This mirrors
the real_session_eval privacy contract: read text only in-process, emit counts.
Privacy invariants (do not violate)
- Local-only. No cloud APIs. Raw transcript text never leaves the machine.
- Stats-only output. The report (markdown + json + optional HTML) contains ONLY
counts and rates — never a transcript substring, never a detected PII value.
- No filenames. Per-file rows are keyed by anonymized ids
own-00,own-01, …
The original path/name is never written or printed. On an unreadable file, only the
index + exception class name is recorded.
- Safe to surface. Because it is counts-only, the aggregate report can be shared
with a cloud agent or pasted into a chat. The PII stays on the machine.
What it reports
n_files, total / mean / min / max document charsspans_by_type(PERSON, EMAIL, PHONE, DATE, …) andspans_by_layer(regex / natasha / llm)overall_redaction_rateplus the per-session redaction-rate distribution
(min / median / mean / max)
- a coarse residual proxy: spans still detectable after redaction — ~0 on a clean RED
corpus, a leakage signal on a GREEN corpus.
Run it
Point it at a folder (recurses, processes every .md/.txt; skips confide's own
.green.md / .stats.json outputs):
python3 skills/audit/scripts/audit.py FOLDER
Options:
--list paths.txt— also/instead audit absolute paths listed one per line.--layers regex,natasha,llm— choose detection layers (default from config).
Use --layers regex for a fully offline, deterministic pass (no models/network).
--out report.md— report path; areport.jsonsibling is written alongside.--html— also write a Tufte-ish dashboard (report.html, counts only).
Writes the markdown + json report (and optional HTML) and prints the aggregate summary —
all counts only.
RED vs GREEN
- RED (raw) corpus: sizes the PII problem before any redaction.
- GREEN (redacted) corpus: the residual proxy and remaining
spans_by_typetell you
whether redaction is holding at scale.
After running
- Report the aggregate summary (file count, span totals by type/layer, redaction-rate
distribution, residual proxy) — never paste PII.
- If residual is non-trivial on a GREEN corpus, point the user at confide:anon to
re-redact and confide:red to probe re-identification risk.
Setup
Layer availability (Natasha, local LLM via Ollama) comes from config — run
confide:setup if they aren't installed. --layers regex always works offline.
想直接用这个技能?
本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。