跳到主要内容
知仓学习社ZHICANG

datasheets

Extract structured specifications from electronic component datasheet PDFs — pinouts, electrical characteristics, peripherals, topology, and feature…

执行命令写文件读文件联网严重 0 · 高危 3aklofas/kicad-happy

它会碰到什么

扫了多少57 个文本文件,563 KB
它会碰到什么执行命令写文件读文件联网
命中总数115 处
命中统计严重 0 · 高 3 · 中 8 · 低 104

这个仓库里自带 3 个测试样本文件(有些技能仓会放故意的恶意样本做演示),它们不计入上面的能力与命中。

逐条看命中(3 条严重或高危)
  • scripts/datasheet_page_selector.py:185exec-spawn
    result = subprocess.run(
  • scripts/datasheet_page_selector.py:221exec-spawn
    result = subprocess.run(
  • scripts/datasheet_page_selector.py:251exec-spawn
    result = subprocess.run(

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

Datasheets Skill

Related Skills

| Skill | Relationship |

|-------|--------------|

| digikey / mouser / lcsc / element14 | Producers — download the PDFs under <project>/datasheets/ that this skill extracts from |

| kicad | Primary consumer — VM-001/PU-001/FS-001/PP-001/LR-001/XT-001 + Phase 4b lookup detectors (AM-001/OV-001/TJ-001/FT-001/EX-001) query extractions via lookup(mpn) for verified-IC knowledge |

| emc | Consumer — switching-frequency, package-Rθ_JA, and operating-voltage data sharpen EMC heuristics |

| spice | Consumer — SPICE model presence + IBIS data feed simulation-readiness checks |

| thermal | Consumer — package Rθ_JA + junction temperature limits drive Tj estimates (TS-001..TJ-001) |

| bom | Indirect — coverage of structured extractions affects BOM verification confidence |

Handoff guidance: This skill is consumer infrastructure. The typical flow is distributor skill downloads PDF → datasheets skill extracts → analyzer skill queries. Use this skill directly when (a) the user asks to extract or verify a specific MPN, (b) an analyzer reports trust_level: low and the gap is per-MPN extraction quality, or (c) a new MPN was added to the BOM and downstream detectors should pick up its verified specs. Don't run this skill in isolation if the user just wants a design review — call it from the kicad workflow at the "Sync datasheets" step instead.

Purpose

Extract structured, machine-readable specifications from component datasheet PDFs and make them available to analyzer skills. Works on whatever PDFs are downloaded under <project>/datasheets/ (downloads are owned by distributor skills like digikey, mouser, lcsc, element14).

Scope

This skill owns:

  • Extraction schemas — canonical JSON structures for per-MPN specs. v1.4 ships 6 JSON Schema Draft 2020-12 schemas under schemas/ (base, pinout, spec_value, regulator, extraction, manifest) plus 5 v1.4 category extensions (diode, transistor, opamp, mcu, crystal). v1.3 cache format (EXTRACTION_VERSION in scripts/datasheet_extract_cache.py) is still read for compat.
  • Typed access layer (v1.4)datasheet_types/ package exposes DatasheetFacts, SpecValue, Pin, Pinout, lookup(), best(), trusted(), has_data(). Recommended for all new consumers.
  • PDF page selection — heuristics to pick pages most likely to contain pinouts, e-chars, applications, SPICE models.
  • Quality scoring — v1.4 uses a three-dimension rubric (pinout completeness, base completeness, category-extension completeness, 0–100 scale). v1.3 5-dimension weighted rubric still applies to legacy caches.
  • Consumer APIsscripts/datasheet_lookup.py for v1.4 typed access; scripts/datasheet_features.py for the v1.3 dict-shaped helpers (get_regulator_features, get_mcu_features, get_pin_function) — the v1.3 helpers dual-read v1.4 caches and translate to v1.3 dict shape for legacy detector code. Sunset planned for v1.6.
  • Verificationdatasheet_verify.py (v1.3, schema-vs-usage cross-check) plus datasheet_verify_v14_extraction (v1.4, power_domain references resolve, recommended ≤ absolute, regulator pin references exist).

Non-goals

  • No PDF downloading. That is owned by distributor skills (digikey, mouser, lcsc, element14).
  • No global library. Each project's extractions live in <project>/datasheets/extracted/. There is no shared cross-project cache.

Cache location

<project>/
  design.kicad_sch
  datasheets/
    TPS61023DRLR.pdf        # downloaded by distributor skills
    extracted/
      manifest.json         # extraction manifest (legacy name: index.json)
      TPS61023DRLR.json     # structured extraction (this skill's output)

Reference guides

  • references/extraction-schema.md — canonical schema, every field defined
  • references/field-extraction-guide.md — how to find each field in datasheets from common vendors (TI, ST, NXP, Espressif, Microchip)
  • references/quality-scoring.md — rubric details, score thresholds
  • references/consumer-api.md — how kicad/emc/spice/thermal consume extractions
  • references/cache-layout.md — v1.4 cache directory convention (per-MPN files, _families/ reservation, staleness rules)

Entry-point scripts

  • scripts/datasheet_extract_cache.py — v1.3 cache manager, resolver, indexer
  • scripts/datasheet_page_selector.py — page selection heuristics (used by both v1.3 and v1.4 pipelines)
  • scripts/datasheet_score.py — v1.3 extraction quality scoring
  • scripts/datasheet_verify.py — cross-check extraction vs schematic usage (v1.3 + v1.4 verify_v14_extraction mode)
  • scripts/datasheet_lookup.pyv1.4 typed lookup(mpn) → DatasheetFacts facade with staleness detection
  • scripts/datasheet_features.py — v1.3 consumer helper API (dual-reads v1.4 caches via _derive_*_v14 translators)
  • scripts/plan_extraction.pyv1.4 orchestration plan generator (Phase 3 extraction pipeline)
  • scripts/merge_results.pyv1.4 per-task result validator + merger
  • datasheet_types/v1.4 typed access layer package (DatasheetFacts, SpecValue, Pin, Pinout, lookup, best, trusted, has_data)

Extraction workflow

Run python3 skills/datasheets/scripts/plan_extraction.py <project> to generate an orchestration plan, then merge_results.py to validate and merge per-task outputs. Full scout→plan→dispatch→merge procedure: [references/extraction-pipeline.md](references/extraction-pipeline.md).

Consuming extractions (v1.4 typed API)

The recommended consumer surface is the typed lookup(mpn, cache_dir=...) facade plus the trust-gating helpers from datasheet_types. Import like:

import sys, pathlib
sys.path.insert(0, str(pathlib.Path(__file__).parent.parent / "datasheets"))
from datasheet_types import lookup, has_data, best, trusted

# Returns Optional[DatasheetFacts]. None on cache miss / stale PDF / low quality.
facts = lookup("TPS61023DRLR", cache_dir=pathlib.Path("datasheets/extracted"))
if facts is None:
    return  # heuristic-only path; no datasheet evidence available

# Field-level trust gating — every SpecValue list runs through has_data() / best() / trusted().
pu_range = facts.base.recommended_pullup_range  # Optional[list[SpecValue]]
if has_data(pu_range):
    # Most-trusted single value (first SpecValue meeting threshold, preserves extractor order).
    rec = best(pu_range, min_confidence="medium")  # Optional[SpecValue]
    if rec is not None and rec.min is not None:
        ...  # use rec.min, rec.max, rec.typ, rec.unit, rec.evidence.{page,section,confidence}

# All SpecValues at threshold (for multi-value fields like absolute_max).
hi_conf = trusted(facts.base.absolute_max.get("VDD", []), min_confidence="high")

Defensive patterns (mirrors kicad/SKILL.md § "Probing Analyzer JSON"):

  • lookup() returns None on cache miss, stale PDF (PDF newer than extraction), or quality score below the configured floor. Always guard with if facts is None: return.
  • Category extensions are optional on DatasheetFacts. facts.regulator is None when the part isn't in the regulator category — check before dereferencing.
  • SpecValue lists can be None (field not extracted), [] (extracted but empty), or list[SpecValue]. has_data() collapses the first two to False; pair with best() / trusted() for confidence gating.
  • SpecValue.min / .max / .typ are each Optional[float]. A SpecValue carrying only typ (no range) makes > / < comparisons against .min / .max raise TypeError — guard with explicit is not None chains on every numeric access.
  • confidence is one of "low" / "medium" / "high". Calling best() / trusted() with any other string raises ValueError.

v1.3 compat shim

Legacy detectors still call get_regulator_features(mpn) / get_mcu_features(mpn) / get_pin_function(mpn, pin) from scripts/datasheet_features.py. These dual-read v1.4 caches and translate to the v1.3 dict shape. Sunset planned for v1.6 — new code should use lookup() directly.

When to trigger this skill

  • Immediately after downloading datasheets via sync_datasheets_digikey.py, sync_datasheets_lcsc.py, or equivalent. Without extraction, IC-aware checks (VM-001 rail voltage, PS-001 power-good, PR-004 USB, DP-002 USB speed classification) fall back to heuristics on unknown ICs.
  • Before running analyzers on a new project where datasheets are present but datasheets/extracted/ is empty — the analyzers won't produce the extractions themselves.
  • When a review flags low trust level due to missing manufacturer evidence: extracting the ICs referenced by power regulators, MCUs, and high-speed peripherals typically flips trust_level: lowmixed or high.
  • When a user asks for pin verification ("verify U1 pin names match datasheet") — this skill's cached extraction is the authoritative source.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 1,235
本站分层T1
该仓技能数11
原文件路径skills/datasheets/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 11 个技能