extract-document-data
Extract structured, grounded fields from documents — values cite their page, missing values abstain instead of hallucinating. Use for parsing invoic…
它会碰到什么
这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。
技能内容
Extract Document Data
Extract structured JSON from documents with per-value grounding: every extracted value cites where it came from (page number, confidence), and values that aren't clearly present are reported in not_found rather than hallucinated. Uses the Stipple API (free anonymous tier).
When to use
- Parsing payslips, invoices, bank statements, receipts, or contracts
- Converting unstructured documents to JSON for downstream systems
- Any extraction where hallucinated values are worse than missing values (lending, accounting, compliance)
Instructions
- Get the document. URL or local file path (PDF, PNG, JPEG, DOCX).
- Choose the extraction mode:
- Ad-hoc fields — tell the API exactly which fields you want:
curl -X POST https://www.stipple.sh/v1/extract \
-F "file=@payslip.pdf" \
-F 'fields=[{"name":"employer_name"},{"name":"net_pay"},{"name":"pay_date"}]' \
-H "Authorization: Bearer $STIPPLE_API_KEY"
- Template — use a built-in schema:
payslip,tax_invoice,bank_statement,receipt,contract - Schema-free — omit
fieldsand let the model extract what it finds
- Interpret the response.
{
"mode": "schema_free",
"document_type": "payslip",
"pages_read": 1,
"fields": {
"employer_name": {"value": "Acme Cleaning Pty Ltd", "confidence": 0.95, "page": 1},
"net_pay": {"value": "2845.10", "confidence": 0.97, "page": 1}
},
"not_found": ["ytd_tax"]
}
- Every value carries
confidence(the model's self-report) andpage(grounding) not_found[]lists requested fields the model couldn't find — absences are reported, never guessedpages_readshows how many pages were processed (page limits apply per document)
- Report honestly. This is extraction, not verification — values are what the document shows, not proof it's genuine:
- "Employer: Acme Cleaning Pty Ltd (confidence 0.95, page 1)"
- "ytd_tax: not found in document" — never "ytd_tax: 0" or a guess
- For "is this document genuine?", pair with the
verify-documentskill first
Output format
Payslip fields (grounded, not guessed):
Employer Acme Cleaning Pty Ltd (confidence 0.95, page 1)
Employee J. Citizen (confidence 0.98, page 1)
Net pay 2,845.10 (confidence 0.97, page 1)
Superannuation 268.20 (confidence 0.93, page 1)
not_found: ytd_tax
(absences are reported, never hallucinated)
Limitations and Safety
- Invoices, statements, payslips, and contracts often contain sensitive personal,
financial, or commercial data. Obtain explicit approval before uploading them to
a hosted third party, minimize the submitted content, and confirm current
retention, residency, access, and deletion terms.
- Confidence and page grounding do not prove that an extracted value is correct or
that the source document is authentic. Reconcile consequential values against the
original document and authoritative systems before payment, lending, accounting,
compliance, or legal action.
- Keep the original file and extraction response so a human reviewer can reproduce
and correct disputed fields.
Notes
- Costs 1 credit per page read by the model (minimum 1); free weekly allowance applies
- Templates:
payslip,tax_invoice,bank_statement,receipt,contract— pass as thetemplateform field - Tables are extracted with structure preserved; multi-page documents are processed page by page
- Pairs with
verify-document(run first, for authenticity) — an extracted value from a tampered document is still wrong - Free key at https://www.stipple.sh for metering beyond the anonymous allowance
想直接用这个技能?
本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。
它属于哪个仓库
plugins/agentic-awesome-skills/skills/extract-document-data/SKILL.md同一个仓库里的其他技能
同名技能的其他版本
有 3 个不同仓库或目录里都有叫 extract-document-data 的技能。它们内容并不相同,别混用:
- sickn33/agentic-awesome-skills — Extract structured, grounded fields from documents — values cite their page, missing value
- sickn33/agentic-awesome-skills — Extract structured, grounded fields from documents — values cite their page, missing value