yao-ocr
OCR text recognition expert. ALWAYS invoke this skill when you need to extract text from images or PDFs — including invoices, receipts, ID cards, ba…
它会碰到什么
这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。
技能内容
OCR Tools
Two tools for optical character recognition, supporting both VLM-OCR (vision language models) and traditional OCR APIs (Baidu, Google, Azure, PaddleOCR).
ocr_recognize
Extract text from images or PDF files using OCR.
Basic usage (plain text output):
tai tool ocr_recognize --source /path/to/image.png
With URL:
tai tool ocr_recognize --source https://example.com/document.jpg
Table extraction as Markdown:
tai tool ocr_recognize --source /path/to/table.png --type table --output_format markdown
Invoice structured extraction:
tai tool ocr_recognize --source /path/to/invoice.pdf --type invoice --output_format json
With specific provider:
tai tool ocr_recognize --source /path/to/doc.png --provider baidu
VLM-OCR with custom prompt:
tai tool ocr_recognize --source /path/to/doc.png --provider llm:qwen-ocr --prompt "只提取表格中的金额列"
PDF page range:
tai tool ocr_recognize --source /path/to/report.pdf --pages "1-5" --output_format markdown
| Parameter | Type | Required | Description |
| ------------- | ------ | -------- | -------------------------------------------------------------------------------------------- |
| source | string | yes | Image or PDF file path/URL to recognize |
| provider | string | no | LLM connector ID (llm:xxx) or OCR settings key (baidu/paddleocr/google/azure). Auto-selects if omitted |
| type | string | no | Recognition type (default: general). See type table below |
| output_format | string | no | text (default), json (with coordinates/fields), or markdown (structured) |
| mode | string | no | accurate (default, best quality) or standard (faster) |
| language | string | no | Language hint (ISO 639-1, e.g. en, zh, ja). Auto-detected if omitted |
| prompt | string | no | Custom instruction for VLM-OCR only, appended to system prompt. Ignored by traditional OCR |
| pages | string | no | PDF page range, e.g. 1-5 or 1,3,7. All pages if omitted |
| extra | JSON | no | Provider-specific parameters as a JSON object |
Recognition types
| Type | Description | Best output_format |
| ---------------- | -------------------------- | ------------------ |
| general | General text (default) | text |
| table | Table extraction | markdown |
| handwriting | Handwritten text | text |
| document | Document layout parsing | markdown |
| invoice | Invoice (VAT) | json |
| receipt | Receipt / ticket | json |
| id_card | ID card | json |
| bank_card | Bank card | json |
| license | Business license | json |
| vehicle_license | Vehicle license | json |
| passport | Passport | json |
| license_plate | License plate | json |
If the chosen provider does not support the requested type, it automatically degrades to general and annotates the response metadata with degraded_from. VLM-OCR supports all types via prompt adaptation.
ocr_providers
List available OCR providers and their supported recognition types.
tai tool ocr_providers
Returns a list of providers including VLM-OCR models (from LLM connectors with ocr capability) and traditional API providers (from OCR settings). Each entry includes id, name, type (vlm or traditional), and supported_types.
PDF support
| Provider | PDF | Notes |
| ---------- | --- | ---------------------------------------- |
| Baidu | yes | pdf_file parameter |
| Azure | yes | Document Intelligence native support |
| PaddleOCR | yes | pdf + fileType=0 |
| Google | no | Sync API does not support PDF |
| VLM (llm:) | no | Vision models accept images only |
For providers that do not support PDF, use Baidu, Azure, or PaddleOCR instead.
Multi-page PDF response
Multi-page PDFs are automatically split page-by-page. Instead of printing all text, the tool returns a JSON summary with file paths for each page result:
{
"source": "report.pdf",
"total_pages": 10,
"pages": 3,
"results": [
{"page": 1, "file": ".tool-tmp/ocr-a1b2c3d4/page-1.txt", "preview": "Invoice No: INV-001..."},
{"page": 2, "file": ".tool-tmp/ocr-a1b2c3d4/page-2.txt", "preview": "Invoice No: INV-002..."},
{"page": 3, "file": ".tool-tmp/ocr-a1b2c3d4/page-3.txt", "preview": "Invoice No: INV-003..."}
]
}
To read full content of a specific page, use cat:
cat .tool-tmp/ocr-a1b2c3d4/page-2.txt
Single-page PDFs and images return inline text as usual (no file indirection).
Guidelines
- Use
output_format=text(default) when you just need the text content — simplest for LLM processing - Use
output_format=jsonfor structured types (invoice, id_card, etc.) to get key-value fields - Use
output_format=markdownfor documents and tables to preserve layout - The
promptparameter only works with VLM-OCR providers; traditional OCR ignores it - For structured document types (invoice, receipt, id_card, etc.), prefer
jsonoutput to getfieldswith key-value pairs - Use
ocr_providersfirst to check which providers are available and what types they support - Multi-page PDFs return a JSON summary with temporary file paths; use
cat <file>to read specific pages - Google Vision and VLM providers do not support PDF input directly; use Baidu, Azure, or PaddleOCR for PDF files
想直接用这个技能?
本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。