跳到主要内容
知仓学习社ZHICANG

picking-a-format

Use when choosing an output format for extracted documents — text, markdown, djot, html, or JSON. Maps consumer (LLM, parser, archive) to the right …

不碰外部(只输出文字)无严重或高危命中hashgraph-online/awesome-codex-plugins

它会碰到什么

扫了多少1 个文本文件,4 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

Picking a format

Kreuzberg has two orthogonal format knobs. Get them right up front and the

downstream code stays simple.

| Knob | What it controls | Values | Default |

| ------------------- | ------------------------------------------------- | -------------------------------------- | ---------------- |

| --format | How the CLI prints the result | text, json | text (extract), json (batch) |

| --content-format | How extracted content is rendered inside result | plain, markdown, djot, html | plain |

| --token-reduction | Strip whitespace / boilerplate for LLM contexts | off, light, moderate, aggressive | off |

--format json always returns the full ExtractionResult (content +

metadata + tables + images). --format text prints just content.

--content-format is what shows up inside that content field.

Decision tree

Who consumes the output?
├── LLM (Claude, GPT, Gemini, local) — embed/prompt context
│       --format text --content-format markdown
├── Vector store / RAG indexer
│       --format json --content-format markdown
│       (markdown preserves structure for chunking)
├── Downstream parser that expects machine-readable JSON
│       --format json --content-format plain
│       (cleanest text + structured metadata)
├── Human review / archival
│       --format text --content-format markdown
├── HTML re-rendering / web display
│       --format json --content-format html
├── Lossless intermediate for pandoc / academic tooling
│       --format json --content-format djot
└── Token-budget-constrained pipeline
        --format text --content-format plain
        (drops markup; add --token-reduction moderate for further savings)

Examples

Feed a PDF directly into an LLM:

kreuzberg extract paper.pdf --content-format markdown

Index a corpus into a RAG store with tables and headings preserved:

kreuzberg batch docs/*.pdf --format json --content-format markdown \
  | jq -c '.[] | {path: .metadata.path, content: .content, tables: .tables}'

Strip a file to bare text for a token-tight summarizer:

kreuzberg extract long.pdf \
  --content-format plain \
  --token-reduction moderate

Pull metadata only, ignore content:

kreuzberg extract file.pdf --format json | jq '.metadata'

When in doubt

  • Default to markdown as the content format. It is the best

compromise across LLMs, RAG, and human review, and Kreuzberg has the

most faithful renderer for it.

  • Reach for plain only when downstream cannot tolerate any markup.
  • Reach for djot only if you're already in a djot/pandoc pipeline.
  • Reach for html only when re-rendering for the web.

Token-reduction (orthogonal)

--token-reduction collapses whitespace, strips repeated headers/footers,

and trims boilerplate. It composes with any --content-format:

  • off (default), light, moderate, aggressive, maximum.

Use moderate as a safe starting point for LLM context windows. maximum

is lossy — verify before relying on it.

See references/cli-reference.md for the full flag set and

references/configuration.md for the equivalent output_format and

token_reduction keys in kreuzberg.toml.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 1,027
本站分层T1
该仓技能数1910
原文件路径plugins/kreuzberg-dev/plugins/plugins/kreuzberg/skills/picking-a-format/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 1910 个技能