跳到主要内容
知仓学习社ZHICANG

arrowspace

Spectral vector search using graph Laplacian eigenstructure. Use when cosine/L2 similarity misses latent structure in your embeddings.

不碰外部(只输出文字)无严重或高危命中sickn33/agentic-awesome-skills

它会碰到什么

扫了多少1 个文本文件,4 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

ArrowSpace

Spectral vector search that augments nearest-neighbour search with graph Laplacian features. Computes a Laplacian over the item graph and uses the Rayleigh quotient to produce a λτ (lambda-tau) score per item, enabling search that respects both semantic similarity and structural role.

When to Use This Skill

  • Cosine or L2 similarity misses latent structure in your embeddings
  • You want graph-based retrieval with spectral awareness
  • You need to characterise the spectral properties of an embedding space
  • You are building RAG pipelines where contextual role matters alongside semantic content

How It Works

Step 1: Install and import

pip install arrowspace
from arrowspace import ArrowSpaceBuilder
import numpy as np

Step 2: Prepare your data

Pass an (N, d) float64 NumPy array of embedding vectors:

items = np.array([[0.1, 0.2, 0.3],
                  [0.0, 0.5, 0.1],
                  [0.9, 0.1, 0.0]], dtype=np.float64)

Step 3: Configure graph parameters

graph_params = {"eps": 0.2, "k": 6, "topk": 3, "p": 2.0, "sigma": 1.0}
builder = ArrowSpaceBuilder(items, graph_params=graph_params)
aspace = builder.build()

Step 4: Query

lambdas = aspace.lambdas()           # array indexed by insertion order
sorted_res = aspace.lambdas_sorted()  # (score, index) pairs ascending

Higher λτ values indicate items that are both semantically close and structurally central.

Examples

Example 1: Basic spectral retrieval

items = np.random.randn(100, 64).astype(np.float64)
builder = ArrowSpaceBuilder(items, graph_params={"eps": 0.5, "k": 10, "topk": 5, "p": 2.0, "sigma": None})
aspace = builder.build()
scores = aspace.lambdas()
top_indices = np.argsort(scores)[-5:]

Example 2: Compare spectral vs cosine ranking

from sklearn.metrics.pairwise import cosine_similarity
cos_sim = cosine_similarity(items)
cosine_order = np.argsort(cos_sim[0])[::-1]
spectral_order = np.argsort(aspace.lambdas())[::-1]

Best Practices

  • ✅ Normalise embeddings to unit norm before passing to ArrowSpace
  • ✅ Start with eps proportional to 1/sqrt(dim) and tune from there
  • ✅ Use k between 3 and 25 depending on dataset size (rule: N/50)
  • ✅ Set sigma=None to auto-select kernel width from distance distribution
  • ❌ Don't use with fewer than 10 items (graph structure is not meaningful)
  • ❌ Don't use for real-time streaming data (ArrowSpace is batch-oriented)

Limitations

  • This skill does not replace environment-specific validation, testing, or expert review.
  • ArrowSpace is batch-oriented and not designed for real-time indexing of streaming data.

Common Pitfalls

  • Problem: eps is too small, producing a disconnected graph

Solution: Increase eps, or set it proportional to 1/sqrt(embedding_dim)

  • Problem: k is too large, producing a dense graph with washed-out spectral features

Solution: Keep k ≤ 25 for most datasets

Related Skills

  • vector-database-engineer — General vector database expertise
  • embedding-strategies — Embedding model selection and chunking
  • similarity-search-patterns — Semantic search implementation patterns
  • hybrid-search-implementation — Combined semantic + keyword search

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

同名技能的其他版本

有 3 个不同仓库或目录里都有叫 arrowspace 的技能。它们内容并不相同,别混用: