Fast State-of-the-Art Static Embeddings
-
Updated
Oct 5, 2026 - Python
Fast State-of-the-Art Static Embeddings
Fast Multimodal Semantic Deduplication & Filtering
Pre-train Static Word Embeddings
벼리 — headless hybrid (lexical + semantic) search for a Markdown vault. No app, no GPU, no cloud. Korean/cross-lingual. Apache-2.0.
Fast, local semantic search over web content for AI agents. Hybrid BM25 + potion-retrieval-32M embeddings, cross-page dedup, token-budget mode, MCP server, SearXNG bridge. ~90% fewer tokens than raw web_fetch.
An effort to build a holistic matching engine that supports regex, cucumber expressions, semantic phrases, and combinations of that. Could be useful for Guardrails, Caching systems etc. Can eventually be a symbolic AI engine
One install, three MCPs: graph chat memory + semantic code search + backend. Benchmarked with claude-haiku-4-5: -19% tokens & 23× faster on memory-recall questions.
Reading the profile, not the keywords — a hybrid, CPU-only ranker that surfaces the top-100 candidates for a Senior AI Engineer JD, beating keyword-stuffers and honeypots. Structured JD-fit + contrastive static embeddings + behavioral/honeypot gating. No LLM at inference; reproducible in ~3 min.
Semantic cache for LLM calls with no torch, no server, no API key and no GPU. One file, CPU static embeddings. Benchmarked on 5,000 real question pairs: 70% of rephrasings served from cache, 28% false matches.
Self-learning AI workload router — routes each task to the cheapest model that can actually do it, and gets smarter with every outcome.
Semantic code search MCP server: tree-sitter chunking + local Model2Vec embeddings + BM25/dense hybrid (RRF). Hands an AI only the relevant code (file:line), ~10-100x fewer tokens. Multi-language, no API key.
To associate your repository with the model2vec topic, visit your repo's landing page and select "manage topics."