MCPcopy Create free account
hub / github.com/alibaba/zvec

github.com/alibaba/zvec

Chat with this repo
repository ↗ · DeepWiki ↗ · release v0.7.0 ↗ · + Follow · compare 3 versions
13,616 symbols 45,941 edges 1,151 files ⚖ custom 4,367 documented · 32% updated todayv0.7.0 · 2026-08-24★ 15,51426 open issues

Browse by type

Functions 11,660 Types & classes 1,906 Endpoints 50
What it actually does AI analysis from the code graph
loading…
README

English | 中文

<img src="https://zvec.oss-cn-hongkong.aliyuncs.com/logo/github_logo_1.svg" width="400" alt="zvec logo" />

Code Coverage Main License PyPI Release Python Versions npm Release

alibaba%2Fzvec | Trendshift

🚀 Quickstart | 🏠 Home | 📚 Docs | 📊 Benchmarks | 🔎 DeepWiki | 🎮 Discord | 🐦 X (Twitter)

Zvec is an open-source, in-process vector database — lightweight, lightning-fast, and designed to embed directly into applications. Battle-tested within Alibaba Group, it delivers production-grade, low-latency and scalable similarity search with minimal setup.

[!Important] 🚀 v0.6.0 (July 20, 2026)

  • Group-By Search: Retrieve top-K results per group instead of globally (group-by deduplication) across Flat, HNSW, HNSW-RaBitQ, and sparse indexes.
  • Random Rotation Quantization: Optional random rotation for INT8/INT4 quantization distributes variance evenly across dimensions, significantly boosting recall.
  • Enhanced Full-Text Search: Upgraded FTS pipeline with a Unicode UAX #29 standard tokenizer, UTF-8 / ASCII folding, and a Snowball-based stemmer supporting 34+ languages.
  • Faster & More Robust: Block-max skip speeds up FTS conjunction queries by 22–38%, plus a new DiskANN C API and numerous stability fixes.

👉 Read the Release Notes | View Roadmap 📍

💫 Features

  • Blazing Fast: Searches billions of vectors in milliseconds.
  • Simple, Just Works: Install and start searching in seconds. Pure local, no servers, no config, no fuss.
  • Dense + Sparse Vectors: Support dense and sparse embeddings, multi-vector queries, and a rich selection of vector index types that scale from memory to disk.
  • Full-Text Search (FTS): Native keyword-based full-text search — query string fields with natural-language or structured expressions.
  • Hybrid Search: Fuse vector similarity, full-text search, and structured filters in a single query for precise results.
  • Durable Storage: Write-ahead logging (WAL) guarantees persistence — data is never lost, even on process crash or power failure.
  • Concurrent Access: Multiple processes can read the same collection simultaneously; writes are single-process exclusive.
  • Runs Anywhere: As an in-process library, Zvec runs wherever your code runs — notebooks, servers, CLI tools, or even edge devices.

📦 Installation

Zvec offers official SDKs across multiple languages:

  • Python: pip install zvec (requires 64-bit Python 3.10–3.14)
  • Node.js: npm install @zvec/zvec
  • Go: High-performance Go bindings.
  • Rust: cargo add zvec-rust
  • Dart/Flutter: flutter pub add zvec

Prefer a visual tool? Try Zvec Studio to browse data and debug queries — no code required.

✅ Supported Platforms

  • Linux (x86_64, ARM64)
  • macOS (ARM64)
  • Windows (x86_64)

🛠️ Building from Source

If you prefer to build Zvec from source, please check the Building from Source guide.

⚡ One-Minute Example

import zvec

# Define collection schema
schema = zvec.CollectionSchema(
    name="example",
    vectors=zvec.VectorSchema("embedding", zvec.DataType.VECTOR_FP32, 4),
)

# Create collection
collection = zvec.create_and_open(path="./zvec_example", schema=schema)

# Insert documents
collection.insert([
    zvec.Doc(id="doc_1", vectors={"embedding": [0.1, 0.2, 0.3, 0.4]}),
    zvec.Doc(id="doc_2", vectors={"embedding": [0.2, 0.3, 0.4, 0.1]}),
])

# Search by vector similarity
results = collection.query(
    zvec.Query(field_name="embedding", vector=[0.4, 0.3, 0.3, 0.1]),
    topk=10
)

# Results: list of {'id': str, 'score': float, ...}, sorted by relevance
print(results)

📈 Performance at Scale

Zvec delivers exceptional speed and efficiency, making it ideal for demanding production workloads.

Zvec Performance Benchmarks

For detailed benchmark methodology, configurations, and complete results, please see our Benchmarks documentation.

🤝 Join Our Community

💬 DingTalk 📱 WeChat 🎮 Discord X (Twitter)
DingTalk QR Code WeChat QR Code Discord X (formerly Twitter) Follow
Scan to join Scan to join Click to join Click to follow

❤️ Contributing

We welcome and appreciate contributions from the community! Whether you're fixing a bug, adding a feature, or improving documentation, your help makes Zvec better for everyone.

Check out our Contributing Guide to get started!

Core symbols most depended-on inside this repo

browse all functions →

Shape

Method 8,916
Function 2,744
Class 1,788
Enum 118
Route 50

Languages

C++89%
Python9%
C2%

Modules by API surface

src/binding/c/c_api.cc403 symbols
src/include/zvec/core/framework/index_holder.h253 symbols
src/include/zvec/ailego/encoding/json/mod_json_plus.h194 symbols
src/ailego/encoding/json/mod_json.c182 symbols
src/include/zvec/ailego/buffer/concurrentqueue.h168 symbols
src/include/zvec/ailego/pattern/expected.hpp139 symbols
python/tests/test_embedding.py135 symbols
src/db/sqlengine/antlr/gen/SQLParser.cc131 symbols
src/db/index/segment/segment.cc97 symbols
src/include/zvec/ailego/container/vector.h93 symbols
tests/c/c_api_test.c85 symbols
src/db/sqlengine/antlr/gen/SQLParserBaseListener.h85 symbols

Dependencies from manifests, versioned

numpy1.23 · 1×

For agents

$ claude mcp add zvec \
  -- python -m otcore.mcp_server <graph>

⬇ download graph artifact

Ask about this repo answers extend the page