Open-source picks · Vector Databases
Best Vector Databases
Updated daily. 60 projects, ranked by stars and checked for fork and maintenance activity.
Top Vector Databases
Established projects, still shipping. Tagged vector-database, vector-search on GitHub.
-
For developers, who are building real-time data-driven applications, Redis is the preferred, fastest, and most feature-rich cache, data structure server, and document and vector query engine.
-
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience
-
A lightning-fast search engine API bringing AI-powered hybrid search to your sites and applications.
-
Ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data. 🐳Docker-friendly.⚡Always in sync with Sharepoint, Google Drive, S3, Kafka, PostgreSQL, real-time data APIs, and more.
-
LlamaIndex is the document processing platform for AI
-
Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search
-
📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
-
Qdrant - High-performance, massive-scale Vector Database and Vector Search Engine for the next generation of AI. Also available in the cloud https://cloud.qdrant.io/
-
Open Source AI Platform - AI Chat with advanced features that works with every LLM
-
A modern replacement for Redis and Memcached
-
Cognee is the open-source AI memory platform for agents. Give your AI agents persistent long-term memory across sessions with a self-hosted knowledge graph engine.
-
This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.
-
The #1 AI Harness for Building Resumes, PDFs, Cover Letters & more, locally with 100+ LLMs support.
-
TencentDB Agent Memory is a team-level memory hub for AI Agents - turning conversations, docs, and code into four reusable memory assets (Chat Memory, Skill, LLM-Wiki, Code-Graph) that are governed, shared, and equipped across agents and frameworks.
-
Open Source alternative to Algolia + Pinecone and an Easier-to-Use alternative to ElasticSearch ⚡ 🔍 ✨ Fast, typo tolerant, in-memory fuzzy Search Engine for building delightful search experiences
-
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
-
A vector index built on TurboQuant, written in Rust with Python bindings
-
Weaviate is an open-source vector database that stores both objects and vectors, allowing for the combination of vector search with structured filtering with the fault tolerance and scalability of a cloud-native database.
-
Memory layer for AI Agents. Replace complex RAG pipelines with a serverless, single-file memory layer. Give your agents instant retrieval and long-term memory.
-
LangChain4j is an idiomatic, open-source Java library for building LLM-powered applications on the JVM. It offers a unified API over popular LLM providers and vector stores, and makes implementing tool calling (including MCP support), agents and RAG easy. It integrates seamlessly with enterprise Java frameworks like Quarkus and Spring Boot.
-
💡 All-in-one AI framework for semantic search, LLM orchestration and language model workflows
-
[MLsys2026 Best Paper]: https://arxiv.org/abs/2506.08276. RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on your personal device.
-
Code search MCP for Claude Code. Make entire codebase the context for any coding agent.
-
Developer-friendly OSS embedded retrieval library for multimodal AI. Search More; Manage Less.
-
Refine high-quality datasets and visual AI models
-
🌌 A complete search engine and RAG pipeline in your browser, server or edge network with support for full-text, vector, and hybrid search in less than 2kb.
-
OceanBase is the unified distributed database for the AI era - open-source, multi-model, one engine for your most demanding workloads.
-
Data Agent Ready Warehouse : One for Analytics, Search, AI, Python Sandbox. - rebuilt from scratch. Unified architecture on your S3.
-
One Postgres for your application data, full-text search, vector retrieval, and aggregations. Home of the pg_search extension.
-
Open Source Deep Research Alternative to Reason and Search on Private Data. Written in Python.
-
MariaDB server is a community developed fork of MySQL server. Started by core members of the original MySQL team, MariaDB actively works with outside developers to deliver the most featureful, stable, and sanely licensed open SQL server in the industry.
-
The AI search platform
-
Open-source framework for building agentic apps in JavaScript, Go, Dart, and Python, built and used in production by Google
-
A query and indexing engine for Redis, providing secondary indexing, full-text search, vector similarity search and aggregations.
-
Go Open Source, Distributed, Simple and efficient Search Engine
-
HelixDB is an OLTP graph database with native vector and full-text search built in Rust on Object Storage.
-
MineContext is your proactive context-aware AI partner(Context-Engineering+ChatGPT Pulse)
-
A distributed approximate nearest neighborhood search (ANN) library which provides a high quality vector index build, search and distributed online serving toolkits for large scale vector search scenario.
-
The AI-native database built for LLM applications, providing incredibly fast hybrid search of dense vector, sparse vector, tensor (multi-vector), and full-text.
-
Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.
New and rising Vector Databases
Created in the last 90 days and climbing fast.
-
Local-first search across your workspace, built for humans and AI agents.
-
XERJ is the new way for AI to search data. Its autoindex capability activates agents to know your data without the token waste of grep and sed. One command indexes code, docs, logs and PDFs for search, RAG, security audits and agent memory, using 40x fewer tokens than grep. Elasticsearch compatible, so existing clients just work.
-
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora - with retrieval and evaluation tooling included.
-
Encrypted, fully offline agentic memory. One click install, GUI w/ memory map, all OS and agents. Superior memory creation, storage and retrieval.
-
Markdown-first shared memory vault for Claude Code and Codex with SQLite, Zvec, Git, closeout, and audit
-
A free, self-paced 24-week AI engineering course: Python, machine learning, LLMs, RAG, fine-tuning, agents and MCP, Azure and Vertex and Bedrock, and Databricks. 43 runnable notebooks, one continuous case study. MIT licensed, no signup. By Zorost Intelligence AI Lab.
-
Engineering deterministic, production-grade systems around non-deterministic LLMs - FSM, durable execution, retries, DAGs, agent runtimes, model routing, edge inference, RAG, memory, multi-agent orchestration, security, and observability. 14 runnable proof-of-concept phases.
-
High-performance Knowledge Graph engine for AI, LLMs, and GraphRAG - built for the next generation of intelligent applications.
-
Spotlight-style local search for everything you saved and forgot: GitHub stars, local files, images, and bookmarks. Privacy-first, no full-disk scanning, fully on-device.
-
Drop-in AI memory layer with 2x faster retrieval and 10x lower cost. Fully compatible with Mem0 API. Migrate in 5 minutes without any code changes. Self-host for free.
-
Selfhost modern LLM stacks. Run the whole fleet from your terminal
-
🧠 The memory that dreams - cross-session memory for DeepSeek Harness. Offline & private, auto-consolidates in its sleep (autoDream), visualized in a memory panel.
-
Agentic event-venue operator demo built on MongoDB Atlas. Showcases three-layer memory (long-term, short-term, shared), hybrid retrieval with Voyage AI multimodal embeddings, optional Langfuse traces, and persona-segmented agent decisions during a rain delay scenario at a.
-
Billion-scale embedded vector database built entirely on Parquet and Arrow.
-
Comparison of memory systems for agents
-
RuvNet Brain - a downloadable, source-grounded brain for Claude Code over Reuven Cohen's (rUv's) RuvNet stack: RuVector/RVF, Ruflo, AgentDB, RuLake, SPARC + 21 building blocks. Grounds Claude in real source via one MCP tool (search_ruvnet), so it builds with the stack instead of drifting off it.
-
Local document Q&A in your terminal - powered by on-device LLMs.
-
A zero-to-100 learning path for applied AI engineering - RAG, embeddings, vector search, agents, MCP, and the production engineering around them. 56 pages, built as a searchable site.
-
Repo Explainer - turn any GitHub repo into a visual explainer page. Pipeline + 5 live examples.
-
Git for Agent Memory: Copy-On-Write vector branching for embedded multi-agent memory (83x faster, 3000x smaller snapshots)
How this list is built
Candidates come from the GitHub topics vector-database, vector-search. Star counts and weekly growth come from the official GitHub star history endpoint, so the numbers match what GitHub reports. Archived repositories and repositories carrying more than 20 topics are left out. A project is marked when its star count is high relative to its forks, which says the two moved apart, not why.