Awesome Data Engineering
Section: Databases · An open-source search database for full-text, vector, and hybrid search with real-time indexing and SQL.
Entry
Appears in 4 awesome lists
Full-text search and data analytics, with fast response time for small, medium and big data (alternative to Elasticsearch). GPL-3.0 Docker/deb/C++/K8S
Section: Databases · An open-source search database for full-text, vector, and hybrid search with real-time indexing and SQL.
Section: 5. Retrieval-Augmented Generation (RAG) & Knowledge · Easy to use open source fast database for search. Good alternative to Elasticsearch with SQL-like interface and vector search capabilities.
Section: Search Engines · Full-text search and data analytics, with fast response time for small, medium and big data (alternative to Elasticsearch). GPL-3.0 Docker/deb/C++/K8S
Section: Other · Open-source search database for full-text, vector, and hybrid search with real-time indexing and SQL.
(MIT) provides modules for structured outputs at different levels of abstraction, including output parsers for text completion endpoints, Pydantic programs for mapping prompts to structured outputs using function calling or output parsing, and pre-defined Pydantic programs for specific output types.
Milvus is a cloud-native, open-source vector database built to manage embedding vectors generated by machine learning models and neural networks.
Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search,…
TiDB is built for agentic workloads that grow unpredictably, with ACID guarantees and native support for transactions, analytics, and vector search. No data silos. No noisy neighbors. No infrastructure ceiling.
is a search engine based on the Lucene library. It provides a distributed, multitenant-capable full-text search engine with an HTTP web interface and schema-free JSON documents. Elasticsearch is developed in Java.
General purpose, document-based, distributed database built for modern applications.
Python tool for converting files and office documents to Markdown. Supports PDF, PowerPoint, Word, Excel, images, audio, HTML, and more with OCR and transcription capabilities. MIT licensed.