> SingleStore 是一个分布式 SQL 数据库,用于事务性和分析性工作负载。您可以在 云 或本地运行。
SingleStore 支持向量存储和基于 SQL 的相似性搜索。它包含向量函数,例如 点积_ 和 欧几里得_距离。有关表设计、索引和查询模式,请参阅 使用向量数据 ,请参阅 SingleStore 文档。
您还可以将向量搜索与 基于 Lucene 的全文索引 和文档元数据过滤相结合。根据您的工作负载,您可以预过滤文本或向量,或组合分数(例如,使用加权求和)。
使用以下部分将 SingleStore 连接到 LangChain。
| 类 | 包 | JS 支持 |
|---|---|---|
SingleStoreVectorStore | langchain_singlestore | ✅ |
设置
要访问 SingleStore 向量存储,您需要安装 langchain-singlestore 集成包。 pip install -qU "langchain-singlestore"
初始化
要初始化 SingleStoreVectorStore,您需要一个 Embeddings 对象和 SingleStore 数据库的连接参数。
必需参数
- 嵌入 (
Embeddings):文本嵌入模型。
可选参数
- 距离_策略 (
DistanceStrategy):计算向量距离的策略。默认为DOT_PRODUCT。选项: - -
DOT_PRODUCT:计算两个向量的标量积。 - -
EUCLIDEAN_DISTANCE:计算两个向量之间的欧几里得距离。 - 表_名称 (
str):表的名称。默认为embeddings. - 内容_字段 (
str):存储内容的字段。默认为content. - 元数据_字段 (
str):存储元数据的字段。默认为metadata. - 向量_字段 (
str):存储向量的字段。默认为vector. - id_字段 (
str):存储 ID 的字段。默认为id. - 使用_向量_索引 (
bool):启用向量索引(需要 SingleStore 8.5+)。默认为False. - 向量_索引_名称 (
str):向量索引的名称。如果use_vector_indexisFalse. - 向量_索引_选项 (
dict):向量索引的选项。如果use_vector_indexisFalse. - 向量_大小 (
int):向量的大小。如果use_vector_indexisTrue. - 使用_完整_文本_搜索 (
bool):启用内容全文索引。默认为False.
连接池参数
- 池_大小 (
int):池中活动连接数。默认为5. - 最大_溢出 (
int):超出pool_size的最大连接数。默认为10. - 超时 (
float):连接超时时间(秒)。默认为30.
数据库连接参数
- 主机 (
str):数据库的主机名、IP 或 URL。 - 用户 (
str):数据库用户名。 - 密码 (
str):数据库密码。 - 端口 (
int):数据库端口。默认为3306. - 数据库 (
str):数据库名称。
其他选项
- 纯_python (
bool):启用纯 Python 模式。 - 本地_输入文件 (
bool):允许本地文件上传。 - 字符集 (
str):字符串值的字符集。 - ssl_密钥, **ssl_证书**, **ssl_ca** (
str):SSL 文件路径。 - ssl_禁用 (
bool):禁用 SSL。 - ssl_验证_证书 (
bool):验证服务器证书。 - ssl_验证_身份 (
bool): 验证服务器身份。 - autocommit (
bool): 启用自动提交。 - results_type (
str): 查询结果结构(例如,tuples,dicts).
from langchain_singlestore.vectorstores import SingleStoreVectorStore
os.environ["SINGLESTOREDB_URL"] = "root:pass@localhost:3306/db"
vector_store = SingleStoreVectorStore(embeddings=embeddings)
管理向量存储
该 SingleStoreVectorStore 假设文档的ID是一个整数。以下是管理向量存储的示例。
向向量存储添加项目
您可以按以下方式向向量存储添加文档:
pip install -qU langchain-core
from langchain_core.documents import Document
docs = [
Document(
page_content="""In the parched desert, a sudden rainstorm brought relief,
as the droplets danced upon the thirsty earth, rejuvenating the landscape
with the sweet scent of petrichor.""",
metadata={"category": "rain"},
),
Document(
page_content="""Amidst the bustling cityscape, the rain fell relentlessly,
creating a symphony of pitter-patter on the pavement, while umbrellas
bloomed like colorful flowers in a sea of gray.""",
metadata={"category": "rain"},
),
Document(
page_content="""High in the mountains, the rain transformed into a delicate
mist, enveloping the peaks in a mystical veil, where each droplet seemed to
whisper secrets to the ancient rocks below.""",
metadata={"category": "rain"},
),
Document(
page_content="""Blanketing the countryside in a soft, pristine layer, the
snowfall painted a serene tableau, muffling the world in a tranquil hush
as delicate flakes settled upon the branches of trees like nature's own
lacework.""",
metadata={"category": "snow"},
),
Document(
page_content="""In the urban landscape, snow descended, transforming
bustling streets into a winter wonderland, where the laughter of
children echoed amidst the flurry of snowballs and the twinkle of
holiday lights.""",
metadata={"category": "snow"},
),
Document(
page_content="""Atop the rugged peaks, snow fell with an unyielding
intensity, sculpting the landscape into a pristine alpine paradise,
where the frozen crystals shimmered under the moonlight, casting a
spell of enchantment over the wilderness below.""",
metadata={"category": "snow"},
),
]
vector_store.add_documents(docs)
更新向量存储中的项目
要更新向量存储中的现有文档,请使用以下代码:
updated_document = Document(
page_content="qux", metadata={"source": "https://another-example.com"}
)
vector_store.update_documents(document_id="1", document=updated_document)
从向量存储删除项目
要从向量存储中删除文档,请使用以下代码:
vector_store.delete(ids=["3"])
查询向量存储
一旦您的向量存储已创建并且相关文档已添加,您很可能希望在运行链或代理时对其进行查询。
直接查询
执行简单的相似性搜索可以按如下方式完成:
results = vector_store.similarity_search(query="trees in the snow", k=1)
for doc in results:
print(f"* {doc.page_content} [{doc.metadata}]")
如果您想执行相似性搜索并接收相应的分数,可以运行:
- - 待办事项:编辑并运行代码单元以生成输出
results = vector_store.similarity_search_with_score(query="trees in the snow", k=1)
for doc, score in results:
print(f"* [SIM={score:3f}] {doc.page_content} [{doc.metadata}]")
元数据过滤
SingleStoreDB通过基于元数据字段进行预过滤,使用户能够增强和优化搜索结果,从而提升搜索能力。这项功能使开发者和数据分析师能够微调查询,确保搜索结果精确符合其需求。通过使用特定元数据属性过滤搜索结果,用户可以缩小查询范围,仅关注相关的数据子集。
SingleStoreVectorStore支持使用强大的查询运算符进行简单和高级的元数据过滤。
简单元数据过滤
使用简单的字典式语法进行精确匹配和向后兼容:
# Filter by a single field
query = "trees branches"
docs = vector_store.similarity_search(
query, filter={"category": "snow"}
)
# Filter by multiple fields (implicit AND)
docs = vector_store.similarity_search(
query="landmarks",
filter={"country": "France", "category": "museum"}
)
高级元数据过滤
使用带有运算符的高级过滤器,例如 $eq, $gt, $in, $and, $or,以及更多用于复杂查询:
比较运算符:
# Greater than, less than, and other comparisons
results = vector_store.similarity_search(
query="old structures",
k=10,
filter={"year_built": {"$lt": 1900}} # Built before 1900
)
# Other operators: $eq, $ne, $gt, $gte, $lte
results = vector_store.similarity_search(
query="landmarks",
filter={"year_built": {"$gte": 1800, "$lte": 1950}}
)
集合运算符:
# Check if value is in a list
results = vector_store.similarity_search(
query="landmarks",
k=10,
filter={"country": {"$in": ["France", "UK"]}}
)
# Not in ($nin)
results = vector_store.similarity_search(
query="museums",
filter={"country": {"$nin": ["USA", "Canada"]}}
)
存在性检查:
# Check if a field exists
results = vector_store.similarity_search(
query="heritage sites",
k=10,
filter={"heritage_status": {"$exists": True}}
)
逻辑运算符:
# Combine multiple conditions with $and
results = vector_store.similarity_search(
query="european landmarks",
k=10,
filter={
"$and": [
{"category": "landmark"},
{"year_built": {"$gte": 1800}},
{"country": {"$in": ["France", "UK"]}}
]
}
)
# Use $or for alternative conditions
results = vector_store.similarity_search(
query="cultural sites",
filter={
"$or": [
{"category": "museum"},
{"category": "landmark"}
]
}
)
# Complex nested queries
results = vector_store.similarity_search(
query="cultural sites",
k=10,
filter={
"$or": [
{
"$and": [
{"category": "museum"},
{"country": "France"}
]
},
{
"$and": [
{"category": "landmark"},
{"year_built": {"$lt": 1900}}
]
}
]
}
)
向量索引
通过利用 ANN向量索引来提升您的搜索效率,适用于SingleStore DB 8.5或更高版本。通过在创建向量存储对象时设置 use_vector_index=True ,您可以激活此功能。此外,如果您的向量维度与默认的OpenAI嵌入大小1536不同,请确保相应地指定 vector_size 参数。
搜索策略
SingleStoreDB提供了多种搜索策略,每种都经过精心设计以满足特定的用例和用户偏好。默认的 VECTOR_ONLY 策略利用向量操作(如 dot_product or euclidean_distance )在向量之间直接计算相似度分数,而 TEXT_ONLY 采用基于Lucene的全文本搜索,对以文本为中心的应用特别有利。对于寻求平衡方案的用户, FILTER_BY_TEXT 首先根据文本相似性优化结果,然后进行向量比较,而 FILTER_BY_VECTOR 优先考虑向量相似性,在评估文本相似性之前先筛选结果以获得最佳匹配。值得注意的是,两者 FILTER_BY_TEXT 和 FILTER_BY_VECTOR 都需要全文索引才能运行。此外, WEIGHTED_SUM 作为一种复杂的策略应运而生,通过权衡向量相似性和文本相似性来计算最终相似度分数,但仅使用点积_距离计算,并且也需要全文索引。这些多功能的策略使用户能够根据其独特需求微调搜索,促进高效和精确的数据检索与分析。此外,SingleStoreDB的混合方法,以 FILTER_BY_TEXT, FILTER_BY_VECTOR和 WEIGHTED_SUM 策略为代表,无缝融合向量和基于文本的搜索,以最大化效率和准确性,确保用户能够充分利用平台的强大功能处理各种应用场景。
from langchain_singlestore.vectorstores import DistanceStrategy
docsearch = SingleStoreVectorStore.from_documents(
docs,
embeddings,
distance_strategy=DistanceStrategy.DOT_PRODUCT, # Use dot product for similarity search
use_vector_index=True, # Use vector index for faster search
use_full_text_search=True, # Use full text index
)
vectorResults = docsearch.similarity_search(
"rainstorm in parched desert, rain",
k=1,
search_strategy=SingleStoreVectorStore.SearchStrategy.VECTOR_ONLY,
filter={"category": "rain"},
)
print(vectorResults[0].page_content)
textResults = docsearch.similarity_search(
"rainstorm in parched desert, rain",
k=1,
search_strategy=SingleStoreVectorStore.SearchStrategy.TEXT_ONLY,
)
print(textResults[0].page_content)
filteredByTextResults = docsearch.similarity_search(
"rainstorm in parched desert, rain",
k=1,
search_strategy=SingleStoreVectorStore.SearchStrategy.FILTER_BY_TEXT,
filter_threshold=0.1,
)
print(filteredByTextResults[0].page_content)
filteredByVectorResults = docsearch.similarity_search(
"rainstorm in parched desert, rain",
k=1,
search_strategy=SingleStoreVectorStore.SearchStrategy.FILTER_BY_VECTOR,
filter_threshold=0.1,
)
print(filteredByVectorResults[0].page_content)
weightedSumResults = docsearch.similarity_search(
"rainstorm in parched desert, rain",
k=1,
search_strategy=SingleStoreVectorStore.SearchStrategy.WEIGHTED_SUM,
text_weight=0.2,
vector_weight=0.8,
)
print(weightedSumResults[0].page_content)
通过转换实现检索器功能
您还可以将向量存储转换为检索器,以便在链中更方便地使用。
retriever = vector_store.as_retriever(search_kwargs={"k": 1})
retriever.invoke("trees in the snow")
多模态示例:利用CLIP和OpenClip嵌入
在多模态数据分析领域,集成图像和文本等不同信息类型变得越来越重要。促进这种集成的一个强大工具是 CLIP,这是一个尖端的模型,能够将图像和文本嵌入到共享的语义空间中。通过这样做,CLIP能够通过相似性搜索检索跨不同模态的相关内容。
例如,让我们考虑一个应用场景,我们的目标是有效分析多模态数据。在这个例子中,我们利用OpenClip多模态嵌入的功能,它利用了CLIP的框架。通过OpenClip,我们可以将文本描述与相应的图像无缝嵌入,从而实现全面的分析和检索任务。无论是根据文本查询识别视觉相似的图像,还是查找与特定视觉内容相关的文本段落,OpenClip都能让用户以卓越的效率和准确性探索多模态数据并从中提取洞察。
pip install -U langchain openai lanchain-singlestore langchain-experimental
from langchain_experimental.open_clip import OpenCLIPEmbeddings
from langchain_singlestore.vectorstores import SingleStoreVectorStore
os.environ["SINGLESTOREDB_URL"] = "root:pass@localhost:3306/db"
TEST_IMAGES_DIR = "../../modules/images"
docsearch = SingleStoreVectorStore(OpenCLIPEmbeddings())
image_uris = sorted(
[
os.path.join(TEST_IMAGES_DIR, image_name)
for image_name in os.listdir(TEST_IMAGES_DIR)
if image_name.endswith(".jpg")
]
)
# Add images
docsearch.add_images(uris=image_uris)
用于检索增强生成的示例
有关如何使用此向量存储进行检索增强生成(RAG)的指南,请参阅以下部分:
- - 检索文档
- - 使用LangChain构建RAG应用
- - 代理式RAG
API参考
有关所有SingleStore文档加载器功能和配置的详细文档,请访问github页面: https://github.com/singlestore-labs/langchain-singlestore/