Manticore 为带有连续书写的语言提供内置索引支持(即单词或句子之间不使用空格或其他分隔符的语言)。这让你可以用两种不同的方式处理这些语言的文本:
- 使用 ICU 库进行精确分词。目前仅支持中文。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- Javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) charset_table = 'cont' morphology = 'icu_chinese'POST /cli -d "
CREATE TABLE products(title text, price float) charset_table = 'cont' morphology = 'icu_chinese'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'charset_table' => 'cont',
'morphology' => 'icu_chinese'
]);utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'cont\' morphology = \'icu_chinese\'')await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'cont\' morphology = \'icu_chinese\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'cont\' morphology = \'icu_chinese\'');utilsApi.sql("CREATE TABLE products(title text, price float) charset_table = 'cont' morphology = 'icu_chinese'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) charset_table = 'cont' morphology = 'icu_chinese'", true);utils_api.sql("CREATE TABLE products(title text, price float) charset_table = 'cont' morphology = 'icu_chinese'", Some(true)).await;table products {
charset_table = cont
morphology = icu_chinese
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}- 使用 Jieba 库进行精确分词。与 ICU 类似,它目前仅支持中文。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- Javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) charset_table = 'cont' morphology = 'jieba_chinese'POST /cli -d "
CREATE TABLE products(title text, price float) charset_table = 'cont' morphology = 'jieba_chinese'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'charset_table' => 'cont',
'morphology' => 'jieba_chinese'
]);utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'cont\' morphology = \'jieba_chinese\'')await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'cont\' morphology = \'jieba_chinese\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'cont\' morphology = \'jieba_chinese\'');utilsApi.sql("CREATE TABLE products(title text, price float) charset_table = 'cont' morphology = 'jieba_chinese'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) charset_table = 'cont' morphology = 'jieba_chinese'", true);utils_api.sql("CREATE TABLE products(title text, price float) charset_table = 'cont' morphology = 'jieba_chinese'", Some(true)).await;table products {
charset_table = cont
morphology = jieba_chinese
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}- 使用 N-gram 选项 ngram_len 和 ngram_chars 进行基本支持。
对于每种使用连续书写的语言,都有单独的字符集表(
chinese、korean、japanese、thai),可以使用。或者,您可以使用通用的cont字符集表同时支持所有 CJK 和泰语语言,或者使用cjk字符集仅包括所有 CJK 语言。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) charset_table = 'non_cont' ngram_len = '1' ngram_chars = 'cont'
/* Or, alternatively */
CREATE TABLE products(title text, price float) charset_table = 'non_cont' ngram_len = '1' ngram_chars = 'cjk,thai'POST /cli -d "
CREATE TABLE products(title text, price float) charset_table = 'non_cont' ngram_len = '1' ngram_chars = 'cont'"
/* Or, alternatively */
POST /cli -d "
CREATE TABLE products(title text, price float) charset_table = 'non_cont' ngram_len = '1' ngram_chars = 'cjk,thai'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'charset_table' => 'non_cont',
'ngram_len' => '1',
'ngram_chars' => 'cont'
]);utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'non_cont\' ngram_len = \'1\' ngram_chars = \'cont\'')await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'non_cont\' ngram_len = \'1\' ngram_chars = \'cont\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'non_cont\' ngram_len = \'1\' ngram_chars = \'cont\'');utilsApi.sql("CREATE TABLE products(title text, price float) charset_table = 'non_cont' ngram_len = '1' ngram_chars = 'cont'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) charset_table = 'non_cont' ngram_len = '1' ngram_chars = 'cont'", true);utils_api.sql("CREATE TABLE products(title text, price float) charset_table = 'non_cont' ngram_len = '1' ngram_chars = 'cont'", Some(true)).await;table products {
charset_table = non_cont
ngram_len = 1
ngram_chars = cont
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}此外,还提供了对中文 停用词 的内置支持,别名 zh。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) charset_table = 'chinese' morphology = 'icu_chinese' stopwords = 'zh'POST /cli -d "
CREATE TABLE products(title text, price float) charset_table = 'chinese' morphology = 'icu_chinese' stopwords = 'zh'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'charset_table' => 'chinese',
'morphology' => 'icu_chinese',
'stopwords' => 'zh'
]);utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'chinese\' morphology = \'icu_chinese\' stopwords = \'zh\'')await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'chinese\' morphology = \'icu_chinese\' stopwords = \'zh\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'chinese\' morphology = \'icu_chinese\' stopwords = \'zh\'');utilsApi.sql("CREATE TABLE products(title text, price float) charset_table = 'chinese' morphology = 'icu_chinese' stopwords = 'zh'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) charset_table = 'chinese' morphology = 'icu_chinese' stopwords = 'zh'", true);utils_api.sql("CREATE TABLE products(title text, price float) charset_table = 'chinese' morphology = 'icu_chinese' stopwords = 'zh'", Some(true)).await;table products {
charset_table = chinese
morphology = icu_chinese
stopwords = zh
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}当文本被索引到 Manticore 中时,它会被拆分成词,并进行大小写折叠,这样像 "Abc"、"ABC" 和 "abc" 这样的词会被视为同一个词。
要正确执行这些操作,Manticore 必须知道:
- 源文本的编码(应始终为 UTF-8)
- 哪些字符被视为字母,哪些不是
- 哪些字母应该折叠成其他字母
你可以使用 charset_table 选项按表进行配置。这会指定一个数组,将字母字符映射到它们的大小写折叠形式(或你希望的其他字符)。数组中未出现的字符会被视为非字母,并在该表的索引或搜索过程中被当作分隔符。
默认字符集是 non_cont,它包含 大多数语言。
你也可以定义文本模式替换规则。例如,使用以下规则:
regexp_filter = \**(\d+)\" => \1 inch
regexp_filter = (BLUE|RED) => COLOR
文本 RED TUBE 5" LONG 会被索引为 COLOR TUBE 5 INCH LONG,而 PLANK 2" x 4" 会被索引为 PLANK 2 INCH x 4 INCH。这些规则按指定顺序应用。规则也会应用于查询,因此搜索 BLUE TUBE 实际上搜索的是 COLOR TUBE。
你可以在这里了解更多关于 regexp_filter 的信息。
# default
charset_table = non_cont
# only English and Russian letters
charset_table = 0..9, A..Z->a..z, _, a..z, \
U+410..U+42F->U+430..U+44F, U+430..U+44F, U+401->U+451, U+451
# english charset defined with alias
charset_table = 0..9, english, _
# override the default transliteration to preserve German umlauts and sharp s;
# uppercase variants are mapped to lowercase
charset_table = non_cont, german
charset_table 指定一个数组,将字母字符映射到它们的大小写折叠形式(或你希望的其他字符)。默认字符集是 non_cont,它包含使用 非连续 脚本的大多数语言。
charset_table 是 Manticore 分词流程中的核心组件,它负责从文档文本或查询文本中提取关键词。它控制哪些字符会被接受为有效字符,以及它们应如何转换(例如是否去除大小写)。
默认情况下,每个字符都会映射到 0,这意味着它不被视为有效关键词,而是作为分隔符处理。一旦某个字符出现在表中,它就会被映射到另一个字符(最常见的是映射到它自身或一个小写字母),并被视为有效关键词的一部分。
charset_table 使用以逗号分隔的映射列表来声明字符为有效字符,或将其映射到其他字符。也提供了用于一次映射一段字符范围的简写语法:
- 单字符映射:
A->a。将源字符 'A' 声明为关键词中允许出现的字符,并将其映射到目标字符 'a'(但不会将 'a' 声明为允许字符)。 - 范围映射:
A..Z->a..z。将源范围内的所有字符声明为允许字符,并将它们映射到目标范围。不会将目标范围声明为允许字符。会检查两个范围的长度。 - 单独字符映射:
a。将某个字符声明为允许字符,并将其映射到它自身。等价于单字符映射a->a。 - 单独范围映射:
a..z。将该范围内的所有字符声明为允许字符,并将它们映射到它们自身。等价于范围映射a..z->a..z。 - 棋盘式范围映射:
A..Z/2。将每两个字符映射到第二个字符。例如,A..Z/2等价于A->B, B->B, C->D, D->D, ..., Y->Z, Z->Z。这种映射简写对大小写字母交错排列的 Unicode 块很有帮助。
对于编码从 0 到 32 的字符,以及 127 到 8 位 ASCII 和 Unicode 字符范围内的字符,Manticore 一律将它们视为分隔符。为了避免配置文件编码问题,8 位 ASCII 字符和 Unicode 字符必须以 U+XXX 形式指定,其中 XXX 是十六进制码点。可接受的最小 Unicode 字符编码是 U+0021。
如果默认映射不能满足你的需求,你可以通过再次指定这些字符来重新定义字符映射。例如,如果内置的 non_cont 数组包含字符 Ä 和 ä,并把它们都映射到 ASCII 字符 a,你可以像这样通过添加它们的 Unicode 码点来重新定义这些字符:
charset_table = non_cont,U+00E4,U+00C4
用于区分大小写搜索,或者
charset_table = non_cont,U+00E4,U+00C4->U+00E4
用于不区分大小写搜索。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) charset_table = '0..9, A..Z->a..z, _, a..z, U+410..U+42F->U+430..U+44F, U+430..U+44F, U+401->U+451, U+451'POST /cli -d "
CREATE TABLE products(title text, price float) charset_table = '0..9, A..Z->a..z, _, a..z, U+410..U+42F->U+430..U+44F, U+430..U+44F, U+401->U+451, U+451'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'charset_table' => '0..9, A..Z->a..z, _, a..z, U+410..U+42F->U+430..U+44F, U+430..U+44F, U+401->U+451, U+451'
]);utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'0..9, A..Z->a..z, _, a..z, U+410..U+42F->U+430..U+44F, U+430..U+44F, U+401->U+451, U+451\'')await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'0..9, A..Z->a..z, _, a..z, U+410..U+42F->U+430..U+44F, U+430..U+44F, U+401->U+451, U+451\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'0..9, A..Z->a..z, _, a..z, U+410..U+42F->U+430..U+44F, U+430..U+44F, U+401->U+451, U+451\'');utilsApi.sql("CREATE TABLE products(title text, price float) charset_table = '0..9, A..Z->a..z, _, a..z, U+410..U+42F->U+430..U+44F, U+430..U+44F, U+401->U+451, U+451'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) charset_table = '0..9, A..Z->a..z, _, a..z, U+410..U+42F->U+430..U+44F, U+430..U+44F, U+401->U+451, U+451'", true);utils_api.sql("CREATE TABLE products(title text, price float) charset_table = '0..9, A..Z->a..z, _, a..z, U+410..U+42F->U+430..U+44F, U+430..U+44F, U+401->U+451, U+451'", Some(true)).await;table products {
charset_table = 0..9, A..Z->a..z, _, a..z, \
U+410..U+42F->U+430..U+44F, U+430..U+44F, U+401->U+451, U+451
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}除了字符和映射的定义之外,还可以使用若干内置别名。当前别名如下:
chinesecjkcontenglishgermanjapanesekoreannon_cont(non_cjk)russianthai
german 别名会保留 ä、ö、ü 和 ß,而不是把它们映射为 ASCII 字符,并会将其大写变体映射为小写,包括把 ẞ 映射为 ß。如上所示,将它追加在 non_cont 之后,以保留其余默认字符映射。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) charset_table = '0..9, english, _'POST /cli -d "
CREATE TABLE products(title text, price float) charset_table = '0..9, english, _'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'charset_table' => '0..9, english, _'
]);utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'0..9, english, _\'')await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'0..9, english, _\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'0..9, english, _\'');utilsApi.sql("CREATE TABLE products(title text, price float) charset_table = '0..9, english, _'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) charset_table = '0..9, english, _'", true);utils_api.sql("CREATE TABLE products(title text, price float) charset_table = '0..9, english, _'", Some(true)).await;table products {
charset_table = 0..9, english, _
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}如果你需要在搜索中支持多种语言,逐一为所有语言定义有效字符集和折叠规则会很繁琐。我们通过提供默认的 charset 表 non_cont 和 cont 简化了这件事,分别覆盖使用非连续脚本和连续脚本(中文、日文、韩文、泰文)的语言。在大多数情况下,这些字符集已经足够。
请注意,目前以下语言不受支持:
- 阿萨姆语
- 比什努普里亚语
- 布希德语
- 加洛语
- 苗语
- 霍语
- 科米语
- 大花苗语
- 马巴语
- 迈蒂利语
- 马拉地语
- 门德语
- Mru 语
- Myene 语
- 恩甘贝语
- 奥里亚语
- 桑塔利语
- 信德语
- Sylheti 语
Unicode 语言列表中列出的 其他所有语言默认都受支持。
要同时处理 cont 和 non-cont 语言,请按如下方式在配置文件中设置选项(中文有一个例外):
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) charset_table = 'non_cont' ngram_len = '1' ngram_chars = 'cont'POST /cli -d "
CREATE TABLE products(title text, price float) charset_table = 'non_cont' ngram_len = '1' ngram_chars = 'cont'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'charset_table' => 'non_cont',
'ngram_len' => '1',
'ngram_chars' => 'cont'
]);utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'non_cont\' ngram_len = \'1\' ngram_chars = \'cont\'')await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'non_cont\' ngram_len = \'1\' ngram_chars = \'cont\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) charset_table = \'non_cont\' ngram_len = \'1\' ngram_chars = \'cont\'');utilsApi.sql("CREATE TABLE products(title text, price float) charset_table = 'non_cont' ngram_len = '1' ngram_chars = 'cont'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) charset_table = 'non_cont' ngram_len = '1' ngram_chars = 'cont'", true);utils_api.sql("CREATE TABLE products(title text, price float) charset_table = 'non_cont' ngram_len = '1' ngram_chars = 'cont'", Some(true)).await;table products {
charset_table = non_cont
ngram_len = 1
ngram_chars = cont
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}如果你不需要支持连续脚本语言,可以直接去掉 ngram_len 和 ngram_chars。 选项。有关这些选项的更多信息,请参阅对应文档章节。
如果要将一个字符映射为多个字符,或将多个字符映射为一个字符,regexp_filter 可能会有帮助。
blend_chars = +, &, U+23
blend_chars = +, &->+
混合字符列表。可选,默认为空。
混合字符会同时作为分隔符和有效字符被索引。例如,当 & 被定义为混合字符,并且 AT&T 出现在已索引文档中时,会索引出三个不同的关键词:at&t、at 和 t。
此外,混合字符还会影响索引,使关键词被索引时仿佛这些混合字符根本没有输入过。这个行为在指定 blend_mode = trim_all 时尤为明显。例如,短语 some_thing 在 blend_mode = trim_all 下会被索引为 some、something 和 thing。
使用混合字符时要小心,因为将某个字符定义为混合字符,就意味着它不再是分隔符。
- 因此,如果你把逗号放进
blend_chars,然后搜索dog,cat,它会把它当作单个 tokendog,cat。如果dog,cat没有被索引为dog,cat,而只是保留为dog cat,那么它就无法匹配。 - 因而,这种行为应通过 blend_mode 设置来控制。
通过用空白替换混合字符得到的 token,其位置会按常规分配,普通关键词会像完全没有指定 blend_chars 一样被索引。一个额外的 token 会把混合字符和非混合字符组合起来,并放在起始位置。例如,如果 AT&T company 出现在文本字段的最开头,at 的位置会是 1,t 的位置会是 2,company 的位置会是 3,同时 AT&T 也会被放在位置 1,并与开头的普通关键词合并。因此,查询 AT&T 或仅 AT 都会匹配该文档。短语查询 "AT T" 也会匹配,短语查询 "AT&T company" 同样会匹配。
混合字符可能与查询语法中使用的特殊字符重叠,例如 T-Mobile 或 @twitter。在可能的情况下,查询解析器会把混合字符按混合字符处理。例如,如果 hello @twitter 位于引号内(短语操作符),查询解析器会把 @ 符号当作混合字符处理。不过,如果 @ 符号不在引号内,这个字符就会被当作操作符处理。因此,建议对关键词进行转义。
混合字符可以重新映射,从而把多个不同的混合字符归一化为一个基础形式。这在索引多个具有等价字形的 Unicode 码位变体时很有用。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) blend_chars = '+, &, U+23, @->_'POST /cli -d "
CREATE TABLE products(title text, price float) blend_chars = '+, &, U+23, @->_'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'blend_chars' => '+, &, U+23, @->_'
]);utilsApi.sql('CREATE TABLE products(title text, price float) blend_chars = \'+, &, U+23, @->_\'')await utilsApi.sql('CREATE TABLE products(title text, price float) blend_chars = \'+, &, U+23, @->_\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) blend_chars = \'+, &, U+23, @->_\'');utilsApi.sql("CREATE TABLE products(title text, price float) blend_chars = '+, &, U+23, @->_'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) blend_chars = '+, &, U+23, @->_'", true);utils_api.sql("CREATE TABLE products(title text, price float) blend_chars = '+, &, U+23, @->_'", Some(true)).await;table products {
blend_chars = +, &, U+23, @->_
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}blend_mode = option [, option [, ...]]
option = trim_none | trim_head | trim_tail | trim_both | trim_all | skip_pure
通过 blend_mode 指令启用混合 token 索引模式。
默认情况下,混合字符和非混合字符混合在一起的 token 会整体被索引。例如,当 blend_chars 中同时包含 at 符号和感叹号时,字符串 @dude! 会被索引成两个 token:@dude!(包含所有混合字符)和 dude(不包含任何混合字符)。因此,查询 @dude 不会匹配它。
blend_mode 为这种索引行为增加了灵活性。它接受一个逗号分隔的选项列表,每个选项都指定一种 token 索引变体。
如果指定了多个选项,同一个 token 会索引出多个变体。普通关键词(即把混合字符替换为分隔符后得到的 token)始终会被索引。
可用选项如下:
trim_none- 索引整个 tokentrim_head- 去掉开头的混合字符,并索引结果 tokentrim_tail- 去掉结尾的混合字符,并索引结果 tokentrim_both- 去掉开头和结尾的混合字符,并索引结果 tokentrim_all- 去掉开头、结尾和中间的混合字符,并索引结果 tokenskip_pure- 如果 token 纯粹由混合字符组成,则不索引它
使用上面的 @dude! 示例字符串,设置 blend_mode = trim_head, trim_tail 会得到两个被索引的 token:@dude 和 dude!。使用 trim_both 不会有任何效果,因为去掉两端的混合字符后会得到 dude,而它已经作为普通关键词被索引了。使用 trim_both 索引 @U.S.A.(并假设点号也是混合字符)时,会得到 U.S.A 被索引。最后,skip_pure 允许你忽略完全由混合字符组成的序列。例如,one @@@ two 会被索引为 one two,并可按短语匹配。这在默认情况下并不会发生,因为完全混合的 token 会被索引,并将第二个关键词的位置偏移。
默认行为是索引整个 token,这等同于 blend_mode = trim_none。
请注意,使用混合模式会限制你的搜索范围,即使是默认模式 trim_none,如果你假定 . 是混合字符也是如此:
.dog.在索引时会变成.dog. dog- 而且你无法通过
dog.找到它。
使用更多模式会提高你的关键词匹配到某些内容的概率。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) blend_mode = 'trim_tail, skip_pure' blend_chars = '+, &'POST /cli -d "
CREATE TABLE products(title text, price float) blend_mode = 'trim_tail, skip_pure' blend_chars = '+, &'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'blend_mode' => 'trim_tail, skip_pure',
'blend_chars' => '+, &'
]);utilsApi.sql('CREATE TABLE products(title text, price float) blend_mode = \'trim_tail, skip_pure\' blend_chars = \'+, &\'')await utilsApi.sql('CREATE TABLE products(title text, price float) blend_mode = \'trim_tail, skip_pure\' blend_chars = \'+, &\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) blend_mode = \'trim_tail, skip_pure\' blend_chars = \'+, &\'');utilsApi.sql("CREATE TABLE products(title text, price float) blend_mode = 'trim_tail, skip_pure' blend_chars = '+, &'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) blend_mode = 'trim_tail, skip_pure' blend_chars = '+, &'", true);utils_api.sql("CREATE TABLE products(title text, price float) blend_mode = 'trim_tail, skip_pure' blend_chars = '+, &'", Some(true)).await;table products {
blend_mode = trim_tail, skip_pure
blend_chars = +, &
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}min_word_len = length
min_word_len 是 Manticore 中一个可选的索引配置项,用于指定可被索引的最小词长。默认值是 1,这意味着所有内容都会被索引。
只有不短于该最小长度的词才会被索引。例如,如果 min_word_len 为 4,那么 'the' 不会被索引,而 'they' 会被索引。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) min_word_len = '4'POST /cli -d "
CREATE TABLE products(title text, price float) min_word_len = '4'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'min_word_len' => '4'
]);utilsApi.sql('CREATE TABLE products(title text, price float) min_word_len = \'4\'')await utilsApi.sql('CREATE TABLE products(title text, price float) min_word_len = \'4\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) min_word_len = \'4\'');utilsApi.sql("CREATE TABLE products(title text, price float) min_word_len = '4'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) min_word_len = '4'", true);utils_api.sql("CREATE TABLE products(title text, price float) min_word_len = '4'", Some(true)).await;table products {
min_word_len = 4
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}ngram_len = 1
用于 N-gram 索引的 N-gram 长度。可选,默认值为 0(禁用 n-gram 索引)。已知值为 0 和 1。
N-gram 为未分词文本中的连续脚本语言提供基础支持。使用连续脚本语言进行搜索的问题在于单词之间没有清晰的分隔符。在某些情况下,你可能不想使用基于词典的分词,例如中文所使用的方式。在这些场景下,n-gram 分词也可能效果很好。
启用此功能后,这类语言的文本流(或在 ngram_chars 中定义的其他字符)会按 N-gram 方式索引。例如,如果输入文本是 "ABCDEF"(其中 A 到 F 代表某种语言字符),并且 ngram_len 为 1,那么它会像 "A B C D E F" 一样被索引。目前只支持 ngram_len=1。只有在 ngram_chars 表中列出的字符才会以这种方式拆分;其他字符不受影响。
请注意,如果搜索查询已经分词,也就是单词之间有分隔符,那么在应用侧把单词加上引号并使用扩展模式,即使文本本身没有分词,也能得到正确匹配。例如,假设原始查询是 BC DEF。在应用侧加上引号后,它应该看起来像 "BC" "DEF"(带引号)。这个查询会传给 Manticore,并在内部也拆成 1-gram,结果变成 "B C" "D E F" 查询,仍然带着作为短语匹配操作符的引号。这样即使文本中没有分隔符,也能匹配到文本。
即使搜索查询没有分词,借助短语相关性排序,Manticore 仍应能产生良好结果:它会把更接近短语匹配的结果(在 N-gram 词的情况下,这可能意味着更接近的多字符词匹配)排到前面。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) ngram_chars = 'cont' ngram_len = '1'POST /cli -d "
CREATE TABLE products(title text, price float) ngram_chars = 'cont' ngram_len = '1'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'ngram_chars' => 'cont',
'ngram_len' => '1'
]);utilsApi.sql('CREATE TABLE products(title text, price float) ngram_chars = \'cont\' ngram_len = \'1\'')await utilsApi.sql('CREATE TABLE products(title text, price float) ngram_chars = \'cont\' ngram_len = \'1\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) ngram_chars = \'cont\' ngram_len = \'1\'');utilsApi.sql("CREATE TABLE products(title text, price float) ngram_chars = 'cont' ngram_len = '1'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) ngram_chars = 'cont' ngram_len = '1'", true);utils_api.sql("CREATE TABLE products(title text, price float) ngram_chars = 'cont' ngram_len = '1'", Some(true)).await;table products {
ngram_chars = cont
ngram_len = 1
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}ngram_chars = cont
ngram_chars = cont, U+3000..U+2FA1F
N-gram 字符列表。可选,默认为空。
要与 ngram_len 配合使用,该列表定义了哪些字符的连续序列会被提取为 N-gram。由其他字符组成的词不会受到 N-gram 索引功能的影响。其值格式与 charset_table 相同。N-gram 字符不能出现在 charset_table 中。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) ngram_chars = 'U+3000..U+2FA1F' ngram_len = '1'POST /cli -d "
CREATE TABLE products(title text, price float) ngram_chars = 'U+3000..U+2FA1F' ngram_len = '1'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'ngram_chars' => 'U+3000..U+2FA1F',
'ngram_len' => '1'
]);utilsApi.sql('CREATE TABLE products(title text, price float) ngram_chars = \'U+3000..U+2FA1F\' ngram_len = \'1\'')await utilsApi.sql('CREATE TABLE products(title text, price float) ngram_chars = \'U+3000..U+2FA1F\' ngram_len = \'1\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) ngram_chars = \'U+3000..U+2FA1F\' ngram_len = \'1\'');utilsApi.sql("CREATE TABLE products(title text, price float) ngram_chars = 'U+3000..U+2FA1F' ngram_len = '1'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) ngram_chars = 'U+3000..U+2FA1F' ngram_len = '1'", true);utils_api.sql("CREATE TABLE products(title text, price float) ngram_chars = 'U+3000..U+2FA1F' ngram_len = '1'", Some(true)).await;table products {
ngram_chars = U+3000..U+2FA1F
ngram_len = 1
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}你也可以像示例中那样为我们的默认 N-gram 表使用别名。在大多数情况下,这已经足够。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) ngram_chars = 'cont' ngram_len = '1'POST /cli -d "
CREATE TABLE products(title text, price float) ngram_chars = 'cont' ngram_len = '1'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'ngram_chars' => 'cont',
'ngram_len' => '1'
]);utilsApi.sql('CREATE TABLE products(title text, price float) ngram_chars = \'cont\' ngram_len = \'1\'')await utilsApi.sql('CREATE TABLE products(title text, price float) ngram_chars = \'cont\' ngram_len = \'1\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) ngram_chars = \'cont\' ngram_len = \'1\'');utilsApi.sql("CREATE TABLE products(title text, price float) ngram_chars = 'cont' ngram_len = '1'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) ngram_chars = 'cont' ngram_len = '1'", true);utils_api.sql("CREATE TABLE products(title text, price float) ngram_chars = 'cont' ngram_len = '1'", Some(true)).await;table products {
ngram_chars = cont
ngram_len = 1
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}ignore_chars = U+AD
被忽略的字符列表。可选,默认为空。
这在某些字符(例如软连字符 U+00AD)不应仅仅被视为分隔符,而应被完全忽略时很有用。例如,如果 - 只是不在 charset_table 中,那么文本 "abc-def" 会被索引为 "abc" 和 "def" 两个关键词。相反,如果把 - 加入 ignore_chars 列表,同样的文本会被索引为单个 "abcdef" 关键词。
语法与 charset_table 相同,但这里只允许声明字符,不允许映射它们。另外,被忽略的字符不能出现在 charset_table 中。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) ignore_chars = 'U+AD'POST /cli -d "
CREATE TABLE products(title text, price float) ignore_chars = 'U+AD'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'ignore_chars' => 'U+AD'
]);utilsApi.sql('CREATE TABLE products(title text, price float) ignore_chars = \'U+AD\'')await utilsApi.sql('CREATE TABLE products(title text, price float) ignore_chars = \'U+AD\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) ignore_chars = \'U+AD\'');utilsApi.sql("CREATE TABLE products(title text, price float) ignore_chars = 'U+AD'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) ignore_chars = 'U+AD'", true);utils_api.sql("CREATE TABLE products(title text, price float) ignore_chars = 'U+AD'", Some(true)).await;table products {
ignore_chars = U+AD
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}bigram_index = {none|all|first_freq|both_freq|second_numeric|second_has_digit}
双字索引模式。可选,默认是 none。
双字索引是一种加速短语搜索的功能。在索引时,它会把所有或部分相邻词对的文档列表存入索引。之后在搜索时,这个列表可被用来显著加速短语或子短语匹配。
bigram_index 控制具体词对的选择。已知模式如下:
all,索引每一个词对first_freq,仅索引那些第一个词在高频词列表中的词对(见 bigram_freq_words)。例如,设置bigram_freq_words = the, in, i, a时,索引 "alone in the dark" 文本会把 "in the" 和 "the dark" 这两个词对作为 bigram 存储,因为它们都以高频关键词(分别是 "in" 或 "the")开头,但 "alone in" 不会被索引,因为 "in" 在该词对中是第二个词。both_freq,仅索引两个词都属于高频词的词对。继续沿用同一个例子,在这种模式下索引 "alone in the dark" 时,只会把 "in the"(从搜索角度看最差的那个)作为 bigram 存储,而其他词对都不会被索引。second_numeric,仅索引第二个 token 只包含 ASCII 数字的词对。例如,xt 806会匹配,但xt rt9600和xt v2不会。second_has_digit,仅索引第二个 token 至少包含一个 ASCII 数字的词对。例如,xt 806、xt rt9600和xt v2会匹配,但xt abc不会。
对于大多数用例,both_freq 是最佳模式,但实际效果可能因场景而异。
需要注意的是,bigram_index 只在分词层面起作用,不会考虑 morphology、wordforms 或 stopwords 等转换。这意味着它创建的 token 非常直接,这会让短语搜索更精确、更严格。虽然这可以提高短语匹配的准确性,但也会降低系统识别词形变化或词语不同写法的能力。
与数字相关的模式只使用 ASCII 数字(0-9)。它们不会把 +、- 或 Unicode 数字视为数字。检查还会使用当前分词器路径生成的 token 文本,不会做额外的标点归一化。
使用 bigram_delimiter 可以控制符合条件的 bigram 是作为内部带分隔符的 token 存储,还是作为像 iphone17 这样的粘连 token 存储,或者两种形式都存。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) bigram_freq_words = 'the, a, you, i' bigram_index = 'both_freq'POST /cli -d "
CREATE TABLE products(title text, price float) bigram_freq_words = 'the, a, you, i' bigram_index = 'both_freq'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'bigram_freq_words' => 'the, a, you, i',
'bigram_index' => 'both_freq'
]);utilsApi.sql('CREATE TABLE products(title text, price float) bigram_freq_words = \'the, a, you, i\' bigram_index = \'both_freq\'')await utilsApi.sql('CREATE TABLE products(title text, price float) bigram_freq_words = \'the, a, you, i\' bigram_index = \'both_freq\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) bigram_freq_words = \'the, a, you, i\' bigram_index = \'both_freq\'');utilsApi.sql("CREATE TABLE products(title text, price float) bigram_freq_words = 'the, a, you, i' bigram_index = 'both_freq'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) bigram_freq_words = 'the, a, you, i' bigram_index = 'both_freq'", true);utils_api.sql("CREATE TABLE products(title text, price float) bigram_freq_words = 'the, a, you, i' bigram_index = 'both_freq'", Some(true)).await;table products {
bigram_index = both_freq
bigram_freq_words = the, a, you, i
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}bigram_delimiter = {true|none|both}
双字 token 存储模式。可选,默认是 true。
bigram_delimiter 控制由 bigram_index 选出的符合条件的 bigram 存储哪种 token 形式:
true,只存储内部带分隔符的 bigram token。这是当前默认行为。none,只存储粘连 token 形式,例如iphone17。both,同时存储内部带分隔符的形式和粘连形式。
搜索行为取决于所选模式:
- 使用
true时,短语优化会把符合条件的短语词对改写为内部带分隔符的 token - 使用
none时,短语优化会把符合条件的短语词对改写为粘连 token,例如"iphone 17"会变成iphone17 - 使用
both时,短语优化会被跳过,短语查询仍保持普通短语查询,但因为粘连形式也被存储了,粘连 token 搜索仍然可以匹配
bigram_delimiter 只改变已存储 token 的形态。它并不决定哪些词对符合条件;这仍由 bigram_index 控制。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) bigram_index = 'all' bigram_delimiter = 'none'POST /cli -d "
CREATE TABLE products(title text, price float) bigram_index = 'all' bigram_delimiter = 'none'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'bigram_index' => 'all',
'bigram_delimiter' => 'none'
]);utilsApi.sql('CREATE TABLE products(title text, price float) bigram_index = \'all\' bigram_delimiter = \'none\'')await utilsApi.sql('CREATE TABLE products(title text, price float) bigram_index = \'all\' bigram_delimiter = \'none\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) bigram_index = \'all\' bigram_delimiter = \'none\'');utilsApi.sql("CREATE TABLE products(title text, price float) bigram_index = 'all' bigram_delimiter = 'none'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) bigram_index = 'all' bigram_delimiter = 'none'", true);utils_api.sql("CREATE TABLE products(title text, price float) bigram_index = 'all' bigram_delimiter = 'none'", Some(true)).await;table products {
bigram_index = all
bigram_delimiter = none
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}bigram_freq_words = the, a, you, i
在索引 bigram 时被视为“高频”的关键词列表。可选,默认为空。
某些 bigram 索引模式(见 bigram_index)需要一个高频关键词列表。这些词不要与停用词混淆。停用词会在索引和搜索时被完全移除。高频关键词只被 bigram 用来判断是否要索引当前词对。
bigram_freq_words 允许你定义这样一组关键词。
只有在 first_freq 和 both_freq 中才需要这个选项。
以下模式中它必须保持为空:
noneallsecond_numericsecond_has_digit
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) bigram_freq_words = 'the, a, you, i' bigram_index = 'first_freq'POST /cli -d "
CREATE TABLE products(title text, price float) bigram_freq_words = 'the, a, you, i' bigram_index = 'first_freq'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'bigram_freq_words' => 'the, a, you, i',
'bigram_index' => 'first_freq'
]);utilsApi.sql('CREATE TABLE products(title text, price float) bigram_freq_words = \'the, a, you, i\' bigram_index = \'first_freq\'')await utilsApi.sql('CREATE TABLE products(title text, price float) bigram_freq_words = \'the, a, you, i\' bigram_index = \'first_freq\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) bigram_freq_words = \'the, a, you, i\' bigram_index = \'first_freq\'');utilsApi.sql("CREATE TABLE products(title text, price float) bigram_freq_words = 'the, a, you, i' bigram_index = 'first_freq'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) bigram_freq_words = 'the, a, you, i' bigram_index = 'first_freq'", true);utils_api.sql("CREATE TABLE products(title text, price float) bigram_freq_words = 'the, a, you, i' bigram_index = 'first_freq'", Some(true)).await;table products {
bigram_freq_words = the, a, you, i
bigram_index = first_freq
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}dict = {keywords|keywords_32k|crc}
字典类型由三个已知值之一标识:keywords、keywords_32k 或 crc。该设置是可选的;keywords 是默认值。
dict=keywords 和 dict=keywords_32k 是词典。词典会在索引中存储原始关键词文本,并在搜索时执行通配符扩展。dict=crc 则存储关键词校验值。
dict=keywords 是默认的词典。
dict=keywords_32k 是一个可启用的词典,支持 32 KiB 关键词 token。它在普通表和 RT 表中都支持最长 32768 字节的规范化 token。超过该限制的 token 会被跳过并发出警告,而不会被作为截断词条索引。当启用了常规的精确、前缀或内联前缀设置时,它支持精确查找、前缀查找和内联查找。有关 42 字节常规 token 限制以及 32768 字节 keywords_32k 限制的详细信息,请参阅词长限制。
keywords_32k 适用于较长的机器生成值,例如哈希、生成的 ID、消息标识符以及长的类邮箱 token。
以下功能目前还不支持 dict=keywords_32k:
CALL SUGGEST和CALL QSUGGEST不能在使用dict=keywords_32k的表上工作。- Percolate 表不能使用
dict=keywords_32k。 - 摘要和高亮仍使用常规 token 限制。最长 42 字节的 token 可以被高亮;更长的
keywords_32ktoken 会在摘要/高亮处理中被跳过。 indextool --dumpdict目前还不能转储dict=keywords_32k词典。
CRC 词典不会在索引中存储原始关键词文本。相反,它们在搜索和索引过程中都会用一个控制和数值(使用 FNV64 计算)替换关键词。这个值在索引内部使用。此方法有两个缺点:
- 首先,不同关键词对之间存在控制和碰撞的风险。这个风险会随着索引中唯一关键词数量的增加而增长。不过,这一问题相对较小,因为在一个包含 10 亿条目的词典中,单次 FNV64 碰撞的概率大约是 1/16,也就是 6.25%。由于典型的人类自然语言只有 100 万到 1000 万种词形,大多数词典的关键词数量会远少于 10 亿。
- 其次,更关键的是,用控制和执行子串搜索并不直接。Manticore 通过预先将所有可能的子串索引为独立关键词来解决这个问题(见 min_prefix_len、min_infix_len 指令)。这种方法还有一个额外优势,就是能以最快的方式匹配子串。然而,预先索引所有子串会显著增加索引大小(通常会增加 3 到 10 倍甚至更多),并进一步影响索引时间,使大索引上的子串搜索相当不实用。
词典方式解决了这两个问题。它把关键词存储在索引中,并在搜索时执行通配符扩展。例如,搜索前缀 test* 时,系统可能会基于词典内容在内部扩展为 test|tests|testing 查询。这个扩展过程对应用完全不可见,唯一的例外是:所有匹配关键词的独立每关键词统计信息现在也会被报告。
对于子串(内联)搜索,可以使用扩展通配符。? 和 % 等特殊字符兼容子串(内联)搜索(例如,t?st*、run%、*abc*)。请注意,通配符操作符适用于 dict=keywords 和 dict=keywords_32k,而 REGEX 操作符仅适用于 dict=keywords。
对于正常大小的 token,使用词典索引大约比常规的非子串索引慢 1.1 倍到 1.3 倍,但仍明显快于子串索引(无论是前缀还是内联)。索引大小应只比标准的非子串表略大,总差异大约为 1..10%。常规关键词搜索的耗时在这三种索引类型之间应当几乎相同或完全一致(CRC 非子串、CRC 子串、词典)。子串搜索时间会随实际匹配到该子串的关键词数量而明显波动(也就是搜索词会扩展成多少关键词)。可匹配关键词的最大数量受 expansion_limit 指令限制。
总之,词典和 CRC 词典为子串搜索提供了两种不同的取舍。你可以选择牺牲索引时间和索引大小,以获得最快的最坏情况搜索性能(CRC 词典);或者尽量减少对索引时间的影响,但在前缀扩展到大量关键词时牺牲最坏情况搜索时间(词典)。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) dict = 'keywords'POST /cli -d "
CREATE TABLE products(title text, price float) dict = 'keywords'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'dict' => 'keywords'
]);utilsApi.sql('CREATE TABLE products(title text, price float) dict = \'keywords\'')await utilsApi.sql('CREATE TABLE products(title text, price float) dict = \'keywords\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) dict = \'keywords\'');utilsApi.sql("CREATE TABLE products(title text, price float) dict = 'keywords'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) dict = 'keywords'", true);utils_api.sql("CREATE TABLE products(title text, price float) dict = 'keywords'", Some(true)).await;table products {
dict = keywords
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}embedded_limit = size
嵌入式 exceptions、wordforms 或 stop words 文件大小限制。可选,默认是 16K。
创建表时,上面提到的这些文件可以与表一起保存到外部,也可以直接嵌入表中。大小低于 embedded_limit 的文件会存储到表里。对于更大的文件,只会存储文件名。这样也简化了将表文件迁移到其他机器的过程;你可能只需要复制一个文件。
对于较小的文件,这种嵌入方式减少了表所依赖的外部文件数量,有助于维护。但与此同时,把一个 100 MB 的 wordforms 词典嵌入一个很小的 delta 表里显然没有意义。所以需要一个大小阈值,而 embedded_limit 就是这个阈值。
- CONFIG
table products {
embedded_limit = 32K
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}global_idf = /path/to/global.idf
包含全局(集群范围)关键词 IDF 的文件路径。可选,默认为空(使用本地 IDF)。
在多表集群中,不同表之间每个关键词的频率很可能不同。这意味着当排序函数使用基于 TF-IDF 的值时,例如 BM25 系列因子,结果的排序可能会因所在集群节点不同而略有差异。
修复这个问题最简单的方法是创建并使用一个全局频率字典,简称全局 IDF 文件。该指令允许你指定该文件的位置。建议使用 .idf 扩展名,但不是强制要求。当给定表指定了 IDF 文件,并且 OPTION global_idf 设置为 1 时,引擎将使用 global_idf 文件中的关键词频率和集合文档数,而不是仅使用本地表的数据。这样,IDF 以及依赖它们的值就能在整个集群中保持一致。
IDF 文件可以在多个表之间共享。即使有很多表引用同一个 IDF 文件,searchd 也只会加载该文件的一份副本。如果 IDF 文件内容发生变化,可以通过 SIGHUP 加载新内容。
你可以使用 indextool 工具生成 .idf 文件:先用 --dumpdict dict.txt --stats 开关转储词典,再用 --buildidf 将其转换为 .idf 格式,然后用 --mergeidf 合并集群中的所有 .idf 文件。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) global_idf = '/usr/local/manticore/var/global.idf'POST /cli -d "
CREATE TABLE products(title text, price float) global_idf = '/usr/local/manticore/var/global.idf'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'global_idf' => '/usr/local/manticore/var/global.idf'
]);utilsApi.sql('CREATE TABLE products(title text, price float) global_idf = \'/usr/local/manticore/var/global.idf\'')await utilsApi.sql('CREATE TABLE products(title text, price float) global_idf = \'/usr/local/manticore/var/global.idf\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) global_idf = \'/usr/local/manticore/var/global.idf\'');utilsApi.sql("CREATE TABLE products(title text, price float) global_idf = '/usr/local/manticore/var/global.idf'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) global_idf = '/usr/local/manticore/var/global.idf'", true);utils_api.sql("CREATE TABLE products(title text, price float) global_idf = '/usr/local/manticore/var/global.idf'", Some(true)).await;table products {
global_idf = /usr/local/manticore/var/global.idf
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}hitless_words = {all|path/to/file}
无位置词列表。可选,允许的值是 'all' 或一个列表文件名。
默认情况下,Manticore 全文索引不仅会为每个给定关键词存储匹配文档列表,还会存储其在文档中的位置列表(称为 hitlist)。Hitlist 支持短语、邻近、严格顺序以及其他高级搜索类型,同时也支持短语邻近排序。不过,某些高频关键词的 hitlist(这些词由于某种原因即使很常见也不能被停用)可能会非常大,因此在查询时处理起来会很慢。另外,在某些场景下,我们可能只关心布尔关键词匹配,而从不需要基于位置的搜索操作符(例如短语匹配)或短语排序。
hitless_words 允许你创建完全不包含位置信息(hitlist)的索引,或者仅对特定关键词跳过位置信息。
无位置索引通常会比相应的常规全文索引占用更少空间(通常可预期约 1.5 倍)。索引和搜索速度都应更快,但代价是缺少位置查询和排序支持。
如果在位置查询中使用(例如短语查询),这些无位置词会从查询中移除,并作为不带位置的操作数使用。例如,如果 "hello" 和 "world" 是无位置词,而 "simon" 和 "says" 不是,那么短语查询 "simon says hello world" 会被转换为 ("simon says" & hello & world),从而匹配文档中任意位置的 "hello" 和 "world",以及作为精确短语的 "simon says"。
如果一个位置查询只包含无位置词,那么它会生成一个空的短语节点,因此整个查询将返回空结果并给出警告。如果整个词典都是无位置的(使用 all),那么在相应索引上只能使用布尔匹配。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) hitless_words = 'all'POST /cli -d "
CREATE TABLE products(title text, price float) hitless_words = 'all'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'hitless_words' => 'all'
]);utilsApi.sql('CREATE TABLE products(title text, price float) hitless_words = \'all\'')await utilsApi.sql('CREATE TABLE products(title text, price float) hitless_words = \'all\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) hitless_words = \'all\'');utilsApi.sql("CREATE TABLE products(title text, price float) hitless_words = 'all'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) hitless_words = 'all'", true);utils_api.sql("CREATE TABLE products(title text, price float) hitless_words = 'all'", Some(true)).await;table products {
hitless_words = all
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}hitless_words_list = 'word1; word2; ...'
hitless_words_list 设置允许你在 CREATE TABLE 语句中直接指定无位置词。它仅支持 RT 模式。
这些值必须用分号(;)分隔。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
CREATE TABLE products(title text, price float) hitless_words_list = 'hello; world'POST /cli -d "
CREATE TABLE products(title text, price float) hitless_words_list = 'hello; world'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'hitless_words_list' => 'hello; world'
]);utilsApi.sql('CREATE TABLE products(title text, price float) hitless_words_list = \'hello; world\'')await utilsApi.sql('CREATE TABLE products(title text, price float) hitless_words_list = \'hello; world\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) hitless_words_list = \'hello; world\'');utilsApi.sql("CREATE TABLE products(title text, price float) hitless_words_list = 'hello; world'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) hitless_words_list = 'hello; world'", true);utils_api.sql("CREATE TABLE products(title text, price float) hitless_words_list = 'hello; world'", Some(true)).await;index_field_lengths = {0|1}
启用将字段长度(每文档和每索引平均值)计算并存储到全文索引中。可选,默认是 0(不计算也不存储)。
当 index_field_lengths 设置为 1 时,Manticore 会:
- 为每个全文字段创建一个对应的长度属性,名称相同但带有
__len后缀 - 为每个文档计算字段长度(按关键词计数)并存储到相应属性中
- 计算每个索引的平均值。这些长度属性会采用特殊的 TOKENCOUNT 类型,但其值实际上是普通的 32 位整数,并且通常可以直接访问。
表达式排序器中的 BM25A() 和 BM25F() 函数基于这些长度,并且需要启用 index_field_lengths。历史上,Manticore 使用的是一个简化版、裁剪版的 BM25,与完整函数不同,它不考虑文档长度。现在也支持完整版本的 BM25,以及它面向多个字段的扩展,称为 BM25F。它们分别需要每文档长度和每字段长度。因此才有了这个额外指令。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) index_field_lengths = '1'POST /cli -d "
CREATE TABLE products(title text, price float) index_field_lengths = '1'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'index_field_lengths' => '1'
]);utilsApi.sql('CREATE TABLE products(title text, price float) index_field_lengths = \'1\'')await utilsApi.sql('CREATE TABLE products(title text, price float) index_field_lengths = \'1\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) index_field_lengths = \'1\'');utilsApi.sql("CREATE TABLE products(title text, price float) index_field_lengths = '1'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) index_field_lengths = '1'", true);utils_api.sql("CREATE TABLE products(title text, price float) index_field_lengths = '1'", Some(true)).await;table products {
index_field_lengths = 1
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}index_token_filter = my_lib.so:custom_blend:chars=@#&
全文索引的索引时 token 过滤器。可选,默认为空。
index_token_filter 指令指定了一个可选的全文索引时 token 过滤器。该指令用于创建自定义分词器,以便根据自定义规则生成 token。这个过滤器由 indexer 在将源数据索引到普通表时创建,或由 RT 表在处理 INSERT 或 REPLACE 语句时创建。插件使用如下格式定义:library name:plugin name:optional string of settings。例如,my_lib.so:custom_blend:chars=@#&。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) index_token_filter = 'my_lib.so:custom_blend:chars=@#&'POST /cli -d "
CREATE TABLE products(title text, price float) index_token_filter = 'my_lib.so:custom_blend:chars=@#&'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'index_token_filter' => 'my_lib.so:custom_blend:chars=@#&'
]);utilsApi.sql('CREATE TABLE products(title text, price float) index_token_filter = \'my_lib.so:custom_blend:chars=@#&\'')await utilsApi.sql('CREATE TABLE products(title text, price float) index_token_filter = \'my_lib.so:custom_blend:chars=@#&\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) index_token_filter = \'my_lib.so:custom_blend:chars=@#&\'');utilsApi.sql("CREATE TABLE products(title text, price float) index_token_filter = 'my_lib.so:custom_blend:chars=@#&'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) index_token_filter = 'my_lib.so:custom_blend:chars=@#&'", true);utils_api.sql("CREATE TABLE products(title text, price float) index_token_filter = 'my_lib.so:custom_blend:chars=@#&'", Some(true)).await;table products {
index_token_filter = my_lib.so:custom_blend:chars=@#&
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}overshort_step = {0|1}
超短(短于 min_word_len)关键词的位置增量。可选,允许的值是 0 和 1,默认值是 1。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) overshort_step = '1'POST /cli -d "
CREATE TABLE products(title text, price float) overshort_step = '1'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'overshort_step' => '1'
]);utilsApi.sql('CREATE TABLE products(title text, price float) overshort_step = \'1\'')utilsApi.sql('CREATE TABLE products(title text, price float) overshort_step = \'1\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) overshort_step = \'1\'');utilsApi.sql("CREATE TABLE products(title text, price float) overshort_step = '1'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) overshort_step = '1'", true);utils_api.sql("CREATE TABLE products(title text, price float) overshort_step = '1'", Some(true)).await;table products {
overshort_step = 1
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}phrase_boundary = ., ?, !, U+2026 # horizontal ellipsis
短语边界字符列表。可选,默认为空。
该列表控制哪些字符会被视为短语边界,以便调整词的位置,并通过邻近搜索实现短语级搜索模拟。语法与 charset_table 类似,但不允许映射,且边界字符不得与其他内容重叠。
在短语边界处,会向当前词位置额外加上一个词位置增量(由 phrase_boundary_step 指定)。这使得可以通过邻近查询实现短语级搜索:不同短语中的词将被保证彼此相距超过 phrase_boundary_step;因此,在该距离内的邻近搜索就等同于短语级搜索。
只有当此类字符后面跟着一个分隔符时,才会触发短语边界条件;这样可以避免像 S.T.A.L.K.E.R 这样的缩写或 URL 被当作多个短语。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) phrase_boundary = '., ?, !, U+2026' phrase_boundary_step = '10'POST /cli -d "
CREATE TABLE products(title text, price float) phrase_boundary = '., ?, !, U+2026' phrase_boundary_step = '10'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'phrase_boundary' => '., ?, !, U+2026',
'phrase_boundary_step' => '10'
]);utilsApi.sql('CREATE TABLE products(title text, price float) phrase_boundary = \'., ?, !, U+2026\' phrase_boundary_step = \'10\'')await utilsApi.sql('CREATE TABLE products(title text, price float) phrase_boundary = \'., ?, !, U+2026\' phrase_boundary_step = \'10\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) phrase_boundary = \'., ?, !, U+2026\' phrase_boundary_step = \'10\'');utilsApi.sql("CREATE TABLE products(title text, price float) phrase_boundary = '., ?, !, U+2026' phrase_boundary_step = '10'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) phrase_boundary = '., ?, !, U+2026' phrase_boundary_step = '10'", true);utils_api.sql("CREATE TABLE products(title text, price float) phrase_boundary = '., ?, !, U+2026' phrase_boundary_step = '10'", Some(true)).await;table products {
phrase_boundary = ., ?, !, U+2026
phrase_boundary_step = 10
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}phrase_boundary_step = 100
短语边界词位置增量。可选,默认是 0。
在短语边界处,当前词位置会额外增加这个数值。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) phrase_boundary_step = '100' phrase_boundary = '., ?, !, U+2026'POST /cli -d "
CREATE TABLE products(title text, price float) phrase_boundary_step = '100' phrase_boundary = '., ?, !, U+2026'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'phrase_boundary_step' => '100',
'phrase_boundary' => '., ?, !, U+2026'
]);utilsApi.sql('CREATE TABLE products(title text, price float) phrase_boundary_step = \'100\' phrase_boundary = \'., ?, !, U+2026\'')await utilsApi.sql('CREATE TABLE products(title text, price float) phrase_boundary_step = \'100\' phrase_boundary = \'., ?, !, U+2026\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) phrase_boundary_step = \'100\' phrase_boundary = \'., ?, !, U+2026\'');utilsApi.sql("CREATE TABLE products(title text, price float) phrase_boundary_step = '100' phrase_boundary = '., ?, !, U+2026'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) phrase_boundary_step = '100' phrase_boundary = '., ?, !, U+2026'", true);utils_api.sql("CREATE TABLE products(title text, price float) phrase_boundary_step = '100' phrase_boundary = '., ?, !, U+2026'", Some(true)).await;table products {
phrase_boundary_step = 100
phrase_boundary = ., ?, !, U+2026
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}# index '13"' as '13inch'
regexp_filter = \b(\d+)\" => \1inch
# index 'blue' or 'red' as 'color'
regexp_filter = (blue|red) => color
用于过滤字段和查询的正则表达式。该指令是可选的,支持多值,默认是空正则表达式列表。Manticore Search 使用的正则引擎是 Google 的 RE2,它以速度快和安全性高而闻名。关于 RE2 支持的语法细节,可以查看 RE2 语法指南。
在某些应用中,例如商品搜索,产品、型号或属性可能有很多不同的表达方式。例如,iPhone 3gs 和 iPhone 3 gs(甚至 iPhone3 gs)很可能指的是同一款产品。另一个例子是笔记本屏幕尺寸的不同表达方式,例如 13-inch、13 inch、13" 或 13in。
正则表达式提供了一种机制,可以为这类情况指定定制规则。在第一个例子里,你或许可以用 wordforms 文件处理少量 iPhone 型号,但在第二个例子里,最好指定规则,把 "13-inch" 和 "13in" 归一化为同一种形式。
regexp_filter 中列出的正则表达式会按列出的顺序,在尽可能早的阶段应用,也就是在任何其他处理之前(包括 exceptions),甚至在分词之前。也就是说,正则表达式会在索引时作用于原始源字段,在搜索时作用于原始查询文本。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) regexp_filter = '(blue|red) => color'POST /cli -d "
CREATE TABLE products(title text, price float) regexp_filter = '(blue|red) => color'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'regexp_filter' => '(blue|red) => color'
]);utilsApi.sql('CREATE TABLE products(title text, price float) regexp_filter = \'(blue|red) => color\'')await utilsApi.sql('CREATE TABLE products(title text, price float) regexp_filter = \'(blue|red) => color\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) regexp_filter = \'(blue|red) => color\'');utilsApi.sql("CREATE TABLE products(title text, price float) regexp_filter = '(blue|red) => color'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) regexp_filter = '(blue|red) => color'", true);utils_api.sql("CREATE TABLE products(title text, price float) regexp_filter = '(blue|red) => color'", Some(true)).await;table products {
# index '13"' as '13inch'
regexp_filter = \b(\d+)\" => \1inch
# index 'blue' or 'red' as 'color'
regexp_filter = (blue|red) => color
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}通配符搜索是一种常见的文本搜索类型。在 Manticore 中,它是在词典级别执行的。默认情况下,普通表和 RT 表都使用一种名为 dict 的词典类型。在这种模式下,单词会按原样存储,因此启用通配符不会影响表的大小。执行通配符搜索时,系统会在词典中查找该通配词的所有可能展开形式。当展开结果很多,或者展开结果对应的命中列表很大时,这种展开在查询时的计算开销会成为问题,尤其是在中缀的情况下,因为通配符会同时加在单词的开头和结尾。为避免这类问题,可以使用 expansion_limit。
注意:更改已填充表的索引时通配符设置,例如
min_prefix_len或min_infix_len,只会影响更改之后被索引的文档。要将新设置应用到现有文档,请重新索引它们。
min_prefix_len = length
此设置确定要索引和搜索的最小单词前缀长度。默认情况下,它设置为 0,这意味着不允许前缀。
前缀允许通过 wordstart* 通配符进行通配符搜索。
例如,如果单词 "example" 使用 min_prefix_len=3 索引,则可以通过搜索 "exa"、"exam"、"examp" 和 "exampl" 以及完整的单词来找到它。
请注意,在 dict=crc 模式下,min_prefix_len 会影响全文索引的大小,因为每个单词展开形式都会被额外存储。
Manticore 可以区分完美匹配的单词和前缀匹配,并在满足以下条件时将前者排名更高:
- dict=keywords(默认开启)或 dict=keywords_32k
- index_exact_words=1(默认关闭),
- expand_keywords=1(默认也关闭)
请注意,在 dict=crc 模式下,或者在禁用上述任一选项时,都无法区分前缀和完整单词,因此无法对完全匹配的单词赋予更高排名。
当设置了 最小后缀长度 为正数时,最小前缀长度始终被视为 1。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) min_prefix_len = '3'POST /cli -d "
CREATE TABLE products(title text, price float) min_prefix_len = '3'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'min_prefix_len' => '3'
]);utilsApi.sql('CREATE TABLE products(title text, price float) min_prefix_len = \'3\'')await utilsApi.sql('CREATE TABLE products(title text, price float) min_prefix_len = \'3\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) min_prefix_len = \'3\'');utilsApi.sql("CREATE TABLE products(title text, price float) min_prefix_len = '3'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) min_prefix_len = '3'", true);utils_api.sql("CREATE TABLE products(title text, price float) min_prefix_len = '3'", Some(true)).await;table products {
min_prefix_len = 3
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}min_infix_len = length
min_infix_len 设置确定要索引和搜索的最小后缀前缀长度。它是可选的,默认值为 0,这意味着不允许后缀。最小允许的非零值为 2。
启用后,后缀允许使用如 start*、*end、*middle* 等术语模式进行通配符搜索。它还允许您禁用太短且搜索成本过高的通配符。
如果满足以下条件,Manticore 可以区分完美匹配的单词和后缀匹配,并将前者排名更高:
- dict=keywords(默认开启)或 dict=keywords_32k
- index_exact_words=1(默认关闭),
- expand_keywords=1(默认也关闭)
请注意,在 dict=crc 模式下,或者在禁用上述任一选项时,都无法区分中缀和完整单词,因此无法对完全匹配的单词赋予更高排名。
后缀通配符搜索的查询时间可能会有很大差异,这取决于子字符串实际扩展成多少关键词。像 *in* 或 *ti* 这样的短而频繁的音节可能会扩展成太多关键词,所有这些都需要匹配和处理。因此,为了启用子字符串搜索,通常会将 min_infix_len 设置为 2。为了限制来自太短通配符的搜索的影响,可能会将其设置得更高。
后缀必须至少为两个字符长,出于性能原因,*a* 类似的通配符是不允许的。
当 min_infix_len 设置为正数时, minimum prefix length 会被视为 1。对于 dict 词形中缀和前缀不能同时启用。对于 dict 以及通过 prefix_fields 声明前缀的其他字段,禁止在两个列表中声明同一个字段。
在 dict=keywords 或 dict=keywords_32k 下,除了通配符 * 之外,还可以使用另外两个通配符字符:
?可匹配任何(一个)字符:t?st将匹配test,但不匹配teast%可匹配零个或一个字符:tes%将匹配tes或test,但不匹配testing
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) min_infix_len = '3'POST /cli -d "
CREATE TABLE products(title text, price float) min_infix_len = '3'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'min_infix_len' => '3'
]);utilsApi.sql('CREATE TABLE products(title text, price float) min_infix_len = \'3\'')await utilsApi.sql('CREATE TABLE products(title text, price float) min_infix_len = \'3\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) min_infix_len = \'3\'');utilsApi.sql("CREATE TABLE products(title text, price float) min_infix_len = '3'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) min_infix_len = '3'", true);utils_api.sql("CREATE TABLE products(title text, price float) min_infix_len = '3'", Some(true)).await;table products {
min_infix_len = 3
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}prefix_fields = field1[, field2, ...]
prefix_fields 设置用于在 dict=crc 模式下限制前缀索引到特定的全文字段。默认情况下,所有字段都在前缀模式下进行索引,但因为前缀索引可能会影响索引和搜索性能,所以可能需要将其限制在某些字段上。
要将前缀索引限制到特定字段,请使用 prefix_fields 设置,后跟逗号分隔的字段名列表。如果未设置 prefix_fields,则所有字段将在前缀模式下进行索引。
- CONFIG
table products {
prefix_fields = title, name
min_prefix_len = 3
dict = crcinfix_fields = field1[, field2, ...]
infix_fields 设置允许您指定一个列表的全文字段,以限制中缀索引。此设置仅适用于 dict=crc,是可选的,默认情况下所有字段都在中缀模式下进行索引。 此设置类似于 prefix_fields,但它允许您将中缀索引限制到特定字段。
- CONFIG
table products {
infix_fields = title, name
min_infix_len = 3
dict = crcmax_substring_len = length
max_substring_len 指令设置为前缀或中缀搜索建立索引的最大子串长度。该设置是可选的,默认值为 0(这意味着所有可能的子串都会被建立索引)。它仅适用于 dict=crc。
默认情况下, dict=crc 的子串索引会把 所有 可能的子串都作为独立关键词建立索引,这可能会导致全文索引过大。因此,max_substring_len 指令允许你跳过过长的子串,因为这些子串很可能永远不会被搜索到。
例如,一个包含 10,000 篇博客文章的测试表在不同设置下的磁盘空间使用量如下:
- 基线:6.4 MB(没有子字符串)
- min_prefix_len = 3:24.3 MB(3.8 倍)
- min_prefix_len = 3,max_substring_len = 8:22.2 MB(3.5 倍)
- min_prefix_len = 3,max_substring_len = 6:19.3 MB(3.0 倍)
- min_infix_len = 3:94.3 MB(14.7 倍)
- min_infix_len = 3,max_substring_len = 8:84.6 MB(13.2 倍)
- min_infix_len = 3,max_substring_len = 6:70.7 MB(11.0 倍)
因此,限制最大子字符串长度可以节省 10-15% 的表大小。
在使用 dict=keywords 或 dict=keywords_32k 模式时,子串索引由词典处理,而不是通过预先建立 CRC 子串索引。因此,该指令不适用,并且在这种情况下会被刻意禁止。不过,如果需要,你仍然可以在应用代码中限制要搜索的子串长度。
- CONFIG
table products {
max_substring_len = 12
min_infix_len = 3
dict = crcexpand_keywords = {0|1|exact|star}
此设置扩展关键字及其精确形式和/或星号。支持的值包括:
- 1 - 扩展为精确形式和带有星号的形式。例如,
running将变为(running | *running* | =running) exact- 仅增加关键字的精确形式。例如,running将变为(running | =running)star- 通过在其周围添加星号来增加关键字。例如,running将变为(running | *running*)此设置是可选的,默认值为 0(关键字不扩展)。
启用 expand_keywords 特性的表中的查询内部会进行扩展:如果表是在前缀或中缀索引启用的情况下构建的,每个关键字都会被替换为其自身和相应的前缀或中缀(带星号)的析取。如果表是在启用了词干提取和 index_exact_words 的情况下构建的,精确形式也会被添加。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) expand_keywords = '1'POST /cli -d "
CREATE TABLE products(title text, price float) expand_keywords = '1'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'expand_keywords' => '1'
]);utilsApi.sql('CREATE TABLE products(title text, price float) expand_keywords = \'1\'')await utilsApi.sql('CREATE TABLE products(title text, price float) expand_keywords = \'1\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) expand_keywords = \'1\'');utilsApi.sql("CREATE TABLE products(title text, price float) expand_keywords = '1'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) expand_keywords = '1'", true);utils_api.sql("CREATE TABLE products(title text, price float) expand_keywords = '1'", Some(true)).await;table products {
expand_keywords = 1
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}扩展查询自然会花费更长时间完成,但可能会提高搜索质量,因为具有精确形式匹配的文档通常应该比具有词干或中缀匹配的文档排名更高。
请注意,现有的查询语法无法模拟这种扩展,因为内部扩展工作在关键字级别,并且会在短语或群集操作符内扩展关键字(这通过查询语法是不可能实现的)。请参阅示例以及 expand_keywords 如何影响搜索结果权重,以及如何通过不需要添加星号的方式找到“runsy”:
- expand_keywords_enabled
- expand_keywords_disabled
mysql> create table t(f text) min_infix_len='2' expand_keywords='1' morphology='stem_en';
Query OK, 0 rows affected, 1 warning (0.00 sec)
mysql> insert into t values(1,'running'),(2,'runs'),(3,'runsy');
Query OK, 3 rows affected (0.00 sec)
mysql> select *, weight() from t where match('runs');
+------+---------+----------+
| id | f | weight() |
+------+---------+----------+
| 2 | runs | 1560 |
| 1 | running | 1500 |
| 3 | runsy | 1500 |
+------+---------+----------+
3 rows in set (0.01 sec)
mysql> drop table t;
Query OK, 0 rows affected (0.01 sec)
mysql> create table t(f text) min_infix_len='2' expand_keywords='exact' morphology='stem_en';
Query OK, 0 rows affected, 1 warning (0.00 sec)
mysql> insert into t values(1,'running'),(2,'runs'),(3,'runsy');
Query OK, 3 rows affected (0.00 sec)
mysql> select *, weight() from t where match('running');
+------+---------+----------+
| id | f | weight() |
+------+---------+----------+
| 1 | running | 1590 |
| 2 | runs | 1500 |
+------+---------+----------+
2 rows in set (0.00 sec)mysql> create table t(f text) min_infix_len='2' morphology='stem_en';
Query OK, 0 rows affected, 1 warning (0.00 sec)
mysql> insert into t values(1,'running'),(2,'runs'),(3,'runsy');
Query OK, 3 rows affected (0.00 sec)
mysql> select *, weight() from t where match('runs');
+------+---------+----------+
| id | f | weight() |
+------+---------+----------+
| 1 | running | 1500 |
| 2 | runs | 1500 |
+------+---------+----------+
2 rows in set (0.00 sec)
mysql> drop table t;
Query OK, 0 rows affected (0.01 sec)
mysql> create table t(f text) min_infix_len='2' morphology='stem_en';
Query OK, 0 rows affected, 1 warning (0.00 sec)
mysql> insert into t values(1,'running'),(2,'runs'),(3,'runsy');
Query OK, 3 rows affected (0.00 sec)
mysql> select *, weight() from t where match('running');
+------+---------+----------+
| id | f | weight() |
+------+---------+----------+
| 1 | running | 1500 |
| 2 | runs | 1500 |
+------+---------+----------+
2 rows in set (0.00 sec)expansion_limit = number
单个通配符展开的最大关键词数。详情见 这里。
停用词是在索引和搜索过程中被忽略的词语,通常是因为它们的频率高且对搜索结果的贡献较低。
Manticore Search 默认会对停用词应用 stemming,这可能导致不理想的结果,但可以使用 stopwords_unstemmed 来关闭此功能。
小型停用词文件存储在表头中,嵌入文件的大小有限制,由 embedded_limit 选项定义。
停用词不会被索引,但会影响关键词的位置。例如,如果 "the" 是一个停用词,文档 1 包含短语 "in office",而文档 2 包含短语 "in the office",搜索 "in office" 作为精确短语时,只会返回第一个文档,即使第二个文档中的 "the" 被跳过作为停用词。此行为可以通过 stopword_step 指令进行修改。
stopwords=path/to/stopwords/file[ path/to/another/file ...]
stopwords 设置是可选的,默认为空。它允许您指定一个或多个停用词文件的路径,用空格分隔。所有文件都将被加载。在实时模式下,仅允许使用绝对路径。
停用词文件格式是简单的纯文本,使用 UTF-8 编码。文件数据将根据 charset_table 设置进行分词,因此您可以使用与索引数据相同的分隔符。
当 ngram_len 索引处于活动状态时,由 ngram_chars 中字符组成的停用词本身会被分词为 N-gram。因此,每个单独的 N-gram 都会成为单独的停用词。例如,使用 ngram_len=1 和合适的 ngram_chars,停用词 test 将被解释为 t、e、s、t 四个不同的停用词。
停用词文件可以手动或半自动创建。indexer 提供了一种模式,可以创建表的频率字典,按关键词频率排序。该字典中的顶级关键词通常可以用作停用词。有关详细信息,请参阅 --buildstops 和 --buildfreqs 开关。该字典中的顶级关键词通常可以用作停用词。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) stopwords = '/usr/local/manticore/data/stopwords.txt /usr/local/manticore/data/stopwords-ru.txt /usr/local/manticore/data/stopwords-en.txt'POST /cli -d "
CREATE TABLE products(title text, price float) stopwords = '/usr/local/manticore/data/stopwords.txt stopwords-ru.txt stopwords-en.txt'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'stopwords' => '/usr/local/manticore/data/stopwords.txt stopwords-ru.txt stopwords-en.txt'
]);utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'/usr/local/manticore/data/stopwords.txt /usr/local/manticore/data/stopwords-ru.txt /usr/local/manticore/data/stopwords-en.txt\'')await utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'/usr/local/manticore/data/stopwords.txt /usr/local/manticore/data/stopwords-ru.txt /usr/local/manticore/data/stopwords-en.txt\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'/usr/local/manticore/data/stopwords.txt /usr/local/manticore/data/stopwords-ru.txt /usr/local/manticore/data/stopwords-en.txt\'');utilsApi.sql("CREATE TABLE products(title text, price float) stopwords = '/usr/local/manticore/data/stopwords.txt /usr/local/manticore/data/stopwords-ru.txt /usr/local/manticore/data/stopwords-en.txt'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) stopwords = '/usr/local/manticore/data/stopwords.txt /usr/local/manticore/data/stopwords-ru.txt /usr/local/manticore/data/stopwords-en.txt'", true);utils_api.sql("CREATE TABLE products(title text, price float) stopwords = '/usr/local/manticore/data/stopwords.txt /usr/local/manticore/data/stopwords-ru.txt /usr/local/manticore/data/stopwords-en.txt'", Some(true)).await;table products {
stopwords = /usr/local/manticore/data/stopwords.txt
stopwords = stopwords-ru.txt stopwords-en.txt
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}或者,您可以使用 Manticore 自带的默认停用词文件。目前有 50 种语言的停用词可用。以下是它们的完整别名列表:
- af - 阿非利卡语
- ar - 阿拉伯语
- bg - 保加利亚语
- bn - 孟加拉语
- ca - 加泰罗尼亚语
- ckb- 库尔德语
- cz - 捷克语
- da - 丹麦语
- de - 德语
- el - 希腊语
- en - 英语
- eo - 世界语
- es - 西班牙语
- et - 爱沙尼亚语
- eu - 巴斯克语
- fa - 波斯语
- fi - 芬兰语
- fr - 法语
- ga - 爱尔兰语
- gl - 加利西亚语
- hi - 印地语
- he - 希伯来语
- hr - 克罗地亚语
- hu - 匈牙利语
- hy - 亚美尼亚语
- id - 印度尼西亚语
- it - 意大利语
- ja - 日语
- ko - 韩语
- la - 拉丁语
- lt - 立陶宛语
- lv - 拉脱维亚语
- mr - 马拉地语
- nl - 荷兰语
- no - 挪威语
- pl - 波兰语
- pt - 葡萄牙语
- ro - 罗马尼亚语
- ru - 俄语
- sk - 斯洛伐克语
- sl - 斯洛文尼亚语
- so - 索马里语
- st - 索托语
- sv - 瑞典语
- sw - 斯瓦希里语
- th - 泰语
- tr - 土耳其语
- yo - 约鲁巴语
- zh - 中文
- zu - 祖鲁语
例如,要在配置文件中使用意大利语的停用词,请添加以下行:
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) stopwords = 'it'POST /cli -d "
CREATE TABLE products(title text, price float) stopwords = 'it'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'stopwords' => 'it'
]);utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'it\'')await utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'it\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'it\'');utilsApi.sql("CREATE TABLE products(title text, price float) stopwords = 'it'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) stopwords = 'it'", true);utils_api.sql("CREATE TABLE products(title text, price float) stopwords = 'it'", Some(true)).await;table products {
stopwords = it
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}如果需要使用多种语言的停用词,应列出所有别名,用逗号(RT 模式)或空格(plain 模式)分隔:
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) stopwords = 'en, it, ru'POST /cli -d "
CREATE TABLE products(title text, price float) stopwords = 'en, it, ru'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'stopwords' => 'en, it, ru'
]);utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'en, it, ru\'')await utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'en, it, ru\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'en, it, ru\'');utilsApi.sql("CREATE TABLE products(title text, price float) stopwords = 'en, it, ru'", true);utilsApi.sql("CREATE TABLE products(title text, price float) stopwords = 'en, it, ru'", true);utils_api.sql("CREATE TABLE products(title text, price float) stopwords = 'en, it, ru'", Some(true)).await;table products {
stopwords = en it ru
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}stopwords_list = 'value1; value2; ...'
设置 stopwords_list 允许您直接在 CREATE TABLE 语句中指定停用词。它仅支持在 RT 模式 中。
值必须用分号(;)分隔。如果需要使用分号作为字面字符,必须用反斜杠(\;)转义。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
CREATE TABLE products(title text, price float) stopwords_list = 'a; the'POST /cli -d "
CREATE TABLE products(title text, price float) stopwords_list = 'a; the'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'stopwords_list' => 'a; the'
]);utilsApi.sql('CREATE TABLE products(title text, price float) stopwords_list = \'a; the\'')await utilsApi.sql('CREATE TABLE products(title text, price float) stopwords_list = \'a; the\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) stopwords_list = \'a; the\'');utilsApi.sql("CREATE TABLE products(title text, price float) stopwords_list = 'a; the'", true);utilsApi.sql("CREATE TABLE products(title text, price float) stopwords_list = 'a; the'", true);utils_api.sql("CREATE TABLE products(title text, price float) stopwords_list = 'a; the'", Some(true)).await;stopword_step={0|1}
在 停用词 上的 position_increment 设置是可选的,允许的值为 0 和 1,默认值为 1。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) stopwords = 'en' stopword_step = '1'POST /cli -d "
CREATE TABLE products(title text, price float) stopwords = 'en' stopword_step = '1'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'stopwords' => 'en, it, ru',
'stopword_step' => '1'
]);utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'en\' stopword_step = \'1\'')await utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'en\' stopword_step = \'1\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'en\' stopword_step = \'1\'');utilsApi.sql("CREATE TABLE products(title text, price float) stopwords = \'en\' stopword_step = \'1\'", true);utilsApi.sql("CREATE TABLE products(title text, price float) stopwords = \'en\' stopword_step = \'1\'", true);utils_api.sql("CREATE TABLE products(title text, price float) stopwords = \'en\' stopword_step = \'1\'", Some(true)).await;table products {
stopwords = en
stopword_step = 1
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}stopwords_unstemmed={0|1}
在词干提取之前或之后应用停用词。可选,默认值为 0(在词干提取之后应用停用词过滤器)。
默认情况下,停用词本身会被词干提取,然后应用于词干提取(或任何其他形态处理)后的标记。这意味着当 stem(token) 等于 stem(stopword) 时,标记会被停止。这种默认行为可能导致意外结果,当标记被错误地词干提取为被停止的根时。例如,“Andes”可能会被词干提取为“and”,所以当“and”是停用词时,“Andes”也会被跳过。
但是,您可以通过启用 stopwords_unstemmed 指令来更改此行为。启用后,停用词会在词干提取之前应用(因此应用于原始词形),当标记等于停用词时,标记会被跳过。
- SQL
- JSON
- PHP
- Python
- Python-asyncio
- javascript
- Java
- C#
- Rust
- CONFIG
CREATE TABLE products(title text, price float) stopwords = 'en' stopwords_unstemmed = '1'POST /cli -d "
CREATE TABLE products(title text, price float) stopwords = 'en' stopwords_unstemmed = '1'"$index = new \Manticoresearch\Index($client);
$index->setName('products');
$index->create([
'title'=>['type'=>'text'],
'price'=>['type'=>'float']
],[
'stopwords' => 'en, it, ru',
'stopwords_unstemmed' => '1'
]);utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'en\' stopwords_unstemmed = \'1\'')await utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'en\' stopwords_unstemmed = \'1\'')res = await utilsApi.sql('CREATE TABLE products(title text, price float) stopwords = \'en\' stopwords_unstemmed = \'1\'');utilsApi.sql("CREATE TABLE products(title text, price float) stopwords = \'en\' stopwords_unstemmed = \'1\'", true);utilsApi.Sql("CREATE TABLE products(title text, price float) stopwords = \'en\' stopwords_unstemmed = \'1\'", true);utils_api.sql("CREATE TABLE products(title text, price float) stopwords = \'en\' stopwords_unstemmed = \'1\'", Some(true)).await;table products {
stopwords = en
stopwords_unstemmed = 1
type = rt
path = tbl
rt_field = title
rt_attr_uint = price
}