
作者来自 Elastic David Pilato本教程将Apache Lucene作为进程内搜索索引嵌入用于搜索领域 Bean —— 这里使用的是 Rekordbox 风格音乐库中的Track记录。同样的模式也适用于任何 Java Bean。Lucene 作为一个派生缓存位于你的对象旁边而不是事实来源。将数据库作为系统的权威记录将 Bean 映射为Document执行搜索将命中的 id 与原始列表关联起来并在成功写入后重新构建或 upsert Lucene。这篇文章只介绍映射 —— 先介绍 analyzer然后介绍字段。将 Lucene 添加到 Maven一个项目两个 artifact使用相同版本。实现时请在 Maven Central 上查找最新的稳定 Lucene 版本本系列使用10.5.1。!-- Index, search, documents, queries -- dependency groupIdorg.apache.lucene/groupId artifactIdlucene-core/artifactId version10.5.1/version /dependency !-- Tokenizers / filters -- dependency groupIdorg.apache.lucene/groupId artifactIdlucene-analysis-common/artifactId version10.5.1/version /dependencyLucene 是纯 Java它可以被打包到 fat-jar 中无需任何本地库。从你现有的 Bean 开始补充现有 Bean 的 Java 代码翻译并补全后续示例内容说明 fat-jar 和本地库public record Track( String id, String title, Artist artist, Genre genre, MusicalKey key, double bpm, int ratingStars, int year // … album, comment, paths, … ) {}为需要查找文档的内容建立索引完整的 Bean 保存在其他地方并在搜索后通过 id 进行关联。稳定的 id—Track.id用于 upsert 和删除。全文搜索— 用户输入的字符串title、artist。过滤 / 范围查询— 精确的 keyword 或数值genre、rating、bpm、year。选择 analyzerAround The World经过 StandardTokenizer → LowerCaseFilter → ASCIIFoldingFilter。对于TextFieldanalyzer 会在索引时运行并且应该与查询时生成的 token 保持一致补充查询时的 analyzer 示例说明各个 token filter 的作用完善 Java 代码示例Analyzer analyzer new Analyzer() { Override protected TokenStreamComponents createComponents(String fieldName) { Tokenizer source new StandardTokenizer(); TokenStream filter new LowerCaseFilter(source); filter new ASCIIFoldingFilter(filter); return new TokenStreamComponents(source, filter); } }; // Analyze a text TokenStream ts analyzer.tokenStream(title, Around The World);不进行 stemmingartist 名称保持完整不使用停用词Around The World仍然可以被搜索。ASCII folding 会将café/François转换为cafe/francoisCafé del Mar — Around The World (François Kevorkian Mix)— tokenizer → lowercase → ASCII folding重音符号会在最后一个阶段进行折叠。最终的 token 会以排序后的形式进入索引around、cafe、del……——就像一本书最后的索引一样。字母顺序让人们无需阅读每一页就能快速找到某个词条Lucene 采用了相同的思路因此查找时可以直接跳转到所需的 term而不是扫描整个词典。在输入和输出时使用相同的analyzer。将 Bean 映射为 LuceneDocument补充完整 Bean 映射示例统一术语与格式风格澄清 analyzer 的使用时机选择一条 trackLucene 会存储一个可用于搜索的DocumentTextField / StringField / 数值字段。模式示例Lucene 类型分析后的文本title、artistTextField精确 keywordid、genre.rawStringField数值bpm、rating、yearDoubleField/IntFieldTextField会进行 token 化用于搜索。StringField不会进行 token 化用于 id、过滤条件。数值字段用于范围过滤和排序——暂时还不用于直方图。存储你需要用来呈现匹配结果的数据Field.Store.YES无论如何都要存储 id。这就是一个可用于搜索的DocumentDocument doc new Document(); // stored join key back to the Track bean doc.add(new StringField(id, 172523747, Store.YES)); // title: TextField is analyzed (MUST). .raw keeps the original for display. .raw.normalized is the exact FILTER. doc.add(new TextField(title, Around The World, Store.YES)); doc.add(new StringField(title.raw, Around The World, Store.YES)); doc.add(new StringField(title.raw.normalized, around the world, Store.YES)); // artist: TextField is analyzed (MUST). .raw keeps the original for display. .raw.normalized is the exact FILTER. doc.add(new TextField(artist, Daft Punk, Store.YES)); doc.add(new StringField(artist.raw, Daft Punk, Store.YES)); doc.add(new StringField(artist.raw.normalized, daft punk, Store.YES)); // genre: analyzed text keyword FILTER (.raw.normalized) doc.add(new TextField(genre, Club, Store.YES)); doc.add(new StringField(genre.raw, Club, Store.YES)); doc.add(new StringField(genre.raw.normalized, club, Store.YES)); // numeric range / sort. numericValue() is IEEE 754 bits; read storedValue().getDoubleValue() doc.add(new DoubleField(bpm, 121.29, Store.YES)); // Camelot key — exact FILTER / MUST_NOT (lowercased) doc.add(new StringField(key.code, 9a, Store.YES)); // rating: numeric filter / sort doc.add(new IntField(rating, 5, Store.YES)); // year: numeric filter / sort doc.add(new IntField(year, 1997, Store.YES)); // album: analyzed free text only — no keyword twin doc.add(new TextField(album, , Store.YES)); // label: analyzed free text only — no keyword twin doc.add(new TextField(label, , Store.YES)); // comment: analyzed free text only — no keyword twin doc.add(new TextField(comment, 09A - Energy 7, Store.YES))完整演示代码位于 GitHublucene-search-tracks。原文Search your beans with Lucene — Mapping | David Pilato