ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

PyText 外部稠密特征(External Dense Features)接入指南:为文本分类模型注入数值型特征

PyText 外部稠密特征(External Dense Features)接入指南:为文本分类模型注入数值型特征 NLP深度学习【免费下载链接】pytextA natural language modeling framework based on PyTorch项目地址https://gitcode.com/gh_mirrors/py/pytext点击查看免费下载导读本文以 PyTextA natural language modeling framework based on PyTorch官方文档 dense.rst 为骨架系统讲解如何在文本分类模型中接入外部稠密特征Dense Features。你将学会稠密特征在数据中的组织方式、通过FloatListTensorizer完成数值化与可选归一化、自定义模型的Config/from_config/arrange_model_inputs/forward四处关键改动以及 PyText 内置DocModel中稠密特征从训练到 TorchScript 导出的完整落地路径。什么是外部稠密特征在自然语言处理建模中文本往往不是唯一的信号。例如要分类一条带有图片的文本可以先由图像模型单独提取图片特征再将这些特征与文本一起喂给分类器。这类文本之外的、由外部系统预先计算好的数值型特征就是本文所说的外部稠密特征External Dense Features。在 PyText 的数据组织方式中稠密特征以一个额外的数据列column存在该列的取值是一个 JSON 格式的浮点数列表list of floats例如[0.64, 0.75, 0.55, 0.24, 0.0, 0.94, ...]。仓库中提供了可直接运行的示例数据train_dense_features_tiny.tsv 与 test_dense_features_tiny.tsv其第 4 列即为 10 维稠密特征alarm/modify_alarm 16:24:datetime,39:57:datetime change my alarm tomorrow to wake me up 30 minutes earlier [0.64840776,0.7575,0.5531,0.2403,0,0.9481,0,0.1538,0.2403,0.3564]可见稠密特征与文本共用同一行样本只是以额外列的形式参与建模。基线无稠密特征的简单分类模型为了直观对比先看一个仅使用文本、不含稠密特征的简单分类器。它只展示模型代码中与本文主题相关的部分class MyModel(Model): class Config(Model.Config): class ModelInput(Model.Config.InputConfig): tokens: TokenTensorizer.Config TokenTensorizer.Config() labels: LabelTensorizer.Config LabelTensorizer.Config() inputs: ModelInput ModelInput() token_embedding: WordEmbedding.Config WordEmbedding.Config() representation: RepresentationBase.Config DocNNRepresentation.Config() decoder: DecoderBase.Config MLPDecoder.Config() output_layer: OutputLayerBase.Config ClassificationOutputLayer.Config() def from_config(cls, config, tensorizers): token_embedding create_module(config.token_embedding, tensorizertensorizers[tokens]) representation create_module(config.representation, embed_dimtoken_embedding.embedding_dim) labels tensorizers[labels].vocab decoder create_module( config.decoder, in_dimrepresentation.representation_dim out_dimlen(labels), ) output_layer create_module(config.output_layer, labelslabels) return cls(token_embedding, representation, decoder, output_layer) def arrange_model_inputs(self, tensor_dict): return (tensor_dict[tokens],) def forward( self, tokens_in: Tuple[torch.Tensor, torch.Tensor], ) - List[torch.Tensor]: word_tokens, seq_lens tokens embedding_out self.embedding(word_tokens) representation_out self.representation(embedding_out, seq_lens) return self.decoder(representation_out)该模型的数据通路为TokenTensorizer将文本转为 token 序列 →WordEmbedding生成词嵌入 →DocNNRepresentation生成文档表示 →MLPDecoder输出分类分数。此时 decoder 的输入维度完全取决于representation.representation_dim。加入稠密特征四处关键改动要启用稠密特征典型做法是让稠密特征直接进入 decoder绕过处理文本的 embedding 与 representation 阶段——即把稠密向量与文本表示向量拼接后送入分类器。下面是同一模型的稠密特征版本改动处已用# --标注class MyModel(Model): class Config(Model.Config): class ModelInput(Model.Config.InputConfig): tokens: TokenTensorizer.Config TokenTensorizer.Config() dense: FloatListTensorizer.Config FloatListTensorizer.Config() # -- labels: LabelTensorizer.Config LabelTensorizer.Config() inputs: ModelInput ModelInput() token_embedding: WordEmbedding.Config WordEmbedding.Config() representation: RepresentationBase.Config DocNNRepresentation.Config() decoder: DecoderBase.Config MLPDecoder.Config() output_layer: OutputLayerBase.Config ClassificationOutputLayer.Config() def from_config(cls, config, tensorizers): token_embedding create_module(config.token_embedding, tensorizertensorizers[tokens]) representation create_module(config.representation, embed_dimtoken_embedding.embedding_dim) dense_dim tensorizers[dense].out_dim # -- labels tensorizers[labels].vocab decoder create_module( config.decoder, in_dimrepresentation.representation_dim dense_dim # -- out_dimlen(labels), ) output_layer create_module(config.output_layer, labelslabels) return cls(token_embedding, representation, decoder, output_layer) def arrange_model_inputs(self, tensor_dict): return (tensor_dict[tokens], tensor_dict[dense]) # -- def forward( self, tokens_in: Tuple[torch.Tensor, torch.Tensor], dense_in: torch.Tensor, # -- ) - List[torch.Tensor]: word_tokens, seq_lens tokens embedding_out self.embedding(word_tokens) representation_out self.representation(embedding_out, seq_lens) representation_out torch.cat((representation_out, dense_in), 1) # -- return self.decoder(representation_out)四处改动逐一说明ModelInput 中新增dense输入声明FloatListTensorizer.Config与 tokens、labels 并列使稠密特征成为模型的标准输入之一。from_config中扩展 decoder 输入维度通过tensorizers[dense].out_dim取稠密特征的维度与representation.representation_dim相加后作为 decoder 的in_dim。这样 decoder 才能同时容纳文本表示与稠密向量。arrange_model_inputs返回稠密张量把tensor_dict[dense]追加进模型输入元组使其参与前向计算。forward中拼接稠密特征用torch.cat((representation_out, dense_in), 1)在特征维度第 1 维上拼接文本表示与稠密向量再送入 decoder。FloatListTensorizer稠密特征的数值化实现稠密特征的读取与批处理由 FloatListTensorizer 完成其Config包含以下关键字段配置项类型默认值说明columnstr必填从数据源中解析稠密特征所在的列名error_checkboolFalse开启后校验每行稠密向量长度是否等于dim不匹配直接断言失败dimOptional[int]None稠密特征的固定维度。配合归一化或错误检查时必须显式指定normalizeboolFalse是否在训练时对稠密特征做标准化(x - mean) / stddevis_inputbool继承自基类是否为模型输入实现细节来自源码 tensorizers.pycolumn_schema声明该列的数据类型为List[float]供数据管道解析。initialize / numberize若开启normalize在初始化阶段逐行统计特征均值与标准差update_meta_data→calculate_feature_stats随后在numberize中对每一行调用normalizer.normalize完成标准化否则normalize为恒等变换。error_check开启时校验len(dense) self.dim用于在训练早期发现特征维度不一致的数据。tensorize将 batch 中的稠密向量pad_and_tensorize为torch.float张量在 fp16 训练下会通过pad_shape显式指定(batch, dim)形状避免多余 padding。需要特别注意的是normalize与error_check都要求dim非空源码中通过assert not normalize or self.dim is not None与assert not self.error_check or self.dim is not None强制保证。若两者均关闭且dim为Nonedim会被置为 0 以正常构造VectorNormalizer。归一化的底层实现位于 VectorNormalizer它按特征维度逐位计算(x - mean) / stddev其中标准差为 0 的特征会除以 1.0 以避免除零。该模块被设计为torch.nn.ModuleScriptModule 友好因此训练时可在 tensorizer 中使用、推理时也可直接嵌入 TorchScript 前向函数——这正是下一节DocModel的做法。配置实战docnn_dense_feat.json 与数据管道仓库在 demo/configs/docnn_dense_feat.json 中提供了开箱即用的完整配置展示稠密特征在 PyText 配置体系中如何被声明{ task: { DocClassificationTask: { features : { dense_feat : { column: dense_feat, dim: 10 } }, data_handler: { columns_to_read: [doc_label, text, dict_feat, dense_feat], train_path: tests/data/train_dense_features_tiny.tsv, eval_path: tests/data/test_dense_features_tiny.tsv, test_path: tests/data/test_dense_features_tiny.tsv } } } }要点解读features 段以键名dense_feat声明稠密特征 tensorizer指定column: dense_feat与dim: 10对应示例数据中每个样本的 10 维浮点列表。columns_to_read数据管道按该列表依次读取列稠密特征列与文本、标签列并列这是外部特征作为额外列加入输入数据在配置层的直接体现。训练/评估/测试路径分别指向 train_dense_features_tiny.tsv 与 test_dense_features_tiny.tsv。对比不含稠密特征的基线配置 demo/configs/docnn.json其TSVDataSource.field_names仅为[label, slots, text]可以清晰看出加入稠密特征后数据源需要新增一个dense_feat列tensorizer 需要新增对应的FloatListTensorizer声明。DocModel 的官方实现稠密特征如何进入 decoder官方内置的文档分类模型 DocModel 正是稠密特征直接进 decoder这一模式的完整实现ModelInput 声明dense: Optional[FloatListTensorizer.Config] Nonedoc_model.py即稠密特征是可选输入不配置时模型退化为纯文本分类。create_decoder 动态扩维doc_model.py当config.inputs.dense存在时in_dim累加config.inputs.dense.dim即in_dim representation_dim dense_dim并记录decoder.num_decoder_modules 1。arrange_model_inputs 条件拼接doc_model.py仅当dense in tensor_dict时才把稠密张量追加进模型输入。导出输入名get_export_input_names在有稠密特征时追加float_vec_valsdoc_model.py。from_config的组装逻辑doc_model.py与文档示例模型一致embedding → representation →create_decoder动态扩维→ output layer其中MLPDecoder是in_dim到out_dim的全连接网络默认使用 ReLU 激活mlp_decoder.pyClassificationOutputLayer则默认使用CrossEntropyLoss并支持标签权重doc_classification_output_layer.py。推理与导出TorchScript 中的稠密特征归一化稠密特征接入后的推理导出同样被官方支持。DocModel.torchscriptifydoc_model.py会依据是否配置densetensorizer 返回两类 ScriptModuleModel无稠密特征输入为texts/tokens前向完成 tokenize、查表、padding 后调用 traced model。ModelWithDenseFeat有稠密特征额外接收dense_feat: Optional[List[List[float]]]并在进入 traced model 前调用self.normalizer.normalize(dense_feat)完成与训练一致的标准化再以torch.float张量传入。这一设计印证了FloatListTensorizer中normalize字段的注释若在训练数据上做了归一化推理数据也应归一化保证训练/推理分布一致。TorchScript 路径通过把VectorNormalizer作为模块成员嵌入导出模型使归一化在推理侧零成本复用。配套的导出配置示例见 demo/configs/docnn_wo_export.json 等文件export_torchscript_path/export_caffe2_path字段可指定导出产物路径适合在训练完成后直接产出可用于线上服务的模型文件。总结与实践建议综合官方文档与仓库源码接入外部稠密特征的最小改动可归纳为四条主线数据侧在 TSV 中新增一列 JSON 浮点列表并在配置的features中声明FloatListTensorizercolumndim。模型侧在ModelInput中加入dense在from_config中把tensorizers[dense].out_dim并入 decoder 的in_dim在arrange_model_inputs与forward中完成张量传递与torch.cat拼接。归一化若训练时开启normalize务必在推理/导出侧使用同一VectorNormalizer保证分布一致error_check可用于维度校验。复用官方实现文本分类任务可直接使用DocModel其可选dense输入、动态扩维 decoder 与 TorchScript 导出链路均已内置无需自行实现。这套模式适用于任何文本 外部数值特征的场景——无论是图像特征、用户画像特征还是规则引擎输出的打分特征只要特征能以固定维度的浮点向量表达即可按本文所述方式接入 PyText 模型。赞分享NLP深度学习【免费下载链接】pytextA natural language modeling framework based on PyTorch项目地址https://gitcode.com/gh_mirrors/py/pytext点击查看免费下载相关推荐D2L.ai特征工程数值特征、类别特征与文本特征处理完整指南D2L.ai特征工程数值特征、类别特征与文本特征处理完整指南 想要在深度学习项目中获得出色的模型性能吗特征工程是关键 D2L.ai 提供了完整的特征工文档教程人工智能深度学习NLP计算机视觉强化学习Rust进阶特征深入理解关联类型与特征继承Rust进阶特征深入理解关联类型与特征继承 在Rust编程语言中特征 trait 是定义共享行为的强大工具。本文将深入探讨Rust特征系统的一些高级概念包文档教程示例工程EfficientLoFTR高效半稠密局部特征匹配EfficientLoFTR高效半稠密局部特征匹配 项目介绍 EfficientLoFTR是一个基于深度学习的半稠密局部特征匹配项目它以类似稀疏特征匹配的速人工智能计算机视觉深度学习科研上一篇如何快速掌握FilesWindows上最强大的现代化文件管理器下一篇B站会员购自动化抢票工具从手动到智能的抢票革命创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表