ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

MMDetection3D ScanNet 数据集使用指南:数据导出、训练管线与评测全流程

MMDetection3D ScanNet 数据集使用指南:数据导出、训练管线与评测全流程 人工智能计算机视觉深度学习自动驾驶【免费下载链接】mmdetection3dOpenMMLabs next-generation platform for general 3D object detection.项目地址https://gitcode.com/gh_mirrors/mm/mmdetection3d点击查看免费下载ScanNet 是 MMDetection3D 中面向室内场景的 LiDAR 点云 3D 检测与语义分割基准数据集。本文以 scannet.md 为主线系统讲解从原始扫描数据导出、RGB 位姿抽取、info 文件生成到检测/分割训练管线、评估指标与在线 Benchmark 提交的完整流程并辅以仓库源码转换脚本、配置基线与数据变换实现验证每个环节的底层原理帮助读者在本地完整复现 ScanNet 任务的训练与测试。ScanNet 在 MMDetection3D 中的定位ScanNet 提供 1513 个室内扫描场景官方划分为 1201 个训练场景、312 个验证场景与 100 个测试场景每个场景包含稠密重建的彩色网格mesh、语义/实例标注与轴对齐的 3D 物体包围盒。MMDetection3D 以原始点云为输入同时支持两类任务LiDAR 3D 物体检测基于ScanNetDatasetmmdet3d/datasets/scannet_dataset.py采用 Depth 坐标系与轴对齐yaw 恒为 0的 3D 包围盒3D 语义分割基于ScanNetSegDatasetmmdet3d/datasets/seg3d_dataset.py通常做 20 类分类19 个有效类别 1 个 ignore 类。从整体流程看数据准备阶段分成两步先用 ScanNet 官方工具把.sens原始数据导出为点云与标注再通过tools/create_data.py生成训练所需的.bin点云/掩码文件与.pklinfo 文件此后的训练、评测与提交则完全由 MMDetection3D 的数据管线与评估模块驱动。数据准备ScanNet 数据准备的完整流程含原始数据下载、目录布局以数据目录下的 README 为准本小节聚焦仓库内可复现的核心步骤。导出 ScanNet 点云数据导出点云数据的目的是把原始 mesh 数据转换为点云并生成语义标签、实例标签与 GT 3D 包围盒三类标注。在data/scannet目录下执行python batch_load_scannet_data.py准备开始前的目录结构如下mmdetection3d ├── mmdet3d ├── tools ├── configs ├── data │ ├── scannet │ │ ├── meta_data │ │ ├── scans │ │ │ ├── scenexxxx_xx │ │ ├── batch_load_scannet_data.py │ │ ├── load_scannet_data.py │ │ ├── scannet_utils.py │ │ ├── README.mdscans目录下共有 1201 个训练场景与 312 个验证场景文件夹每个文件夹以scene0001_01为例包含以下原始文件文件内容scene0001_01_vh_clean_2.ply保存每个顶点坐标与颜色的 mesh 文件mesh 顶点即原始点云数据scene0001_01.aggregation.json聚合文件包含物体 ID、分段 ID 与语义标签scene0001_01_vh_clean_2.0.010000.segs.json分段文件包含分段 ID 与顶点归属scene0001_01.txt元数据文件包含轴对齐矩阵axis-aligned matrix等scene0001_01_vh_clean_2.labels.ply标注文件保存每个顶点的语义类别运行python batch_load_scannet_data.py的导出过程主要包含 3 步将原始文件导出为点云、实例标签、语义标签与包围盒文件对原始点云下采样并过滤无效类别保存点云数据与相关标注文件。load_scannet_data.py中的核心函数export完整逻辑如下def export(mesh_file, agg_file, seg_file, meta_file, label_map_file, output_fileNone, test_modeFalse): # label map file: ./data/scannet/meta_data/scannetv2-labels.combined.tsv # the various label standards in the label map file, e.g. nyu40id label_map scannet_utils.read_label_mapping( label_map_file, label_fromraw_category, label_tonyu40id) # load raw point cloud data, 6-dims feature: XYZRGB mesh_vertices scannet_utils.read_mesh_vertices_rgb(mesh_file) # Load scene axis alignment matrix: a 4x4 transformation matrix # transform raw points in sensor coordinate system to a coordinate system # which is axis-aligned with the length/width of the room lines open(meta_file).readlines() # test set data doesnt have align_matrix axis_align_matrix np.eye(4) for line in lines: if axisAlignment in line: axis_align_matrix [ float(x) for x in line.rstrip().strip(axisAlignment ).split( ) ] break axis_align_matrix np.array(axis_align_matrix).reshape((4, 4)) # perform global alignment of mesh vertices pts np.ones((mesh_vertices.shape[0], 4)) # raw point cloud in homogeneous coordinates, each row: [x, y, z, 1] pts[:, 0:3] mesh_vertices[:, 0:3] # transform raw mesh vertices to aligned mesh vertices pts np.dot(pts, axis_align_matrix.transpose()) # Nx4 aligned_mesh_vertices np.concatenate([pts[:, 0:3], mesh_vertices[:, 3:]], axis1) # Load semantic and instance labels if not test_mode: # each object has one semantic label and consists of several segments object_id_to_segs, label_to_segs read_aggregation(agg_file) # many points may belong to the same segment seg_to_verts, num_verts read_segmentation(seg_file) label_ids np.zeros(shape(num_verts), dtypenp.uint32) object_id_to_label_id {} for label, segs in label_to_segs.items(): label_id label_map[label] for seg in segs: verts seg_to_verts[seg] # each point has one semantic label label_ids[verts] label_id instance_ids np.zeros( shape(num_verts), dtypenp.uint32) # 0: unannotated for object_id, segs in object_id_to_segs.items(): for seg in segs: verts seg_to_verts[seg] # object_id is 1-indexed, i.e. 1,2,3,.,,,.NUM_INSTANCES # each point belongs to one object instance_ids[verts] object_id if object_id not in object_id_to_label_id: object_id_to_label_id[object_id] label_ids[verts][0] # bbox format is [x, y, z, x_size, y_size, z_size, label_id] # [x, y, z] is gravity center of bbox, [x_size, y_size, z_size] is axis-aligned # [label_id] is semantic label id in nyu40id standard # Note: since 3D bbox is axis-aligned, the yaw is 0. unaligned_bboxes extract_bbox(mesh_vertices, object_id_to_segs, object_id_to_label_id, instance_ids) aligned_bboxes extract_bbox(aligned_mesh_vertices, object_id_to_segs, object_id_to_label_id, instance_ids) ... return mesh_vertices, label_ids, instance_ids, unaligned_bboxes, \ aligned_bboxes, object_id_to_label_id, axis_align_matrix几点关键设计值得注意标签映射通过meta_data/scannetv2-labels.combined.tsv中的raw_category → nyu40id映射把原始类别统一到nyu40id标准轴对齐矩阵axisAlignment是一个 4×4 变换矩阵把传感器坐标系下的原始点变换到与房间长宽轴对齐的坐标系测试集数据没有该矩阵故默认置为单位阵包围盒格式[x, y, z, x_size, y_size, z_size, label_id][x, y, z]为包围盒重心尺寸为轴对齐尺寸标签为nyu40id标准下的语义 ID由于 3D 包围盒是轴对齐的yaw 恒为 0实例标注instance_ids中 0 表示未标注物体 ID 从 1 开始递增每个点同时拥有唯一的语义标签与实例归属。每个场景导出完成后若点数过多会对原始点云下采样例如降到 50000 点若该点云同时用于 3D 语义分割任务则不下采样。此外需要过滤掉nyu40id标准之外的无效语义标签以及可选的DONOT CARE类别。最终点云数据、语义标签、实例标签与 GT 包围盒保存为.npy文件。导出 ScanNet RGB 数据可选若计划使用多视角检测如 MVXNet、ImVoteNet 等基于 RGB 图像的多模态方案需要额外导出每个场景的 RGB 图像及对应位姿。此步骤为可选项若不做多视角检测可完全跳过。python extract_posed_images.py1201 个训练场景、312 个验证场景与 100 个测试场景各含一个.sens文件例如data/scannet/scans/scene0001_01/0001_01.sens。运行后该场景的全部图像与位姿被导出到data/scannet/posed_images/scene0001_01具体包含约 300 张xxxxx.jpg图像、300 个xxxxx.txt位姿文件每个均为 4×4 矩阵以及一个单独的intrinsic.txt相机内参文件。典型场景实际包含数千张图像为控制磁盘占用默认导出的结果占用小于 100 GB默认只导出其中 300 张如需导出更多图像可使用--max-images-per-scene参数调整。创建数据集生成 .bin 与 .pkl完成上述导出后运行如下命令生成训练与验证所用的数据文件python tools/create_data.py scannet --root-path ./data/scannet \ --out-dir ./data/scannet --extra-tag scannet该命令由 tools/dataset_converters/indoor_converter.py 中的create_indoor_info_file驱动pkl_prefixscannet分支把上一步导出的点云文件、语义标签文件、实例标签文件进一步保存为.bin格式同时为 train/val/test 生成.pklinfo 文件。从源码看ScanNet 是少数同时具有 train-val-test 三段划分的室内数据集# ScanNet has a train-val-test split train_dataset ScanNetData(root_pathdata_path, splittrain) val_dataset ScanNetData(root_pathdata_path, splitval) test_dataset ScanNetData(root_pathdata_path, splittest) test_filename os.path.join(save_path, f{pkl_prefix}_infos_test.pkl)其中 train/val 调用get_infos(..., has_labelTrue)test 调用get_infos(..., has_labelFalse)tools/dataset_converters/indoor_converter.py。获取单场景 info 的核心函数process_single_scene位于 tools/dataset_converters/scannet_data_utils.py如下def process_single_scene(sample_idx): # save point cloud, instance label and semantic label in .bin file respectively, get info[pts_path], info[pts_instance_mask_path] and info[pts_semantic_mask_path] ... # get annotations if has_label: annotations {} # box is of shape [k, 6 class] aligned_box_label self.get_aligned_box_label(sample_idx) unaligned_box_label self.get_unaligned_box_label(sample_idx) annotations[gt_num] aligned_box_label.shape[0] if annotations[gt_num] ! 0: aligned_box aligned_box_label[:, :-1] # k, 6 unaligned_box unaligned_box_label[:, :-1] classes aligned_box_label[:, -1] # k annotations[name] np.array([ self.label2cat[self.cat_ids2class[classes[i]]] for i in range(annotations[gt_num]) ]) # default names are given to aligned bbox for compatibility # we also save unaligned bbox info with marked names annotations[location] aligned_box[:, :3] annotations[dimensions] aligned_box[:, 3:6] annotations[gt_boxes_upright_depth] aligned_box annotations[unaligned_location] unaligned_box[:, :3] annotations[unaligned_dimensions] unaligned_box[:, 3:6] annotations[ unaligned_gt_boxes_upright_depth] unaligned_box annotations[index] np.arange( annotations[gt_num], dtypenp.int32) annotations[class] np.array([ self.cat_ids2class[classes[i]] for i in range(annotations[gt_num]) ]) axis_align_matrix self.get_axis_align_matrix(sample_idx) annotations[axis_align_matrix] axis_align_matrix # 4x4 info[annos] annotations return info这里保存了aligned与unaligned两套包围盒默认以轴对齐aligned包围盒为基准与评测口径一致同时以带标记名称的 unaligned 字段保留未对齐信息axis_align_matrix一并写入 annos供训练管线中的GlobalAlignment使用。此外源码还处理了多视角数据的读取若posed_images目录存在则加载每场景的内参intrinsics与所有外参extrinsics、图像路径img_paths并且会过滤掉包含非法非有限数值的外参矩阵tools/dataset_converters/scannet_data_utils.py。处理完成后的目录结构如下scannet ├── meta_data ├── batch_load_scannet_data.py ├── load_scannet_data.py ├── scannet_utils.py ├── README.md ├── scans ├── scans_test ├── scannet_instance_data ├── points │ ├── xxxxx.bin ├── instance_mask │ ├── xxxxx.bin ├── semantic_mask │ ├── xxxxx.bin ├── seg_info │ ├── train_label_weight.npy │ ├── train_resampled_scene_idxs.npy │ ├── val_label_weight.npy │ ├── val_resampled_scene_idxs.npy ├── posed_images │ ├── scenexxxx_xx │ │ ├── xxxxxx.txt │ │ ├── xxxxxx.jpg │ │ ├── intrinsic.txt ├── scannet_infos_train.pkl ├── scannet_infos_val.pkl ├── scannet_infos_test.pkl各目录与字段的含义points/xxxxx.bin下采样后的未轴对齐点云。由于 ScanNet 3D 检测任务以轴对齐点云为输入而 3D 语义分割任务使用未对齐点云因此仓库统一存储未对齐点云及其轴对齐变换矩阵检测任务在预处理管线GlobalAlignmentmmdet3d/datasets/transforms/transforms_3d.py中完成对齐instance_mask/xxxxx.bin逐点实例标签取值范围[0, NUM_INSTANCES]0 表示未标注semantic_mask/xxxxx.bin逐点语义标签取值范围[1, 40]即nyu40id标准训练时由管线PointSegClassMapping将nyu40id映射为训练 IDseg_info为语义分割训练生成的信息train_label_weight.npy各语义类别的权重因子。不同类别点数差异悬殊使用标签重加权label re-weighting是提升性能的常见做法train_resampled_scene_idxs.npy逐场景重采样索引。不同房间会根据其点数被多次采样以平衡训练数据posed_images/scenexxxx_xx一组.jpg图像、对应.txt4×4 位姿文件与单个intrinsic.txt内参矩阵scannet_infos_train.pkl训练 info 文件每个场景包含info[lidar_points]包含点云相关信息的字典info[lidar_points][lidar_path]点云数据文件名info[lidar_points][num_pts_feats]点的特征维度info[lidar_points][axis_align_matrix]轴对齐变换矩阵info[pts_semantic_mask_path]语义掩码标注文件名info[pts_instance_mask_path]实例掩码标注文件名info[instances]包含所有标注的字典列表每个字典对应单个实例的完整标注。对第 i 个实例info[instances][i][bbox_3d]6 个数字组成的列表表示深度坐标系下的轴对齐 3D 包围盒顺序为(x, y, z, l, w, h)info[instances][i][bbox_label_3d]3D 包围盒标签scannet_infos_val.pkl验证 info 文件格式与训练文件相同scannet_infos_test.pkl测试 info 文件格式与训练文件几乎相同但缺少标注has_labelFalse。值得一提的是分割任务用的seg_info由ScanNetSegData.get_seg_infostools/dataset_converters/scannet_data_utils.py生成采样索引的生成逻辑是点数越多的场景被采样越多次sample_prob正比于场景点数num_iter由总点数除以num_points8192得到标签权重则采用 PointNet 论文中1.0 / np.log(1.2 x)的反频加权函数。测试集不生成 seg info。训练管线ScanNet 检测与分割任务的数据增强策略差异很大检测任务面向整场景 40000 点采样与全局旋转分割任务则采用室内 patch 裁剪。3D 检测训练管线典型检测训练管线如下与 configs/base/datasets/scannet-3d.py 一致train_pipeline [ dict( typeLoadPointsFromFile, coord_typeDEPTH, shift_heightTrue, load_dim6, use_dim[0, 1, 2]), dict( typeLoadAnnotations3D, with_bbox_3dTrue, with_label_3dTrue, with_mask_3dTrue, with_seg_3dTrue), dict(typeGlobalAlignment, rotation_axis2), dict(typePointSegClassMapping), dict(typePointSample, num_points40000), dict( typeRandomFlip3D, sync_2dFalse, flip_ratio_bev_horizontal0.5, flip_ratio_bev_vertical0.5), dict( typeGlobalRotScaleTrans, rot_range[-0.087266, 0.087266], scale_ratio_range[1.0, 1.0], shift_heightTrue), dict( typePack3DDetInputs, keys[ points, gt_bboxes_3d, gt_labels_3d, pts_semantic_mask, pts_instance_mask ]) ]各变换要点GlobalAlignment利用axis_align_matrix对点云做全局轴对齐。源码实现mmdet3d/datasets/transforms/transforms_3d.py会校验旋转矩阵的合法性行列式为 1且rotation_axis对应轴保持不动此处rotation_axis2即绕 z 轴旋转然后执行旋转与平移。注释特别说明这里不记录对齐变换与GlobalRotScaleTrans不同因为 ScanNet 检测任务直接用对齐后的 GT 包围盒评测无需逆变换回滚PointSegClassMapping把原始语义类别映射到有效类别 ID。源码mmdet3d/datasets/transforms/loading.py通过seg_label_mapping数组做查表映射有效类别映射为0 ~ len(valid_cat_ids)-1其余类别映射为len(valid_cat_ids)检测任务中有效类别为 18 类见ScanNetData.classestools/dataset_converters/scannet_data_utils.py对应nyu40id集合{3,4,5,6,7,8,9,10,11,12,14,16,24,28,33,34,36,39}数据增强PointSample对输入点云下采样num_points40000RandomFlip3D以 0.5 概率在 BEV 水平方向或垂直方向随机翻转点云GlobalRotScaleTrans旋转点云ScanNet 通常取[-5, 5]度代码中的[-0.087266, 0.087266]即 ±5° 的弧度值缩放因子固定为 1.0即不缩放平移量为 0即不平移。配置基线中检测任务还使用了RepeatDataset(times5)重复训练数据、box_type_3dDepth的 Depth 盒子类型与filter_empty_gtFalse允许无 GT 的场景参与训练评测器为IndoorMetricconfigs/base/datasets/scannet-3d.py。3D 语义分割训练管线典型语义分割训练管线如下与 configs/base/datasets/scannet-seg.py 一致train_pipeline [ dict( typeLoadPointsFromFile, coord_typeDEPTH, shift_heightFalse, use_colorTrue, load_dim6, use_dim[0, 1, 2, 3, 4, 5]), dict( typeLoadAnnotations3D, with_bbox_3dFalse, with_label_3dFalse, with_mask_3dFalse, with_seg_3dTrue), dict( typePointSegClassMapping), dict( typeIndoorPatchPointSample, num_pointsnum_points, block_size1.5, ignore_indexlen(class_names), use_normalized_coordFalse, enlarge_size0.2, min_unique_numNone), dict(typeNormalizePointsColor, color_meanNone), dict(typePack3DDetInputs, keys[points, pts_semantic_mask]) ]要点说明PointSegClassMapping有效类别映射为[0, 20)的训练 ID其余类别映射为ignore_index等于20。configs/_base_/datasets/scannet-seg.py中定义了 20 个类别名wall、floor、cabinet、bed、chair、sofa、table、door、window、bookshelf、picture、counter、desk、curtain、refrigerator、showercurtrain、toilet、sink、bathtub、otherfurnitureIndoorPatchPointSample从输入点云裁剪包含固定点数num_points默认 8192的 patch。block_size表示裁剪块的大小ScanNet 通常取1.5米enlarge_size0.2会将采样 patch 扩大为[-block_size/2 - enlarge_size, block_size/2 enlarge_size]作为数据增强ignore_index同时用作 patch 选取准则优先选取含有效标签的 patch。源码实现mmdet3d/datasets/transforms/transforms_3d.py源自 PointNet 的scannet_dataset.py训练时还会减去 patch 中心坐标z 维不做中心化并可选择拼接归一化坐标特征use_normalized_coordNormalizePointsColor将点云 RGB 颜色值除以255归一化源码见 mmdet3d/datasets/transforms/loading.py。注意该管线shift_heightFalse且use_dim[0,1,2,3,4,5]使用 XYZRGB 全部 6 维特征与检测管线仅用 XYZshift_heightTrue形成对比。评估指标3D 物体检测mAPScanNet 检测任务通常使用平均精度均值mAP评估常见口径为mAP0.25与mAP0.5。仓库调用多类别 3D 检测的通用精度/召回计算函数indoor_evalmmdet3d/evaluation/functional/indoor_eval.py先按类别汇总预测框含 score与 GT 框再对每个 IoU 阈值计算各类别 AP 与 AR最终输出mAP_{iou}、mAR_{iou}及按类别细分的 AP/AR 表格。重要说明如导出 ScanNet 数据一节所述ScanNet 所有 GT 3D 包围盒均为轴对齐yaw 为 0因此网络预测的 yaw 目标也为 0后处理阶段采用与旋转无关的轴对齐 3D 非极大值抑制NMS。3D 语义分割mIoU语义分割任务通常使用平均交并比mIoU评估先计算多个类别的 IoU再取平均得到 mIoU具体实现见 mmdet3d/evaluation/functional/seg_eval.py。训练/验证时由SegMetric评测器驱动configs/base/datasets/scannet-seg.py同时可配合Seg3DTTAModel使用测试时增强。测试与提交在线 Benchmark默认情况下代码库在验证集上评估语义分割结果。若要在 ScanNet 在线 Benchmark 上测试模型性能需要在评测脚本中加入--format-only参数并在 configs/base/datasets/scannet-seg.py 中把ann_filedata_root scannet_infos_val.pkl改为ann_filedata_root scannet_infos_test.pkl同时指定txt_prefix为保存测试结果的目录。以 PointNetSSG在 ScanNet 上的推理为例仓库实际提供的配置为configs/pointnet2/pointnet2_ssg_2xb16-cosine-200e_scannet-seg.py另见configs/pointnet2/pointnet2_msg_2xb16-cosine-250e_scannet-seg.py在测试集上执行推理的命令如下./tools/dist_test.sh configs/pointnet2/pointnet2_ssg_2xb16-cosine-200e_scannet-seg.py \ work_dirs/pointnet2_ssg/latest.pth --format-only \ --eval-options txt_prefixwork_dirs/pointnet2_ssg/test_submission其中--format-only只保存预测结果而不在本地评测因为测试集没有公开标注--eval-options txt_prefix...指定结果输出目录预测掩码将以逐场景文件的形式写入该目录work_dirs/pointnet2_ssg/latest.pth训练得到的权重路径请按实际输出目录替换。生成结果后将结果文件夹压缩并上传至 ScanNet 官方评测服务器即可完成提交。若需要在本地验证集上评测去掉--format-only并保持ann_file指向scannet_infos_val.pkl即可。相关仓库文件索引数据集转换入口tools/dataset_converters/indoor_converter.pyScanNet info 与 seg info 生成实现tools/dataset_converters/scannet_data_utils.py检测任务配置基线configs/base/datasets/scannet-3d.py分割任务配置基线configs/base/datasets/scannet-seg.py点云变换GlobalAlignment / PointSample / IndoorPatchPointSample / GlobalRotScaleTransmmdet3d/datasets/transforms/transforms_3d.py加载与标签映射PointSegClassMapping / NormalizePointsColormmdet3d/datasets/transforms/loading.py检测评测实现mmdet3d/evaluation/functional/indoor_eval.py分割评测实现mmdet3d/evaluation/functional/seg_eval.py赞分享人工智能计算机视觉深度学习自动驾驶【免费下载链接】mmdetection3dOpenMMLabs next-generation platform for general 3D object detection.项目地址https://gitcode.com/gh_mirrors/mm/mmdetection3d点击查看免费下载相关推荐MMDetection3D 自定义数据集训练全流程指南MMDetection3D 自定义数据集训练全流程指南 引言 在3D目标检测领域使用公开数据集进行模型训练和测试是常见做法。然而在实际应用中我们经常需要处理人工智能计算机视觉深度学习自动驾驶一步粘完电子课本 PDF 下载tchMaterial-parser 实用指南一步粘完电子课本 PDF 下载tchMaterial parser 实用指南 tchMaterial parser 是面向教师和家长的电子课本 PDF 下载工人工智能计算机视觉深度学习自动驾驶mmdetection3d自定义数据集指南标注格式与训练流程mmdetection3d自定义数据集指南标注格式与训练流程 1. 引言3D检测中的数据挑战 在自动驾驶Autonomous Driving和机器人视觉人工智能计算机视觉深度学习自动驾驶上一篇dbt-jinja 实战在模板回调函数中使用 tokio block_on 同步等待异步操作下一篇Tamagui 条件式 RNGH Press 处理原生端按下事件的双路径架构解析创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表