
人工智能计算机视觉深度学习自动驾驶【免费下载链接】mmdetection3dOpenMMLabs next-generation platform for general 3D object detection.项目地址https://gitcode.com/gh_mirrors/mm/mmdetection3d点击查看免费下载PointRCNN 是首个直接从原始点云进行两阶段 3D 目标检测的经典算法CVPR 2019第一阶段以自底向上的方式生成高质量 3D proposal第二阶段在规范坐标系canonical coordinates中细化 proposal 得到最终检测结果全程仅使用点云、不依赖图像或体素化。本文以 MMDetection3D 仓库中 configs/point_rcnn/README.md 为核心结合 完整训练配置 与对应源码系统讲解 PointRCNN 的算法原理、MMDetection3D 中的模块实现、KITTI 数据集上的配置细节与训练调度并给出可直接复现的完整配置片段与性能参考。PointRCNN 算法核心两阶段纯点云检测框架PointRCNN 的完整框架由两个阶段组成对应 configs/point_rcnn/README.md 中的 Abstract 描述Stage-1自底向上的 3D proposal 生成。与之前从 RGB 图像生成 proposal、或把点云投影到鸟瞰图/体素的做法不同Stage-1 子网络直接对整场景点云做前景/背景分割以自底向上的方式生成少量高质量 3D proposal。分割本身把“哪些点属于物体”的任务前置使 proposal 生成不再依赖固定的先验 anchor。Stage-2规范坐标系中的 proposal 细化。将每个 proposal 内池化得到的点变换到规范坐标系学习更好的局部空间特征再与 Stage-1 学到的逐点全局语义特征融合用于精确的框回归和置信度预测得到最终检测结果。论文在 KITTI 3D 检测基准上的实验表明仅使用点云作为输入该架构相比当时的 SOTA 方法就有显著提升。MMDetection3D 在 configs/point_rcnn/README.md 的 Introduction 中说明仓库实现了 PointRCNN 并在 KITTI 数据集上提供了 checkpoint 与评测结果。MMDetection3D 中的实现从检测器到各子模块检测器主体PointRCNN仓库将整个网络封装为PointRCNN检测器类定义在 mmdet3d/models/detectors/point_rcnn.py继承自通用的两阶段 3D 检测器基类TwoStage3DDetector见 mmdet3d/models/detectors/two_stage.py。其extract_feat方法返回三个关键量供 RPN head 与 RoI head 共享# 来自 mmdet3d/models/detectors/point_rcnn.py 的 extract_feat def extract_feat(self, batch_inputs_dict): points torch.stack(batch_inputs_dict[points]) x self.backbone(points) if self.with_neck: x self.neck(x) return dict( fp_featuresx[fp_features].clone(), # 逐点特征backboneneck 输出 fp_pointsx[fp_xyz].clone(), # 逐点坐标 raw_pointspoints) # 原始点云可见 PointRCNN 的显著特点是全程保持逐点表示backbone 输出的不是体素特征图而是逐点特征fp_features与逐点坐标fp_points这为 Stage-1 的逐点前景分割和 Stage-2 的规范坐标变换提供了基础。完整的模型配置骨架PointRCNN 的模型定义位于 configs/base/models/point_rcnn.py由以下模块组成模块配置类型说明backbonePointNet2SAMSG点级 Set Abstraction MSG 骨干4 层下采样4096→1024→256→64neckPointNetFPNeck特征传播FPneck逐级上采样回传逐点特征rpn_headPointRPNHeadStage-1逐点前景分割 逐点框回归生成 3D proposalroi_headPointRCNNRoIHeadStage-2proposal 内点池化 规范坐标细化train_cfg / test_cfg——两阶段的 NMS、IoU 分配器、采样器与后处理参数其中PointRPNHead定义在 mmdet3d/models/dense_heads/point_rpn_head.pyPointRCNNRoIHead定义在 mmdet3d/models/roi_heads/point_rcnn_roi_head.py其下的细化头PointRCNNBboxHead定义在 mmdet3d/models/roi_heads/bbox_heads/point_rcnn_bbox_head.py。KITTI 数据集配置与数据流水线训练入口配置 point-rcnn_8xb2_kitti-3d-3class.py 继承了四个基础配置_base_ [ ../_base_/datasets/kitti-3d-car.py, # 数据集基础该文件实为 3 类 KITTI 配置 ../_base_/models/point_rcnn.py, # 上述模型骨架 ../_base_/default_runtime.py, # 运行时默认设置 ../_base_/schedules/cyclic-40e.py # 循环余弦调度基础 ]数据与点云范围dataset_type KittiDataset data_root data/kitti/ class_names [Pedestrian, Cyclist, Car] metainfo dict(classesclass_names) point_cloud_range [0, -40, -3, 70.4, 40, 1] # x/y/z 的 [min, max]单位为米 input_modality dict(use_lidarTrue, use_cameraFalse) # 纯 LiDAR 输入 backend_args Noneclass_names与基础数据集配置 kitti-3d-3class.py 一致为 KITTI 的 3 类检测任务行人、骑车人、汽车。point_cloud_range限定了检测范围前方 070.4 m、左右各 40 m、高度 -31 m。input_modality明确use_lidarTrue, use_cameraFalse对应算法“仅使用点云”的核心特性。DB 采样器db_samplerPointRCNN 训练采用基于数据库的目标采样增广GT-DB Sampling将kitti_dbinfos_train.pkl中的物体点云片段粘贴进场景db_sampler dict( data_rootdata_root, info_pathdata_root kitti_dbinfos_train.pkl, rate1.0, preparedict( filter_by_difficulty[-1], filter_by_min_pointsdict(Car5, Pedestrian5, Cyclist5)), sample_groupsdict(Car20, Pedestrian15, Cyclist15), classesclass_names, points_loaderdict( typeLoadPointsFromFile, coord_typeLIDAR, load_dim4, use_dim4, backend_argsbackend_args), backend_argsbackend_args)与基础配置相比PointRCNN 配置把各类别的最小采样点数门槛统一为 5基础配置中行人/骑车人为 10并提高了每个类别的采样组数量Car20、Pedestrian15、Cyclist15基础配置为 12/6/6即每个场景平均粘贴更多物体点云增强前景多样性。训练流水线train_pipeline [ dict(typeLoadPointsFromFile, coord_typeLIDAR, load_dim4, use_dim4, backend_argsbackend_args), # 加载点云x,y,z,intensity dict(typeLoadAnnotations3D, with_bbox_3dTrue, with_label_3dTrue), dict(typePointsRangeFilter, point_cloud_rangepoint_cloud_range), dict(typeObjectRangeFilter, point_cloud_rangepoint_cloud_range), dict(typeObjectSample, db_samplerdb_sampler), # 数据库采样增广 dict(typeRandomFlip3D, flip_ratio_bev_horizontal0.5), dict(typeObjectNoise, num_try100, translation_std[1.0, 1.0, 0.5], global_rot_range[0.0, 0.0], rot_range[-0.78539816, 0.78539816]), # 目标平移/旋转扰动 dict(typeGlobalRotScaleTrans, rot_range[-0.78539816, 0.78539816], scale_ratio_range[0.95, 1.05]), # 全局旋转与缩放 dict(typePointsRangeFilter, point_cloud_rangepoint_cloud_range), dict(typePointSample, num_points16384, sample_range40.0), dict(typePointShuffle), dict(typePack3DDetInputs, keys[points, gt_bboxes_3d, gt_labels_3d]) ]关键点说明固定采样点数PointSample将每个场景采样为 16384 个点采样范围 40 m保证输入点云尺寸一致便于 PointNet2 骨干的批量处理。范围过滤PointsRangeFilter与ObjectRangeFilter分别在点级和目标级裁剪出point_cloud_range内的数据。增广顺序目标噪声、全局旋转缩放之后再执行范围过滤避免增广把点移出有效范围。测试流水线使用MultiScaleFlipAug3D同样执行PointSample采样 16384 点保证推理与训练输入分布一致。模型配置逐模块解析以下关键片段全部来自 configs/base/models/point_rcnn.py可直接对照源码理解。BackbonePointNet2SAMSG PointNetFPNeckbackbonedict( typePointNet2SAMSG, in_channels4, # (x, y, z, intensity) num_points(4096, 1024, 256, 64), # 四级下采样点数 radii((0.1, 0.5), (0.5, 1.0), (1.0, 2.0), (2.0, 4.0)), num_samples((16, 32), (16, 32), (16, 32), (16, 32)), sa_channels(((16, 16, 32), (32, 32, 64)), ((64, 64, 128), (64, 96, 128)), ((128, 196, 256), (128, 196, 256)), ((256, 256, 512), (256, 384, 512))), fps_mods((D-FPS), (D-FPS), (D-FPS), (D-FPS)), fps_sample_range_lists((-1), (-1), (-1), (-1)), aggregation_channels(None, None, None, None), dilated_group(False, False, False, False), out_indices(0, 1, 2, 3), norm_cfgdict(typeBN2d, eps1e-3, momentum0.1), sa_cfgdict(typePointSAModuleMSG, pool_modmax, use_xyzTrue, normalize_xyzFalse)), neckdict( typePointNetFPNeck, fp_channels((1536, 512, 512), (768, 512, 512), (608, 256, 256), (257, 128, 128)))每层 SA 使用双半径 MSG 分支radii中两个半径对应sa_channels中两个分支通道并在第四层输出 512 维逐点特征。PointNetFPNeck通过四层特征传播将高层语义逐级回传最终在所有输入点上得到 128 维特征fp_features实现“逐点全局语义特征”。Stage-1 RPNPointRPNHeadrpn_headdict( typePointRPNHead, num_classes3, enlarge_width0.1, # 扩大 GT 框以圈定“忽略”点区域 pred_layer_cfgdict(in_channels128, cls_linear_channels(256, 256), reg_linear_channels(256, 256)), cls_lossdict(typemmdet.FocalLoss, use_sigmoidTrue, reductionsum, gamma2.0, alpha0.25, loss_weight1.0), bbox_lossdict(typemmdet.SmoothL1Loss, beta1.0 / 9.0, reductionsum, loss_weight1.0), bbox_coderdict( typePointXYZWHLRBBoxCoder, code_size8, # code_size: 中心残差(3) 尺寸回归(3) cos(yaw)(1) sin(yaw)(1) use_mean_sizeTrue, mean_size[[3.9, 1.6, 1.56], [0.8, 0.6, 1.73], [1.76, 0.6, 1.73]]))对照 mmdet3d/models/dense_heads/point_rpn_head.py 中的loss_by_feat网络在每个点上输出num_classes3维语义分类 logits前景类和code_size8维框回归量实现“逐点前景分割 逐点框回归”。正负样本分配get_targets_single通过points_in_boxes判断点是否落在 GT 框内落在框内为正样本落在enlarge_width0.1扩展框外为负样本扩展框内到原始框之间为忽略区域——这正是论文中“忽略邻近点”的设计。PointXYZWHLRBBoxCoder定义于 mmdet3d/models/task_modules/coders/point_xyzwhlr_bbox_coder.py回归目标以类别均值尺寸Car 3.9×1.6×1.56 m 等为锚中心偏移用对角线归一化尺寸用对数残差朝向用 cos/sin 双通道表示。语义损失使用 Focal Lossgamma2.0, alpha0.25缓解前景点与背景点的类别不平衡框回归使用 Smooth L1 Loss并通过正样本 mask 归一化权重。预测阶段predict_by_feat先取语义分数的最大值作为 objectness再经类无关旋转 NMSclass_agnostic_nms训练时iou_thr0.8输出最多 512 个 proposalproposal 的类别由其内点的语义 argmax 决定。Stage-2 RoI 细化PointRCNNRoIHead PointRCNNBboxHeadroi_headdict( typePointRCNNRoIHead, bbox_roi_extractordict( typeSingle3DRoIPointExtractor, roi_layerdict(typeRoIPointPool3d, num_sampled_points512)), bbox_headdict( typePointRCNNBboxHead, num_classes1, loss_bboxdict(typemmdet.SmoothL1Loss, beta1.0 / 9.0, reductionsum, loss_weight1.0), loss_clsdict(typemmdet.CrossEntropyLoss, use_sigmoidTrue, reductionsum, loss_weight1.0), pred_layer_cfgdict(in_channels512, cls_conv_channels(256, 256), reg_conv_channels(256, 256), biasTrue), in_channels5, # 3(xyz) 1(score) 1(depth) mlp_channels[128, 128], num_points(128, 32, -1), # 三级 SA最后一层全部点 radius(0.2, 0.4, 100), num_samples(16, 16, 16), sa_channels((128, 128, 128), (128, 128, 256), (256, 256, 512)), with_corner_lossTrue), depth_normalizer70.0)对照 mmdet3d/models/roi_heads/point_rcnn_roi_head.py 与 point_rcnn_bbox_head.pyStage-2 的流程为特征拼接将 Stage-1 的逐点语义分数point_scores、归一化深度fp_points.norm / 70.0 - 0.5与 backbone 的 128 维逐点特征拼接成 130 维输入对应in_channels5指前 5 维为 xyz、score、depth后续为 backbone 特征。RoI 池化Single3DRoIPointExtractor对每个 proposal 用RoIPointPool3d采样 512 个点。规范坐标变换_get_target_single中将 GT 框中心减去 RoI 中心、朝向角减去 RoI 的 ry再按-(roi_ry)旋转从而把 GT 变换到以 RoI 为原点的规范坐标系并处理朝向相反π 翻转的情况——对应论文中“canonical coordinates 中学习局部空间特征”的核心思想。三级 SA 聚合num_points(128, 32, -1)、radius(0.2, 0.4, 100)三级 set abstraction 逐级聚合局部特征到 512 维。损失类别分支用带cls_pos_thr0.7 / cls_neg_thr0.25的软标签中间 IoU 区间线性插值回归分支在规范坐标下用 Smooth L1此外启用with_corner_lossTrue用 8 个角点的 Huber 损失含朝向翻转的 min 处理见get_corner_loss_lidar约束框的几何形状。训练与测试配置train_cfg / test_cfgtrain_cfgdict( pos_distance_thr10.0, rpndict(rpn_proposaldict(use_rotate_nmsTrue, score_thrNone, iou_thr0.8, nms_pre9000, nms_post512)), rcnndict( assigner[ # Pedestrian / Cyclist / Car 每类一个 Max3DIoUAssigner dict(typeMax3DIoUAssigner, iou_calculatordict(typeBboxOverlaps3D, coordinatelidar), pos_iou_thr0.55, neg_iou_thr0.55, min_pos_iou0.55, ignore_iof_thr-1, match_low_qualityFalse), ...], samplerdict( typeIoUNegPiecewiseSampler, num128, pos_fraction0.5, neg_piece_fractions[0.8, 0.2], neg_iou_piece_thrs[0.55, 0.1], neg_pos_ub-1, add_gt_as_proposalsFalse, return_iouTrue), cls_pos_thr0.7, cls_neg_thr0.25)), test_cfgdict( rpndict(nms_cfgdict(use_rotate_nmsTrue, iou_thr0.85, nms_pre9000, nms_post512, score_thrNone)), rcnndict(use_rotate_nmsTrue, nms_thr0.1, score_thr0.1))Stage-1 训练时 NMS IoU 阈值 0.8、测试时提高到 0.85最多保留 512 个 proposal。Stage-2 采用每类独立分配器三类共用 IoU 阈值 0.55pos_iou_thrneg_iou_thrmin_pos_iou0.55即大于 0.55 为正、否则为负采样器IoUNegPiecewiseSampler每批采样 128 个 RoI、正样本比例 0.5。测试阶段 RoI head 用旋转 NMSscore_thr0.1、nms_thr0.1做最终去重。训练调度cyclic 40e 与 80 轮实际训练PointRCNN 配置继承 cyclic-40e.py 的循环余弦调度思想但做了针对性覆盖见 point-rcnn_8xb2_kitti-3d-3class.pylr 0.001 # 初始学习率注意PointRCNN 覆盖了基础配置的 0.0018 optim_wrapper dict(optimizerdict(lrlr, betas(0.95, 0.85))) train_cfg dict(by_epochTrue, max_epochs80, val_interval2) auto_scale_lr dict(enableFalse, base_batch_size16) # 8 GPU × 2 样本 param_scheduler [ dict(typeCosineAnnealingLR, T_max35, eta_minlr * 10, begin0, end35, by_epochTrue, convert_to_iter_basedTrue), dict(typeCosineAnnealingLR, T_max45, eta_minlr * 1e-4, begin35, end80, by_epochTrue, convert_to_iter_basedTrue), dict(typeCosineAnnealingMomentum, T_max35, eta_min0.85 / 0.95, begin0, end35, by_epochTrue, convert_to_iter_basedTrue), dict(typeCosineAnnealingMomentum, T_max45, eta_min1, begin35, end80, by_epochTrue, convert_to_iter_basedTrue), ]要点双段余弦前 35 轮学习率从 0 升至lr × 10 0.01预热升到峰值后 45 轮从峰值余弦衰减至lr × 1e-4动量调度与之同步从 0.85/0.95 升至 1。数据重复与轮数基础cyclic-40e定义max_epochs40而 PointRCNN 配置的train_dataloader使用RepeatDataset(times2)把 40 轮的数据重复为两遍因此max_epochs80README 表格中的 “cyclic 40e” 指基础调度实际训练为 80 轮。批量与学习率batch_size2、num_workers2单卡 2 样本8 卡共 16 样本base_batch_size16auto_scale_lr默认关闭。KITTI 上的结果与模型README 中提供的结果由 PointNet 骨干、cyclic 40e实际 80 轮调度在 KITTI 3 类任务上取得训练显存约 4.6 GBBackboneClassLr schdMem (GB)Inf time (fps)mAPPointNet配置见 point-rcnn_8xb2_kitti-3d-3class.py3 Classcyclic 40e4.6——70.83注表中 mAP 表示3 类在 moderate 难度设置下的 AP11 结果。模型权重与训练日志的下载信息记录在 configs/point_rcnn/metafile.yml 中。KITTI 3D 检测AP11 指标的逐类详细结果如下类别EasyModerateHardCar89.1378.7278.24Pedestrian65.8159.5752.75Cyclist93.5174.1970.73可以看出PointRCNN 在 KITTI 3D 检测上对三类目标均有较高精度汽车在 Easy 难度 AP 达 89.13骑车人在 Easy 难度 AP 达 93.51行人由于点云稀疏且形态多变精度相对较低moderate 59.57。复现与验证训练、测试与单元测试训练数据准备完成后KITTI 原始数据需先经 tools/create_data.py 转为kitti_infos_train.pkl、kitti_dbinfos_train.pkl等格式单卡训练命令为python tools/train.py configs/point_rcnn/point-rcnn_8xb2_kitti-3d-3class.py8 卡分布式训练可使用 tools/dist_train.sh./tools/dist_train.sh configs/point_rcnn/point-rcnn_8xb2_kitti-3d-3class.py 8测试与评测加载 checkpoint 在 KITTI 验证集上评测python tools/test.py configs/point_rcnn/point-rcnn_8xb2_kitti-3d-3class.py \ checkpoint路径 --eval kitti单元测试仓库在 tests/test_models/test_detectors/test_pointrcnn.py 中提供了TestPointRCNN单元测试它直接以本配置构建模型并验证前向预测结果包含bboxes_3d、scores_3d、labels_3d字段loss 前向产生rpn_bbox_loss、rpn_semantic_lossStage-1以及loss_cls、loss_bbox、loss_cornerStage-2五类损失且均不小于 0。运行方式与仓库其他测试一致python -m pytest tests/test_models/test_detectors/test_pointrcnn.py这从工程角度印证了 Stage-1 分割回归 Stage-2 规范坐标细化 角点损失这一完整链路在 MMDetection3D 框架下的正确性。小结PointRCNN 是两阶段纯点云 3D 检测的代表性方法MMDetection3D 的 point_rcnn 配置目录 提供了从数据流水线、双阶段模型、循环余弦调度到 KITTI 评测结果的完整闭环。本文以 configs/point_rcnn/README.md 为主线结合 模型配置、训练配置 及point_rpn_head.py、point_rcnn_roi_head.py、point_rcnn_bbox_head.py等源码梳理了逐点前景分割、规范坐标变换、点级 RoI 细化与角点损失等关键机制并给出可直接复现的配置、命令与性能基线可作为理解与二次开发两阶段点云检测算法的起点。引用若在研究中使用了 PointRCNN请引用原论文来自 configs/point_rcnn/README.md 的 Citationinproceedings{Shi_2019_CVPR, title {PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud}, author {Shi, Shaoshuai and Wang, Xiaogang and Li, Hongsheng}, booktitle {The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, month {June}, year {2019} }赞分享人工智能计算机视觉深度学习自动驾驶【免费下载链接】mmdetection3dOpenMMLabs next-generation platform for general 3D object detection.项目地址https://gitcode.com/gh_mirrors/mm/mmdetection3d点击查看免费下载相关推荐RetinaNet 目标检测实战指南Focal Loss 与 MMDetection 完整实现解析RetinaNet 目标检测实战指南Focal Loss 与 MMDetection 完整实现解析 本文以 configs/retinanet/README.人工智能计算机视觉深度学习自动驾驶Front-End-Checklist 无障碍指南为 ARIA command 元素提供可访问名称Accessible NamesFront End Checklist 无障碍指南为 ARIA command 元素提供可访问名称Accessible Names 本文基于 Front人工智能计算机视觉深度学习自动驾驶mmdetection3d中的SMOKE单阶段monocular 3D检测mmdetection3d中的SMOKE单阶段monocular 3D检测 引言单目3D检测的挑战与解决方案 在自动驾驶Autonomous Drivin人工智能计算机视觉深度学习自动驾驶上一篇CANN社区活动作品提交平台下一篇CANN ops-blas入门教程从零开始构建你的第一个GEMM应用创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考