ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

oneTBB 嵌套流图(Nested Flow Graphs)实战指南:在节点内部构建并执行子图

oneTBB 嵌套流图(Nested Flow Graphs)实战指南:在节点内部构建并执行子图 并发编程高性能计算【免费下载链接】oneTBBoneAPI Threading Building Blocks (oneTBB)项目地址https://gitcode.com/gh_mirrors/on/oneTBB点击查看免费下载导读oneAPI Threading Building BlocksoneTBB的 Flow Graph 支持在算法嵌套之外的另一层组合能力图的嵌套graph nesting。本文围绕 use_nested_flow_graphs.rst 的核心内容展开讲解如何在一个function_node的 body 中临时构造并执行一个内层依赖图dependence graph或数据流图data flow graph以及当内层图结构在多轮调用间保持不变时如何通过持久复用图来消除重复构造的开销。读完本文你将掌握嵌套流图的两种实现模式、wait_for_all()的正确使用时机以及嵌套图与task_arena、节点优先级组合使用的源码级细节。一、嵌套流图在节点内部再建一张图oneTBB Flow Graph 的组合能力体现在两个层次算法嵌套在一个节点内部调用parallel_for、parallel_reduce等并行算法图嵌套在一个节点的 body 内直接构造并运行另一张独立的graph。嵌套流图适用于把一个大的数据流问题拆分为多个内部子流程的场景例如外层图负责整体调度与扇出/扇入而每个节点内部运行一段独立的子流程子图。每个graph对象都拥有独立的task_group_context见 flow_graph.h 的graph::graph()构造实现其中my_context new ... task_group_context(FLOW_TASKS)因此内层图的取消与异常处理与外层图相互隔离这是嵌套能够成立的结构基础。1.1 在节点中构造并执行内层依赖图以下示例来自 use_nested_flow_graphs.rst外层图g有两个节点a和b节点a收到消息后在 body 内新建一张依赖图h四个continue_node风格的节点n1~n4先扇出后扇入节点b收到消息后则在 body 内新建一张数据流图graph g; function_node int, int a( g, unlimited, []( int i ) - int { graph h; node_t n1( h, { cout n1: i \n; } ); node_t n2( h, { cout n2: i \n; } ); node_t n3( h, { cout n3: i \n; } ); node_t n4( h, { cout n4: i \n; } ); make_edge( n1, n2 ); make_edge( n1, n3 ); make_edge( n2, n4 ); make_edge( n3, n4 ); n1.try_put(continue_msg()); h.wait_for_all(); return i; } ); function_node int, int b( g, unlimited, []( int i ) - int { graph h; function_node int, int m1( h, unlimited, []( int j ) - int { cout m1: j \n; return j; } ); function_node int, int m2( h, unlimited, []( int j ) - int { cout m2: j \n; return j; } ); function_node int, int m3( h, unlimited, []( int j ) - int { cout m3: j \n; return j; } ); function_node int, int m4( h, unlimited, []( int j ) - int { cout m4: j \n; return j; } ); make_edge( m1, m2 ); make_edge( m1, m3 ); make_edge( m2, m4 ); make_edge( m3, m4 ); m1.try_put(i); h.wait_for_all(); return i; } ); make_edge( a, b ); for ( int i 0; i 3; i ) { a.try_put(i); } g.wait_for_all();关键执行流程外层通过a.try_put(i)注入 3 个消息make_edge(a, b)使a的输出作为b的输入。节点a的每次调用都会新建依赖图h、连边、以n1.try_put(continue_msg())启动并调用h.wait_for_all()等待内层图空闲后返回。节点b的每次调用则新建一张由 4 个function_node int, int 构成的数据流图同样以m1.try_put(i)启动、h.wait_for_all()收尾。这里wait_for_all()是必须的因为h是节点 body 的局部变量body 返回时h的析构函数会先调用wait_for_all()再销毁内部节点与上下文见 flow_graph.h 中graph::~graph()的实现wait_for_all(); if (own_context) { ... }。若不等待就返回内层图任务会在析构过程中被强制等待语义上等价但会把等待延迟到析构阶段且无法保证h已经空闲后再复用因此显式调用wait_for_all()是更清晰、可控的写法。二、结构不变时的优化持久复用内层图如果内层图的结构在多次调用之间保持不变那么每次调用都重新构造一次图就是多余的——make_edge、节点注册、arena 重新 attach 都会带来额外开销。此时可以把内层图提升为外层节点可捕获的持久对象只在节点 body 中注入消息。2.1 复用依赖图结构的改进版将b的内层图提升到外层作用域b的 body 通过引用捕获并直接try_putgraph h; function_node int, int m1( h, unlimited, []( int j ) - int { cout m1: j \n; return j; } ); function_node int, int m2( h, unlimited, []( int j ) - int { cout m2: j \n; return j; } ); function_node int, int m3( h, unlimited, []( int j ) - int { cout m3: j \n; return j; } ); function_node int, int m4( h, unlimited, []( int j ) - int { cout m4: j \n; return j; } ); make_edge( m1, m2 ); make_edge( m1, m3 ); make_edge( m2, m4 ); make_edge( m3, m4 ); graph g; function_node int, int a( g, unlimited, []( int i ) - int { graph h; node_t n1( h, { cout n1: i \n; } ); node_t n2( h, { cout n2: i \n; } ); node_t n3( h, { cout n3: i \n; } ); node_t n4( h, { cout n4: i \n; } ); make_edge( n1, n2 ); make_edge( n1, n3 ); make_edge( n2, n4 ); make_edge( n3, n4 ); n1.try_put(continue_msg()); h.wait_for_all(); return i; } ); function_node int, int b( g, unlimited, - int { m1.try_put(i); h.wait_for_all(); // optional since h is not destroyed return i; } ); make_edge( a, b ); for ( int i 0; i 3; i ) { a.try_put(i); } g.wait_for_all();对比第一版注意两处差异b的 lambda 捕获方式由[]变为[]以便引用外层持久图h及其节点b不再在局部构造内层图只执行m1.try_put(i)即可复用已建好的图。2.2wait_for_all()何时可以省略原文档明确指出持久图场景下h.wait_for_all()是可选的。原因如下第一版中h是局部对象离开作用域即析构graph::~graph()内部会执行wait_for_all()所以不等待是不行的改进版中h是外层持久对象b的 body 返回后h并不会被销毁。因此若业务允许b不阻塞等待内层图完成可以直接m1.try_put(i); return i;让内层图在外层图的调度下继续运行。不过要注意权衡省略等待意味着b返回时内层消息可能尚未处理完若后续逻辑依赖内层结果则仍需调用h.wait_for_all()反之若希望最大化重叠执行overlap省略等待能让b的调用立即返回内层图与外层后续节点并行推进。从实现上看graph::wait_for_all()的核心是my_task_arena-execute(... d1::wait(my_wait_context_vertex.get_context(), *my_context))等待线程在阻塞期间会去窃取任务工作The waiting thread will go off and steal work while it is blocked in the wait_for_all见 _flow_graph_impl.h因此无论是否等待等待线程本身都不会空转。三、嵌套图的并发与优先级源码与测试中的佐证3.1 内层图与 task_arena每个graph构造时会prepare_task_arena()优先 attach 到当前活动的task_arena失败时新建一个默认初始化的 arena见 _flow_graph_impl.h。因此内层图默认运行在创建它的线程当前所在的 arena上。若需要让内层图运行在特定 arena可以在构造内层图前用task_arena::execute切换或在图生命周期跨越多个execute调用时通过graph::reset()重新 attach见 flow_graph.h 中prepare_task_arena(/*reinit*/true)的注释说明。3.2 测试用例NestedCase仓库测试 test_flow_graph_priorities.cpp 中的NestedCase直接验证了嵌套图的正确性其要点外层图outer_graph上挂 10 个function_nodeint,int彼此全连接任意两节点之间都建边每个外层节点 bodyOuterBody::operator()内部新建graph inner_graph挂 4 个continue_nodecontinue_msgstart_node、mid_node1、mid_node2、end_node并组成扇出扇入结构内层节点可以携带优先级mid_node1为node_priority_t(5)、end_node为node_priority_t(15)证明嵌套图与节点优先级机制兼容测试通过task_arena控制内外层图是否运行在同一个 arenasame_arena与different_arena两种配置见test_in_arena的INFO输出并遍历不同线程数验证每次运行前通过outer_graph.reset()、inner_graph.reset()重置图状态说明嵌套图可安全地反复启动reset 会依次重置 context、所有注册节点并重新 attach arena见 flow_graph.h。对应的测试注册在 test_flow_graph_priorities.cpp 的TEST_CASE(Nested test case)覆盖内层图在外层节点体内反复构造执行与内外层图分处不同 arena两类场景可作为嵌套流图实现的参考基准。四、实践要点与注意事项局部图必须等待若内层graph是节点 body 内的局部对象离开作用域时析构会强制执行wait_for_all()请在返回前显式调用h.wait_for_all()语义更清晰也避免析构阶段的隐式等待。持久图按需等待若内层图持久存在于外层作用域h.wait_for_all()可以省略以实现内外层图的异步重叠执行需要确定性结果时再等待。捕获方式持久复用要求节点 lambda 以引用捕获[]外层图对象临时构造则可用值捕获[]。上下文隔离每个graph拥有独立task_group_contextFLOW_TASKS类型嵌套图之间取消与异常状态互不干扰wait_for_all()返回前会同步cancelled/caught_exception状态。arena 归属内层图默认 attach 到当前线程所在 arena需要指定 arena 时结合task_arena::execute构造内层图或使用graph::reset()重新 attach。性能取舍仅当内层图结构在多次调用间不变时才值得持久复用若每次调用结构都不同临时构造是唯一选择此时应控制单次构造的节点规模避免构造开销淹没图本身的执行时间。五、延伸阅读图对象与wait_for_all/reserve_wait/release_wait的完整语义flow_graph.hgraph::reset与 arena 重 attach 的实现flow_graph.h嵌套图与节点优先级组合的测试test_flow_graph_priorities.cpp流图基础概念与图对象说明Flow_Graph.rst、Graph_Object.rst将嵌套流图与算法嵌套parallel_for等结合的组合模式Parallelizing_Flow_Graph.rst赞分享并发编程高性能计算【免费下载链接】oneTBBoneAPI Threading Building Blocks (oneTBB)项目地址https://gitcode.com/gh_mirrors/on/oneTBB点击查看免费下载相关推荐oneTBB 流图嵌套并行技巧在节点内部嵌套算法与嵌套流图oneTBB 流图嵌套并行技巧在节点内部嵌套算法与嵌套流图 导读 本文围绕 oneTBBoneAPI Threading Building BlocksF并发编程高性能计算Vue Flow 嵌套节点Nested Nodes完整指南parentNode、extent 与 expandParent 实战Vue Flow 嵌套节点Nested Nodes完整指南parentNode、extent 与 expandParent 实战 本文是 Vue Flow前端UI组件oneTBB 图并行编程指南用 Flow Graph 构建数据流图与依赖图oneTBB 图并行编程指南用 Flow Graph 构建数据流图与依赖图 本篇技术指南围绕 oneAPI Threading Building Blocks并发编程高性能计算上一篇Android抽屉动画效果终极指南MaterialDrawer转场动画实现详解下一篇NetSonar高级技巧如何自定义ping服务与导出网络性能报告创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表