TA的每日心情 | 开心 2020-4-8 10:45 |
|---|
签到天数: 227 天 [LV.7]分神
|
) p# J0 _8 m$ A9 i( [在论文里,这是第3.2.2节的内容+ y* P" W- e" j: F. D/ r& R
( T6 x5 o. ~# u1 n5 D' v7 I* Z1 ^3.2.2. Efficient Implementation of Cross-Node All-to-All Communication
5 m Z, D* u& r! L% b2 i) qIn order to ensure sufficient computational performance for DualPipe, we customize efficient
' j' `% Z7 Q$ y& [2 fcross-node all-to-all communication kernels (including dispatching and combining) to conserve0 \, p% y4 P; W. N! x$ E8 q5 C
the number of SMs dedicated to communication. The implementation of the kernels is codesigned with the MoE gating algorithm and the network topology of our cluster. To be specific,3 ]5 t- C( }6 `, y& L1 M& e
in our cluster, cross-node GPUs are fully interconnected with IB, and intra-node communications6 Q# b4 D) J j+ ]! C" }0 u" \8 H
are handled via NVLink. NVLink offers a bandwidth of 160 GB/s, roughly 3.2 times that of IB K l0 ]4 W4 Q: d
(50 GB/s). To effectively leverage the different bandwidths of IB and NVLink, we limit each( |: K2 s7 o. W Y( w% P
token to be dispatched to at most 4 nodes, thereby reducing IB traffic. For each token, when its
8 w) ~" s z: S+ o' q6 krouting decision is made, it will first be transmitted via IB to the GPUs with the same in-node
3 \$ G- S6 A$ i- F9 i- qindex on its target nodes. Once it reaches the target nodes, we will endeavor to ensure that it is
0 J- m) S/ R$ [! Cinstantaneously forwarded via NVLink to specific GPUs that host their target experts, without
: s% e: `4 k, U9 F9 bbeing blocked by subsequently arriving tokens. In this way, communications via IB and NVLink
0 A/ C0 x3 j- \ Y6 p3 r lare fully overlapped, and each token can efficiently select an average of 3.2 experts per node" [- c4 ] } \6 @# Z
without incurring additional overhead from NVLink. This implies that, although DeepSeek-V34 o, B# M! W- s! N* a2 }
13 |5 f/ ]7 `* V+ h; G- |& N
selects only 8 routed experts in practice, it can scale up this number to a maximum of 13 experts
1 Y; f& ]8 d) S* l" Y! I0 i, l(4 nodes × 3.2 experts/node) while preserving the same communication cost. Overall, under* u& p) C2 [) w% \. V% G/ g: u& \
such a communication strategy, only 20 SMs are sufficient to fully utilize the bandwidths of IB7 c1 w6 B) V/ a' t; T. T: x% A
and NVLink.+ O; F% w7 b! i3 K: R+ ^, O
In detail, we employ the warp specialization technique (Bauer et al., 2014) and partition
* I' y' W2 ~5 i$ i$ W2 C, |20 SMs into 10 communication channels. During the dispatching process, (1) IB sending, (2)0 H$ k5 }. P' B( H% S: }" Y% C+ ^
IB-to-NVLink forwarding, and (3) NVLink receiving are handled by respective warps. The& S$ U3 E/ \- Q: \# J% S6 E) H% s9 X
number of warps allocated to each communication task is dynamically adjusted according to the, s! |+ x' Q/ B& [& t
actual workload across all SMs. Similarly, during the combining process, (1) NVLink sending,
; N8 g) M+ y. F(2) NVLink-to-IB forwarding and accumulation, and (3) IB receiving and accumulation are also3 f8 l: u" R8 h B. z7 V/ W
handled by dynamically adjusted warps. In addition, both dispatching and combining kernels7 c' b# t) p4 U) _6 H! \, ~
overlap with the computation stream, so we also consider their impact on other SM computation y' S% m# h& \1 A1 z: y
kernels. Specifically, we employ customized PTX (Parallel Thread Execution) instructions and
/ x5 [+ C) i" F7 C, F2 r, m6 v1 ?auto-tune the communication chunk size, which significantly reduces the use of the L2 cache
% i! ]" s) \ Z/ U$ d) Kand the interference to other SMs.
4 v6 \: z2 W9 T7 d3 ~( w! l9 J6 d: p
通俗一点说,就是为了实现高效的跨节点全面通信。解决的问题本质上和唐家山老师日志里说的双机对拷的场景差不多。一般来说单机多卡之间用nvlink,多机多卡之间依赖IB网络,但nvlink的速率是IB网络的速率的3.2倍,需要通过一些优化来实现更好的传输策略。这是一整套方案。
& F% O! u8 c0 ~7 u2 y' L& p% q6 j" F. T( P- q& m. e! m
我的理解,使用PTX在其中,是为了更精准的定制线程执行减少通信块分配传输之间的串扰。
- Q5 E7 G) b4 e; J% J9 j7 t. G
" ?; c) I7 K2 c* d. m: E' Y/ ]; v目的不是为了绕cuda,反而是为了让cuda的效率更高。
0 J5 Q5 C; D9 a8 |$ [
5 P3 X3 h. R5 O" b3 Z% t3 S& k类比一下,就好比发现网卡驱动在对拷特定内存块的时候会和应用的线程执行出现串行导致效率降低,而绕开操作系统定义的与网卡驱动的接口,直接使用网卡支持的指令集进行了优化。 |
评分
-
查看全部评分
|