PAPER

Near-Optimal Wafer-Scale Reduce

AI Systems and Hardware OLPP DPS
2024年04月24日
高性能计算(HPC)应用的关键基石是高效的Reduce和AllReduce通信集合。我们首次系统地研究了Cerebras Wafer-Scale Engine(WSE)上的Reduce和AllReduce。该架构已被证明在机器学习工作负载和FFT等其他计算问题上实现了前所未有的性能。我们引入了一个性能模型来估计WSE上算法的执行时间,并在广泛的输入大小范围内进行了实验验证我们的预测。除了现有的实现,我们还设计和实现了几个专门针对该架构的新算法。此外,我们为Reduce操作在WSE上的运行时间建立了下限。基于我们的模型,我们自动生成了代码,在整个输入大小范围内实现了接近最优的性能。实验表明,我们的新Reduce和AllReduce算法比当前供应商解决方案的性能提高了最多3.27倍。此外,我们的模型预测的性能误差不到4%。所提出的通信集合扩大了可以受益于WSE高吞吐量的HPC应用的范围。我们的模型驱动方法展示了一种有纪律的方法,可以引领在晶片级别的架构上进一步的算法进步。
Efficient Reduce and AllReduce communication collectives are a critical cornerstone of high-performance computing (HPC) applications. We present the first systematic investigation of Reduce and AllReduce on the Cerebras Wafer-Scale Engine (WSE). This architecture has been shown to achieve unprecedented performance both for machine learning workloads and other computational problems like FFT. We introduce a performance model to estimate the execution time of algorithms on the WSE and validate our predictions experimentally for a wide range of input sizes. In addition to existing implementations, we design and implement several new algorithms specifically tailored to the architecture. Moreover, we establish a lower bound for the runtime of a Reduce operation on the WSE. Based on our model, we automatically generate code that achieves near-optimal performance across the whole range of input sizes. Experiments demonstrate that our new Reduce and AllReduce algorithms outperform the current vendor solution by up to 3.27x. Additionally, our model predicts performance with less than 4% error. The proposed communication collectives increase the range of HPC applications that can benefit from the high throughput of the WSE. Our model-driven methodology demonstrates a disciplined approach that can lead the way to further algorithmic advancements on wafer-scale architectures.
许愿