Automatic Generation of the HPC Challenge’s Global FFT Benchmark for BlueGene/P Franz Franchetti Yevgen Voroneko Gheorghe Almasi 10.1184/R1/6468404.v1 https://kilthub.cmu.edu/articles/journal_contribution/Automatic_Generation_of_the_HPC_Challenge_s_Global_FFT_Benchmark_for_BlueGene_P/6468404 <p>We present the automatic synthesis of the HPC Challenge’s Global FFT, a large 1D FFT across a whole supercomputer system. We extend the Spiral system to synthesize specialized single-node FFT libraries that combine a data layout transformation with the actual on-node FFT computation to improve the network performance through enabling all-to-all collectives. We run our optimized Global FFT benchmark on up to 128k cores (32 racks) of ANL’s BlueGene/P “Intrepid” and achieved 6.4 Tflop/s, outperforming ANL’s 2008 HPC Challenge Class I Global FFT run (5 Tflop/s). Our code was part of IBM’s winning 2010 HPC Challenge Class II submission. Further, we show first single-thread results on BlueGene/Q.</p> 2012-07-01 00:00:00 Electrical & Computer Engineering