TY - GEN
T1 - Productivity, portability, performance: data-centric python
AU - Ziogas, Alexandros Nikolaos
AU - Schneider, Timo
AU - Nun, Tal Ben
AU - Calotoiu, Alexandru
AU - Matteis, Tiziano De
AU - Licht, Johannes De Fine
AU - Lavarini, Luca
AU - Hoefler, Torsten
PY - 2021
Y1 - 2021
N2 - Python has become the de facto language for scientific computing. Programming in Python is highly productive, mainly due to its rich science-oriented software ecosystem built around the NumPy module. As a result, the demand for Python support in High Performance Computing (HPC) has skyrocketed. However, the Python language itself does not necessarily offer high performance. In this work, we present a workflow that retains Pythons high productivity while achieving portable performance across different architectures the workflows key features are HPC-oriented language extensions and a set of automatic optimizations powered by a data-centric intermediate representation. We show performance results and scaling across CPU, GPU, FPGA, and the Piz Daint supercomputer (up to 23,328 cores), with 2.47x and 3.75x speedups over previous-best solutions, first-ever Xilinx and Intel FPGA results of annotated Python, and up to 93.16% scaling efficiency on 512 nodes.
AB - Python has become the de facto language for scientific computing. Programming in Python is highly productive, mainly due to its rich science-oriented software ecosystem built around the NumPy module. As a result, the demand for Python support in High Performance Computing (HPC) has skyrocketed. However, the Python language itself does not necessarily offer high performance. In this work, we present a workflow that retains Pythons high productivity while achieving portable performance across different architectures the workflows key features are HPC-oriented language extensions and a set of automatic optimizations powered by a data-centric intermediate representation. We show performance results and scaling across CPU, GPU, FPGA, and the Piz Daint supercomputer (up to 23,328 cores), with 2.47x and 3.75x speedups over previous-best solutions, first-ever Xilinx and Intel FPGA results of annotated Python, and up to 93.16% scaling efficiency on 512 nodes.
UR - https://www.scopus.com/pages/publications/85119391282
U2 - 10.1145/3458817.3476176
DO - 10.1145/3458817.3476176
M3 - Conference contribution
T3 - International Conference for High Performance Computing, Networking, Storage and Analysis, SC
SP - 1
EP - 13
BT - Proceedings of SC 2021
PB - IEEE Computer Society
T2 - 33rd International Conference for High Performance Computing, Networking, Storage and Analysis: Science and Beyond, SC 2021
Y2 - 14 November 2021 through 19 November 2021
ER -