Zur Hauptnavigation wechseln Zur Suche wechseln Zum Hauptinhalt wechseln

Porting the MPI-parallelised les model PALM to multi-GPU systems and many integrated core processors - an experience report

  • Helge Knoop*
  • , Tobias Gronemeier
  • , Matthias Sühring
  • , Peter Steinbach
  • , Matthias Noack
  • , Florian Wende
  • , Thomas Steinke
  • , Christoph Knigge
  • , Siegfried Raasch
  • , Klaus Ketelsen
  • *Korrespondierende*r Autor*in für diese Arbeit

Publikation: Beitrag in FachzeitschriftArtikelForschungPeer-Review

Abstract

The computational power and availability of graphics processing units (GPUs) and many integrated core (MIC) processors on high performance computing (HPC) systems is rapidly evolving. However, HPC applications need to be ported to take advantage of such hardware. This paper is a report on our experience of porting the MPI+OpenMP parallelised large-eddy simulation model (PALM) to multi-GPU as well as to MIC processor environments using OpenACC and OpenMP. PALM is written in Fortran, entails 140 kLOC and runs on HPC farms of up to 43,200 cores. The main porting challenges are the size and complexity of PALM, its inconsistent modularisation and no unit-tests. We report the methods used to identify performance issues as well as our experiences with state-of-the-art profiling tools. Moreover, we outline the required porting steps, describe the problems and bottlenecks we encountered and present separate performance tests for both architectures. We however, do not provide benchmark information.

OriginalspracheEnglisch
Seiten (von - bis)297-309
Seitenumfang13
FachzeitschriftInternational Journal of Computational Science and Engineering
Jahrgang17
Ausgabenummer3
DOIs
PublikationsstatusVeröffentlicht - 27 Okt. 2018

ASJC Scopus Sachgebiete

  • Software
  • Modellierung und Simulation
  • Hardware und Architektur
  • Computational Mathematics
  • Theoretische Informatik und Mathematik

Dieses zitieren