Large-Scale Distributed Training Applied to Generative Adversarial Networks for Calorimeter Simulation

Vlimant Jean-Roch; Pantaleo Felice; Pierini Maurizio; Loncar Vladimir; Vallecorsa Sofia; Anderson Dustin; Nguyen Thong; Zlokapa Alexander

doi:10.1051/epjconf/201921406025

EPJ Web of Conferences (Jan 2019)

Large-Scale Distributed Training Applied to Generative Adversarial Networks for Calorimeter Simulation

Vlimant Jean-Roch,
Pantaleo Felice,
Pierini Maurizio,
Loncar Vladimir,
Vallecorsa Sofia,
Anderson Dustin,
Nguyen Thong,
Zlokapa Alexander

Affiliations

Vlimant Jean-Roch
Pantaleo Felice
Pierini Maurizio
Loncar Vladimir
Vallecorsa Sofia
Anderson Dustin
Nguyen Thong
Zlokapa Alexander

DOI: https://doi.org/10.1051/epjconf/201921406025
Journal volume & issue: Vol. 214
p. 06025

Abstract

Read online

In recent years, several studies have demonstrated the benefit of using deep learning to solve typical tasks related to high energy physics data taking and analysis. In particular, generative adversarial networks are a good candidate to supplement the simulation of the detector response in a collider environment. Training of neural network models has been made tractable with the improvement of optimization methods and the advent of GP-GPU well adapted to tackle the highly-parallelizable task of training neural nets. Despite these advancements, training of large models over large data sets can take days to weeks. Even more so, finding the best model architecture and settings can take many expensive trials. To get the best out of this new technology, it is important to scale up the available network-training resources and, consequently, to provide tools for optimal large-scale distributed training. In this context, our development of a new training workflow, which scales on multi-node/multi-GPU architectures with an eye to deployment on high performance computing machines is described. We describe the integration of hyper parameter optimization with a distributed training framework using Message Passing Interface, for models defined in keras [12] or pytorch [13]. We present results on the speedup of training generative adversarial networks trained on a data set composed of the energy deposition from electron, photons, charged and neutral hadrons in a fine grained digital calorimeter.

Published in EPJ Web of Conferences

ISSN: 2100-014X (Online)
Publisher: EDP Sciences
Country of publisher: France
LCC subjects: Science: Physics
Website: http://www.epj-conferences.org/

About the journal