Distributed Learning of CNNs on Heterogeneous CPU/GPU Architectures

Marques, José; Falcao, Gabriel; Alexandre, Luís

Publication

Distributed Learning of CNNs on Heterogeneous CPU/GPU Architectures

2018Journal article

dc.contributor.author	Marques, José
dc.contributor.author	Falcao, Gabriel
dc.contributor.author	Alexandre, Luís
dc.date.accessioned	2020-01-09T09:45:41Z
dc.date.available	2020-01-09T09:45:41Z
dc.date.issued	2018
dc.description.abstract	Convolutional Neural Networks (CNNs) have shown to be powerful classi cation tools in tasks that range from check reading to medical diagnosis, reaching close to human perception, and in some cases surpassing it. However, the problems to solve are becoming larger and more complex, which translates to larger CNNs, leading to longer training times\|the computational complex part\|that not even the adoption of Graphics Processing Units (GPUs) could keep up to. This problem is partially solved by using more processing units and distributed training methods that are o ered by several frameworks dedicated to neural network training, such as Ca e, Torch or TensorFlow. However, these techniques do not take full advantage of the possible parallelization o ered by CNNs and the cooperative use of heterogeneous devices with di erent processing capabilities, clock speeds, memory size, among others. This paper presents a new method for the parallel training of CNNs that can be considered as a particular instantiation of model parallelism, where only the convolutional layer is distributed. In fact, the convolutions processed during training (forward and backward propagation included) represent from 60-90% of global processing time. The paper analyzes the in uence of network size, bandwidth, batch size, number of devices, including their processing capabilities, and other parameters. Results show that this technique is capable of diminishing the training time without a ecting the classi cation performance for both CPUs and GPUs. For the CIFAR-10 dataset, using a CNN with two convolutional layers, and 500 and 1500 kernels, respectively, best speedups achieve 3:28 using four CPUs and 2:45 with three GPUs. Modern imaging datasets, larger and more complex than CIFAR-10 will certainly require more than 60-90% of processing time calculating convolutions, and speedups will tend to increase accordingly.	pt_PT
dc.description.version	info:eu-repo/semantics/publishedVersion	pt_PT
dc.identifier.doi	10.1080/08839514.2018.1508814	pt_PT
dc.identifier.uri	http://hdl.handle.net/10400.6/8141
dc.language.iso	eng	pt_PT
dc.peerreviewed	no	pt_PT
dc.title	Distributed Learning of CNNs on Heterogeneous CPU/GPU Architectures	pt_PT
dc.type	journal article
dspace.entity.type	Publication
oaire.citation.endPage	844	pt_PT
oaire.citation.issue	9-10	pt_PT
oaire.citation.startPage	822	pt_PT
oaire.citation.title	Applied Artificial Intelligence	pt_PT
oaire.citation.volume	32	pt_PT
person.familyName	Falcao
person.familyName	Alexandre
person.givenName	Gabriel
person.givenName	Luís
person.identifier	1483922
person.identifier.ciencia-id	251F-BD6A-8DF9
person.identifier.ciencia-id	2014-0F06-A3E3
person.identifier.orcid	0000-0001-9805-6747
person.identifier.orcid	0000-0002-5133-5025
person.identifier.rid	P-9142-2014
person.identifier.rid	E-8770-2013
person.identifier.scopus-author-id	17433774200
person.identifier.scopus-author-id	8847713100
rcaap.rights	openAccess	pt_PT
rcaap.type	article	pt_PT
relation.isAuthorOfPublication	f9be499e-6059-41dc-983e-5fe9022ea0db
relation.isAuthorOfPublication	131ec6eb-b61a-4f27-953f-12e948a43a96
relation.isAuthorOfPublication.latestForDiscovery	131ec6eb-b61a-4f27-953f-12e948a43a96

Files

Original bundle

Now showing 1 - 1 of 1

Name:: 1712.02546(1).pdf
Size:: 826.73 KB
Format:: Adobe Portable Document Format

Download

License bundle

Now showing 1 - 1 of 1

Name:: license.txt
Size:: 1.71 KB
Format:: Item-specific license agreed upon to submission
Description:

Download

Collections

FE - DI | Documentos por Auto-Depósito
ICI - IT-UBI | Documentos por Auto-Depósito