Performance analysis of convolution algorithms for deep learning on edge processors
Abstract:
We provide a complete performance comparison of two realizations of the convolution, based on the lowering approach and a blocked variant of the direct convolution algorithm. The theoretical analysis focuses on the conventional, high performance implementation of the general matrix multiplication (gemm), which is the key computational kernel underneath these two algorithms. The study leverages a simulator calibrated for the GAP8 edge processor and exploits the determinism of the memory system in this type of architectures to deliver accurate predictions of the arithmetic and data transfer costs.
Año de publicación:
2022
Keywords:
Fuente:
googleTipo de documento:
Article
Estado:
Acceso abierto
Áreas de conocimiento:
- Aprendizaje profundo
- Algoritmo
- Algoritmo
Áreas temáticas de Dewey:
- Física aplicada
- Ciencias de la computación
- Métodos informáticos especiales
Objetivos de Desarrollo Sostenible:
- ODS 9: Industria, innovación e infraestructura
- ODS 17: Alianzas para lograr los objetivos
- ODS 8: Trabajo decente y crecimiento económico