Parallel GEMM-based convolution for deep learning on multicore RISC-V processors
Abstract:
We address the efficient implementation of the convolution operator on the GAP8 parallel ultra-low power platform (PULP), a heterogeneous multi-core processor equipped with a fabric controller (FC); a cluster of eight compute cores; and a four-level memory hierarchy with scratchpads instead of conventional, hardware-assisted cache memories. Our solution for this platform transforms the convolution into a general matrix–matrix multiplication (gemm) via the lowering approach, demonstrating that it is possible to attain reasonable performance on the GAP8 by carefully adapting techniques such as tiling and loop parallelism, which are mainstream in the multi-threaded, cache-aware realization of gemm.
Año de publicación:
2024
Keywords:
- Convolutional layers
- Deep learning
- Edge processors
- Performance Analysis
Fuente:
scopus
googleTipo de documento:
Article
Estado:
Acceso abierto
Áreas de conocimiento:
- Aprendizaje profundo
- Ciencias de la computación
- Ciencias de la computación
Áreas temáticas de Dewey:
- Métodos informáticos especiales
- Ciencias de la computación
- Física aplicada
Objetivos de Desarrollo Sostenible:
- ODS 9: Industria, innovación e infraestructura
- ODS 7: Energía asequible y no contaminante
- ODS 8: Trabajo decente y crecimiento económico