Parallel GEMM-based convolution for deep learning on multicore RISC-V processors


Abstract:

We address the efficient implementation of the convolution operator on the GAP8 parallel ultra-low power platform (PULP), a heterogeneous multi-core processor equipped with a fabric controller (FC); a cluster of eight compute cores; and a four-level memory hierarchy with scratchpads instead of conventional, hardware-assisted cache memories. Our solution for this platform transforms the convolution into a general matrix–matrix multiplication (gemm) via the lowering approach, demonstrating that it is possible to attain reasonable performance on the GAP8 by carefully adapting techniques such as tiling and loop parallelism, which are mainstream in the multi-threaded, cache-aware realization of gemm.

Año de publicación:

2024

Keywords:

  • Convolutional layers
  • Deep learning
  • Edge processors
  • Performance Analysis

Fuente:

scopusscopus
googlegoogle

Tipo de documento:

Article

Estado:

Acceso abierto

Áreas de conocimiento:

  • Aprendizaje profundo
  • Ciencias de la computación
  • Ciencias de la computación

Áreas temáticas de Dewey:

  • Métodos informáticos especiales
  • Ciencias de la computación
  • Física aplicada
Procesado con IAProcesado con IA

Objetivos de Desarrollo Sostenible:

  • ODS 9: Industria, innovación e infraestructura
  • ODS 7: Energía asequible y no contaminante
  • ODS 8: Trabajo decente y crecimiento económico
Procesado con IAProcesado con IA