Performance analysis of convolution algorithms for deep learning on edge processors


Abstract:

We provide a complete performance comparison of two realizations of the convolution, based on the lowering approach and a blocked variant of the direct convolution algorithm. The theoretical analysis focuses on the conventional, high performance implementation of the general matrix multiplication (gemm), which is the key computational kernel underneath these two algorithms. The study leverages a simulator calibrated for the GAP8 edge processor and exploits the determinism of the memory system in this type of architectures to deliver accurate predictions of the arithmetic and data transfer costs.

Año de publicación:

2022

Keywords:

    Fuente:

    googlegoogle

    Tipo de documento:

    Article

    Estado:

    Acceso abierto

    Áreas de conocimiento:

    • Aprendizaje profundo
    • Algoritmo
    • Algoritmo

    Áreas temáticas de Dewey:

    • Física aplicada
    • Ciencias de la computación
    • Métodos informáticos especiales
    Procesado con IAProcesado con IA

    Objetivos de Desarrollo Sostenible:

    • ODS 9: Industria, innovación e infraestructura
    • ODS 17: Alianzas para lograr los objetivos
    • ODS 8: Trabajo decente y crecimiento económico
    Procesado con IAProcesado con IA

    Contribuidores: