Integrate new pretranspose_b_array with extra fused transpose of B

This patch fuses the transposition taking place in Acl with the transformations done in arm_gemm (called pretranspose_b_array) if the underlying kernel and transform supports it. This should improve start-up time (as it's for constant Rhs matrices) and memory footprint. The transformations in arm_gemm are kernel specific. The Rhs matrix is transformed into certain layouts to improve the performance.

Resolves: COMPMID-6595

Change-Id: Id2932dd966e59f903c279417bebcea83d9a42464
Signed-off-by: Gunes Bayir <gunes.bayir@arm.com>
Reviewed-on: https://review.mlplatform.org/c/ml/ComputeLibrary/+/11144
Tested-by: Arm Jenkins <bsgcomp@arm.com>
Reviewed-by: Viet-Hoa Do <viet-hoa.do@arm.com>
Comments-Addressed: Arm Jenkins <bsgcomp@arm.com>
Benchmark: Arm Jenkins <bsgcomp@arm.com>
diff --git a/filelist.json b/filelist.json
index dcf3204..d44a721 100644
--- a/filelist.json
+++ b/filelist.json
@@ -1592,6 +1592,7 @@
               "src/core/NEON/kernels/arm_gemm/gemm_quint8.cpp",
               "src/core/NEON/kernels/arm_gemm/gemm_uint16.cpp",
               "src/core/NEON/kernels/arm_gemm/gemm_uint8.cpp",
+              "src/core/NEON/kernels/arm_gemm/interleave-8way.cpp",
               "src/core/NEON/kernels/arm_gemm/interleave_indirect.cpp",
               "src/core/NEON/kernels/arm_gemm/mergeresults-fp16.cpp",
               "src/core/NEON/kernels/arm_gemm/mergeresults.cpp",