arm: vp9itxfm: Optimize 16x16 and 32x32 idct dc by unrolling

Multimedia / Libav - Martin Storsjö [martin.st] - 10 February 2017 17:31 EST

This work is sponsored by, and copyright, Google.

Before: Cortex A7 A8 A9 A53
vp9_inv_dct_dct_16x16_sub1_add_neon: 273.0 189.5 211.7 235.8
vp9_inv_dct_dct_32x32_sub1_add_neon: 752.0 459.2 862.2 553.9 After:
vp9_inv_dct_dct_16x16_sub1_add_neon: 226.5 145.0 225.1 171.8
vp9_inv_dct_dct_32x32_sub1_add_neon: 721.2 415.7 727.6 475.0

a76bf8c arm: vp9itxfm: Optimize 16x16 and 32x32 idct dc by unrolling
libavcodec/arm/vp9itxfm_neon.S | 54 ++++++++++++++++++++++++++++--------------
1 file changed, 36 insertions(+), 18 deletions(-)

Upstream: git.libav.org


  • Share