about summary refs log tree commit diff
path: root/sysdeps/arm/armv7/multiarch/ifunc-impl-list.c
diff options
context:
space:
mode:
authorPrakhar Bahuguna <prakhar.bahuguna@arm.com>2017-06-27 15:43:50 +0000
committerJoseph Myers <joseph@codesourcery.com>2017-06-27 15:43:50 +0000
commitf8f72bc0c3da8ba039e6a1ed670ca576120b1f85 (patch)
tree83b3438aea7f6425cf94c5f97cbbca1d62797683 /sysdeps/arm/armv7/multiarch/ifunc-impl-list.c
parenta37b5daa6bc7fbcbbc229b2549a161fa15023f41 (diff)
downloadglibc-f8f72bc0c3da8ba039e6a1ed670ca576120b1f85.tar.gz
glibc-f8f72bc0c3da8ba039e6a1ed670ca576120b1f85.tar.xz
glibc-f8f72bc0c3da8ba039e6a1ed670ca576120b1f85.zip
[ARM] Optimise memchr for NEON-enabled processors
This patch provides an optimised implementation of memchr using NEON
instructions to improve its performance, especially with longer search regions.
This gave an improvement in performance against the Thumb2+DSP optimised code,
with more significant gains for larger inputs. The NEON code also wins in cases
where the input is small (less than 8 bytes) by defaulting to a simple
byte-by-byte search. This avoids the overhead imposed by filling two quadword
registers from memory.

	* sysdeps/arm/armv7/multiarch/Makefile: Add memchr_neon to
	sysdep_routines.
	* sysdeps/arm/armv7/multiarch/ifunc-impl-list.c: Add define for
	__memchr_neon.
	Add ifunc definitions for __memchr_neon and __memchr_noneon.
	* sysdeps/arm/armv7/multiarch/memchr.S: New file.
	* sysdeps/arm/armv7/multiarch/memchr_impl.S: Likewise.
	* sysdeps/arm/armv7/multiarch/memchr_neon.S: Likewise.

Testing done: Ran regression tests for arm-none-linux-gnueabihf as well as a
full toolchain bootstrap. Benchmark tests were ran on ARMv7-A and ARMv8-A
hardware targets.
Diffstat (limited to 'sysdeps/arm/armv7/multiarch/ifunc-impl-list.c')
-rw-r--r--sysdeps/arm/armv7/multiarch/ifunc-impl-list.c5
1 files changed, 5 insertions, 0 deletions
diff --git a/sysdeps/arm/armv7/multiarch/ifunc-impl-list.c b/sysdeps/arm/armv7/multiarch/ifunc-impl-list.c
index b8094fd393..8f33156317 100644
--- a/sysdeps/arm/armv7/multiarch/ifunc-impl-list.c
+++ b/sysdeps/arm/armv7/multiarch/ifunc-impl-list.c
@@ -34,6 +34,7 @@ __libc_ifunc_impl_list (const char *name, struct libc_ifunc_impl *array,
   bool use_neon = true;
 #ifdef __ARM_NEON__
 # define __memcpy_neon	memcpy
+# define __memchr_neon	memchr
 #else
   use_neon = (GLRO(dl_hwcap) & HWCAP_ARM_NEON) != 0;
 #endif
@@ -52,5 +53,9 @@ __libc_ifunc_impl_list (const char *name, struct libc_ifunc_impl *array,
 #endif
 	      IFUNC_IMPL_ADD (array, i, memcpy, 1, __memcpy_arm));
 
+  IFUNC_IMPL (i, name, memchr,
+	      IFUNC_IMPL_ADD (array, i, memchr, use_neon, __memchr_neon)
+	      IFUNC_IMPL_ADD (array, i, memchr, 1, __memchr_noneon));
+
   return i;
 }