5.4.1
Fix comment
Add initial support for Intel AVX512F
Add TFloat data type for neural network
bugfix of FMA port to FAST_FLOAT
Add missing check for __ARM_NEON
reformat code (files with tabs)
Facilitate vectorization for generic build (#4223)
Implement DotProductSSE() for FAST_FLOAT
Remove unused variable assignments
Prepare using float instead of double for LSTM calculations
Detect availability of AVX512-VNNI