Universal intrinsics#

Topics#

Detailed Description#

“Universal intrinsics” is a types and functions set intended to simplify vectorization of code on different platforms. Currently a few different SIMD extensions on different architectures are supported.

OpenCV Universal Intrinsics support the following instruction sets:

  • 128 bit registers of various types support is implemented for a wide range of architectures including

    • x86(SSE/SSE2/SSE4.2),

    • ARM(NEON): 64-bit float (64F) requires AArch64,

    • PowerPC(VSX),

    • MIPS(MSA),

    • LoongArch(LSX),

    • RISC-V(RVV 0.7.1): Fixed-length implementation,

    • WASM: 64-bit float (64F) is not supported,

  • 256 bit registers are supported on

    • x86(AVX2),

    • LoongArch (LASX),

  • 512 bit registers are supported on

    • x86(AVX512),

  • Vector Length Agnostic (VLA) registers are supported on

    • RISC-V(RVV 1.0)

    • ARM(SVE/SVE2): Powered by Arm KleidiCV integration (OpenCV 4.11+),

In case when there is no SIMD extension available during compilation, fallback C++ implementation of intrinsics will be chosen and code will work as expected although it could be slower.

Types#

There are several types representing packed values vector registers, each type is implemented as a structure based on a one SIMD register.

Exact bit length(and value quantity) of listed types is compile time deduced and depends on architecture SIMD capabilities chosen as available during compilation of the library. All the types contains nlanes enumeration to check for exact value quantity of the type.

In case the exact bit length of the type is important it is possible to use specific fixed length register types.

There are several types representing 128-bit registers.

There are several types representing 256-bit registers.

Note

256 bit registers at the moment implemented for AVX2 SIMD extension only, if you want to use this type directly, don’t forget to check the CV_SIMD256 preprocessor definition:

#if CV_SIMD256
//...
#endif

There are several types representing 512-bit registers.

Load and store operations#

These operations allow to set contents of the register explicitly or by loading it from some memory block and to save contents of the register to memory block.

There are variable size register load operations that provide result of maximum available size depending on chosen platform capabilities.

  • Constructors: from memory,

  • Other create methods: vx_setall_s8, vx_setall_u8, …, vx_setzero_u8, vx_setzero_s8, …

  • Memory load operations: vx_load, vx_load_aligned, vx_load_low, vx_load_halves,

  • Memory operations with expansion of values: vx_load_expand, vx_load_expand_q

Also there are fixed size register load/store operations.

For 128 bit registers

For 256 bit registers(check CV_SIMD256 preprocessor definition)

For 512 bit registers(check CV_SIMD512 preprocessor definition)

Store to memory operations are similar across different platform capabilities: v_store, v_store_aligned, v_store_high, v_store_low

Value reordering#

These operations allow to reorder or recombine elements in one or multiple vectors.

Arithmetic, bitwise and comparison operations#

Element-wise binary and unary operations.

Reduce and mask#

Most of these operations return only one value.

Other math#

Conversions#

Different type conversions and casts:

Matrix operations#

In these operations vectors represent matrix rows/columns: v_dotprod, v_dotprod_fast, v_dotprod_expand, v_dotprod_expand_fast, v_matmul, v_transpose4x4

Usability#

Most operations are implemented only for some subset of the available types, following matrices shows the applicability of different operations to the types.

Regular integers:

Operations\Types

uint 8

int 8

uint 16

int 16

uint 32

int 32

load, store

x

x

x

x

x

x

interleave

x

x

x

x

x

x

expand

x

x

x

x

x

x

expand_low

x

x

x

x

x

x

expand_high

x

x

x

x

x

x

expand_q

x

x

add, sub

x

x

x

x

x

x

add_wrap, sub_wrap

x

x

x

x

mul_wrap

x

x

x

x

mul

x

x

x

x

x

x

mul_expand

x

x

x

x

x

compare

x

x

x

x

x

x

shift

x

x

x

x

dotprod

x

x

dotprod_fast

x

x

dotprod_expand

x

x

x

x

x

dotprod_expand_fast

x

x

x

x

x

logical

x

x

x

x

x

x

min, max

x

x

x

x

x

x

absdiff

x

x

x

x

x

x

absdiffs

x

x

reduce

x

x

x

x

x

x

mask

x

x

x

x

x

x

pack

x

x

x

x

x

x

pack_u

x

x

pack_b

x

unpack

x

x

x

x

x

x

extract

x

x

x

x

x

x

rotate (lanes)

x

x

x

x

x

x

cvt_flt32

x

cvt_flt64

x

transpose4x4

x

x

reverse

x

x

x

x

x

x

extract_n

x

x

x

x

x

x

broadcast_element

x

x

Big integers:

Operations\Types

uint 64

int 64

load, store

x

x

add, sub

x

x

shift

x

x

logical

x

x

reverse

x

x

extract

x

x

rotate (lanes)

x

x

cvt_flt64

x

extract_n

x

x

Floating point:

Operations\Types

float 32

float 64

load, store

x

x

interleave

x

add, sub

x

x

mul

x

x

div

x

x

compare

x

x

min, max

x

x

absdiff

x

x

reduce

x

mask

x

x

unpack

x

x

cvt_flt32

x

cvt_flt64

x

sqrt, abs

x

x

float math

x

x

transpose4x4

x

extract

x

x

rotate (lanes)

x

x

reverse

x

x

extract_n

x

x

broadcast_element

x

exp

x

x

log

x

x

sin, cos

x

x

Classes#

Name

Description

struct cv::v_reg

View details

Enumerations#

enum cv {
    simd128_width = 16,
    simd256_width = 32,
    simd512_width = 64,
    simdmax_width = simd512_width
}

View details

Pack boolean values#

Return

Name

Description

template<int n>
v_reg< uchar, 2 *n >

cv::v_pack_b(
const v_reg< ushort, n > & a,
const v_reg< ushort, n > & b )

! For 16-bit boolean values

template<int n>
v_reg< uchar, 4 *n >

cv::v_pack_b(
const v_reg< unsigned, n > & a,
const v_reg< unsigned, n > & b,
const v_reg< unsigned, n > & c,
const v_reg< unsigned, n > & d )

template<int n>
v_reg< uchar, 8 *n >

cv::v_pack_b(
const v_reg< uint64, n > & a,
const v_reg< uint64, n > & b,
const v_reg< uint64, n > & c,
const v_reg< uint64, n > & d,
const v_reg< uint64, n > & e,
const v_reg< uint64, n > & f,
const v_reg< uint64, n > & g,
const v_reg< uint64, n > & h )

Typedef Documentation#

v_float32x16#

typedef v_reg< float, 16 > cv::v_float32x16

#include <opencv2/core/hal/intrin_cpp.hpp>

Sixteen 32-bit floating point values (single precision)

v_float32x4#

typedef v_reg< float, 4 > cv::v_float32x4

#include <opencv2/core/hal/intrin_cpp.hpp>

Four 32-bit floating point values (single precision)

v_float32x8#

typedef v_reg< float, 8 > cv::v_float32x8

#include <opencv2/core/hal/intrin_cpp.hpp>

Eight 32-bit floating point values (single precision)

v_float64x2#

typedef v_reg< double, 2 > cv::v_float64x2

#include <opencv2/core/hal/intrin_cpp.hpp>

Two 64-bit floating point values (double precision)

v_float64x4#

typedef v_reg< double, 4 > cv::v_float64x4

#include <opencv2/core/hal/intrin_cpp.hpp>

Four 64-bit floating point values (double precision)

v_float64x8#

typedef v_reg< double, 8 > cv::v_float64x8

#include <opencv2/core/hal/intrin_cpp.hpp>

Eight 64-bit floating point values (double precision)

v_int16x16#

typedef v_reg< short, 16 > cv::v_int16x16

#include <opencv2/core/hal/intrin_cpp.hpp>

Sixteen 16-bit signed integer values.

v_int16x32#

typedef v_reg< short, 32 > cv::v_int16x32

#include <opencv2/core/hal/intrin_cpp.hpp>

Thirty two 16-bit signed integer values.

v_int16x8#

typedef v_reg< short, 8 > cv::v_int16x8

#include <opencv2/core/hal/intrin_cpp.hpp>

Eight 16-bit signed integer values.

v_int32x16#

typedef v_reg< int, 16 > cv::v_int32x16

#include <opencv2/core/hal/intrin_cpp.hpp>

Sixteen 32-bit signed integer values.

v_int32x4#

typedef v_reg< int, 4 > cv::v_int32x4

#include <opencv2/core/hal/intrin_cpp.hpp>

Four 32-bit signed integer values.

v_int32x8#

typedef v_reg< int, 8 > cv::v_int32x8

#include <opencv2/core/hal/intrin_cpp.hpp>

Eight 32-bit signed integer values.

v_int64x2#

typedef v_reg< int64, 2 > cv::v_int64x2

#include <opencv2/core/hal/intrin_cpp.hpp>

Two 64-bit signed integer values.

v_int64x4#

typedef v_reg< int64, 4 > cv::v_int64x4

#include <opencv2/core/hal/intrin_cpp.hpp>

Four 64-bit signed integer values.

v_int64x8#

typedef v_reg< int64, 8 > cv::v_int64x8

#include <opencv2/core/hal/intrin_cpp.hpp>

Eight 64-bit signed integer values.

v_int8x16#

typedef v_reg< schar, 16 > cv::v_int8x16

#include <opencv2/core/hal/intrin_cpp.hpp>

Sixteen 8-bit signed integer values.

v_int8x32#

typedef v_reg< schar, 32 > cv::v_int8x32

#include <opencv2/core/hal/intrin_cpp.hpp>

Thirty two 8-bit signed integer values.

v_int8x64#

typedef v_reg< schar, 64 > cv::v_int8x64

#include <opencv2/core/hal/intrin_cpp.hpp>

Sixty four 8-bit signed integer values.

v_uint16x16#

typedef v_reg< ushort, 16 > cv::v_uint16x16

#include <opencv2/core/hal/intrin_cpp.hpp>

Sixteen 16-bit unsigned integer values.

v_uint16x32#

typedef v_reg< ushort, 32 > cv::v_uint16x32

#include <opencv2/core/hal/intrin_cpp.hpp>

Thirty two 16-bit unsigned integer values.

v_uint16x8#

typedef v_reg< ushort, 8 > cv::v_uint16x8

#include <opencv2/core/hal/intrin_cpp.hpp>

Eight 16-bit unsigned integer values.

v_uint32x16#

typedef v_reg< unsigned, 16 > cv::v_uint32x16

#include <opencv2/core/hal/intrin_cpp.hpp>

Sixteen 32-bit unsigned integer values.

v_uint32x4#

typedef v_reg< unsigned, 4 > cv::v_uint32x4

#include <opencv2/core/hal/intrin_cpp.hpp>

Four 32-bit unsigned integer values.

v_uint32x8#

typedef v_reg< unsigned, 8 > cv::v_uint32x8

#include <opencv2/core/hal/intrin_cpp.hpp>

Eight 32-bit unsigned integer values.

v_uint64x2#

typedef v_reg< uint64, 2 > cv::v_uint64x2

#include <opencv2/core/hal/intrin_cpp.hpp>

Two 64-bit unsigned integer values.

v_uint64x4#

typedef v_reg< uint64, 4 > cv::v_uint64x4

#include <opencv2/core/hal/intrin_cpp.hpp>

Four 64-bit unsigned integer values.

v_uint64x8#

typedef v_reg< uint64, 8 > cv::v_uint64x8

#include <opencv2/core/hal/intrin_cpp.hpp>

Eight 64-bit unsigned integer values.

v_uint8x16#

typedef v_reg< uchar, 16 > cv::v_uint8x16

#include <opencv2/core/hal/intrin_cpp.hpp>

Sixteen 8-bit unsigned integer values.

v_uint8x32#

typedef v_reg< uchar, 32 > cv::v_uint8x32

#include <opencv2/core/hal/intrin_cpp.hpp>

Thirty two 8-bit unsigned integer values.

v_uint8x64#

typedef v_reg< uchar, 64 > cv::v_uint8x64

#include <opencv2/core/hal/intrin_cpp.hpp>

Sixty four 8-bit unsigned integer values.

Enumeration Type Documentation#

enum#

#include <opencv2/core/hal/intrin_cpp.hpp>

Enumerator:

simd128_width

simd256_width

simd512_width

simdmax_width

Function Documentation#

v_pack_b() [1/3]#

template<int n>
inline v_reg< uchar, 2 *n > cv::v_pack_b(
const v_reg< ushort, n > & a,
const v_reg< ushort, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

! For 16-bit boolean values

Scheme:

a  {0xFFFF 0 0 0xFFFF 0 0xFFFF 0xFFFF 0}
## b  {0xFFFF 0 0xFFFF 0 0 0xFFFF 0 0xFFFF}
{
   0xFF 0 0 0xFF 0 0xFF 0xFF 0
   0xFF 0 0xFF 0 0 0xFF 0 0xFF
}

v_pack_b() [2/3]#

template<int n>
inline v_reg< uchar, 4 *n > cv::v_pack_b(
const v_reg< unsigned, n > & a,
const v_reg< unsigned, n > & b,
const v_reg< unsigned, n > & c,
const v_reg< unsigned, n > & d )

#include <opencv2/core/hal/intrin_cpp.hpp>

This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts. For 32-bit boolean values

Scheme:

a  {0xFFFF.. 0 0 0xFFFF..}
b  {0 0xFFFF.. 0xFFFF.. 0}
c  {0xFFFF.. 0 0xFFFF.. 0}
## d  {0 0xFFFF.. 0 0xFFFF..}
{
   0xFF 0 0 0xFF 0 0xFF 0xFF 0
   0xFF 0 0xFF 0 0 0xFF 0 0xFF
}

v_pack_b() [3/3]#

template<int n>
inline v_reg< uchar, 8 *n > cv::v_pack_b(
const v_reg< uint64, n > & a,
const v_reg< uint64, n > & b,
const v_reg< uint64, n > & c,
const v_reg< uint64, n > & d,
const v_reg< uint64, n > & e,
const v_reg< uint64, n > & f,
const v_reg< uint64, n > & g,
const v_reg< uint64, n > & h )

#include <opencv2/core/hal/intrin_cpp.hpp>

This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts. For 64-bit boolean values

Scheme:

a  {0xFFFF.. 0}
b  {0 0xFFFF..}
c  {0xFFFF.. 0}
d  {0 0xFFFF..}

e  {0xFFFF.. 0}
f  {0xFFFF.. 0}
g  {0 0xFFFF..}
## h  {0 0xFFFF..}
{
   0xFF 0 0 0xFF 0xFF 0 0 0xFF
   0xFF 0 0xFF 0 0 0xFF 0 0xFF
}

v256_cleanup()#

inline void cv::v256_cleanup()

#include <opencv2/core/hal/intrin_cpp.hpp>

v256_load()#

template<typename _Tp>
inline v_reg< _Tp, simd256_width/sizeof(_Tp)> cv::v256_load(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load 256-bit length register contents from memory.

Note

Returned type will be detected from passed pointer type, for example uchar ==> cv::v_uint8x32, int ==> cv::v_int32x8, etc.

Check CV_SIMD256 preprocessor definition prior to use. Use vx_load version to get maximum available register length result

Alignment requirement: if CV_STRONG_ALIGNMENT=1 then passed pointer must be aligned (sizeof(lane type) should be enough). Do not cast pointer types without runtime check for pointer alignment (like uchar* => int*).

Parameters

  • ptr — pointer to memory block with data

Returns — register object

Here is the call graph for this function:

cv::v256_load Node1 cv::v256_load Node2 cv::isAligned Node1->Node2

cv::v256_load Node1 cv::v256_load Node2 cv::isAligned Node1->Node2

v256_load_aligned()#

template<typename _Tp>
inline v_reg< _Tp, simd256_width/sizeof(_Tp)> cv::v256_load_aligned(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load register contents from memory (aligned)

similar to cv::v256_load, but source memory block should be aligned (to 32-byte boundary in case of SIMD256, 64-byte - SIMD512, etc)

Note

Check CV_SIMD256 preprocessor definition prior to use. Use vx_load_aligned version to get maximum available register length result

Here is the call graph for this function:

cv::v256_load_aligned Node1 cv::v256_load_aligned Node2 cv::isAligned Node1->Node2

cv::v256_load_aligned Node1 cv::v256_load_aligned Node2 cv::isAligned Node1->Node2

v256_load_expand() [1/2]#

template<typename _Tp>
inline v_reg< typename V_TypeTraits< _Tp >::w_type, simd256_width/sizeof(typename V_TypeTraits< _Tp >::w_type)> cv::v256_load_expand(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load register contents from memory with double expand.

Same as cv::v256_load, but result pack type will be 2x wider than memory type.

short buf[8] = {1, 2, 3, 4, 5, 6, 7, 8}; // type is int16
v_int32x8 r = v256_load_expand(buf); // r = {1, 2, 3, 4, 5, 6, 7, 8} - type is int32

For 8-, 16-, 32-bit integer source types.

Note

Check CV_SIMD256 preprocessor definition prior to use. Use vx_load_expand version to get maximum available register length result

Here is the call graph for this function:

cv::v256_load_expand Node1 cv::v256_load_expand Node2 cv::isAligned Node1->Node2

cv::v256_load_expand Node1 cv::v256_load_expand Node2 cv::isAligned Node1->Node2

v256_load_expand() [2/2]#

inline v_reg< float, simd256_width/sizeof(float)> cv::v256_load_expand(const hfloat * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

v256_load_expand_q()#

template<typename _Tp>
inline v_reg< typename V_TypeTraits< _Tp >::q_type, simd256_width/sizeof(typename V_TypeTraits< _Tp >::q_type)> cv::v256_load_expand_q(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load register contents from memory with quad expand.

Same as cv::v256_load_expand, but result type is 4 times wider than source.

char buf[8] = {1, 2, 3, 4, 5, 6, 7, 8}; // type is int8
v_int32x8 r = v256_load_expand_q(buf); // r = {1, 2, 3, 4, 5, 6, 7, 8} - type is int32

For 8-bit integer source types.

Note

Check CV_SIMD256 preprocessor definition prior to use. Use vx_load_expand_q version to get maximum available register length result

Here is the call graph for this function:

cv::v256_load_expand_q Node1 cv::v256_load_expand_q Node2 cv::isAligned Node1->Node2

cv::v256_load_expand_q Node1 cv::v256_load_expand_q Node2 cv::isAligned Node1->Node2

v256_load_halves()#

template<typename _Tp>
inline v_reg< _Tp, simd256_width/sizeof(_Tp)> cv::v256_load_halves(
const _Tp * loptr,
const _Tp * hiptr )

#include <opencv2/core/hal/intrin_cpp.hpp>

Load register contents from two memory blocks.

int lo[4] = { 1, 2, 3, 4 }, hi[4] = { 5, 6, 7, 8 };
v_int32x8 r = v256_load_halves(lo, hi);

Note

Check CV_SIMD256 preprocessor definition prior to use. Use vx_load_halves version to get maximum available register length result

Parameters

  • loptr — memory block containing data for first half (0..n/2)

  • hiptr — memory block containing data for second half (n/2..n)

Here is the call graph for this function:

cv::v256_load_halves Node1 cv::v256_load_halves Node2 cv::isAligned Node1->Node2

cv::v256_load_halves Node1 cv::v256_load_halves Node2 cv::isAligned Node1->Node2

v256_load_low()#

template<typename _Tp>
inline v_reg< _Tp, simd256_width/sizeof(_Tp)> cv::v256_load_low(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load 128-bits of data to lower part (high part is undefined).

int lo[4] = { 1, 2, 3, 4 };
v_int32x8 r = v256_load_low(lo);

Note

Check CV_SIMD256 preprocessor definition prior to use. Use vx_load_low version to get maximum available register length result

Parameters

  • ptr — memory block containing data for first half (0..n/2)

Here is the call graph for this function:

cv::v256_load_low Node1 cv::v256_load_low Node2 cv::isAligned Node1->Node2

cv::v256_load_low Node1 cv::v256_load_low Node2 cv::isAligned Node1->Node2

v512_cleanup()#

inline void cv::v512_cleanup()

#include <opencv2/core/hal/intrin_cpp.hpp>

v512_load()#

template<typename _Tp>
inline v_reg< _Tp, simd512_width/sizeof(_Tp)> cv::v512_load(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load 512-bit length register contents from memory.

Note

Returned type will be detected from passed pointer type, for example uchar ==> cv::v_uint8x64, int ==> cv::v_int32x16, etc.

Check CV_SIMD512 preprocessor definition prior to use. Use vx_load version to get maximum available register length result

Alignment requirement: if CV_STRONG_ALIGNMENT=1 then passed pointer must be aligned (sizeof(lane type) should be enough). Do not cast pointer types without runtime check for pointer alignment (like uchar* => int*).

Parameters

  • ptr — pointer to memory block with data

Returns — register object

Here is the call graph for this function:

cv::v512_load Node1 cv::v512_load Node2 cv::isAligned Node1->Node2

cv::v512_load Node1 cv::v512_load Node2 cv::isAligned Node1->Node2

v512_load_aligned()#

template<typename _Tp>
inline v_reg< _Tp, simd512_width/sizeof(_Tp)> cv::v512_load_aligned(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load register contents from memory (aligned)

similar to cv::v512_load, but source memory block should be aligned (to 64-byte boundary in case of SIMD512, etc)

Note

Check CV_SIMD512 preprocessor definition prior to use. Use vx_load_aligned version to get maximum available register length result

Here is the call graph for this function:

cv::v512_load_aligned Node1 cv::v512_load_aligned Node2 cv::isAligned Node1->Node2

cv::v512_load_aligned Node1 cv::v512_load_aligned Node2 cv::isAligned Node1->Node2

v512_load_expand() [1/2]#

template<typename _Tp>
inline v_reg< typename V_TypeTraits< _Tp >::w_type, simd512_width/sizeof(typename V_TypeTraits< _Tp >::w_type)> cv::v512_load_expand(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load register contents from memory with double expand.

Same as cv::v512_load, but result pack type will be 2x wider than memory type.

short buf[8] = {1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16}; // type is int16
v_int32x16 r = v512_load_expand(buf); // r = {1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16} - type is int32

For 8-, 16-, 32-bit integer source types.

Note

Check CV_SIMD512 preprocessor definition prior to use. Use vx_load_expand version to get maximum available register length result

Here is the call graph for this function:

cv::v512_load_expand Node1 cv::v512_load_expand Node2 cv::isAligned Node1->Node2

cv::v512_load_expand Node1 cv::v512_load_expand Node2 cv::isAligned Node1->Node2

v512_load_expand() [2/2]#

inline v_reg< float, simd512_width/sizeof(float)> cv::v512_load_expand(const hfloat * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

v512_load_expand_q()#

template<typename _Tp>
inline v_reg< typename V_TypeTraits< _Tp >::q_type, simd512_width/sizeof(typename V_TypeTraits< _Tp >::q_type)> cv::v512_load_expand_q(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load register contents from memory with quad expand.

Same as cv::v512_load_expand, but result type is 4 times wider than source.

char buf[16] = {1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16}; // type is int8
v_int32x16 r = v512_load_expand_q(buf); // r = {1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16} - type is int32

For 8-bit integer source types.

Note

Check CV_SIMD512 preprocessor definition prior to use. Use vx_load_expand_q version to get maximum available register length result

Here is the call graph for this function:

cv::v512_load_expand_q Node1 cv::v512_load_expand_q Node2 cv::isAligned Node1->Node2

cv::v512_load_expand_q Node1 cv::v512_load_expand_q Node2 cv::isAligned Node1->Node2

v512_load_halves()#

template<typename _Tp>
inline v_reg< _Tp, simd512_width/sizeof(_Tp)> cv::v512_load_halves(
const _Tp * loptr,
const _Tp * hiptr )

#include <opencv2/core/hal/intrin_cpp.hpp>

Load register contents from two memory blocks.

int lo[4] = { 1, 2, 3, 4, 5, 6, 7, 8 }, hi[4] = { 9, 10, 11, 12, 13, 14, 15, 16 };
v_int32x16 r = v512_load_halves(lo, hi);

Note

Check CV_SIMD512 preprocessor definition prior to use. Use vx_load_halves version to get maximum available register length result

Parameters

  • loptr — memory block containing data for first half (0..n/2)

  • hiptr — memory block containing data for second half (n/2..n)

Here is the call graph for this function:

cv::v512_load_halves Node1 cv::v512_load_halves Node2 cv::isAligned Node1->Node2

cv::v512_load_halves Node1 cv::v512_load_halves Node2 cv::isAligned Node1->Node2

v512_load_low()#

template<typename _Tp>
inline v_reg< _Tp, simd512_width/sizeof(_Tp)> cv::v512_load_low(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load 256-bits of data to lower part (high part is undefined).

int lo[8] = { 1, 2, 3, 4, 5, 6, 7, 8 };
v_int32x16 r = v512_load_low(lo);

Note

Check CV_SIMD512 preprocessor definition prior to use. Use vx_load_low version to get maximum available register length result

Parameters

  • ptr — memory block containing data for first half (0..n/2)

Here is the call graph for this function:

cv::v512_load_low Node1 cv::v512_load_low Node2 cv::isAligned Node1->Node2

cv::v512_load_low Node1 cv::v512_load_low Node2 cv::isAligned Node1->Node2

v_absdiff() [1/3]#

template<typename _Tp, int n>
inline v_reg< typename V_TypeTraits< _Tp >::abs_type, n > cv::v_absdiff(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Add values without saturation.

For 8- and 16-bit integer values.

Subtract values without saturation

For 8- and 16-bit integer values.

Multiply values without saturation

For 8- and 16-bit integer values.

Absolute difference

Returns \( |a - b| \) converted to corresponding unsigned type. Example:

v_int32x4 a, b; // {1, 2, 3, 4} and {4, 3, 2, 1}
v_uint32x4 c = v_absdiff(a, b); // result is {3, 1, 1, 3}

For 8-, 16-, 32-bit integer source types.

v_absdiff() [2/3]#

template<int n>
inline v_reg< double, n > cv::v_absdiff(
const v_reg< double, n > & a,
const v_reg< double, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.

For 64-bit floating point values

v_absdiff() [3/3]#

template<int n>
inline v_reg< float, n > cv::v_absdiff(
const v_reg< float, n > & a,
const v_reg< float, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.

For 32-bit floating point values

v_absdiffs()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_absdiffs(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Saturating absolute difference.

Returns \( saturate(|a - b|) \) . For 8-, 16-bit signed integer source types.

Here is the call graph for this function:

cv::v_absdiffs Node1 cv::v_absdiffs Node2 cv::saturate_cast Node1->Node2

cv::v_absdiffs Node1 cv::v_absdiffs Node2 cv::saturate_cast Node1->Node2

v_add()#

template<typename _Tp, int n>
v_reg< _Tp, n > cv::v_add(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Add values.

For all types.

v_and()#

template<typename _Tp, int n>
v_reg< _Tp, n > cv::v_and(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Bitwise AND.

Only for integer types.

v_broadcast_element()#

template<int i, typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_broadcast_element(const v_reg< _Tp, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Broadcast i-th element of vector.

Scheme:

{ v[0] v[1] v[2] ... v[SZ] } => { v[i], v[i], v[i] ... v[i] }

Restriction: 0 <= i < nlanes Supported types: 32-bit integers and floats (s32/u32/f32)

v_ceil() [1/2]#

template<int n>
inline v_reg< int, n *2 > cv::v_ceil(const v_reg< double, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.

Here is the call graph for this function:

cv::v_ceil Node1 cv::v_ceil Node2 cvCeil Node1->Node2

cv::v_ceil Node1 cv::v_ceil Node2 cvCeil Node1->Node2

v_ceil() [2/2]#

template<int n>
inline v_reg< int, n > cv::v_ceil(const v_reg< float, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Ceil elements.

Ceil each value. Input type is float vector ==> output type is int vector.

Note

Only for floating point types.

Here is the call graph for this function:

cv::v_ceil Node1 cv::v_ceil Node2 cvCeil Node1->Node2

cv::v_ceil Node1 cv::v_ceil Node2 cvCeil Node1->Node2

v_check_all()#

template<typename _Tp, int n>
inline bool cv::v_check_all(const v_reg< _Tp, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Check if all packed values are less than zero.

Unsigned values will be casted to signed: uchar 254 => char -2.

v_check_any()#

template<typename _Tp, int n>
inline bool cv::v_check_any(const v_reg< _Tp, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Check if any of packed values is less than zero.

Unsigned values will be casted to signed: uchar 254 => char -2.

v_cleanup()#

inline void cv::v_cleanup()

#include <opencv2/core/hal/intrin_cpp.hpp>

v_combine_high()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_combine_high(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Combine vector from last elements of two vectors.

Scheme:

  {A1 A2 A3 A4}
  {B1 B2 B3 B4}
---------------
  {A3 A4 B3 B4}

For all types except 64-bit.

v_combine_low()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_combine_low(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Combine vector from first elements of two vectors.

Scheme:

  {A1 A2 A3 A4}
  {B1 B2 B3 B4}
---------------
  {A1 A2 B1 B2}

For all types except 64-bit.

v_cvt_f32() [1/3]#

template<int n>
inline v_reg< float, n *2 > cv::v_cvt_f32(const v_reg< double, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Convert lower half to float.

Supported input type is cv::v_float64.

v_cvt_f32() [2/3]#

template<int n>
inline v_reg< float, n *2 > cv::v_cvt_f32(
const v_reg< double, n > & a,
const v_reg< double, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Convert to float.

Supported input type is cv::v_float64.

v_cvt_f32() [3/3]#

template<int n>
inline v_reg< float, n > cv::v_cvt_f32(const v_reg< int, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Convert to float.

Supported input type is cv::v_int32.

v_cvt_f64() [1/3]#

template<int n>
v_reg< double,(n/2)> cv::v_cvt_f64(const v_reg< float, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Convert lower half to double.

Supported input type is cv::v_float32.

v_cvt_f64() [2/3]#

template<int n>
v_reg< double, n/2 > cv::v_cvt_f64(const v_reg< int, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Convert lower half to double.

Supported input type is cv::v_int32.

v_cvt_f64() [3/3]#

template<int n>
v_reg< double, n > cv::v_cvt_f64(const v_reg< int64, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Convert to double.

Supported input type is cv::v_int64.

v_cvt_f64_high() [1/2]#

template<int n>
v_reg< double,(n/2)> cv::v_cvt_f64_high(const v_reg< float, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Convert to double high part of vector.

Supported input type is cv::v_float32.

v_cvt_f64_high() [2/2]#

template<int n>
v_reg< double,(n/2)> cv::v_cvt_f64_high(const v_reg< int, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Convert to double high part of vector.

Supported input type is cv::v_int32.

v_div()#

template<typename _Tp, int n>
v_reg< _Tp, n > cv::v_div(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Divide values.

For floating types only.

v_dotprod() [1/2]#

template<typename _Tp, int n>
inline v_reg< typename V_TypeTraits< _Tp >::w_type, n/2 > cv::v_dotprod(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Dot product of elements.

Multiply values in two registers and sum adjacent result pairs.

Scheme:

  {A1 A2 ...} // 16-bit
x {B1 B2 ...} // 16-bit
-------------
{A1B1+A2B2 ...} // 32-bit

v_dotprod() [2/2]#

template<typename _Tp, int n>
inline v_reg< typename V_TypeTraits< _Tp >::w_type, n/2 > cv::v_dotprod(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b,
const v_reg< typename V_TypeTraits< _Tp >::w_type, n/2 > & c )

#include <opencv2/core/hal/intrin_cpp.hpp>

Dot product of elements.

Same as cv::v_dotprod, but add a third element to the sum of adjacent pairs. Scheme:

  {A1 A2 ...} // 16-bit
x {B1 B2 ...} // 16-bit
-------------
  {A1B1+A2B2+C1 ...} // 32-bit

v_dotprod_expand() [1/4]#

template<typename _Tp, int n>
inline v_reg< typename V_TypeTraits< _Tp >::q_type, n/4 > cv::v_dotprod_expand(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Dot product of elements and expand.

Multiply values in two registers and expand the sum of adjacent result pairs.

Scheme:

  {A1 A2 A3 A4 ...} // 8-bit
x {B1 B2 B3 B4 ...} // 8-bit
-------------
  {A1B1+A2B2+A3B3+A4B4 ...} // 32-bit

v_dotprod_expand() [2/4]#

template<typename _Tp, int n>
inline v_reg< typename V_TypeTraits< _Tp >::q_type, n/4 > cv::v_dotprod_expand(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b,
const v_reg< typename V_TypeTraits< _Tp >::q_type, n/4 > & c )

#include <opencv2/core/hal/intrin_cpp.hpp>

Dot product of elements.

Same as cv::v_dotprod_expand, but add a third element to the sum of adjacent pairs. Scheme:

  {A1 A2 A3 A4 ...} // 8-bit
x {B1 B2 B3 B4 ...} // 8-bit
-------------
  {A1B1+A2B2+A3B3+A4B4+C1 ...} // 32-bit

v_dotprod_expand() [3/4]#

template<int n>
inline v_reg< double, n/2 > cv::v_dotprod_expand(
const v_reg< int, n > & a,
const v_reg< int, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Here is the call graph for this function:

cv::v_dotprod_expand Node1 cv::v_dotprod_expand Node2 cv::v_cvt_f64 Node1->Node2 Node3 cv::v_cvt_f64_high Node1->Node3 Node4 cv::v_fma Node1->Node4 Node5 cv::v_mul Node1->Node5

cv::v_dotprod_expand Node1 cv::v_dotprod_expand Node2 cv::v_cvt_f64 Node1->Node2 Node3 cv::v_cvt_f64_high Node1->Node3 Node4 cv::v_fma Node1->Node4 Node5 cv::v_mul Node1->Node5

v_dotprod_expand() [4/4]#

template<int n>
inline v_reg< double, n/2 > cv::v_dotprod_expand(
const v_reg< int, n > & a,
const v_reg< int, n > & b,
const v_reg< double, n/2 > & c )

#include <opencv2/core/hal/intrin_cpp.hpp>

Here is the call graph for this function:

cv::v_dotprod_expand Node1 cv::v_dotprod_expand Node2 cv::v_cvt_f64 Node1->Node2 Node3 cv::v_cvt_f64_high Node1->Node3 Node4 cv::v_fma Node1->Node4

cv::v_dotprod_expand Node1 cv::v_dotprod_expand Node2 cv::v_cvt_f64 Node1->Node2 Node3 cv::v_cvt_f64_high Node1->Node3 Node4 cv::v_fma Node1->Node4

v_dotprod_expand_fast() [1/4]#

template<typename _Tp, int n>
inline v_reg< typename V_TypeTraits< _Tp >::q_type, n/4 > cv::v_dotprod_expand_fast(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Fast Dot product of elements and expand.

Multiply values in two registers and expand the sum of adjacent result pairs.

Same as cv::v_dotprod_expand, but it may perform unorder sum between result pairs in some platforms, this intrinsic can be used if the sum among all lanes is only matters and also it should be yielding better performance on the affected platforms.

Here is the call graph for this function:

cv::v_dotprod_expand_fast Node1 cv::v_dotprod_expand_fast Node2 cv::v_dotprod_expand Node1->Node2

cv::v_dotprod_expand_fast Node1 cv::v_dotprod_expand_fast Node2 cv::v_dotprod_expand Node1->Node2

v_dotprod_expand_fast() [2/4]#

template<typename _Tp, int n>
inline v_reg< typename V_TypeTraits< _Tp >::q_type, n/4 > cv::v_dotprod_expand_fast(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b,
const v_reg< typename V_TypeTraits< _Tp >::q_type, n/4 > & c )

#include <opencv2/core/hal/intrin_cpp.hpp>

Fast Dot product of elements.

Same as cv::v_dotprod_expand_fast, but add a third element to the sum of adjacent pairs.

Here is the call graph for this function:

cv::v_dotprod_expand_fast Node1 cv::v_dotprod_expand_fast Node2 cv::v_dotprod_expand Node1->Node2

cv::v_dotprod_expand_fast Node1 cv::v_dotprod_expand_fast Node2 cv::v_dotprod_expand Node1->Node2

v_dotprod_expand_fast() [3/4]#

template<int n>
inline v_reg< double, n/2 > cv::v_dotprod_expand_fast(
const v_reg< int, n > & a,
const v_reg< int, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Here is the call graph for this function:

cv::v_dotprod_expand_fast Node1 cv::v_dotprod_expand_fast Node2 cv::v_dotprod_expand Node1->Node2

cv::v_dotprod_expand_fast Node1 cv::v_dotprod_expand_fast Node2 cv::v_dotprod_expand Node1->Node2

v_dotprod_expand_fast() [4/4]#

template<int n>
inline v_reg< double, n/2 > cv::v_dotprod_expand_fast(
const v_reg< int, n > & a,
const v_reg< int, n > & b,
const v_reg< double, n/2 > & c )

#include <opencv2/core/hal/intrin_cpp.hpp>

Here is the call graph for this function:

cv::v_dotprod_expand_fast Node1 cv::v_dotprod_expand_fast Node2 cv::v_dotprod_expand Node1->Node2

cv::v_dotprod_expand_fast Node1 cv::v_dotprod_expand_fast Node2 cv::v_dotprod_expand Node1->Node2

v_dotprod_fast() [1/2]#

template<typename _Tp, int n>
inline v_reg< typename V_TypeTraits< _Tp >::w_type, n/2 > cv::v_dotprod_fast(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Fast Dot product of elements.

Same as cv::v_dotprod, but it may perform unorder sum between result pairs in some platforms, this intrinsic can be used if the sum among all lanes is only matters and also it should be yielding better performance on the affected platforms.

Here is the call graph for this function:

cv::v_dotprod_fast Node1 cv::v_dotprod_fast Node2 cv::v_dotprod Node1->Node2

cv::v_dotprod_fast Node1 cv::v_dotprod_fast Node2 cv::v_dotprod Node1->Node2

v_dotprod_fast() [2/2]#

template<typename _Tp, int n>
inline v_reg< typename V_TypeTraits< _Tp >::w_type, n/2 > cv::v_dotprod_fast(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b,
const v_reg< typename V_TypeTraits< _Tp >::w_type, n/2 > & c )

#include <opencv2/core/hal/intrin_cpp.hpp>

Fast Dot product of elements.

Same as cv::v_dotprod_fast, but add a third element to the sum of adjacent pairs.

Here is the call graph for this function:

cv::v_dotprod_fast Node1 cv::v_dotprod_fast Node2 cv::v_dotprod Node1->Node2

cv::v_dotprod_fast Node1 cv::v_dotprod_fast Node2 cv::v_dotprod Node1->Node2

v_expand()#

template<typename _Tp, int n>
inline void cv::v_expand(
const v_reg< _Tp, n > & a,
v_reg< typename V_TypeTraits< _Tp >::w_type, n/2 > & b0,
v_reg< typename V_TypeTraits< _Tp >::w_type, n/2 > & b1 )

#include <opencv2/core/hal/intrin_cpp.hpp>

Expand values to the wider pack type.

Copy contents of register to two registers with 2x wider pack type. Scheme:

 int32x4     int64x2 int64x2
{A B C D} ==> {A B} , {C D}

v_expand_high()#

template<typename _Tp, int n>
inline v_reg< typename V_TypeTraits< _Tp >::w_type, n/2 > cv::v_expand_high(const v_reg< _Tp, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Expand higher values to the wider pack type.

Same as cv::v_expand_low, but expand higher half of the vector instead.

Scheme:

 int32x4     int64x2
{A B C D} ==> {C D}

v_expand_low()#

template<typename _Tp, int n>
inline v_reg< typename V_TypeTraits< _Tp >::w_type, n/2 > cv::v_expand_low(const v_reg< _Tp, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Expand lower values to the wider pack type.

Same as cv::v_expand, but return lower half of the vector.

Scheme:

 int32x4     int64x2
{A B C D} ==> {A B}

v_extract()#

template<int s, typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_extract(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Vector extract.

Scheme:

  {A1 A2 A3 A4}
## {B1 B2 B3 B4}
shift = 1  {A2 A3 A4 B1}
shift = 2  {A3 A4 B1 B2}
shift = 3  {A4 B1 B2 B3}

Restriction: 0 <= shift < nlanes

Usage:

v_int32x4 a, b, c;
c = v_extract<2>(a, b);

For all types.

v_extract_n()#

template<int s, typename _Tp, int n>
inline _Tp cv::v_extract_n(const v_reg< _Tp, n > & v)

#include <opencv2/core/hal/intrin_cpp.hpp>

Vector extract.

Scheme: Return the s-th element of v. Restriction: 0 <= s < nlanes

Usage:

v_int32x4 a;
int r;
r = v_extract_n<2>(a);

For all types.

v_floor() [1/2]#

template<int n>
inline v_reg< int, n *2 > cv::v_floor(const v_reg< double, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.

Here is the call graph for this function:

cv::v_floor Node1 cv::v_floor Node2 cvFloor Node1->Node2

cv::v_floor Node1 cv::v_floor Node2 cvFloor Node1->Node2

v_floor() [2/2]#

template<int n>
inline v_reg< int, n > cv::v_floor(const v_reg< float, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Floor elements.

Floor each value. Input type is float vector ==> output type is int vector.

Note

Only for floating point types.

Here is the call graph for this function:

cv::v_floor Node1 cv::v_floor Node2 cvFloor Node1->Node2

cv::v_floor Node1 cv::v_floor Node2 cvFloor Node1->Node2

v_fma()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_fma(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b,
const v_reg< _Tp, n > & c )

#include <opencv2/core/hal/intrin_cpp.hpp>

Multiply and add.

Returns \( a*b + c \) For floating point types and signed 32bit int only.

v_interleave_pairs()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_interleave_pairs(const v_reg< _Tp, n > & vec)

#include <opencv2/core/hal/intrin_cpp.hpp>

v_interleave_quads()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_interleave_quads(const v_reg< _Tp, n > & vec)

#include <opencv2/core/hal/intrin_cpp.hpp>

v_invsqrt()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_invsqrt(const v_reg< _Tp, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Inversed square root.

Returns \( 1/sqrt(a) \) For floating point types only.

v_load()#

template<typename _Tp>
inline v_reg< _Tp, simd128_width/sizeof(_Tp)> cv::v_load(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load register contents from memory.

Note

Returned type will be detected from passed pointer type, for example uchar ==> cv::v_uint8x16, int ==> cv::v_int32x4, etc.

Use vx_load version to get maximum available register length result

Alignment requirement: if CV_STRONG_ALIGNMENT=1 then passed pointer must be aligned (sizeof(lane type) should be enough). Do not cast pointer types without runtime check for pointer alignment (like uchar* => int*).

Parameters

  • ptr — pointer to memory block with data

Returns — register object

Here is the call graph for this function:

cv::v_load Node1 cv::v_load Node2 cv::isAligned Node1->Node2

cv::v_load Node1 cv::v_load Node2 cv::isAligned Node1->Node2

v_load_aligned()#

template<typename _Tp>
inline v_reg< _Tp, simd128_width/sizeof(_Tp)> cv::v_load_aligned(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load register contents from memory (aligned)

similar to cv::v_load, but source memory block should be aligned (to 16-byte boundary in case of SIMD128, 32-byte - SIMD256, etc)

Note

Use vx_load_aligned version to get maximum available register length result

Here is the call graph for this function:

cv::v_load_aligned Node1 cv::v_load_aligned Node2 cv::isAligned Node1->Node2

cv::v_load_aligned Node1 cv::v_load_aligned Node2 cv::isAligned Node1->Node2

v_load_deinterleave() [1/3]#

template<typename _Tp, int n>
inline void cv::v_load_deinterleave(
const _Tp * ptr,
v_reg< _Tp, n > & a,
v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Load and deinterleave (2 channels)

Load data from memory deinterleave and store to 2 registers. Scheme:

{A1 B1 A2 B2 ...} ==> {A1 A2 ...}, {B1 B2 ...}

For all types except 64-bit.

Here is the call graph for this function:

cv::v_load_deinterleave Node1 cv::v_load_deinterleave Node2 cv::isAligned Node1->Node2

cv::v_load_deinterleave Node1 cv::v_load_deinterleave Node2 cv::isAligned Node1->Node2

v_load_deinterleave() [2/3]#

template<typename _Tp, int n>
inline void cv::v_load_deinterleave(
const _Tp * ptr,
v_reg< _Tp, n > & a,
v_reg< _Tp, n > & b,
v_reg< _Tp, n > & c )

#include <opencv2/core/hal/intrin_cpp.hpp>

Load and deinterleave (3 channels)

Load data from memory deinterleave and store to 3 registers. Scheme:

{A1 B1 C1 A2 B2 C2 ...} ==> {A1 A2 ...}, {B1 B2 ...}, {C1 C2 ...}

For all types except 64-bit.

Here is the call graph for this function:

cv::v_load_deinterleave Node1 cv::v_load_deinterleave Node2 cv::isAligned Node1->Node2

cv::v_load_deinterleave Node1 cv::v_load_deinterleave Node2 cv::isAligned Node1->Node2

v_load_deinterleave() [3/3]#

template<typename _Tp, int n>
inline void cv::v_load_deinterleave(
const _Tp * ptr,
v_reg< _Tp, n > & a,
v_reg< _Tp, n > & b,
v_reg< _Tp, n > & c,
v_reg< _Tp, n > & d )

#include <opencv2/core/hal/intrin_cpp.hpp>

Load and deinterleave (4 channels)

Load data from memory deinterleave and store to 4 registers. Scheme:

{A1 B1 C1 D1 A2 B2 C2 D2 ...} ==> {A1 A2 ...}, {B1 B2 ...}, {C1 C2 ...}, {D1 D2 ...}

For all types except 64-bit.

Here is the call graph for this function:

cv::v_load_deinterleave Node1 cv::v_load_deinterleave Node2 cv::isAligned Node1->Node2

cv::v_load_deinterleave Node1 cv::v_load_deinterleave Node2 cv::isAligned Node1->Node2

v_load_expand() [1/2]#

template<typename _Tp>
inline v_reg< typename V_TypeTraits< _Tp >::w_type, simd128_width/sizeof(typename V_TypeTraits< _Tp >::w_type)> cv::v_load_expand(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load register contents from memory with double expand.

Same as cv::v_load, but result pack type will be 2x wider than memory type.

short buf[4] = {1, 2, 3, 4}; // type is int16
v_int32x4 r = v_load_expand(buf); // r = {1, 2, 3, 4} - type is int32

For 8-, 16-, 32-bit integer source types.

Note

Use vx_load_expand version to get maximum available register length result

Here is the call graph for this function:

cv::v_load_expand Node1 cv::v_load_expand Node2 cv::isAligned Node1->Node2

cv::v_load_expand Node1 cv::v_load_expand Node2 cv::isAligned Node1->Node2

v_load_expand() [2/2]#

inline v_reg< float, simd128_width/sizeof(float)> cv::v_load_expand(const hfloat * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

v_load_expand_q()#

template<typename _Tp>
inline v_reg< typename V_TypeTraits< _Tp >::q_type, simd128_width/sizeof(typename V_TypeTraits< _Tp >::q_type)> cv::v_load_expand_q(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load register contents from memory with quad expand.

Same as cv::v_load_expand, but result type is 4 times wider than source.

char buf[4] = {1, 2, 3, 4}; // type is int8
v_int32x4 r = v_load_expand_q(buf); // r = {1, 2, 3, 4} - type is int32

For 8-bit integer source types.

Note

Use vx_load_expand_q version to get maximum available register length result

Here is the call graph for this function:

cv::v_load_expand_q Node1 cv::v_load_expand_q Node2 cv::isAligned Node1->Node2

cv::v_load_expand_q Node1 cv::v_load_expand_q Node2 cv::isAligned Node1->Node2

v_load_halves()#

template<typename _Tp>
inline v_reg< _Tp, simd128_width/sizeof(_Tp)> cv::v_load_halves(
const _Tp * loptr,
const _Tp * hiptr )

#include <opencv2/core/hal/intrin_cpp.hpp>

Load register contents from two memory blocks.

int lo[2] = { 1, 2 }, hi[2] = { 3, 4 };
v_int32x4 r = v_load_halves(lo, hi);

Note

Use vx_load_halves version to get maximum available register length result

Parameters

  • loptr — memory block containing data for first half (0..n/2)

  • hiptr — memory block containing data for second half (n/2..n)

Here is the call graph for this function:

cv::v_load_halves Node1 cv::v_load_halves Node2 cv::isAligned Node1->Node2

cv::v_load_halves Node1 cv::v_load_halves Node2 cv::isAligned Node1->Node2

v_load_low()#

template<typename _Tp>
inline v_reg< _Tp, simd128_width/sizeof(_Tp)> cv::v_load_low(const _Tp * ptr)

#include <opencv2/core/hal/intrin_cpp.hpp>

Load 64-bits of data to lower part (high part is undefined).

int lo[2] = { 1, 2 };
v_int32x4 r = v_load_low(lo);

Note

Use vx_load_low version to get maximum available register length result

Parameters

  • ptr — memory block containing data for first half (0..n/2)

Here is the call graph for this function:

cv::v_load_low Node1 cv::v_load_low Node2 cv::isAligned Node1->Node2

cv::v_load_low Node1 cv::v_load_low Node2 cv::isAligned Node1->Node2

v_lut() [1/5]#

template<typename _Tp>
inline v_reg< _Tp, simd128_width/sizeof(_Tp)> cv::v_lut(
const _Tp * tab,
const int * idx )

#include <opencv2/core/hal/intrin_cpp.hpp>

v_lut() [2/5]#

template<int n>
inline v_reg< double, n/2 > cv::v_lut(
const double * tab,
const v_reg< int, n > & idx )

#include <opencv2/core/hal/intrin_cpp.hpp>

v_lut() [3/5]#

template<int n>
inline v_reg< float, n > cv::v_lut(
const float * tab,
const v_reg< int, n > & idx )

#include <opencv2/core/hal/intrin_cpp.hpp>

v_lut() [4/5]#

template<int n>
inline v_reg< int, n > cv::v_lut(
const int * tab,
const v_reg< int, n > & idx )

#include <opencv2/core/hal/intrin_cpp.hpp>

v_lut() [5/5]#

template<int n>
inline v_reg< unsigned, n > cv::v_lut(
const unsigned * tab,
const v_reg< int, n > & idx )

#include <opencv2/core/hal/intrin_cpp.hpp>

v_lut_deinterleave() [1/2]#

template<int n>
inline void cv::v_lut_deinterleave(
const double * tab,
const v_reg< int, n *2 > & idx,
v_reg< double, n > & x,
v_reg< double, n > & y )

#include <opencv2/core/hal/intrin_cpp.hpp>

v_lut_deinterleave() [2/2]#

template<int n>
inline void cv::v_lut_deinterleave(
const float * tab,
const v_reg< int, n > & idx,
v_reg< float, n > & x,
v_reg< float, n > & y )

#include <opencv2/core/hal/intrin_cpp.hpp>

v_lut_pairs()#

template<typename _Tp>
inline v_reg< _Tp, simd128_width/sizeof(_Tp)> cv::v_lut_pairs(
const _Tp * tab,
const int * idx )

#include <opencv2/core/hal/intrin_cpp.hpp>

v_lut_quads()#

template<typename _Tp>
inline v_reg< _Tp, simd128_width/sizeof(_Tp)> cv::v_lut_quads(
const _Tp * tab,
const int * idx )

#include <opencv2/core/hal/intrin_cpp.hpp>

v_magnitude()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_magnitude(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Magnitude.

Returns \( sqrt(a^2 + b^2) \) For floating point types only.

v_matmul()#

template<int n>
inline v_reg< float, n > cv::v_matmul(
const v_reg< float, n > & v,
const v_reg< float, n > & a,
const v_reg< float, n > & b,
const v_reg< float, n > & c,
const v_reg< float, n > & d )

#include <opencv2/core/hal/intrin_cpp.hpp>

Matrix multiplication.

Scheme:

{A0 A1 A2 A3}   |V0|
{B0 B1 B2 B3}   |V1|
{C0 C1 C2 C3}   |V2|
## {D0 D1 D2 D3} x |V3|
{R0 R1 R2 R3}, where:
R0 = A0V0 + B0V1 + C0V2 + D0V3,
R1 = A1V0 + B1V1 + C1V2 + D1V3
...

v_matmuladd()#

template<int n>
inline v_reg< float, n > cv::v_matmuladd(
const v_reg< float, n > & v,
const v_reg< float, n > & a,
const v_reg< float, n > & b,
const v_reg< float, n > & c,
const v_reg< float, n > & d )

#include <opencv2/core/hal/intrin_cpp.hpp>

Matrix multiplication and add.

Scheme:

{A0 A1 A2 A3}   |V0|   |D0|
{B0 B1 B2 B3}   |V1|   |D1|
{C0 C1 C2 C3} x |V2| + |D2|
====================   |D3|
{R0 R1 R2 R3}, where:
R0 = A0V0 + B0V1 + C0V2 + D0,
R1 = A1V0 + B1V1 + C1V2 + D1
...

v_mul()#

template<typename _Tp, int n>
v_reg< _Tp, n > cv::v_mul(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Multiply values.

For 16- and 32-bit integer types and floating types.

v_mul_expand()#

template<typename _Tp, int n>
inline void cv::v_mul_expand(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b,
v_reg< typename V_TypeTraits< _Tp >::w_type, n/2 > & c,
v_reg< typename V_TypeTraits< _Tp >::w_type, n/2 > & d )

#include <opencv2/core/hal/intrin_cpp.hpp>

Multiply and expand.

Multiply values two registers and store results in two registers with wider pack type. Scheme:

  {A B C D} // 32-bit
x {E F G H} // 32-bit
---------------
{AE BF}         // 64-bit
        {CG DH} // 64-bit

Example:

v_uint32x4 a, b; // {1,2,3,4} and {2,2,2,2}
v_uint64x2 c, d; // results
v_mul_expand(a, b, c, d); // c, d = {2,4}, {6, 8}

Implemented only for 16- and unsigned 32-bit source types (v_int16x8, v_uint16x8, v_uint32x4).

v_mul_hi()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_mul_hi(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Multiply and extract high part.

Multiply values two registers and store high part of the results. Implemented only for 16-bit source types (v_int16x8, v_uint16x8). Returns \( a*b >> 16 \)

v_muladd()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_muladd(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b,
const v_reg< _Tp, n > & c )

#include <opencv2/core/hal/intrin_cpp.hpp>

A synonym for v_fma.

Here is the call graph for this function:

cv::v_muladd Node1 cv::v_muladd Node2 cv::v_fma Node1->Node2

cv::v_muladd Node1 cv::v_muladd Node2 cv::v_fma Node1->Node2

v_not()#

template<typename _Tp, int n>
v_reg< _Tp, n > cv::v_not(const v_reg< _Tp, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Bitwise NOT.

Only for integer types.

v_not_nan() [1/2]#

template<int n>
inline v_reg< double, n > cv::v_not_nan(const v_reg< double, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

v_not_nan() [2/2]#

template<int n>
inline v_reg< float, n > cv::v_not_nan(const v_reg< float, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Less-than comparison.

For all types except 64-bit integer values.

Greater-than comparison

For all types except 64-bit integer values.

Less-than or equal comparison

For all types except 64-bit integer values.

Greater-than or equal comparison

For all types except 64-bit integer values.

Equal comparison

Not equal comparison

v_or()#

template<typename _Tp, int n>
v_reg< _Tp, n > cv::v_or(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Bitwise OR.

Only for integer types.

v_pack_store()#

template<int n>
inline void cv::v_pack_store(
hfloat * ptr,
const v_reg< float, n > & v )

#include <opencv2/core/hal/intrin_cpp.hpp>

v_pack_triplets()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_pack_triplets(const v_reg< _Tp, n > & vec)

#include <opencv2/core/hal/intrin_cpp.hpp>

v_popcount()#

template<typename _Tp, int n>
inline v_reg< typename V_TypeTraits< _Tp >::abs_type, n > cv::v_popcount(const v_reg< _Tp, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Count the 1 bits in the vector lanes and return result as corresponding unsigned type.

Scheme:

{A1 A2 A3 ...} => {popcount(A1), popcount(A2), popcount(A3), ...}

For all integer types.

Here is the call graph for this function:

cv::v_popcount Node1 cv::v_popcount Node2 cv::v_reinterpret_as_u8 Node1->Node2

cv::v_popcount Node1 cv::v_popcount Node2 cv::v_reinterpret_as_u8 Node1->Node2

v_recombine()#

template<typename _Tp, int n>
inline void cv::v_recombine(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b,
v_reg< _Tp, n > & low,
v_reg< _Tp, n > & high )

#include <opencv2/core/hal/intrin_cpp.hpp>

Combine two vectors from lower and higher parts of two other vectors.

low = cv::v_combine_low(a, b);
high = cv::v_combine_high(a, b);

v_reduce_sad()#

template<typename _Tp, int n>
inline V_TypeTraits< typenameV_TypeTraits< _Tp >::abs_type >::sum_type cv::v_reduce_sad(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Sum absolute differences of values.

Scheme:

{A1 A2 A3 ...} {B1 B2 B3 ...} => sum{ABS(A1-B1),abs(A2-B2),abs(A3-B3),...}

For all types except 64-bit types.

v_reduce_sum()#

template<typename _Tp, int n>
inline V_TypeTraits< _Tp >::sum_type cv::v_reduce_sum(const v_reg< _Tp, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Element shift left among vector.

For all type

Element shift right among vector

For all type

Sum packed values

Scheme:

{A1 A2 A3 ...} => sum{A1,A2,A3,...}

v_reduce_sum4()#

template<int n>
inline v_reg< float, n > cv::v_reduce_sum4(
const v_reg< float, n > & a,
const v_reg< float, n > & b,
const v_reg< float, n > & c,
const v_reg< float, n > & d )

#include <opencv2/core/hal/intrin_cpp.hpp>

Sums all elements of each input vector, returns the vector of sums.

Scheme:

result[0] = a[0] + a[1] + a[2] + a[3]
result[1] = b[0] + b[1] + b[2] + b[3]
result[2] = c[0] + c[1] + c[2] + c[3]
result[3] = d[0] + d[1] + d[2] + d[3]

v_reverse()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_reverse(const v_reg< _Tp, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Vector reverse order.

Reverse the order of the vector Scheme:

REG {A1 ... An} ==> REG {An ... A1}

For all types.

v_round() [1/3]#

template<int n>
inline v_reg< int, n *2 > cv::v_round(const v_reg< double, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.

Here is the call graph for this function:

cv::v_round Node1 cv::v_round Node2 cvRound Node1->Node2

cv::v_round Node1 cv::v_round Node2 cvRound Node1->Node2

v_round() [2/3]#

template<int n>
inline v_reg< int, n *2 > cv::v_round(
const v_reg< double, n > & a,
const v_reg< double, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.

Here is the call graph for this function:

cv::v_round Node1 cv::v_round Node2 cvRound Node1->Node2

cv::v_round Node1 cv::v_round Node2 cvRound Node1->Node2

v_round() [3/3]#

template<int n>
inline v_reg< int, n > cv::v_round(const v_reg< float, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Round elements.

Rounds each value. Input type is float vector ==> output type is int vector.

Note

Only for floating point types.

Here is the call graph for this function:

cv::v_round Node1 cv::v_round Node2 cvRound Node1->Node2

cv::v_round Node1 cv::v_round Node2 cvRound Node1->Node2

v_scan_forward()#

template<typename _Tp, int n>
inline int cv::v_scan_forward(const v_reg< _Tp, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Get first negative lane index.

Returned value is an index of first negative lane (undefined for input of all positive values) Example:

v_int32x4 r; // set to {0, 0, -1, -1}
int idx = v_heading_zeros(r); // idx = 2

v_select()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_select(
const v_reg< _Tp, n > & mask,
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Per-element select (blend operation)

Return value will be built by combining values a and b using the following scheme: result[i] = mask[i] ? a[i] : b[i];

Note

  • 0: select element from b

  • 0xff/0xffff/etc: select element from a (fully compatible with bitwise-based operator)

v_signmask()#

template<typename _Tp, int n>
inline int cv::v_signmask(const v_reg< _Tp, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Get negative values mask.

Deprecated

v_signmask depends on a lane count heavily and therefore isn’t universal enough

Returned value is a bit mask with bits set to 1 on places corresponding to negative packed values indexes. Example:

v_int32x4 r; // set to {-1, -1, 1, 1}
int mask = v_signmask(r); // mask = 3 <== 00000000 00000000 00000000 00000011

v_sincos()#

template<typename _Tp, int n>
inline void cv::v_sincos(
const v_reg< _Tp, n > & x,
v_reg< _Tp, n > & s,
v_reg< _Tp, n > & c )

#include <opencv2/core/hal/intrin_cpp.hpp>

Natural logarithm \( \log(x) \) of elements.

Only for floating point types. Core implementation steps:

  1. Decompose Input: Use binary representation to decompose the input into mantissa part \( m \) and exponent part \( e \). Such that \( \log(x) = \log(m \cdot 2^e) = \log(m) + e \cdot \ln(2) \).

  2. Adjust Mantissa and Exponent Parts: If the mantissa is less than \( \sqrt{0.5} \), adjust the exponent and mantissa to ensure the mantissa is in the range \( (\sqrt{0.5}, \sqrt{2}) \) for better approximation.

  3. Polynomial Approximation for \( \log(m) \): The closer the \( m \) is to 1, the more accurate the result.

  • For float16 and float32, use a Taylor Series with 9 terms.

  • For float64, use Pade Polynomials Approximation with 6 terms.

  1. Combine Results: Add the two parts together to get the final result.

Note

The precision of the calculation depends on the implementation and the data type of the input.

Similar to the behavior of std::log(), \( \ln(0) = -\infty \).

Error function.

Note

Support FP32 precision for now.

Compute sine \( sin(x) \) and cosine \( cos(x) \) of elements at the same time

Only for floating point types. Core implementation steps:

  1. Input Normalization: Scale the periodicity from 2π to 4 and reduce the angle to the range \( [0, \frac{\pi}{4}] \) using periodicity and trigonometric identities.

  2. Polynomial Approximation for \( sin(x) \) and \( cos(x) \):

  • For float16 and float32, use a Taylor series with 4 terms for sine and 5 terms for cosine.

  • For float64, use a Taylor series with 7 terms for sine and 8 terms for cosine.

  1. Select Results: select and convert the final sine and cosine values for the original input angle.

Note

The precision of the calculation depends on the implementation and the data type of the input vector.

v_sqr_magnitude()#

template<typename _Tp, int n>
inline v_reg< _Tp, n > cv::v_sqr_magnitude(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Square of the magnitude.

Returns \( a^2 + b^2 \) For floating point types only.

v_store() [1/2]#

template<typename _Tp, int n>
inline void cv::v_store(
_Tp * ptr,
const v_reg< _Tp, n > & a )

#include <opencv2/core/hal/intrin_cpp.hpp>

Store data to memory.

Store register contents to memory. Scheme:

REG {A B C D} ==> MEM {A B C D}

Pointer can be unaligned.

Here is the call graph for this function:

cv::v_store Node1 cv::v_store Node2 cv::isAligned Node1->Node2

cv::v_store Node1 cv::v_store Node2 cv::isAligned Node1->Node2

v_store() [2/2]#

template<typename _Tp, int n>
inline void cv::v_store(
_Tp * ptr,
const v_reg< _Tp, n > & a,
hal::StoreMode )

#include <opencv2/core/hal/intrin_cpp.hpp>

Here is the call graph for this function:

cv::v_store Node1 cv::v_store Node2 cv::isAligned Node1->Node2 Node3 cv::v_store Node1->Node3 Node3->Node2

cv::v_store Node1 cv::v_store Node2 cv::isAligned Node1->Node2 Node3 cv::v_store Node1->Node3 Node3->Node2

v_store_aligned() [1/2]#

template<typename _Tp, int n>
inline void cv::v_store_aligned(
_Tp * ptr,
const v_reg< _Tp, n > & a )

#include <opencv2/core/hal/intrin_cpp.hpp>

Store data to memory (aligned)

Store register contents to memory. Scheme:

REG {A B C D} ==> MEM {A B C D}

Pointer should be aligned by 16-byte boundary.

Here is the call graph for this function:

cv::v_store_aligned Node1 cv::v_store_aligned Node2 cv::isAligned Node1->Node2 Node3 cv::v_store Node1->Node3 Node3->Node2

cv::v_store_aligned Node1 cv::v_store_aligned Node2 cv::isAligned Node1->Node2 Node3 cv::v_store Node1->Node3 Node3->Node2

v_store_aligned() [2/2]#

template<typename _Tp, int n>
inline void cv::v_store_aligned(
_Tp * ptr,
const v_reg< _Tp, n > & a,
hal::StoreMode )

#include <opencv2/core/hal/intrin_cpp.hpp>

Here is the call graph for this function:

cv::v_store_aligned Node1 cv::v_store_aligned Node2 cv::isAligned Node1->Node2 Node3 cv::v_store Node1->Node3 Node3->Node2

cv::v_store_aligned Node1 cv::v_store_aligned Node2 cv::isAligned Node1->Node2 Node3 cv::v_store Node1->Node3 Node3->Node2

v_store_aligned_nocache()#

template<typename _Tp, int n>
inline void cv::v_store_aligned_nocache(
_Tp * ptr,
const v_reg< _Tp, n > & a )

#include <opencv2/core/hal/intrin_cpp.hpp>

Here is the call graph for this function:

cv::v_store_aligned_nocache Node1 cv::v_store_aligned _nocache Node2 cv::isAligned Node1->Node2 Node3 cv::v_store Node1->Node3 Node3->Node2

cv::v_store_aligned_nocache Node1 cv::v_store_aligned _nocache Node2 cv::isAligned Node1->Node2 Node3 cv::v_store Node1->Node3 Node3->Node2

v_store_high()#

template<typename _Tp, int n>
inline void cv::v_store_high(
_Tp * ptr,
const v_reg< _Tp, n > & a )

#include <opencv2/core/hal/intrin_cpp.hpp>

Store data to memory (higher half)

Store higher half of register contents to memory. Scheme:

REG {A B C D} ==> MEM {C D}

Here is the call graph for this function:

cv::v_store_high Node1 cv::v_store_high Node2 cv::isAligned Node1->Node2

cv::v_store_high Node1 cv::v_store_high Node2 cv::isAligned Node1->Node2

v_store_interleave() [1/3]#

template<typename _Tp, int n>
inline void cv::v_store_interleave(
_Tp * ptr,
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b,
const v_reg< _Tp, n > & c,
const v_reg< _Tp, n > & d,
hal::StoreMode = hal::STORE_UNALIGNED )

#include <opencv2/core/hal/intrin_cpp.hpp>

Interleave and store (4 channels)

Interleave and store data from 4 registers to memory. Scheme:

{A1 A2 ...}, {B1 B2 ...}, {C1 C2 ...}, {D1 D2 ...} ==> {A1 B1 C1 D1 A2 B2 C2 D2 ...}

For all types except 64-bit.

Here is the call graph for this function:

cv::v_store_interleave Node1 cv::v_store_interleave Node2 cv::isAligned Node1->Node2

cv::v_store_interleave Node1 cv::v_store_interleave Node2 cv::isAligned Node1->Node2

v_store_interleave() [2/3]#

template<typename _Tp, int n>
inline void cv::v_store_interleave(
_Tp * ptr,
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b,
const v_reg< _Tp, n > & c,
hal::StoreMode = hal::STORE_UNALIGNED )

#include <opencv2/core/hal/intrin_cpp.hpp>

Interleave and store (3 channels)

Interleave and store data from 3 registers to memory. Scheme:

{A1 A2 ...}, {B1 B2 ...}, {C1 C2 ...} ==> {A1 B1 C1 A2 B2 C2 ...}

For all types except 64-bit.

Here is the call graph for this function:

cv::v_store_interleave Node1 cv::v_store_interleave Node2 cv::isAligned Node1->Node2

cv::v_store_interleave Node1 cv::v_store_interleave Node2 cv::isAligned Node1->Node2

v_store_interleave() [3/3]#

template<typename _Tp, int n>
inline void cv::v_store_interleave(
_Tp * ptr,
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b,
hal::StoreMode = hal::STORE_UNALIGNED )

#include <opencv2/core/hal/intrin_cpp.hpp>

Interleave and store (2 channels)

Interleave and store data from 2 registers to memory. Scheme:

{A1 A2 ...}, {B1 B2 ...} ==> {A1 B1 A2 B2 ...}

For all types except 64-bit.

Here is the call graph for this function:

cv::v_store_interleave Node1 cv::v_store_interleave Node2 cv::isAligned Node1->Node2

cv::v_store_interleave Node1 cv::v_store_interleave Node2 cv::isAligned Node1->Node2

v_store_low()#

template<typename _Tp, int n>
inline void cv::v_store_low(
_Tp * ptr,
const v_reg< _Tp, n > & a )

#include <opencv2/core/hal/intrin_cpp.hpp>

Store data to memory (lower half)

Store lower half of register contents to memory. Scheme:

REG {A B C D} ==> MEM {A B}

Here is the call graph for this function:

cv::v_store_low Node1 cv::v_store_low Node2 cv::isAligned Node1->Node2

cv::v_store_low Node1 cv::v_store_low Node2 cv::isAligned Node1->Node2

v_sub()#

template<typename _Tp, int n>
v_reg< _Tp, n > cv::v_sub(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Subtract values.

For all types.

v_transpose4x4()#

template<typename _Tp, int n>
inline void cv::v_transpose4x4(
v_reg< _Tp, n > & a0,
const v_reg< _Tp, n > & a1,
const v_reg< _Tp, n > & a2,
const v_reg< _Tp, n > & a3,
v_reg< _Tp, n > & b0,
v_reg< _Tp, n > & b1,
v_reg< _Tp, n > & b2,
v_reg< _Tp, n > & b3 )

#include <opencv2/core/hal/intrin_cpp.hpp>

Transpose 4x4 matrix.

Scheme:

a0  {A1 A2 A3 A4}
a1  {B1 B2 B3 B4}
a2  {C1 C2 C3 C4}
## a3  {D1 D2 D3 D4}
b0  {A1 B1 C1 D1}
b1  {A2 B2 C2 D2}
b2  {A3 B3 C3 D3}
b3  {A4 B4 C4 D4}

v_trunc() [1/2]#

template<int n>
inline v_reg< int, n *2 > cv::v_trunc(const v_reg< double, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.

v_trunc() [2/2]#

template<int n>
inline v_reg< int, n > cv::v_trunc(const v_reg< float, n > & a)

#include <opencv2/core/hal/intrin_cpp.hpp>

Truncate elements.

Truncate each value. Input type is float vector ==> output type is int vector.

Note

Only for floating point types.

v_xor()#

template<typename _Tp, int n>
v_reg< _Tp, n > cv::v_xor(
const v_reg< _Tp, n > & a,
const v_reg< _Tp, n > & b )

#include <opencv2/core/hal/intrin_cpp.hpp>

Bitwise XOR.

Only for integer types.

v_zip()#

template<typename _Tp, int n>
inline void cv::v_zip(
const v_reg< _Tp, n > & a0,
const v_reg< _Tp, n > & a1,
v_reg< _Tp, n > & b0,
v_reg< _Tp, n > & b1 )

#include <opencv2/core/hal/intrin_cpp.hpp>

Interleave two vectors.

Scheme:

  {A1 A2 A3 A4}
  {B1 B2 B3 B4}
---------------
  {A1 B1 A2 B2} and {A3 B3 A4 B4}

For all types except 64-bit.

Macro Definition Documentation#

OPENCV_HAL_HAVE_PACK_STORE_BFLOAT16#

#define OPENCV_HAL_HAVE_PACK_STORE_BFLOAT16

#include <opencv2/core/hal/intrin_cpp.hpp>

Value:

1

OPENCV_HAL_MATH_HAVE_EXP#

#define OPENCV_HAL_MATH_HAVE_EXP

#include <opencv2/core/hal/intrin_cpp.hpp>

Value:

1

Square root of elements.

Only for floating point types.

Exponential \( e^x \) of elements

Only for floating point types. Core implementation steps:

  1. Decompose Input: Convert the input to \( 2^{x \cdot \log_2e} \) and split its exponential into integer and fractional parts: \( x \cdot \log_2e = n + f \), where \( n \) is the integer part and \( f \) is the fractional part.

  2. Compute \( 2^n \): Calculated by shifting the bits.

  3. Adjust Fractional Part: Compute \( f \cdot \ln2 \) to convert the fractional part to base \( e \). \( C1 \) and \( C2 \) are used to adjust the fractional part.

  4. Polynomial Approximation for \( e^{f \cdot \ln2} \): The closer the fractional part is to 0, the more accurate the result.

  • For float16 and float32, use a Taylor Series with 6 terms.

  • For float64, use Pade Polynomials Approximation with 4 terms.

  1. Combine Results: Multiply the two parts together to get the final result: \( e^x = 2^n \cdot e^{f \cdot \ln2} \).

Note

The precision of the calculation depends on the implementation and the data type of the input vector.