# SIMDRegister - How do I do the equivalent of

**URL:** <https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188>\
**Category:** General JUCE discussion\
**Created:** [June 19, 2018, 2:52pm UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188 "2018-06-19T14:52:25Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![nammick](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/nammick/32/2169_2.png) [@nammick](https://forum.juce.com/u/nammick)\
**Post date:** [June 19, 2018, 2:52pm UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/1 "2018-06-19T14:52:25Z")

</div>

I have got my head around Intel Intrinsics but would like to how to do the equivalent of simple operations with SIMDRegister

Take:

```auto
    __m128 v = _mm_setr_ps(0,2.2,1.3,19.9);
    __m128 p = _mm_set1_ps(2.3);
    __m128 u = _mm_add_ps(v , p);
    
    float e[4];
    _mm_store_ps(e,u);
    
    DBG(e[1]);

```

How would I replicate with SIMDRegister?

---

<div class="post-metadata">

**Author:** ![ujam](https://avatars.discourse-cdn.com/v4/letter/u/2bfe46/32.png) [@ujam](https://forum.juce.com/u/ujam)\
**Post date:** [June 20, 2018, 7:00am UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/2 "2018-06-20T07:00:07Z")

</div>

Not the answer to your question, but I think you either need to force good alignment of e[] or use \_mm\_storeu\_ps (store unaligned)

---

<div class="post-metadata">

**Author:** ![nammick](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/nammick/32/2169_2.png) [@nammick](https://forum.juce.com/u/nammick)\
**Post date:** [June 20, 2018, 7:32am UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/3 "2018-06-20T07:32:17Z")

</div>

In code I’m pulling stuff in from a float array and initialising it with align as

```auto
alignas(32) float * samples = nullptr;

```

Though from reading up the bleeding edge case is

```auto
samples = static_cast<float*>(std::aligned_alloc(32, numChannels * numDomains * numSamples * 4));

```

Which I may also start to use.

I have played around with creating wrapper class more inline with the context I want to use but if juce handles it out the box then no point reinventing wheel. I just haven’t got my head around the underlying structure and aligned or not the above is just for a like for like context.

Checking memory alignment I am using this assertion

```auto
jassert((unsigned long)std::addressof(samples) % 32 == 0);

```

---

<div class="post-metadata">

**Author:** ![martinrobinson-2](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/martinrobinson-2/32/630_2.png) [@martinrobinson-2](https://forum.juce.com/u/martinrobinson-2)\
**Post date:** [June 20, 2018, 8:04am UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/4 "2018-06-20T08:04:18Z")

</div>

Something like this:

```
using SIMDFloat = SIMDRegister<float>;

alignas (16) float vraw[] = { 0.0f, 2.2f, 1.3f, 19.9f };
 
SIMDFloat v = SIMDFloat::fromRawArray (vraw);
SIMDFloat p (2.3f);
SIMDFloat u = v + p;

alignas (16) float eraw[4];

u.copyToRawArray (eraw);

DBG (eraw[1]);
```

---

<div class="post-metadata">

**Author:** ![fr810](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/fr810/32/841_2.png) [@fr810](https://forum.juce.com/u/fr810)\
**Post date:** [June 20, 2018, 8:08am UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/5 "2018-06-20T08:08:32Z")

</div>

Depending on which compiler you use, you can also specify a native SIMD literal with an initialiser list. This allows you to do the following:

```
using SIMDFloat = SIMDRegister<float>;

SIMDFloat v = SIMDFloat::fromNative ({ 0.0f, 2.2f, 1.3f, 19.9f });

```

Sadly, not all compilers support this.

---

<div class="post-metadata">

**Author:** ![jules](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@jules](https://forum.juce.com/u/jules)\
**Post date:** [June 20, 2018, 8:22am UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/6 "2018-06-20T08:22:04Z")

</div>

Just for the record, although people are tolerent of quite messy looking SIMD code because so much of it looks like that, it doesn’t have to be so bad if you use a more modern style, e.g. the example above

```auto
using SIMDFloat = SIMDRegister<float>;
alignas (16) float vraw[] = { 0.0f, 2.2f, 1.3f, 19.9f };
SIMDFloat v = SIMDFloat::fromRawArray (vraw);
SIMDFloat p (2.3f);
SIMDFloat u = v + p;
alignas (16) float eraw[4];
u.copyToRawArray (eraw);
DBG (eraw[1]);

```

could be written as simply as this:

```auto
auto v = juce::dsp::SIMDRegister<float>::fromNative ({ 0.0f, 2.2f, 1.3f, 19.9f });
auto u = v + 2.3f;
DBG (u.get(1));

```

…or even as a one-liner:

```auto
DBG ((juce::dsp::SIMDRegister<float>::fromNative ({ 0.0f, 2.2f, 1.3f, 19.9f })
       + 2.3f).get(1));

```

(although that’s a bit too terse for even my taste…)

---

<div class="post-metadata">

**Author:** ![martinrobinson-2](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/martinrobinson-2/32/630_2.png) [@martinrobinson-2](https://forum.juce.com/u/martinrobinson-2)\
**Post date:** [June 20, 2018, 8:44am UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/7 "2018-06-20T08:44:39Z")

</div>

Or even better

DBG (4.5f);

😜

---

<div class="post-metadata">

**Author:** ![nammick](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/nammick/32/2169_2.png) [@nammick](https://forum.juce.com/u/nammick)\
**Post date:** [June 20, 2018, 8:52am UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/8 "2018-06-20T08:52:50Z")

</div>

Ha… (t’was supposed to mundane example)

Cheers for the examples. I was trying to draw some parallels from the XSIMD lib which makes sense to me. I assume I can use SIMDFloat 8 floats AVX aswell.

Thanks again

---

<div class="post-metadata">

**Author:** ![fr810](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/fr810/32/841_2.png) [@fr810](https://forum.juce.com/u/fr810)\
**Post date:** [June 20, 2018, 8:57am UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/9 "2018-06-20T08:57:17Z")

</div>

> [@IanKnowles](#):
>
> I assume I can use SIMDFloat 8 floats AVX aswell.

Yes you can you just need to enable avx2 at compile-time (for example `-mavx2` for clang and gcc).

---

<div class="post-metadata">

**Author:** ![nammick](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/nammick/32/2169_2.png) [@nammick](https://forum.juce.com/u/nammick)\
**Post date:** [June 20, 2018, 8:57am UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/10 "2018-06-20T08:57:35Z")

</div>

is there any cost to accessing with .get(x)?

I had wrapped something like that up myself but handy that its in there

---

<div class="post-metadata">

**Author:** ![jules](https://avatars.discourse-cdn.com/v4/letter/j/41988e/32.png) [@jules](https://forum.juce.com/u/jules)\
**Post date:** [June 20, 2018, 8:58am UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/11 "2018-06-20T08:58:16Z")

</div>

There’s always _some_ cost, but I assume it’ll be pretty low.

---

<div class="post-metadata">

**Author:** ![fr810](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/fr810/32/841_2.png) [@fr810](https://forum.juce.com/u/fr810)\
**Post date:** [June 20, 2018, 8:58am UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/12 "2018-06-20T08:58:45Z")

</div>

It will boil down to a single machine instruction in release mode.

---

<div class="post-metadata">

**Author:** ![ncthom](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/ncthom/32/3721_2.png) [@ncthom](https://forum.juce.com/u/ncthom)\
**Post date:** [June 22, 2018, 6:29pm UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/13 "2018-06-22T18:29:29Z")

</div>

This thread has been really illuminating for me; I was very confused about the SIMDRegister class before this, so thank you! I’d like to tack on another “How do I do X” on the thread since it seems the previous question has an answer:

```auto
SIMDRegister<float> WavetableOscillator::tick(SIMDRegister<float> phase) {
   int index = static_cast<int>(phase * kTableSize)
   int rightIndex = (index + 1 ) & kTableMask;
   float frac = phase * (float) kTableSize - (float) index;
   return lerp(frac, m_table[index], m_table[rightIndex]);
}

```

This is a contrived example, but the point stands: the line `int index = static_cast<int>(phase * kTableSize)` will fail to compile because you can’t cast an `__mm128` to `int`, for example. But if I have this sort of procedure where I want to cast some float type to int for maybe a table lookup with linear interpolation, can it still be vectorized with `SIMDRegister`?

---

<div class="post-metadata">

**Author:** ![nammick](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/nammick/32/2169_2.png) [@nammick](https://forum.juce.com/u/nammick)\
**Post date:** [June 22, 2018, 7:44pm UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/14 "2018-06-22T19:44:09Z")

</div>

My heads been in the vectorising space the last week also trying to implement some really heavy Poly and Linear phase filter banks.

From what I can you don’t want to mix operations SIMDRegister and SIMDRegister with operators as it will be more costly than just plain code. There’s no functions in intel intrinsics to do this + I don’t think you can cast one type to another.

Is kTableSize just an int as in the code there seems to be no benefit to having phase as a SIMD object. I know that SIMDRegister has got the get overload and the [] shorthand.

```auto
int index = static_cast<int>(phase[0] * kTableSize);
// or
int index = static_cast<int>(phase.get(0) * kTableSize);

```

would work…

Other thing I did note aswell using juce SIMD classes compared to others was that

```auto
juce::dsp::SIMDRegister<float> f(2.0);
auto g = f * 8.0f; // works
auto h = 8.0f * f; // doesn't

```

Would be good to have operators for both left hand and right hand side as it can take a little rearranging to get code in the right form which may not be as easy to understand what a formula or also does.

I will let Fabian or Jules answer your question as I’m only levelled up to about a green belt with SIMD and still got quite a bit to learn.

---

<div class="post-metadata">

**Author:** ![chrisboy2000](https://avatars.discourse-cdn.com/v4/letter/c/a3d4f5/32.png) [@chrisboy2000](https://forum.juce.com/u/chrisboy2000)\
**Post date:** [June 22, 2018, 8:24pm UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/15 "2018-06-22T20:24:53Z")

</div>

I also made the experience that you need at least a black belt in SIMD tuning to get stuff like linear interpolation faster than it’s naive implementation.

The thing with linear interpolation is that it can have any stride factor resulting in unpredictable memory access, which is the bottleneck. The raw mathematics are so trivial that it doesn’t matter whether they are SIMDed or not.

But I also started using this class recently and I got the best performance increases by replacing repetitive calls to FloatVectorOperations functions with tight loops that operate on the hot data SIMD-style:

```cpp
// This
FloatVectorOperations::multiply(data, 2.0f, numSamples);
FloatVectorOperations::add(data, otherData, numSamples);
FloatVectorOperations::multiply(data, -1.0f, numSamples);

// becomes something like this
using SSEType = dsp::SIMDRegister<float>;

int numLoop = numSamples / (SSEType::RegisterSize);

while(--numLoop >= 0)
{
    auto a = SSEType::fromRawArray(data);
    auto b = SSEType::fromRawArray(otherData);
    a *= b * SSEType::expand(-1,0f);
    a.copyToRawArray(data);
    data += SSEType::RegisterSize;
    otherData += SSEType::RegisterSize;
}

```

Another thing I noticed regarding the SIMDRegister class is that it requires the AVX2 compiler flag in order to use the AVX register size. This is a pretty tight requirement since they are a lot of CPUs that don’t have AVX2 (I am typing this without AVX for example), but AVX should be the lowest common denominator by now (I think all CPUs since 2011 support it and who uses anything older for serious audio stuff).

Are there some things missing in the AVX instruction set or is there another reason for this decision?

```cpp
#ifdef __AVX2__
#include "native/juce_avx_SIMDNativeOps.h"
#else

```

---

<div class="post-metadata">

**Author:** ![ncthom](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/ncthom/32/3721_2.png) [@ncthom](https://forum.juce.com/u/ncthom)\
**Post date:** [June 22, 2018, 8:54pm UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/16 "2018-06-22T20:54:06Z")

</div>

> [@IanKnowles](#):
>
> Is kTableSize just an int as in the code there seems to be no benefit to having phase as a SIMD object. I know that SIMDRegister has got the get overload and the shorthand.
> 
> ```auto
> int index = static_cast<int>(phase[0] * kTableSize);
> // or
> int index = static_cast<int>(phase.get(0) * kTableSize);
> 
> ```

So `kTableSize` in this example is just an int. Essentially what I’m trying to do (which is how I got to this example) is to process LFO modulators in parallel (i.e. each channel in a stereo channel pair can have its own LFO modulator). I’d like to interleave the stereo buffer and then operate on both the L/R channels via SIMDRegister. I think most of my code should be set up well for that, but I have 2 cases like the above where I need the float type to map to an array lookup.

---

<div class="post-metadata">

**Author:** ![nammick](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/nammick/32/2169_2.png) [@nammick](https://forum.juce.com/u/nammick)\
**Post date:** [June 23, 2018, 6:43am UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/17 "2018-06-23T06:43:01Z")

</div>

I’m not sure you can use SIMD in this way. If you using SIMD object as a LUT then you can only have one lookup per SIMD object. You can’t index individual rows across multiple SIMD objects and mix them into one object in a single operation.

Now I get what you doing though I would say just bite the bullet and process each channel seperately and feed in each value to two new SIMD objects which you can then lerp between as this is going to be the most expensive computationally anyway.

```auto
   auto index = (phase * kTableSize)
   auto rightIndex = (index + 1 ) & kTableMask;

   alignas(16) float frac[4];
   frac[0] = phase * (float) kTableSize - (float) index[0];
   ...
   frac[3] = phase * (float) kTableSize - (float) index[3];

   auto fracSimd = SIMDRegister<float>::fromRawArray(frac);

// do the same with m_table
   alignas(16) float m_table_left[4];
alignas(16) float m_table_right[4];
...
... // fill table values
...
auto m_table_left_Simd = SIMDRegister<float>::fromRawArray(m_table_left);
auto m_table_right_Simd = SIMDRegister<float>::fromRawArray(m_table_right);

   return lerp(frac, m_table_left_Simd, m_table_right_Simd);

```

Then profile it to make sure your getting the boost

---

<div class="post-metadata">

**Author:** ![martinrobinson-2](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/martinrobinson-2/32/630_2.png) [@martinrobinson-2](https://forum.juce.com/u/martinrobinson-2)\
**Post date:** [June 23, 2018, 9:00am UTC](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/18 "2018-06-23T09:00:39Z")

</div>

Have you seen Angus’ talk from ADC2017?
