# Simplest way to use SIMD for basic float multiplication/addition?

**URL:** <https://forum.juce.com/t/simplest-way-to-use-simd-for-basic-float-multiplication-addition/59617>\
**Category:** General JUCE discussion\
**Created:** [January 22, 2024, 8:04am UTC](https://forum.juce.com/t/simplest-way-to-use-simd-for-basic-float-multiplication-addition/59617 "2024-01-22T08:04:28Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![mikej](https://avatars.discourse-cdn.com/v4/letter/m/a9a28c/32.png) [@mikej](https://forum.juce.com/u/mikej)\
**Post date:** [January 22, 2024, 8:04am UTC](https://forum.juce.com/t/simplest-way-to-use-simd-for-basic-float-multiplication-addition/59617/1 "2024-01-22T08:04:28Z")

</div>

I read the tutorial [here](https://docs.juce.com/master/tutorial_simd_register_optimisation.html) but I unfortunately couldn’t get anything usable out of it.

I am trying to just simply improve the optimization of one of my more computationally expensive synths. I must do an enormous amount of multiplications/additions per sample as it is doing some advanced physical modeling.

As I understand it, SIMD allows you to add/multiply up to 4 floats or doubles for the cost of one, right? Is it system dependent how many you can do at once?

What is the absolute simplest code that might allow me to do this hypothetically with SIMD (random data here for example):

```auto
//ALREADY EXISTING DATA IN VARIABLES TO WORK WITH
float set1A = 20;
float set1B = 10;
float set1C = 2;
float set1D = 8;

float set2A = 5;
float set2B = 30;
float set2C = 80;
float set2D = 6;

//NEED TO REPLACE WITH SIMD TO GO FASTER AND ASSIGN RESULTS TO THOSE VARIABLE NAMES:
float resultA = set1A * set2A;
float resultB = set1B * set2B;
float resultC = set1C * set2C;
float resultD = set1D * set2D;

```

Surely there is some easy way to take advantage of the SIMD function to replace the multiplication operations and not make life too hard?

For example, according to [this](https://learn.microsoft.com/en-us/dotnet/standard/simd) in .NET it would be as simple as putting the data into a Vector4 and using the dot function, followed by reassigning your output variables from that:

```auto
Vector4 set1 = new Vector4(set1A, set1B, set1C, set1D);
Vector4 set2 = new Vector4(set2A, set2B, set2C, set2D);
Vector4 result = Vector4.Dot(set1,set2);

float resultA = result.W;
float resultB = result.X;
float resultC = result.Y;
float resultD = result.Z;

```

This is obvious, simple, easy, and intuitive. Whereas I can’t make any sense of the JUCE equivalent, if one exists.

If the JUCE method is inherently cumbersome, I wonder if it is worth implementing a third party library like [this one](https://github.com/CRefice/ssm) or is there any other you would recommend?

Thanks for any help.

---

<div class="post-metadata">

**Author:** ![kerfuffle](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/kerfuffle/32/11805_2.png) [@kerfuffle](https://forum.juce.com/u/kerfuffle)\
**Post date:** [January 22, 2024, 9:35am UTC](https://forum.juce.com/t/simplest-way-to-use-simd-for-basic-float-multiplication-addition/59617/2 "2024-01-22T09:35:11Z")

</div>

The JUCE equivalent is `juce::SIMDRegister<float>` instead of your `Vector4`. Instead of `Vector4.Dot()` you simply keep using the `*` operator.

And yes, it’s system-dependent how many floats / doubles you can do at once. There are different instruction sets (SSE, AVX, etc for Intel, Neon for Arm). But a SIMD abstraction like JUCE’s `SIMDRegister` hides that from you.

---

<div class="post-metadata">

**Author:** ![mikej](https://avatars.discourse-cdn.com/v4/letter/m/a9a28c/32.png) [@mikej](https://forum.juce.com/u/mikej)\
**Post date:** [January 22, 2024, 9:47am UTC](https://forum.juce.com/t/simplest-way-to-use-simd-for-basic-float-multiplication-addition/59617/3 "2024-01-22T09:47:29Z")

</div>

> [@mikej](#):
>
> h implementing a third party library like [this one](https://github.com/CRefice/ssm) or is there any other you would recommend?

Thanks. I didn’t realize SIMDRegister like that has 4 floats in it. Based on your tip, I found a simple example of how to use it here:

> [@SIMDRegister - How do I do the equivalent of](https://forum.juce.com/t/simdregister-how-do-i-do-the-equivalent-of/28188/6):
>
> Just for the record, although people are tolerent of quite messy looking SIMD code because so much of it looks like that, it doesn’t have to be so bad if you use a more modern style, e.g. the example above using SIMDFloat = SIMDRegister\<float\>; alignas (16) float vraw[] = { 0.0f, 2.2f, 1.3f, 19.9f }; SIMDFloat v = SIMDFloat::fromRawArray (vraw); SIMDFloat p (2.3f); SIMDFloat u = v + p; alignas (16) float eraw[4]; u.copyToRawArray (eraw); DBG (eraw[1]); could be written as simply as this: auto…

I will try that soon.

The funny thing about all this is it requires you to:

- create a new (aligned) array of the primitives (floats)
- put that array into a new class to get it into the SIMD format.
- do this twice if you need to multiply two sets of data
- multiply the SIMD types (only step saving operations)
- create a new (aligned) array to copy the data to or run a get function for each

It seems surprising to me this is any better than just multiplying but I guess multiplying is still much more expensive than shuffling all these primitives around.

Also, in the example given, is the `alignas (16)` necessary? Eg.

```auto
alignas (16) float eraw[4];
u.copyToRawArray (eraw);

```

What would happen if you didn’t have the `alignas`? Would it just fail?

I notice Jules’ code skips all this completely. Is it likely less/more or the same efficiency to use:

```auto
float val0 = u.get(0);
float val1 = u.get(1);
float val2 = u.get(2);
float val3 = u.get(3);

```

Thanks for any further thoughts.

---

<div class="post-metadata">

**Author:** ![kerfuffle](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/kerfuffle/32/11805_2.png) [@kerfuffle](https://forum.juce.com/u/kerfuffle)\
**Post date:** [January 22, 2024, 10:31am UTC](https://forum.juce.com/t/simplest-way-to-use-simd-for-basic-float-multiplication-addition/59617/4 "2024-01-22T10:31:05Z")

</div>

IIRC, if you do an aligned load on an unaligned array, then it will either crash or be slower than it needs to be, depending on the CPU that’s being used.

If memory is allocated on the heap it might already be properly aligned (again depending on the CPU) but on the stack there’s no guarantee how your local variables are aligned, hence the `alignas`.

Note that this same shuffling of data also happens in the .NET Vector4 classes, it’s just less obvious.

---

<div class="post-metadata">

**Author:** ![mikej](https://avatars.discourse-cdn.com/v4/letter/m/a9a28c/32.png) [@mikej](https://forum.juce.com/u/mikej)\
**Post date:** [January 22, 2024, 6:24pm UTC](https://forum.juce.com/t/simplest-way-to-use-simd-for-basic-float-multiplication-addition/59617/5 "2024-01-22T18:24:44Z")

</div>

Thanks. I guess more specifically then I’m wondering about these two methods:

**Option 1:**

```auto
using SIMDFloat = SIMDRegister<float>;
alignas (16) float vraw[] = { 0.0f, 2.2f, 1.3f, 19.9f };
SIMDFloat v = SIMDFloat::fromRawArray (vraw);
SIMDFloat p (2.3f);
SIMDFloat u = v + p;
alignas (16) float eraw[4];
u.copyToRawArray (eraw);

```

**Option 2:**

```auto
auto v = juce::dsp::SIMDRegister<float>::fromNative ({ 0.0f, 2.2f, 1.3f, 19.9f });
auto u = v + 2.3f;
DBG (u.get(1));

```

In #1, we align the array before putting it in, and receive it out as an aligned array. I must presume then that `fromNative` creates an aligned array out of the nonaligned data supplied internally then?

And the `get` function either gets it from an internally aligned basic array that is again generated or from the `SIMDRegister` directly perhaps?

Probably these must be equivalent also in some way. Thanks agin.

---

<div class="post-metadata">

**Author:** ![kerfuffle](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.juce.com/kerfuffle/32/11805_2.png) [@kerfuffle](https://forum.juce.com/u/kerfuffle)\
**Post date:** [January 22, 2024, 7:21pm UTC](https://forum.juce.com/t/simplest-way-to-use-simd-for-basic-float-multiplication-addition/59617/6 "2024-01-22T19:21:38Z")

</div>

`fromNative ({ 0.0f, 2.2f, 1.3f, 19.9f })` directly creates a `__m128` object (on Intel SSE), which is automatically aligned already.

`fromRawArray()` does a `load` operation of some memory array, which is not necessarily aligned, into a `__m128` object.

It might be useful to learn how to use compiler intrinsics for SIMD directly, just so you get a sense of how this works under the hood.
