Lesson 45 of 55
9 mins readPython SIMD Vectorization & @simd
In Plain English
The @simd macro gives the Julia compiler permission to vectorize inner loops, executing floating-point arithmetic on multiple data items simultaneously in hardware registers.
Deep Dive: How It Works
SIMD: Single Instruction, Multiple Data allows the CPU to process 4 to 8 floating point numbers in a single clock cycle.
@simd for: Allows the compiler to reorder loop operations and emit AVX2/AVX-512 instructions.
@inbounds @simd: Combining @inbounds with @simd produces inner loop assembly with minimal latency.
Core Rules to Remember

Hardware Vector Registers: Execute arithmetic operations on 4-8 floats per clock cycle.

@simd @inbounds Combo: The standard idiom for maximum inner loop performance in numerical algorithms.
Live Interactive Example
Hit Run Code to see it liveSIMD Accelerated Dot Product
Python 3.12
1
2
3
4
5
6
7
8
9
10
11
12
13
Output Console
Click "Run Code" to view the rendered output.
How it works: 1*2 + 2*0.5 + 3*1 + 4*3 = 2 + 1 + 3 + 12 = 18.0.
Your Turn: Micro Challenge
No pressure! Edit the starter code below and test your solution with instant feedback.
Micro Exercise
Sum with @simd Loop
Define `function simd_sum(arr::Vector{Float64}) s = 0.0; @simd for x in arr s += x end; return s end`.
Compute `simd_sum([10.0, 20.0, 30.0])` and print `"Total: $res"`.
1
2
3
4
5
6
7
8
9
10
Sandbox Output
Click "Run & Check" to test your solution.
Finished reading and practicing?
Mark this lesson as completed to update your course progress.