SIMD helper types, conversions, and operations for Nim with AVX2/SSE/NEON backends.
simd/: core SIMD types, conversions, and operations (including generic traits/helpers and SIMD iterators with masks).matrices/: matrix-oriented SIMD helpers.sequences/: SIMD-aware sequence utilities, including GF(256) field arithmetic.isa/: instruction-set declarations imported by the modules that need them.evaluation/tests/: unit tests for core behavior.
There is no build-flag facade and no when switch per function. Two ordinary
Nim mechanisms do the whole job:
The compiler drops what you never call. Nim emits no C for an unreferenced proc, so importing the package does not drag it into your binary:
| what you import and use | binary |
|---|---|
| nothing (baseline program) | 42,248 B |
simd_nexus/sequences/gf256, one call |
43,416 B |
all of simd_nexus, one gf256 call |
43,808 B |
all of simd_nexus, gf256 + SIMD ops + streams + GPU |
51,664 B |
Using one function out of the package costs about 1.5 KB, not the package.
Each module declares the instruction sets its own intrinsics need, by
importing isa/x86 (SSSE3, SSE4.1) or isa/x86_avx2 (AVX2, which implies the
rest). Because those flags travel with the import, the granularity is automatic:
import simd_nexus/sequences/gf256 -> -mssse3 -msse4.1
import simd_nexus -> -mssse3 -msse4.1 -mavx2
Reaching for one 128-bit byte-stream helper does not hand you an AVX2-only
binary as a side effect of touching the package. Putting these flags in the
repo's nim.cfg instead would only work while SIMD-Nexus is the project being
built — every outside importer got a gcc target specific option mismatch.
nimble testConsumer compiles throwaway consumers from outside the repo so
that failure mode cannot come back unnoticed.
sequences/gf256 multiplies whole byte buffers by a field coefficient, which is
the inner loop of every Reed-Solomon or network coding scheme. A coefficient is
split into two 16-entry tables once, then each byte costs two table lookups and
one xor — and a single SIMD shuffle does 16 (SSSE3, NEON) or 32 (AVX2) lanes of
that at a time.
import simd_nexus
var
parity = newSeq[uint8](1024)
data = newSeq[uint8](1024)
t = gf256Tables(0x1b'u8) # build once
gf256MulAdd(parity, data, t) # parity[i] ^= data[i] * 0x1b
gf256AddInto(parity, data) # parity[i] ^= data[i]The lane width is picked at compile time: AVX2 under -d:simdNexusEnableAvx2,
SSSE3 on any amd64/i386 build, vqtbl1q_u8 on aarch64, and a plain byte
loop everywhere else. Every backend is checked against the scalar reference at
every buffer length from 0 to 200, so tails and lane boundaries stay honest.
import simd_nexus
let x: i32x4 = [1'u32, 2, 3, 4].asM128i()
let y = x + x
echo y[0]
let v = loadU32x4[M128i]([1'u32, 2'u32, 3'u32, 4'u32])
let r = rotl32(v, 8)
echo storeU32x4[M128i](r)[0]
for (i, mask) in simdRangeU32[M128i](0'u32, 6):
let idxs = storeU32x4[M128i](i)
let masks = storeU32x4[M128i](mask)
echo idxs, " ", masksSIMD-Nexus now exposes a higher-level GPU dispatch surface that keeps user code
Nim-native. You write procs/operators over GpuArray[T]; SIMD-Nexus builds the
operation graph and dispatches it to the selected device. A CPU fallback device
is always present, so the API is testable without OpenCL drivers.
import simd_nexus
proc myGpuFunc(x, y: GpuArray[float32]): GpuOp[float32] =
result = (x + y)
var
dataA = @[1.0'f32, 2.0'f32, 3.0'f32]
dataB = @[4.0'f32, 5.0'f32, 6.0'f32]
gpu = getGpu()[0]
a = dispatch(dataA, gpu)
b = dispatch(dataB, gpu)
out = dispatch(myGpuFunc(a, b), gpu)
echo out.toSeq()Arrays work the same way as sequences, and can be converted into a chosen GPU element type when needed:
import simd_nexus
var
gpu = getGpu()[0]
raw: array[4, int32] = [1'i32, 2, 3, 4]
asI32 = toGpuArray(raw, gpu)
asF32 = toGpuArrayAs(raw, float32, gpu)
alsoF32 = dispatch(raw, float32, gpu)Supported high-level operations:
| API | Purpose |
|---|---|
getGpu() |
Return OpenCL GPUs when enabled plus a CPU fallback. |
selectGpu(i) |
Select a device by ordinal, with fallback. |
toGpuArray(arrayOrSeq, gpu) |
Move host data into a GpuArray preserving element type. |
toGpuArrayAs(arrayOrSeq, T, gpu) |
Convert host data into GpuArray[T]. |
dispatch(arrayOrSeq, gpu) |
Move host data into a GpuArray. |
dispatch(arrayOrSeq, T, gpu) |
Convert host data into GpuArray[T]. |
dispatch(op, gpu) |
Execute a numeric operation. |
+ - * / |
Elementwise numeric operations over GpuArray. |
scale, relu, sigmoid, tanhAct |
Basic NN/vector transforms. |
dot, sum |
Reductions returning GpuScalar. |
Backend behavior:
GpuArray / GpuOp
|
v
dispatch(op, selectedGpu)
|
+--> OpenCL f32 elementwise path when compiled with -d:simdNexusOpenCL
|
+--> CPU fallback for all supported numeric ops
The OpenCL path is intentionally generated from Nim-side operation kinds for
basic numeric kernels. Users do not write OpenCL C for +, -, *, /,
scale, relu, sigmoid, or tanhAct.
Dense layers, neural-network orchestration, and particle swarm optimization live
in Lineage-GeneticProgramming. SIMD-Nexus stays focused on device dispatch,
type conversion, primitive vector ops, and low-level kernels.
- Prefer clarity and modularity over micro-optimizations.
- Keep modules layered (helpers/types at top levels; deeper modules depend upward).
- Avoid nested functions; build helpers and call them from high-level procs.
- Use concise parameter names based on meaning; document each parameter with
##. - Declare variables at the top of procs; initialize immediately when possible.
- Add SIMD helpers in shared locations to avoid near-duplicate implementations.
- Keep
agents/PROGRESS.mdupdated with commit message, features, and recent work notes. - Generic nimble tasks (autopush, switch, applyNightly, updateSubmodules, …) come from the shared Nimble-Tasks submodule; only repo-specific tasks live in
simd_nexus.nimble. - Exclude
builds/and*.exein.gitignore.
- Keep functions short and avoid nesting. Prefer small helpers that are called by high-level procs.
Example:
proc myFunc1(): void =
...
proc myFunc2(): void =
...
proc highLevelFunc(): int =
myFunc1()
myFunc2()- Call functions as
funcX(param1, param2)orparam1.funcX(param2). - Avoid the
funcX:block call syntax unless absolutely needed and explain why in a comment above it.
- Parameter names use the first letter of what they represent.
- Explain parameter meaning directly below the function declaration with
##doc comments. - Arrays, sequences, openArrays, and tables end with
s. - State objects that will be mutated use
s(ors0,s1,s2, ...). - Math-heavy functions use
a,b,corx,y,z(thenx1,x2, ...). - Arrays/lists in math functions use uppercase letters like
A,B,CorX,Y,Z. tis reserved for temporary variables inside functions.i,j,kare indices;l,m,nare lengths.- Use
whilefor complex loops andforfor simple one-call loops. - If a function has only one parameter, you may use its first letter unless it collides with index identifiers.
- It is OK to assign to a temporary variable and set
resultat the end for clarity.
Example:
proc myProc(a, b: uint8): uint8 =
var
veryImportantNumber: uint8
veryImportantNumber = callSomeOtherFunc(a, b)
veryImportantNumber = veryImportantNumber + callYetAnotherFunc(a)
result = veryImportantNumber- Declare variables at the start of the proc, not mid-block or inside loops.
- Always indent
var,let,const, andtypewhen declaring multiple values. - Use
constwhenever possible; otherwise usevar, and assign immediately if the value is known.
- The actual project belongs in
src. Create it if missing. - Submodules can live outside
src. - Every repo keeps its handoff notes in
agents/PROGRESS.mdnext tosrc/. - Every module (
.nimfile) must have a description at the top explaining what it does. - Organize modules by dependency levels (helpers/types at top; deeper modules depend upward).
Example structure:
src/utils.nim
src/types.nim
src/level1/module1.nim
src/level1/module2.nim
src/level1/level2/module3.nim
- If you write three similar helpers across modules, move them into
utilsand overload or use generics (when/case) instead.
- Update the README when making bigger project changes.
- Keep this full conventions section at the bottom of the README.
- Add a
toolsfolder when needed (submodule builders or other pre-compile utilities). - Keep tests below
evaluation/tests/, benchmarks belowevaluation/benchmarks/. - After changing code or dependencies, run tests and fix errors.
- If you need an entirely different project as a dependency, ask before starting a new sibling repo.
- Prefer Nim and nimble only. Do not add Python, bash, or PowerShell build tools.
- The shared cNimWrapper repo can be used to generate bindings for C libraries when needed.
- Generic helper functions may be added to
Fylgia-Utils(https://github.com/siriuslee69/fylgia-utils).
- Do not write pre-compile-time import statements that prevent nimsuggest from checking functions.
- Track current commit message, features planned/implemented/in progress, and recent changes/problems.
- Include tasks for tests and builders.
- Include an
autopushtask that reads the commit message fromagents/PROGRESS.md.
- Add
builds/and*.exeto.gitignore.
- Libraries do not need a frontend (at most a CLI).
- Avoid frontend/backend splits in library repos.
- Symptom:
nimble testornimble buildfails with Nimble metadata write errors under%USERPROFILE%\\.nimble(for examplenimbledata2.json). Workaround: run direct Nim commands such asnim c -r evaluation/tests/test_basic.nimandnim c src/simd_nexus.nim, then record the environment issue inagents/PROGRESS.md. - Symptom: stale
*.exefiles appear insrc/orevaluation/tests/after local runs. Workaround: remove generated binaries before committing;*.exeis intentionally ignored and only Nim sources should be tracked. - Symptom:
nimble autopushuses a generic commit message. Workaround: ensureagents/PROGRESS.mdcontains a line starting withCommit Message:and rerunnimble autopush.