Arch-Specific SIMD in Go 1.27: Using the archsimd Package
You can pick every intrinsic correctly and still write SIMD that loses to the scalar loop. A bounds check the compiler couldn’t eliminate. A vector parked in a struct and spilled to the stack. A missing CPU feature check that pushes the compiler into conservative lowering. Instruction selection is the easy part.
Go 1.26 shipped an experimental amd64 SIMD API. Go 1.27 extends it to arm64 (NEON) and wasm (WebAssembly 128-bit SIMD), all behind the simd/archsimd package and GOEXPERIMENT=simd. The official documentation on the Go blog covers the API surface. This post is about everything around it.
If you want the portable story first, we covered that in writing portable SIMD in Go with the experimental simd package. archsimd sits one level below simd. It’s the intrinsics layer. You drop down to it when you need an instruction that doesn’t exist on every architecture, or when the portable version would emulate something your CPU does in one shot.
What archsimd gives you
SIMD means one instruction operating on a wide register holding many elements. archsimd models those registers as distinct Go struct types: Float32x4, Int32x8, Uint8x64. The name tells you the element type and the lane count, and the operations are methods on those types.
The naming rule is worth learning once, because it saves you hundreds of doc lookups later. Semantically identical operations get the same method name everywhere:
// All of these are element-wise addition.
// The receiver type selects the operation and vector shape;
// the target and available CPU features determine the instruction.
a.Add(b) // Float32x4
c.Add(d) // Int32x8
e.Add(f) // Uint8x64 on amd64
One method name instead of eighteen. Compare that to _mm512_slli_epi64, which archsimd just calls ShiftAllLeft.
The flip side: when two architectures have instructions that look the same but differ in edge cases, Go gives them different names on purpose. 16-byte table lookup on amd64 (VPSHUFB) zeros the output when the index is negative and wraps modulo 16 otherwise, so it’s called PermuteOrZero. NEON’s VTBL and wasm’s i8x16.swizzle zero the output for any out-of-range index, so those are LookupOrZero. Silently different behaviour is worse than a compile error, and this is one of the better decisions in the package.
Many 256-bit and 512-bit amd64 instructions also work inside 128-bit lanes rather than across the whole register. Those methods carry a Grouped suffix (InterleaveLoGrouped, PermuteOrZeroGrouped), so the lane boundary shows up at the call site instead of hiding in a doc comment.
Loading, storing, and the tail problem
Methods need a receiver, which doesn’t work when you’re creating a vector from a slice. Loads are package-level functions named after the target type:
v := archsimd.LoadFloat32x4(s) // s []float32, needs 4 elements
v.Store(dst) // writes 4 elements back
The Part variants are the ones that earn their keep. Real slices are rarely a multiple of your vector width, and the tail is where SIMD code usually turns into a mess:
v, n := archsimd.LoadFloat32x4Part(s) // zero-fills remaining lanes
// n = how many were real
v.StorePart(dst[:n]) // writes only the valid lanes
StorePart uses the length of its destination slice, so passing dst[:n] is what limits the write to the elements loaded. That pairing removes the scalar cleanup loop you’d otherwise have to write. For fixed-size arrays there’s LoadFloat32x4Array(*[4]float32) and StoreArray, and BroadcastFloat32x4(v float32) splats a scalar across all lanes.
Masks replace branches
You can’t branch per-lane, so comparisons produce a Mask value:
m := x.Greater(y) // Mask32x4
z := x.Masked(m) // zero where m is false
w := x.IfElse(m, y) // x where true, y where false
IfElse replaces Merge from the Go 1.26 experiment. The mask types are deliberately opaque, because the hardware representation varies wildly: AVX-512 has dedicated k registers, AVX2 and NEON use full-width bitmasks, SVE has predicate registers.
The compiler peepholes mask operations into the surrounding instruction. On AVX-512, x.Add(y).Masked(m) becomes a single zero-masked VPADD. On AVX2 or NEON it lowers to a bitwise AND or a blend. The Go source is identical either way, which is the whole point.
Type reinterpretation without the quadratic mess
Go 1.26 had As<Type> methods, one per type pair. That doesn’t scale, and it doesn’t generalise to scalable vectors at all. Go 1.27 replaces it with composable zero-cost conversions:
// x is archsimd.Uint8x16
f := x.ReshapeToUint32s(). // Uint8x16 -> Uint32x4 (same register width)
BitsToFloat32() // Uint32x4 -> Float32x4
Three rules cover everything. ToBits() reinterprets a signed or float vector as unsigned of the same element width. BitsToInt32() / BitsToFloat32() go back. ReshapeToUint<W>s() changes element width within the same register width, and only exists on unsigned vectors, which is what forces you through ToBits() first and keeps the combinatorics linear instead of quadratic.
Three Go patterns that decide your performance
Easy to skip, expensive to get wrong.
Guard everything with a CPU feature check. The compiler doesn’t know what machine your binary will land on. Call an AVX-512 instruction on a CPU without it and you get SIGILL:
if !archsimd.X86.AVX512GFNI() {
slowFallback(dst, src)
return
}
Feature checks double as optimisation hints. Outside an if archsimd.X86.AVX512() block, the compiler has to lower x.Add(y).IfElse(m, z) conservatively as a VPADD plus a VPBLEND. Inside the block, it knows merge-masking exists and fuses both into one instruction. Guarding for safety and guarding for codegen turn out to be the same guard. On arm64, NEON is always available, and archsimd.ARM64.PMULL(), SVE(), SVE2() cover the optional extensions.
Write the loop bound so the prove pass can kill your bounds checks. This matters more than it looks:
// Good: no bounds checks in the loop body.
for i = 0; i < len(src)-v.Len()+1; i += v.Len() {
v = archsimd.LoadUint8x64(src[i : i+v.Len()])
// ...
}
// Bad: i + v.Len() could overflow int, so prove gives up.
for i = 0; i+v.Len() <= len(src); i += v.Len() {
In the second form, i + v.Len() could theoretically overflow to a negative int if len(src) were near math.MaxInt. prove can’t rule that out, so it can’t prove i stays in bounds, so you pay for a bounds check every iteration. Subtracting instead of adding removes the overflow case entirely. The two loops are the same loop; one of them just talks to the compiler.
Avoid putting vectors in structs or arrays on hot paths. Go’s ABI and SSA backend currently place large composite types in memory rather than registers (#24416). A SIMD vector is 16 to 64 bytes, so spilling one to the stack costs far more than spilling an int. That’s why the Go team’s transpose example takes eight separate Int32x8 parameters instead of an [8]archsimd.Int32x8:
func Transpose8(a0, a1, a2, a3, a4, a5, a6, a7 archsimd.Int32x8) (
b0, b1, b2, b3, b4, b5, b6, b7 archsimd.Int32x8)
The signature is unwieldy, but it keeps the vectors in registers. The same concern applies to local variables, not only parameters. Register promotion for composite types is being worked on, but it didn’t land in 1.27, so benchmark before wrapping vectors in composite values on a hot path.
When dropping to archsimd pays off
The clearest case is an instruction with no portable equivalent. On amd64, the GFNI extension exposes GaloisFieldAffineTransform:
func (x Uint8x16) GaloisFieldAffineTransform(A Uint64x2, b uint8) Uint8x16
Each element of A is an 8x8 bit matrix packed row-by-row into a uint64. The result is z[i] = (A[i/8] * x[i]) + b, where multiply is AND and add is XOR. Any permutation of the 8 bits in a byte can be written as such a matrix, which turns one instruction into a general byte-level bit manipulator.
Reversing bit order within every byte, the SIMD version of bits.Reverse8, normally needs nibble lookup tables, masks, shifts and ORs. With a 512-bit GFNI vector it’s a multiply by the anti-diagonal identity matrix 0x8040201008040201, handling 64 bytes per instruction. The portable simd API does not expose that architecture-specific operation, and emulating it would defeat the purpose of using it.
The second case is shape mismatch. amd64 permutation instructions are lane-grouped in a way that doesn’t map cleanly onto a portable intersection API, so an 8x8 transpose written against simd can pick up emulation overhead. Written directly in archsimd with InterleaveLoGrouped, ConcatPermuteScalarsGrouped and ConcatPermute128Scalars, it can map directly to the hardware operations.
Prefer simd unless you need those architecture-specific operations. Dropping to intrinsics costs you a fallback path, a feature check and a second implementation to keep in sync. Types such as Uint8x64 exist only on amd64, so code using them must live in an architecture-specific file (for example, a _amd64.go file) with the fallback in another file. The runtime feature check then chooses between implementations on amd64 CPUs with different instruction support.
Running it
Everything is behind the experiment flag:
GOEXPERIMENT=simd go test simd/archsimd/...
You can cross-test other targets without leaving your machine:
# wasm SIMD via a WASI runtime
GOOS=wasip1 GOARCH=wasm GOEXPERIMENT=simd go test simd/archsimd/...
# amd64 AVX2 on Apple Silicon via Rosetta
GOARCH=amd64 GOEXPERIMENT=simd go test simd/archsimd/...
The parent tracking issue is #73787. For Go 1.28 the team is working on arm64 SVE and SVE2 with width-agnostic scalable types like archsimd.Float32s, plus riscv64, ppc64, s390x and loong64.
Names will keep moving while this is experimental. Merge became IfElse; As<Type> became the ToBits/Reshape family. If you’re writing against it now, keep every archsimd call behind a function boundary you control. Renames are the cheapest breakage to absorb, and the most likely one coming.