Go Blog Details Arch-Specific SIMD for amd64, arm64 and WebAssembly in Go 1.27, Still Behind GOEXPERIMENT=simd
Go 1.27 extends the experimental simd/archsimd package from amd64 to arm64 and WebAssembly, replacing assembly for vector code, per a Go blog post.
Overview
The Go team has published a detailed look at the architecture-specific SIMD API that ships experimentally in Go 1.27. According to the Go blog, “Since Go 1.26, we have supported an experimental amd64 SIMD API. Now in Go 1.27, we support two more architectures: arm64 and wasm.” The post, written by Junyang Shao and David Chase and dated 2 October 2026, says the APIs are enabled by setting GOEXPERIMENT=simd at build time and live in the simd/archsimd package.
SIMD, short for Single Instruction, Multiple Data, lets one instruction operate on several values held in a wide register. The Go blog says that until now, “anyone who wanted to use SIMD in Go had to endure the friction of writing assembly language.” The release of Go 1.27 itself was previously reported by The Machine Herald.
What We Know
Architecture coverage
According to the Go blog, the amd64 support has the largest API surface, covering AVX, AVX2 and many AVX-512 extensions. The arm64 support currently covers NEON, with SVE and some SVE2 described as already on the way, and the wasm support is described as mostly complete. Matrix extensions such as AMX and SME are not supported yet, because the team has not settled on how to represent matrices in Go.
A companion post on platform-independent SIMD explains how the layers relate: it says Go 1.26 introduced a SIMD API for amd64, and Go 1.27 added APIs for arm64 (specifically NEON) and wasm, and that Go 1.27 also introduces a portable, size-agnostic simd package. The archsimd package is the lower-level layer. Per the Go blog, it is “essentially the intrinsics layer in other languages”, and users can reach exotic single-architecture operations by moving from simd to archsimd.
API design
The Go blog says archsimd does not mirror hardware instruction names directly. Instead of the intrinsic _mm512_slli_epi64, the Go API calls the operation ShiftAllLeft. Vector types are distinct structs such as Float32x4, Int32x8 or Uint8x16, with operations defined as methods, and mask types such as Mask32x4 hide the differences in how architectures encode masks. The package supports only fixed-width vector extensions for now; the post says scalable extensions such as arm64 SVE and RVV will have different struct types.
Where instructions look alike but behave differently, the post gives them separate names. Its example is a 16-byte table lookup: the amd64 form is called PermuteOrZero, while the arm64 NEON and wasm equivalents are called LookupOrZero. On amd64, methods that work within 128-bit lanes carry a Grouped suffix, such as InterleaveLoGrouped.
The Go 1.27 API also changes from the Go 1.26 experiment. The blog says the earlier As<Type> reinterpretation methods are replaced by composable conversions such as ToBits() and ReshapeToUint<W>s(), and that x.IfElse(m, y) replaces Merge from Go 1.26.
Guidance for users
The post warns that code must be guarded with runtime feature checks such as archsimd.X86.AVX512GFNI() or archsimd.ARM64.SVE(). Without the check, according to the Go blog, a program may crash with SIGILL on hardware that lacks the instruction. The post adds that the checks also act as compiler hints: inside an archsimd.X86.AVX512() block, the compiler can fuse an add-and-merge sequence into a single merge-masked AVX-512 instruction.
The authors also note a performance limitation. Go’s ABI and SSA backend currently place large composite types in memory rather than registers, a known issue tracked as #24416, so wrapping vectors in large arrays or structs in hot loops can hurt performance. Efforts to improve this have “not yet landed in Go 1.27”, per the post.
To try the package, the post gives the command GOEXPERIMENT=simd go test simd/archsimd/..., and notes that wasm can be cross-tested with GOOS=wasip1 GOARCH=wasm.
What We Don’t Know
- The Go blog does not give a date for when the API will leave experimental status or be available without
GOEXPERIMENT=simd. - The post does not publish benchmark figures comparing
archsimdcode with hand-written assembly. - Support for SVE and SVE2 is described as in progress for Go 1.28 and beyond, but the post gives no release schedule. It also lists
riscv64,ppc64,s390xandloong64as architectures the team plans to add, again without dates.
Analysis
The two blog posts together show Go taking a two-layer approach: a portable simd package for write-once code, and archsimd for operations that exist on only one architecture. By the post’s own account, the goal is to support as many SIMD instructions across architectures as possible. Because the whole effort remains behind a GOEXPERIMENT flag and the API changed between Go 1.26 and 1.27, programs that depend on it should expect further changes. The authors invite feedback on usability and performance, and point to issue #73787 as the parent issue for the project.