QWEN3.5-2B · WEBGPU

The first hybrid gated-DeltaNet + gated-attention LLM in the browser —
hand-written WGSL kernels only, no ONNX, no transformers.js.
~82 tok/s decode · ~475 tok/s prefill (Apple M4) · GPTQ int4, 1.07 GB · Chrome/Edge 125+