Skip to content

[Bug]: BitNet FFN uses SILU instead of ReLU² — perplexity 99.8 vs 17.1 on every CPU backend #588

Description

@hmeiland

Model: microsoft/BitNet-b1.58-2B-4T (ggml-model-i2_s.gguf, ggml-model-tl1.gguf)
Where: 3rdparty/llama.cpp submodule → src/models/bitnet.cpp (the pinned
isHuangXin/llama.cpp@390c3077), also present in current ggml-org/llama.cpp master.

TL;DR

The BitNet FFN sub-layer applies SILU to the gate activation, but BitNet
b1.58 uses squared ReLU (relu(x)²). This is a one-line graph bug
(LLM_FFN_SILU should be LLM_FFN_RELU_SQR). It does not crash — the model
runs and emits plausible-looking text — but it is numerically mis-calibrated,
which shows up unmistakably in perplexity:

Backend / kernel Pure upstream (SILU) With ReLU² fix Official reference
x86 / I2_S (Intel) 99.8178 ± 0.906 17.1086 ± 0.129 17.1090 ± 0.128
x86 / I2_S (AMD) 99.8178 ± 0.906 (bit-identical to Intel) (identical by construction)
ARM / TL1 78.6877 ± 0.730 14.9641 ± 0.114 (correct band)

The Intel fixed value lands on the official 17.1090 to 4 decimal places
(Δ = 0.0004). The fix restores correctness on two independent kernel families
(scalar I2_S and LUT-based TL1) and across both x86 vendors, confirming this
is a graph-level bug, not a kernel/hardware issue.

Root cause

src/models/bitnet.cpp, in the per-layer FFN build:

        cur = build_ffn(cur,
                model.layers[il].ffn_up,   NULL, model.layers[il].ffn_up_s,
                model.layers[il].ffn_gate, NULL, model.layers[il].ffn_gate_s,
                NULL,                      NULL, NULL,
                NULL,
                LLM_FFN_SILU, LLM_FFN_PAR, il);   // ← BUG: BitNet b1.58 uses ReLU², not SILU
        cb(cur, "ffn_sub_out", il);

The BitNet b1.58 architecture specifies squared ReLU in the FFN
(consistent with the reference model config). SILU here mis-scales activations
into the sub-norm, and the error compounds over 30 layers.

Fix (one line)

-                LLM_FFN_SILU, LLM_FFN_PAR, il);
+                LLM_FFN_RELU_SQR, LLM_FFN_PAR, il);

LLM_FFN_RELU_SQR already exists in the enum (src/llama-graph.h), so this is
a pure activation-op swap with no new code. Throughput is unchanged (activation
op, not GEMM/LUT).

(For the origin fork Eddie-Wang1120/llama.cpp@bitnet, the same bug appears in
the pre-build_ffn monolithic form at llama.cpp:11905 as a literal
cur = ggml_silu(ctx0, cur); — replace with
cur = ggml_sqr(ctx0, ggml_relu(ctx0, cur));.)

Reproduction (clean-room, pure upstream)

Three fresh Azure VMs (Intel x86, AMD x86, ARM Ampere), pure upstream
microsoft/BitNet with recursive submodules, official 2B-4T model, standard
llama-perplexity over the same corpus:

git clone --recursive https://github.com/microsoft/BitNet.git && cd BitNet
python setup_env.py -md models/BitNet-b1.58-2B-4T -q i2_s   # (or -q tl1 on ARM)
./build/bin/llama-perplexity -m models/BitNet-b1.58-2B-4T/ggml-model-i2_s.gguf \
    -f <wikitext-2-raw/wiki.test.raw> -c 2048
# -> PPL ≈ 99.82 (x86 I2_S)  /  ≈ 78.69 (ARM TL1)   [BROKEN]

Apply the one-line fix in 3rdparty/llama.cpp/src/models/bitnet.cpp, rebuild,
re-run:

# -> PPL 17.1086 (x86 I2_S)  /  14.9641 (ARM TL1)   [RESTORED — matches official 17.1090]

Both broken values are far outside the correct band; both fixes land inside it.
The broken PPL differs by kernel (99.82 vs 78.69) because the mis-calibrated
activation interacts with each kernel's numerics differently — but the
brokenness itself is universal and deterministic.

Why this hasn't been pinned before

This is distinct from the existing garbage-output reports:

Unlike those, this bug reproduces on every backend and both x86 vendors,
does not crash or emit a constant token, and only reveals itself as a
perplexity regression (~6× worse) with plausible-but-degraded text. That is
likely why the vague quality complaints — #216 ("local model doesn't have the
same quality as online"), #348, #106, #115 — were never root-caused: nobody
measured perplexity. This bug is a strong candidate explanation for those.

Suggested routing

The fix belongs in the pinned llama.cpp (isHuangXin/llama.cpp
release-bitnet-embedding-0.6b-270m, and canonically ggml-org/llama.cpp
master). Once merged upstream, microsoft/BitNet needs only a
3rdparty/llama.cpp submodule bump. Happy to open the upstream PR (patch
verified to apply clean at the pinned commit) and link it here.

Environment

  • Model: microsoft/BitNet-b1.58-2B-4T official GGUF (I2_S, TL1)
  • llama.cpp submodule: isHuangXin/llama.cpp@390c3077
  • Build: 355c0c4d1 (9916), clang-18, cmake
  • VMs: Intel x86 (AVX-512), AMD x86 (AVX-512), ARM Ampere Altra — westus3
  • Cross-vendor: Intel & AMD produced bit-identical broken PPL (99.8178)

Upstream fix: the canonical fix has been opened against ggml-org/llama.cpp master: ggml-org/llama.cpp#25885 (LLM_FFN_SILULLM_FFN_RELU_SQR in src/models/bitnet.cpp). Recommend re-syncing the pinned llama.cpp submodule once it lands.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions