Peak local-inference fork: Qwen3.8-27B @ 196K + native MTP on single 16GB GPU (RTX 5070 Ti). Gated benchmark suite (needle/toolcall/coherence/speed). den-legacy branch holds prior engine.
-
Updated
Sep 5, 2026 - C++
Peak local-inference fork: Qwen3.8-27B @ 196K + native MTP on single 16GB GPU (RTX 5070 Ti). Gated benchmark suite (needle/toolcall/coherence/speed). den-legacy branch holds prior engine.
To associate your repository with the nvfp4-native topic, visit your repo's landing page and select "manage topics."