Skip to content

ADFA-5188 | Enable KV cache quantization and flash attention - #76

Open
jatezzz wants to merge 4 commits into
fix/ADFA-5187-dynamic-n-ctxfrom
fix/ADFA-5188-kv-cache-quantization
Open

ADFA-5188 | Enable KV cache quantization and flash attention#76
jatezzz wants to merge 4 commits into
fix/ADFA-5187-dynamic-n-ctxfrom
fix/ADFA-5188-kv-cache-quantization

Conversation

@jatezzz

@jatezzz jatezzz commented Aug 21, 2026

Copy link
Copy Markdown
Contributor

Description

This PR implements KV cache quantization (q8_0) and enables flash attention in the llama.cpp context parameters to drastically improve memory efficiency and generation speed.

  • What: Configures the context parameters type_k and type_v to use GGML_TYPE_Q8_0 and enables flash attention (LLAMA_FLASH_ATTN_TYPE_AUTO).
  • How: These settings are exposed through the configure-before-load pattern. If the native context creation fails (e.g., the model's head width is incompatible with the quantized block size), it gracefully falls back to the previous defaults: f16 KV cache and flash attention disabled.
  • Why: Storing the KV cache as q8_0 halves its byte size, allowing for much longer conversations before context is dropped and drastically reducing mid-generation crashes on 4–6 GB RAM devices. Flash attention mitigates generation latency on longer contexts.

Details

Logs confirming n_ctx initialization, cache type used (q8_0 vs f16), and fallback activations.

Flash Attention Enabled

Screenshot 2026-08-21 at 12 16 48 PM

Q8_0 KV cache

Screenshot 2026-08-21 at 1 01 13 PM

Ticket

ADFA-5188

Observation

This implementation works in tandem with dynamic n_ctx sizing and should be validated alongside ADFA-5187, as the memory measurement relies on both features working concurrently.

#75 Needs to be merged first

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.

Tip: disable this comment in your organization's Code Review settings.

@jatezzz jatezzz changed the title feat(ai-agent-local): store the KV cache as q8_0 to double the context ADFA-5188 | Enable KV cache quantization and flash attention Aug 21, 2026
@jatezzz
jatezzz force-pushed the fix/ADFA-5188-kv-cache-quantization branch from 24429e7 to 55b74ac Compare August 26, 2026 17:14
The cache was always f16. A load now asks for q8_0 where the model's head widths divide into whole blocks, costing 34 bytes per 32 elements instead of 64 and so buying about 1.88x the context from the same RAM budget. Flash attention is requested as AUTO, since llama.cpp needs it for a quantized value cache; if a context is refused anyway, the native side retries at f16 at a shorter length.
@jatezzz
jatezzz force-pushed the fix/ADFA-5188-kv-cache-quantization branch from 55b74ac to c45cce1 Compare August 26, 2026 20:20

@hal-eisen-adfa hal-eisen-adfa left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Four things worth a look before this merges — details inline. In short: the native fallback allocates about the same number of bytes as the attempt that just failed, a failed new_context leaks the whole model, the KV budget doesn't account for the weights that get mmap'd right after, and there's a 16 KB page-size ABI change riding along unannounced.

Comment thread ai-agent-local/llama-impl/src/main/cpp/llama-android.cpp
Comment thread ai-agent-local/llama-impl/src/main/java/android/llama/cpp/LLamaAndroid.kt Outdated
Comment thread ai-agent-local/llama-impl/build.gradle.kts
Charge the model weights against the budget the quantized cache is sized from, and floor clamp_context at DEFAULT_N_CTX so the f16 fallback cannot drop a short-context model below the context it always had.
The f16 retry was sized out of the same RAM budget as the q8_0 attempt, so it asked the allocator for roughly the bytes that had just failed, and was byte-identical when quantization was off. Adds a second retry at the DEFAULT_N_CTX floor, the only lever that answers memory pressure, and frees the model and any partial handles when a load step fails instead of pinning the weights for the life of the IDE process.
@jatezzz
jatezzz requested a review from hal-eisen-adfa August 27, 2026 16:14
Keep the pure ModelContextResolver.resolve that takes an already-parsed header, and have it answer with the q8_0 cache type and the f16 fallback size. The clamped KV-budget subtractions now guard the quantized path too, which is the one that can actually grow the context.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants