[Blog]: Add llm-d v0.9 release blog post - #472
Open
chcost wants to merge 10 commits into
Open
Conversation
Signed-off-by: Carlos H. Andrade Costa <chcost@us.ibm.com>
…laim, add links - Add SGLang as first-class backend as its own section per review feedback - Mention agentic serving guide and SGLang expansion in intro - Fix incorrect claim that TensorRT-LLM is the first non-vLLM engine - Add missing guide links for DeepSeek-V4 Signed-off-by: Carlos H. Andrade Costa <chcost@us.ibm.com>
Signed-off-by: Carlos H. Andrade Costa <chcost@us.ibm.com>
chcost
requested review from
Gregory-Pereira,
clubanderson,
davidgs,
jjasghar,
petecheslock and
robertgshaw2-redhat
as code owners
August 13, 2026 02:17
❌ Deploy Preview for llm-d failed. Why did it fail? →
|
- Add FMA, ModelExpress P2P, multi-model routing (IPP), P2P KV cache sharing - Move RL section to last, frame time-slicing as incubation - Revise What Is Next to be high-level directional Signed-off-by: Carlos H. Andrade Costa <chcost@us.ibm.com>
…it paths Signed-off-by: Carlos H. Andrade Costa <chcost@us.ibm.com>
Signed-off-by: Carlos H. Andrade Costa <chcost@us.ibm.com>
Signed-off-by: Carlos H. Andrade Costa <chcost@us.ibm.com>
… teams Signed-off-by: Carlos H. Andrade Costa <chcost@us.ibm.com>
Signed-off-by: Carlos H. Andrade Costa <chcost@us.ibm.com>
Signed-off-by: Carlos H. Andrade Costa <chcost@us.ibm.com>
Contributor
Author
|
/assign @ahg-g |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Draft release blog post for llm-d v0.9 ("Hardened for Scale").
Covers: coordinator HA, plugin lifecycle, flow control maturation, RDMA-free disaggregation, KEDA autoscaling, GPU utilization-aware routing, multimodal/diffusion support, OTel observability, RL architecture direction, and new well-lit paths for GB200/GB300/XPU/ROCm/TPU.
Pending:
release-0.9branch once cut