feat(backend/python): build vllm from source for AMD consumer/RDNA GPUs#10998
Draft
walcz-de wants to merge 5 commits into
Draft
feat(backend/python): build vllm from source for AMD consumer/RDNA GPUs#10998walcz-de wants to merge 5 commits into
walcz-de wants to merge 5 commits into
Conversation
…1 / Strix Halo) ROCm 7.14 adds native gfx1151 (Strix Halo / RDNA 3.5) runtime support. The prebuilt ROCm images pinned rocm/dev-ubuntu-24.04:7.2.1, which predates gfx1151 support, so the published -gpu-rocm-hipblas-* images do not run on that hardware even though gfx1151 is already in AMDGPU_TARGETS. Bump the pin to rocm/dev-ubuntu-24.04:7.14.0-full everywhere it is referenced: - .github/backend-matrix.yml (all hipblas backend build entries) - .github/workflows/base-images.yml (base-grpc-rocm-amd64 gRPC cache is rebuilt on 7.14 so it stays ABI-compatible with the new runtime base) - .github/workflows/image.yml + image-pr.yml (core hipblas image) - .agents/*.md examples + backend/Dockerfile.base-grpc-builder comment Verified on gfx1151 / Strix Halo: a full ROCm 7.14 stack (LocalAI + llama.cpp hipblas, vLLM, embeddings/reranker) runs natively, no HSA_OVERRIDE_GFX_VERSION needed. Other AMD architectures (gfx908/90a/942/1030/1100/1200/1201) build on the same base and are covered by CI; I do not have that hardware to test at runtime. Signed-off-by: stefanwalcz <stefan.walcz@walcz.de>
…gister lib paths)
ROCm 7.x moved to 'TheRock' packaging. Bumping the base to
rocm/dev-ubuntu-24.04:7.14.0-full surfaces two build breakages, fixed here in
ALL THREE spots that installed the legacy ROCm dev metapackages:
1) Legacy hipblas-dev / hipblaslt-dev / rocblas-dev metapackages were removed
(consolidated + arch-split: amdrocm-blas<ver>-gfx*, amdrocm-blas-dev, ...).
'apt-get install' of the old names fails. The -full base already ships the
BLAS dev libs+headers, so the install is dropped in:
- Dockerfile (requirements-drivers) -> core image
- .docker/install-base-deps.sh (section 6) -> C++ backend builder
- backend/Dockerfile.python (hipblas block) -> Python backends (vllm,
sglang, transformers, diffusers, kokoro, ...)
2) TheRock scatters the ROCm libs across /opt/rocm/lib,
/opt/rocm/lib/rocm_sysdeps/lib and /opt/rocm/llvm/lib with no ld.so.conf.d
entry, so the dynamic linker can't resolve them and the built backends fail
at RUNTIME ('... cannot open shared object file') — a failure CI never sees
because it only surfaces when a backend loads on an AMD GPU. Fixed by
registering all three lib dirs in /etc/ld.so.conf.d/rocm.conf before ldconfig
in all three blocks.
Validated end-to-end on gfx1151 / Strix Halo (build + GPU inference, llama-cpp).
Ref: https://rocm.docs.amd.com/en/latest/about/transition-guide-TheRock.html
Signed-off-by: stefanwalcz <stefan.walcz@walcz.de>
Apply the same ROCm 7.14 dependency + linker-path migration to backend/Dockerfile.golang: the removed legacy hipblas-dev / hipblaslt-dev / rocblas-dev metapackages are no longer installed (the rocm/dev-ubuntu-*:*-full base already ships them), and TheRock's scattered ROCm lib dirs are registered in /etc/ld.so.conf.d/rocm.conf before ldconfig so cgo-based Go backends resolve the ROCm runtime libraries at build/link and runtime. Identified by the LocalAI maintenance bot; re-authored under my identity per DCO. Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> Signed-off-by: stefanwalcz <stefan.walcz@walcz.de>
…users/kokoro The community whl/rocm7.0 torch wheel these backends install does not enumerate consumer/RDNA AMD GPUs (torch.cuda.device_count() == 0 on e.g. Strix Halo / gfx1151), so they silently fall back to CPU or fail to load. AMD publishes a stable ROCm 7.14 torch for essentially every AMD arch on its multi-arch index, selected per GPU via the torch[device-gfx<arch>] extra. Install torch from that index in an isolated step: the index returns 403 for packages it doesn't serve (accelerate/transformers/...), which uv treats as fatal, so only the torch family is pulled there and everything else resolves from PyPI. The arch comes from the build via a new AMDGPU_TARGETS ARG/ENV on backend/Dockerfile.python, defaulting to gfx1151. Only gfx1151 was validated on real hardware (transformers generates text, diffusers runs SD-1.5, kokoro runs TTS); the mechanism is arch-generic for any AMD GPU AMD ships a device wheel for. Signed-off-by: stefanwalcz <stefan.walcz@walcz.de>
vllm's ROCm install pulls a prebuilt wheel from wheels.vllm.ai/rocm, but that wheel exists only for CDNA data-center arches (gfx942/gfx950) and is pinned to a rocm7.2.3 torch that can't enumerate consumer/RDNA GPUs (device_count == 0 on e.g. gfx1151). No prebuilt vllm wheel exists for these GPUs anywhere, so build it from source against AMD's ROCm 7.14 torch -- reproducing the recipe AMD's own rocm/vllm:*_rdna_* image uses. - torch[device-gfx<arch>] + rocm-sdk-devel (hipcc + the HIP CMake packages the runtime SDK lacks) from AMD's multi-arch index; ROCM_PATH / CMAKE_PREFIX_PATH resolved from the SDK's own `rocm-sdk path` CLI -- self-contained, no system /opt/rocm needed. - vLLM detects ROCm at runtime via `import amdsmi`; register the SDK's in-place amd_smi with a .pth. A pip copy breaks its relative libamd_smi.so lookup, and forcing the lib via LD_LIBRARY_PATH shadows torch's own ROCm runtime and zeroes device_count. - run.sh exposes the bundled amdclang as CC/CXX: Triton JIT-compiles kernels at first inference and the runtime image ships hipcc but not cc/gcc, so without it inference dies with "Failed to find C compiler". Mirrors the CPU toolchain block just above it. - arch from AMDGPU_TARGETS (added in the multi-arch-torch change this stacks on), defaulting to gfx1151. Validated on gfx1151: builds from a single `make`, loads natively via LocalAI (no external gRPC) and generates. Only gfx1151 was hardware-tested; other AMD arches use the same mechanism but are unverified. Depends on mudler#10978 (ROCm 7.14 base) and the multi-arch-torch change it stacks on. Signed-off-by: stefanwalcz <stefan.walcz@walcz.de>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
vllm's ROCm install pulls a prebuilt wheel from
wheels.vllm.ai/rocm, but that wheel exists only for CDNA data-center arches (gfx942/gfx950) and is pinned to a rocm7.2.3 torch that can't enumerate consumer/RDNA GPUs (device_count == 0on e.g. gfx1151). There is no prebuilt vllm wheel for these GPUs anywhere — which is why AMD ships a wholerocm/vllm:*_rdna_*container. This builds vllm from source against AMD's ROCm 7.14 torch, reproducing that container's recipe inside LocalAI's normalDockerfile.pythonbuild (FROM-scratch, portable venv — no external gRPC / no 73 GB image).How
torch[device-gfx<arch>]+rocm-sdk-devel(hipcc + the HIP CMake packages the runtime_rocm_sdk_corelacks) from AMD's multi-arch index;ROCM_PATH/CMAKE_PREFIX_PATHresolved from the SDK's ownrocm-sdk pathCLI — self-contained, no system/opt/rocmneeded.import amdsmi; register the SDK's in-placeamd_smiwith a.pth. A pip copy breaks its relativelibamd_smi.solookup, and forcing the lib viaLD_LIBRARY_PATHshadows torch's own ROCm runtime and zeroesdevice_count.run.shexposes the bundledamdclangasCC/CXX: Triton JIT-compiles kernels at first inference and the runtime image shipshipccbut notcc/gcc, so without it inference dies with "Failed to find C compiler". Mirrors the CPU-toolchain block just above it.AMDGPU_TARGETS, defaulting togfx1151.Testing
Hardware-validated on gfx1151: builds from a single
make, is loaded natively by LocalAI from/backends/vllm(no external gRPC), and generates. Only gfx1151 was tested on real hardware; other AMD arches use the same mechanism but are unverified.Dependency
Stacks on #10978 (ROCm 7.14 base) and #10997 (multi-arch torch /
AMDGPU_TARGETS). Draft until those land; reviewable change is the top commit.