Skip to content
#

fp4

Here are 29 public repositories matching this topic...

An LLM server for a single RTX 5090, built for agent workloads: tool calls, long conversations, reasoning, and many requests at once. One of the fastest engines on this card, at batch 1 and at dozens of concurrent streams, with the numbers in the repo.

  • Updated Sep 5, 2026
  • Cuda

LoRA fine-tune and serve NVFP4 models on one DGX Spark (GB10, 128 GB UMA): text backbones via generic-family onboarding, plus VLMs (vision tower, or LLM+tower jointly via --train-target both) validated end-to-end on Pixtral and Nemotron-Omni. Fused Triton dequant; runtime-LoRA and merge serving.

  • Updated Aug 1, 2026
  • Python

Python implementations for multi-precision quantization in computer vision and sensor fusion workloads, targeting the XR-NPE Mixed-Precision SIMD Neural Processing Engine. The code includes visual inertial odometry (VIO), object classification, and eye gaze extraction code in FP4, FP8, Posit4, Posit8, and BF16 formats.

  • Updated Aug 17, 2025
  • Jupyter Notebook

Add this topic to your repo

To associate your repository with the fp4 topic, visit your repo's landing page and select "manage topics."

Learn more