Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
-
Updated
Sep 10, 2026 - Python
Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
Anomaly detection related books, papers, videos, and toolboxes. Last update late 2025 for LLM and VLM works!
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
Official repository for VisionZip (CVPR 2025)
Official implementation of paper "VLM³: Vision Language Models Are Native 3D Learners".
[CVPR'24] HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models
Scala client for OpenAI API and other major LLM providers
[Survey Paper] This repository collects research papers of large Foundation Models for Scenario Generation and Analysis in Autonomous Driving. The repository will be continuously updated to track the latest update.
[CVPR 2024] Official implementation of "ViTamin: Designing Scalable Vision Models in the Vision-language Era"
Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders [Technical Report]
Official Repository of OmniCaptioner
[NeurIPS 2024] AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
[CVPR2025] SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
[ACL 2025 🔥] A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding
Hub for researchers exploring VLMs and Multimodal Learning:)
[NeurIPS'24] Official PyTorch Implementation of Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
To associate your repository with the vlms topic, visit your repo's landing page and select "manage topics."