A high-throughput and memory-efficient inference and serving engine for LLMs
-
Updated
Sep 11, 2026 - Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Never stop coding. Free MIT AI gateway: one endpoint, 352 providers (150+ free), 1200+ models Kimi, Claude, GPT, Gemini, GLM, DeepSeek, MiniMax. Works with Claude Code, Codex, Cursor, OpenCode, Cline & Copilot. Quota-aware auto-fallback, RTK+Caveman compression saves 15-95% tokens, MCP/A2A, Desktop/PWA. Built by 550+ contributors
The unified workspace where open-source models get things done for you.
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
✨ AI Coding, Vim Style
Command Code AI — the best coding agent for open models.
Visible multi-agent CLI workspace for mixing Codex, Claude, Gemini, Kimi, Qwen, Cursor, Copilot, Pi, OpenCode, and other AI coding agents
External-model router for Codex with guided Kimi OAuth/API, DeepSeek, safe migration, and rollback.
Intercept and inspect Coding Agent API traffic from Claude Code, Codex CLI, Gemini CLI, Cursor CLI, OpenCode, Kimi/Kimi Code, Pi, and Hermes in a local trace viewer.
Find, benchmark and install in CLI 170+ FREE coding LLM models across 15+ providers in real time
Delegate a coding task to a separate coding agent CLI, review the diff, land the commit yourself — one per implementer.
U-Claw 虾盘 — OpenClaw AI 助手离线安装 U 盘:一键装好 Claude Code/Codex/OpenClaw,无需联网即插即用;新品 U-King AI 装机管家详见 u-king.org,另有远程支持与定制 AI 开发。
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
Web, Desktop & Mobile client for Codex, Claude Code, OpenCode, Kimi, Augment Code, Qwen, fully end-to-end encrypted
A macOS menu bar application that monitors AI coding assistant usage quotas. Keep track of your Claude, Codex, Antigravity ,and Gemini usage at a glance.
AI-native agent harness for coding workflows by python: multi-model LLM orchestration, stateful sessions, tool governance, traceable delivery, and provider routing for GPT, Claude, DeepSeek, Qwen, Kimi, GLM, and MiniMax.
Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere
To associate your repository with the kimi topic, visit your repo's landing page and select "manage topics."