Skip to content

Repository files navigation

Model Router

A model router (AI routing gateway): it accepts requests and routes them to a suitable backend model, either by rules or by an AI decision. It speaks both the OpenAI chat-completions protocol and the Anthropic Messages protocol, and converts between them, so either kind of client can reach either kind of model. The backend is not limited to Azure AI Foundry, and each model can be bound to a different connection. It ships with an Azure-portal-styled React console: GitHub sign-in, API key management, usage statistics, full-chain traces, and configuration management.

Quick start (Docker)

docker run -d --name model-router \
  -p 8000:8000 -v mr-data:/data --restart unless-stopped \
  ghcr.io/satomic/model-router:latest

Nothing to prepare: the configuration is created from the template on first start, and the single /data volume holds all of it (the configuration, the sign-in state, the keys and the traces), so an upgrade is just a new image over the same volume.

Open http://localhost:8000/ and sign in as the local super administrator, admin / admin1234, which forces a password change first. Configure a backend connection from the console, create an API key on the "API keys" page, and point your client at it:

Field OpenAI-compatible client Anthropic-compatible client
Base URL http://localhost:8000/v1 http://localhost:8000
API Key mr_... as Authorization: Bearer mr_... as x-api-key
Model any model name registered under "Routing configuration" the same, or auto

Volumes, port mapping, upgrades and reverse proxies: Docker deployment.

Running from source

.venv\Scripts\Activate.ps1
pip install -r requirements.txt
cd frontend; npm ci; npm run build; cd ..   # FastAPI serves the built console from /
uvicorn app.main:app --host 0.0.0.0 --port 8000

data/config.yaml is created from config.example.yaml on first start here too, and data/ is the same single directory the container mounts. Running on your own machine, the first visit can also use the setup wizard to enter a GitHub OAuth Client ID / Secret. It is offered only to requests from 127.0.0.1, which is why a container uses the local administrator instead.

What it does

  • Routes by rules or by an AI decision model, then adapts parameters per model (reasoning models, the Responses API) before calling the backend.
  • Speaks both protocols on the way in and on the way out. /v1/chat/completions and /v1/messages are two doors onto the same router, and a connection can be Azure OpenAI, OpenAI-compatible or Anthropic-compatible. The four combinations all work, streaming included, so an Anthropic-style client can be answered by an Azure deployment and the reverse.
  • Scopes each API key independently. One key can be limited to a set of models, or to every model of chosen interface types, always as an intersection with what its owner is allowed, and the scope can be narrowed later without reissuing the key.
  • One user interaction is one routing decision and one trace. An agentic client such as GitHub Copilot answers a single question with a loop of HTTP requests; an x-interaction-id holds the model constant across that loop and folds every turn into a single trace record, instead of re-routing the same prompt N times.
  • Attributes every call to a real user. Copilot BYOK passes no identity, so user_id comes from the owner of the API key and cannot be forged by the client.
  • Gates who may create a key on GitHub Enterprise / organization / Enterprise Team membership, answered from a local cache where it can be trusted.
  • Curates the model list per user, team and organization. Named model groups are granted per scope and resolve as a union, and every user has a page showing exactly what they may call and which grant made it available.
  • Records the full chain: request, routing decision, backend call, response, and per-turn tool calls, readable in the console as a collapsible JSON tree.

Documentation

Document Contents
Operations guide / 操作指南 the console, screen by screen: what an administrator configures, then what a standard user does
Architecture and data flow start here: diagrams of the components, the request path, and every routing strategy
Docker deployment the recommended path: the image, port mapping, the data volume, upgrades, reverse proxies
Getting started running from source, frontend development, console languages
Sign-in and authentication the GitHub OAuth App, API keys, the local super administrator, the permission matrix
Backend connections providers, per-model bindings, non-Foundry OpenAI-compatible and Anthropic-compatible endpoints, protocol conversion
Router logic the request flow, interaction stickiness, rule and AI routing, the editable decision prompt
Configuration config.yaml, the console's configuration pages, hot reload
Access control the key-creation policy and the local GitHub structure/member cache
Model policy model groups, and which models each user / team / organization may use
API every endpoint
Full-chain logging the trace format, turns, and how the listing stays cheap at scale
Verification scripts the verify/ suite and the frontend gates

Layout

app/         FastAPI backend: routing, providers, auth, key policy, traces
frontend/    React + Vite console (built output is served by FastAPI from /)
docs/        the documents listed above
verify/      end-to-end verification scripts
Dockerfile   multi-stage build: the console is built in a discarded Node stage
data/        ALL persistent state -- config.yaml, sessions, keys, traces -- gitignored

Credentials never enter the repository: the whole of data/ (which is where config.yaml lives) and .env are gitignored, and config.example.yaml is the committed template with placeholders only.

About

an AI routing gateway that sits in front of any set of model endpoints and turns model selection from a vendor decision into a configuration we own and can audit.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages