diff --git a/AGENTS.md b/AGENTS.md index c74f8c6a..a9080d98 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -58,10 +58,19 @@ make ci # Full CI suite - `provider/core.Provider` is the lower-level provider contract; keep its method set stable and use it across feature packages - Streaming tests need careful goroutine management -- `go.work` here should stay minimal; the parent `graycode-eco/go.work` - connects this independent `flux` checkout beside Rho for local development. +- `go.work` here should stay minimal; the `graycode-eco` workspace connects + this independent `flux` checkout beside Rho for local development. Do not add extra local `replace` directives here without coordinating with the parent workspace. +- There is usually **no** `go.work` in the parent folder, and that is expected. + A parent `go.work` breaks every sibling it does not list, and exporting + `GOWORK` is worse: it is inherited by child `go` processes, so rho's own + tests that shell out to `go test` in temporary projects fail. To exercise + rho against this checkout, use a gitignored module overlay in rho instead — + `go test -modfile=go.local.mod ./...`. See rho/AGENTS.md + "Workspace workflow" for the setup. Without that overlay, rho compiles + against the *published* flux from the module cache, so a green rho suite + does not exercise local flux changes at all. ## Naming Conventions diff --git a/CHANGELOG.md b/CHANGELOG.md index adf5b99d..a372e974 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -15,12 +15,73 @@ Format: [Keep a Changelog](https://keepachangelog.com/en/1.0.0/) · Versioning: credential lookup is required for the replicated route. - Client-owned OpenAI-compatible provider registration through `FluxClient.RegisterCustomProvider`. +- Streams end with exactly one terminal event. When the caller's context is + cancelled or its deadline passes, `provider/core` stream wrappers emit a + terminal `cancelled` event and the engine emits `engine.EventCancelled` + (with `ErrorInfo` and the route) before `Err()` reports `ErrorCancelled`. + Hosts that switch on event types should handle `cancelled`. Cancelling + the context releases the provider request immediately, but the terminal is + delivered like any other event, so callers must still read the stream to + the end or call `Close` (as `StreamResult`, `EventStreamer` and + `engine.Stream` now document). +- Terminal `error`/`cancelled` stream events leaving `provider/core` always + carry a `StreamErrorInfo` (`Kind`/`Retryable`), inferred from the provider + message when the adapter set none. Existing `ErrorInfo` is never + overwritten, and a cancellation never reports as an internal fault. +- Responses and stream events report the route that actually served the + request: `ResolvedRoute.DeploymentID` and `Attempts` after a deployment + failover, a `route_changed` event per deployment attempt, and the route on + engine events. ### Fixed - Circuit breakers now admit at most one concurrent half-open probe and do not reserve probes during route filtering. +- `DeploymentRouter` records a circuit-breaker failure only for errors that + describe the deployment's health (5xx, 529, transport failures). Caller + cancellation, rate limits and 4xx request errors no longer take a healthy + deployment out of rotation. +- Upstream timeouts and provider messages that merely contain "cancelled" are + no longer mistaken for the caller's cancellation. Only the caller's own + context decides that: `DeploymentRouter` fails over from an upstream + timeout (including net/http `Client.Timeout`) and counts it against the + deployment, and the engine reports it as a retryable + `ErrorProviderUnavailable` instead of `ErrorCancelled`. +- Engine stream errors keep the provider's classification: a stream error's + `ErrorInfo.Kind` now maps to `ErrorRateLimited`, `ErrorAuthentication`, + `ErrorContextExceeded` or `ErrorInvalidRequest` (content filtering + included), carries `Retryable`, and names the route's provider and model, + instead of every stream failure becoming a non-retryable + `ErrorProviderUnavailable`. +- `DeploymentRouter` streams no longer end at the first non-output event + after output. A `usage`, `ttft` or `provider_block` event between content + events (Anthropic reports output usage before `message_stop`) used to end + the deployment stream, dropping the remaining content and `done` and + surfacing a truncation error. +- `DeploymentRouter` treats warning-marked diagnostic `error` events (for + example a reasoning-only response) as non-fatal, like `provider/core` and + the engine already did, instead of ending the stream or failing over. +- SSE parsing is bounded per event. `core.ParseSSEStream` capped each line at + 2 MiB but accumulated an event's `data:` lines without limit, and the + Concentrate Responses reader bounded neither lines nor events, so a hostile + or broken endpoint could grow client memory indefinitely. Both now stop + with a stream error once one event exceeds `core.SSEMaxEventBytes` + (16 MiB). +- The response cache key now covers the whole request. It hashed only the + model, system prompt, temperature and message text/tool data, so with + caching enabled a reply produced for one tool set, `max_tokens`, stop + sequence, sampling or thinking setting, response format, image or caller + was served to a different request. Keys now hash every `ChatOptions` and + message field under a versioned prefix, and requests that cannot be + encoded bypass the cache. +- A `provider/core` stream cancelled while its consumer was behind (the + forwarder blocked delivering an event) now still ends with the terminal + `cancelled` event instead of closing silently. ### Changed +- `core.CoordinateStreamResult` returns a stream unchanged when it is already + coordinated under the same context and request ID, so layers that re-wrap + an adapter's stream with the caller's context (`Router`, `ProtocolRouter`) + no longer add a goroutine and buffer per layer. - Removed process-global custom gateway and dynamic provider registration, the ambient `OPENAI_API_BASE` auto-registration path, and no-op API-key prefix inference. Custom gateways and endpoints now require explicit, diff --git a/docs/architecture/HOST-ENGINE-BOUNDARY.md b/docs/architecture/HOST-ENGINE-BOUNDARY.md index 15f0f36d..ba452775 100644 --- a/docs/architecture/HOST-ENGINE-BOUNDARY.md +++ b/docs/architecture/HOST-ENGINE-BOUNDARY.md @@ -153,6 +153,24 @@ additive and must be ignored safely. Flux emits tool requests; Rho authorizes and executes tools, appends results to its history, and begins the next model turn. +Every stream ends exactly once: + +- `done` — success; `Err()` stays nil. +- `cancelled` — the request context was cancelled or its deadline passed. The + event carries `ErrorInfo` (`canceled` or `timeout`) and the route, and + `Err()` then reports `ErrorCancelled` wrapping the context's error. Hosts + should treat it as the user's cancellation, not as a failure, and must not + emit a second terminal for it. +- no terminal event — `Next` returns false and `Err()` reports the provider + failure with the code from its `ErrorInfo` (`rate_limited`, + `authentication_failed`, `context_exceeded`, `invalid_request`, or + `provider_unavailable`) and `Retryable`. An upstream timeout is such a + retryable failure, never `cancelled`. + +Cancelling the context releases the provider request at once, but the +terminal event waits for the host to read it, so hosts must still read to the +end or call `Close`. + ## Readiness Preflight has two explicit modes: diff --git a/docs/guides/DYNAMIC-MODEL-DISCOVERY.md b/docs/guides/DYNAMIC-MODEL-DISCOVERY.md index f088d71e..f0f2997b 100644 --- a/docs/guides/DYNAMIC-MODEL-DISCOVERY.md +++ b/docs/guides/DYNAMIC-MODEL-DISCOVERY.md @@ -244,7 +244,7 @@ the safe credential/gateway reports without reading either file directly. | catalog cache missing/corrupt | report unavailable; do not claim bootstrap readiness | | live list fails or selected model is absent | fail live preflight | | custom gateway URL contains embedded data | reject configuration | -| stream caller exits | close/cancel the Engine stream | +| stream caller exits | close the Engine stream (cancelling its context alone leaves the terminal event waiting for a reader) | Provider-specific friendly error formatting remains Flux-owned; Rho decides where and how to display it. diff --git a/engine/classify.go b/engine/classify.go index ce82991c..08a9b08f 100644 --- a/engine/classify.go +++ b/engine/classify.go @@ -21,6 +21,9 @@ func classify(operation string, route Route, err error) error { switch { case errors.Is(err, context.Canceled), errors.Is(err, context.DeadlineExceeded): code = ErrorCancelled + case errors.Is(err, core.ErrStreamTruncated): + code = ErrorProviderUnavailable + retryable = true default: var providerErr *core.FluxError if errors.As(err, &providerErr) { diff --git a/engine/continuation.go b/engine/continuation.go index ece1083f..bfde6fca 100644 --- a/engine/continuation.go +++ b/engine/continuation.go @@ -40,6 +40,7 @@ func streamWithContinuation(ctx context.Context, provider core.Provider, message var segment strings.Builder hadToolCall := false var terminal core.FluxStreamEvent + segmentLoop: for event := range current.Events { switch event.Type { case "content": @@ -52,7 +53,13 @@ func streamWithContinuation(ctx context.Context, provider core.Provider, message } case "done": terminal = event - continue + break segmentLoop + case "error": + if event.Warning == "" { + _ = emitEngineEvent(streamCtx, out, event) + current.Close() + return + } } if !emitEngineEvent(streamCtx, out, event) { current.Close() @@ -60,12 +67,12 @@ func streamWithContinuation(ctx context.Context, provider core.Provider, message } } current.Close() + if terminal.Type == "" { + return + } needsContinuation := terminal.StopReason == "max_tokens" || terminal.StopReason == "length" if !needsContinuation || hadToolCall || totalOutput >= maxTotalTokens || attempt >= maxContinuations { - if terminal.Type == "" { - terminal = core.FluxStreamEvent{Type: "done", StopReason: terminal.StopReason, RequestID: requestID} - } _ = emitEngineEvent(streamCtx, out, terminal) return } diff --git a/engine/contract_e2e_test.go b/engine/contract_e2e_test.go index 824242ab..29b08dfe 100644 --- a/engine/contract_e2e_test.go +++ b/engine/contract_e2e_test.go @@ -176,6 +176,43 @@ func TestEngineContinuationPreservesEventsAndConversationShape(t *testing.T) { } } +type terminalErrorProvider struct{} + +func (*terminalErrorProvider) Name() string { return "terminal-error" } +func (*terminalErrorProvider) Ping(context.Context) error { return nil } +func (*terminalErrorProvider) Chat(context.Context, []core.FluxMessage, core.ChatOptions) (*core.FluxResponse, error) { + return nil, nil +} + +func (*terminalErrorProvider) StreamChat(context.Context, []core.FluxMessage, core.ChatOptions) (*core.StreamResult, error) { + events := make(chan core.FluxStreamEvent, 2) + events <- core.FluxStreamEvent{Type: "error", Error: "connection reset"} + events <- core.FluxStreamEvent{Type: "done", StopReason: "stop"} + close(events) + return llm.NewStreamResult(events, "request-error", nil), nil +} + +func TestEngineContinuationDoesNotAppendDoneAfterFatalError(t *testing.T) { + source, err := streamWithContinuation( + context.Background(), &terminalErrorProvider{}, + []core.FluxMessage{{Role: "user", Content: "hello"}}, + core.ChatOptions{Model: "model"}, + Limits{MaxContinuations: 1, MaxTotalOutputTokens: 100}, + ) + if err != nil { + t.Fatal(err) + } + defer source.Close() + + var events []core.FluxStreamEvent + for event := range source.Events { + events = append(events, event) + } + if len(events) != 1 || events[0].Type != "error" || events[0].Error != "connection reset" { + t.Fatalf("events = %+v, want one fatal error", events) + } +} + func firstCatalogModelID(cat catalog.Catalog) string { for id := range cat.Models { return id diff --git a/engine/convert.go b/engine/convert.go index 2cc36585..0d3a0668 100644 --- a/engine/convert.go +++ b/engine/convert.go @@ -85,14 +85,42 @@ func cloneStringMap(in map[string]string) map[string]string { return out } +func cloneRoute(route *Route) *Route { + if route == nil { + return nil + } + cloned := *route + return &cloned +} + +// mergeRoute keeps the route the provider reported and fills any blank +// identity fields from the route the engine planned. +func mergeRoute(actual *Route, planned Route) *Route { + if actual == nil { + return cloneRoute(&planned) + } + merged := *actual + if merged.Provider == "" { + merged.Provider = planned.Provider + } + if merged.Model == "" { + merged.Model = planned.Model + } + if !merged.DeploymentRouting { + merged.DeploymentRouting = planned.DeploymentRouting + } + return &merged +} + // fromClientResponse attaches the resolved route to a client response. The -// engine and the client both speak the canonical contract response type, so -// this only sets the route the engine selected. +// engine and the client both speak the canonical contract response type. A +// route the provider already reported (for example the deployment that served +// the request after a failover) wins; the planned route only fills blanks. func fromClientResponse(resp *core.FluxResponse, route Route) *GenerateResponse { if resp == nil { - return &GenerateResponse{Route: &route} + return &GenerateResponse{Route: cloneRoute(&route)} } - resp.Route = &route + resp.Route = mergeRoute(resp.Route, route) return resp } diff --git a/engine/convert_test.go b/engine/convert_test.go index 0c4b2802..14be9b92 100644 --- a/engine/convert_test.go +++ b/engine/convert_test.go @@ -264,6 +264,23 @@ func TestFromClientResponse_AttachesRoute(t *testing.T) { } } +func TestFromClientResponse_PreservesActualRoute(t *testing.T) { + resp := &core.FluxResponse{ + Content: "hello", + Route: &Route{ + Provider: "router", Model: "planned/model", DeploymentID: "deployment-2", Attempts: 2, + }, + } + planned := Route{Provider: "planned", Model: "planned/model", DeploymentRouting: true} + out := fromClientResponse(resp, planned) + if out.Route == nil || out.Route.DeploymentID != "deployment-2" || out.Route.Attempts != 2 || !out.Route.DeploymentRouting { + t.Fatalf("route = %+v, want actual route preserved", out.Route) + } + if out.Route.Provider != "router" || out.Route.Model != "planned/model" { + t.Fatalf("route identity = %+v, want provider route preserved", out.Route) + } +} + func TestFromClientResponse_NilResponse(t *testing.T) { route := Route{Provider: "test", Model: "test/model"} out := fromClientResponse(nil, route) diff --git a/engine/engine_test.go b/engine/engine_test.go index 30593b2f..62722c81 100644 --- a/engine/engine_test.go +++ b/engine/engine_test.go @@ -240,6 +240,77 @@ func TestStreamFatalErrorEventStillTerminal(t *testing.T) { } } +func TestStreamSilentSourceCloseIsTruncated(t *testing.T) { + sourceEvents := make(chan core.FluxStreamEvent) + close(sourceEvents) + + ctx, cancel := context.WithCancel(context.Background()) + stream := newStream(ctx, cancel, llm.NewStreamResult(sourceEvents, "", nil), Route{Provider: "mock", Model: "mock/model"}) + defer stream.Close() + + var events []Event + for stream.Next() { + events = append(events, stream.Event()) + } + err := stream.Err() + if !IsCode(err, ErrorProviderUnavailable) { + t.Fatalf("error = %v, want provider_unavailable", err) + } + if !errors.Is(err, core.ErrStreamTruncated) { + t.Fatalf("error = %v, want stream truncation cause", err) + } + if len(events) != 1 || events[0].Type != EventRouteSelected { + t.Fatalf("events = %+v, want only route_selected", events) + } +} + +func TestStreamContextCancellationEmitsTerminalAndErr(t *testing.T) { + sourceEvents := make(chan core.FluxStreamEvent) + ctx, cancel := context.WithCancel(context.Background()) + stream := newStream(ctx, cancel, llm.NewStreamResult(sourceEvents, "request-cancel", nil), Route{Provider: "mock", Model: "mock/model"}) + + if !stream.Next() { + t.Fatal("expected route event") + } + cancel() + + if !stream.Next() { + t.Fatal("expected cancellation event") + } + event := stream.Event() + if event.Type != EventCancelled || event.ErrorInfo == nil || event.ErrorInfo.Kind != llm.ErrKindCanceled { + t.Fatalf("event = %+v, want cancellation terminal", event) + } + if stream.Next() { + t.Fatal("unexpected event after cancellation terminal") + } + if err := stream.Err(); !IsCode(err, ErrorCancelled) { + t.Fatalf("error = %v, want cancelled", err) + } + _ = stream.Close() +} + +func TestStreamCancelledBeforeFirstEventStillTerminates(t *testing.T) { + // forward races the route_selected emit against the already-cancelled + // context; every run must still end with the cancelled terminal. + for i := 0; i < 200; i++ { + ctx, cancel := context.WithCancel(context.Background()) + cancel() + stream := newStream(ctx, cancel, llm.NewStreamResult(make(chan core.FluxStreamEvent), "request-early", nil), Route{Provider: "mock", Model: "mock/model"}) + var last Event + for stream.Next() { + last = stream.Event() + } + if last.Type != EventCancelled { + t.Fatalf("run %d: last event = %+v, want cancelled terminal", i, last) + } + if err := stream.Err(); !IsCode(err, ErrorCancelled) { + t.Fatalf("run %d: error = %v, want cancelled", i, err) + } + _ = stream.Close() + } +} + func TestSnapshotPublishesCapabilities(t *testing.T) { compiled := &catalog.CompiledCatalog{ ModelsByID: map[string]catalog.Model{ diff --git a/engine/stream.go b/engine/stream.go index fbff3e1e..ed128385 100644 --- a/engine/stream.go +++ b/engine/stream.go @@ -2,28 +2,44 @@ package engine import ( "context" + "encoding/json" + "errors" "sync" + "github.com/GrayCodeAI/flux/llm" "github.com/GrayCodeAI/flux/provider/core" ) // Stream is a normalized, pull-based event stream. Next must not be called // concurrently. Close is idempotent and may be called from another goroutine. +// +// Callers must call Close, or read until Next returns false. Cancelling the +// request context ends the stream with an EventCancelled terminal and releases +// the provider request at once, but the goroutine delivering that terminal +// waits until it is read or Close is called. type Stream struct { ctx context.Context cancel context.CancelFunc source *core.StreamResult route Route events chan Event + closed chan struct{} - mu sync.Mutex - current Event - err error - once sync.Once + mu sync.Mutex + current Event + err error + once sync.Once + releaseOnce sync.Once } func newStream(ctx context.Context, cancel context.CancelFunc, source *core.StreamResult, route Route) *Stream { - s := &Stream{ctx: ctx, cancel: cancel, source: source, route: route, events: make(chan Event, 32)} + if ctx == nil { + ctx = context.Background() + } + s := &Stream{ + ctx: ctx, cancel: cancel, source: source, route: route, + events: make(chan Event, 32), closed: make(chan struct{}), + } go s.forward() return s } @@ -69,6 +85,18 @@ func (s *Stream) Close() error { return nil } s.once.Do(func() { + if s.closed != nil { + close(s.closed) + } + s.release() + }) + return nil +} + +// release cancels the request and closes the provider stream without +// closing the consumer side. +func (s *Stream) release() { + s.releaseOnce.Do(func() { if s.cancel != nil { s.cancel() } @@ -76,30 +104,60 @@ func (s *Stream) Close() error { s.source.Close() } }) - return nil } func (s *Stream) forward() { defer close(s.events) defer s.Close() - if !s.emit(Event{Type: EventRouteSelected, Route: &s.route}) { + if !s.emit(Event{Type: EventRouteSelected, Route: cloneRoute(&s.route)}) { + // The context ended before the first event; the stream must still + // finish with the cancelled terminal instead of closing silently. + if !s.isClosed() { + s.finishCancelled(s.source.RequestID) + } return } for { select { case <-s.ctx.Done(): - s.setError(classify("stream", s.route, s.ctx.Err())) + if !s.isClosed() { + s.finishCancelled(s.source.RequestID) + } return case event, ok := <-s.source.Events: if !ok { + if s.ctx.Err() == nil { + s.setError(classify("stream", s.route, core.ErrStreamTruncated)) + } else if !s.isClosed() { + s.finishCancelled(s.source.RequestID) + } + return + } + if event.RequestID == "" && s.source != nil { + event.RequestID = s.source.RequestID + } + if isTerminalFailure(event) && s.ctx.Err() != nil { + // The caller cancelled or the request deadline passed, and the + // provider reported the consequence. Only the engine's own + // context decides this: an upstream timeout carries the same + // timeout kind but is a provider failure. + if !s.isClosed() { + s.finishCancelled(event.RequestID) + } return } - normalized, err := normalizeEvent(event) + normalized, err := normalizeEvent(event, s.route) if err != nil { s.setError(classify("stream", s.route, err)) return } if !s.emit(normalized) { + if !s.isClosed() && s.ctx.Err() != nil { + s.finishCancelled(normalized.RequestID) + } + return + } + if normalized.Type == EventDone { return } } @@ -115,17 +173,66 @@ func (s *Stream) emit(event Event) bool { } } +// finishCancelled records the caller's cancellation as the terminal error and +// emits the terminal cancelled event. The provider request is released first, +// so a host that cancels without draining or closing the stream strands only +// this goroutine, not the upstream connection. +func (s *Stream) finishCancelled(requestID string) { + err := s.ctx.Err() + if err == nil { + return + } + s.setError(classify("stream", s.route, err)) + s.release() + s.emitCancellation(err, requestID) +} + +func (s *Stream) emitCancellation(err error, requestID string) { + if err == nil || s.isClosed() { + return + } + kind := llm.ErrKindCanceled + if errors.Is(err, context.DeadlineExceeded) { + kind = llm.ErrKindTimeout + } + event := Event{ + Type: EventCancelled, + Error: err.Error(), + RequestID: requestID, + ErrorInfo: &llm.StreamErrorInfo{Kind: kind}, + Route: cloneRoute(&s.route), + } + select { + case s.events <- event: + case <-s.closed: + } +} + +func (s *Stream) isClosed() bool { + if s.closed == nil { + return false + } + select { + case <-s.closed: + return true + default: + return false + } +} + func (s *Stream) setError(err error) { s.mu.Lock() s.err = err s.mu.Unlock() } -func normalizeEvent(event core.FluxStreamEvent) (Event, error) { +func normalizeEvent(event core.FluxStreamEvent, route Route) (Event, error) { out := Event{ Content: event.Content, Thinking: event.Thinking, RequestID: event.RequestID, + Error: event.Error, Warning: event.Warning, ErrorInfo: cloneErrorInfo(event.ErrorInfo), Usage: fromClientUsage(event.Usage), StopReason: event.StopReason, - TTFTms: event.TTFTms, + TTFTms: event.TTFTms, TTFT: event.TTFT, Route: cloneRoute(event.Route), + ProviderBlock: cloneProviderBlock(event.ProviderBlock), } if out.TTFTms == 0 { out.TTFTms = event.TTFT @@ -147,20 +254,105 @@ func normalizeEvent(event core.FluxStreamEvent) (Event, error) { out.Type = EventTTFT case "continuation": out.Type = EventContinuation + case "cancelled", "canceled": + // forward handles the caller's own cancellation before normalizing, so + // a cancellation reaching here came from a context the provider owns + // (for example an adapter-side timeout): an upstream failure. + return Event{}, &Error{ + Code: ErrorProviderUnavailable, Operation: "stream", Provider: route.Provider, Model: route.Model, + Message: event.Error, Retryable: true, + } case "error": if event.Warning != "" { // Non-fatal health diagnostic (e.g. a reasoning-only response): // client/core marks these with Warning so they can be surfaced // without terminating the stream — the terminal done/usage event // follows. Forward as a warning event; do not set Err()/stop. - return Event{Type: EventWarning, Warning: event.Warning}, nil + return Event{ + Type: EventWarning, RequestID: event.RequestID, Error: event.Error, + ErrorInfo: cloneErrorInfo(event.ErrorInfo), Warning: event.Warning, + Route: cloneRoute(event.Route), + }, nil + } + if event.ErrorInfo != nil && event.ErrorInfo.Kind == llm.ErrKindTruncated { + return Event{}, core.ErrStreamTruncated } - return Event{}, &Error{Code: ErrorProviderUnavailable, Operation: "stream", Message: event.Error} + return Event{}, streamEventError(event, route) default: out.Type = event.Type } if event.ToolCall != nil { - out.ToolCall = &ToolCall{ID: event.ToolCall.ID, Name: event.ToolCall.Name, Arguments: event.ToolCall.Arguments} + out.ToolCall = cloneToolCall(event.ToolCall) } return out, nil } + +// streamEventError converts a fatal provider stream event into the engine's +// typed error. It prefers the event's ErrorInfo over its message text so hosts +// can tell an authentication failure from a rate limit or an outage. +func streamEventError(event core.FluxStreamEvent, route Route) *Error { + err := &Error{ + Code: ErrorProviderUnavailable, Operation: "stream", Provider: route.Provider, Model: route.Model, + Message: event.Error, + } + info := event.ErrorInfo + if info == nil { + return err + } + err.Retryable = info.Retryable + switch info.Kind { + case llm.ErrKindRateLimited: + err.Code = ErrorRateLimited + case llm.ErrKindAuth: + err.Code = ErrorAuthentication + case llm.ErrKindContextExceeded: + err.Code = ErrorContextExceeded + case llm.ErrKindInvalidRequest, llm.ErrKindContentFiltered: + err.Code = ErrorInvalidRequest + case llm.ErrKindCanceled, llm.ErrKindTimeout: + // Not the caller's cancellation (forward checked its context): an + // upstream timeout or cancel, worth retrying. + err.Retryable = true + } + return err +} + +func cloneErrorInfo(info *llm.StreamErrorInfo) *llm.StreamErrorInfo { + if info == nil { + return nil + } + cloned := *info + return &cloned +} + +func cloneProviderBlock(block *llm.ProviderBlock) *llm.ProviderBlock { + if block == nil { + return nil + } + cloned := *block + cloned.Data = append([]byte(nil), block.Data...) + return &cloned +} + +func cloneToolCall(call *llm.ToolCall) *llm.ToolCall { + if call == nil { + return nil + } + cloned := *call + if call.RawArguments != nil { + cloned.RawArguments = append([]byte(nil), call.RawArguments...) + } + if call.ProviderMetadata != nil { + cloned.ProviderMetadata = make(map[string]json.RawMessage, len(call.ProviderMetadata)) + for key, value := range call.ProviderMetadata { + cloned.ProviderMetadata[key] = append(json.RawMessage(nil), value...) + } + } + return &cloned +} + +// isTerminalFailure reports whether a provider event ends the stream with a +// failure. Warning-marked errors are non-fatal diagnostics. +func isTerminalFailure(event core.FluxStreamEvent) bool { + return event.Type == "cancelled" || event.Type == "canceled" || event.Type == "error" && event.Warning == "" +} diff --git a/engine/stream_lifecycle_test.go b/engine/stream_lifecycle_test.go new file mode 100644 index 00000000..7b1dceb4 --- /dev/null +++ b/engine/stream_lifecycle_test.go @@ -0,0 +1,267 @@ +package engine + +import ( + "context" + "errors" + "sync" + "testing" + "time" + + "github.com/GrayCodeAI/flux/llm" + "github.com/GrayCodeAI/flux/provider/core" +) + +var lifecycleRoute = Route{Provider: "mock", Model: "mock/model"} + +// drainEngineStream reads until Next reports false, failing if it stalls. +func drainEngineStream(t *testing.T, stream *Stream) []Event { + t.Helper() + done := make(chan []Event, 1) + go func() { + var events []Event + for stream.Next() { + events = append(events, stream.Event()) + } + done <- events + }() + select { + case events := <-done: + return events + case <-time.After(5 * time.Second): + t.Fatal("stream did not finish") + return nil + } +} + +func engineEventTypes(events []Event) []string { + types := make([]string, len(events)) + for i, event := range events { + types[i] = event.Type + } + return types +} + +func streamOf(ctx context.Context, cancel context.CancelFunc, events ...core.FluxStreamEvent) *Stream { + source := make(chan core.FluxStreamEvent, len(events)) + for _, event := range events { + source <- event + } + close(source) + return newStream(ctx, cancel, llm.NewStreamResult(source, "request-1", nil), lifecycleRoute) +} + +func TestStreamProviderFailureWithLiveContextIsNotCancellation(t *testing.T) { + tests := []struct { + name string + event core.FluxStreamEvent + }{ + {"upstream timeout error", core.FluxStreamEvent{ + Type: "error", Error: "context deadline exceeded (Client.Timeout exceeded while reading body)", + ErrorInfo: &llm.StreamErrorInfo{Kind: llm.ErrKindTimeout, Retryable: true}, + }}, + {"upstream canceled error", core.FluxStreamEvent{ + Type: "error", Error: "context canceled", + ErrorInfo: &llm.StreamErrorInfo{Kind: llm.ErrKindCanceled}, + }}, + {"upstream cancelled event", core.FluxStreamEvent{ + Type: "cancelled", Error: "context canceled", + ErrorInfo: &llm.StreamErrorInfo{Kind: llm.ErrKindCanceled}, + }}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + ctx, cancel := context.WithCancel(context.Background()) + stream := streamOf(ctx, cancel, tt.event) + defer stream.Close() + + events := drainEngineStream(t, stream) + for _, event := range events { + if event.Type == EventCancelled { + t.Fatalf("events = %v; a provider failure must not look like the caller's cancellation", engineEventTypes(events)) + } + } + err := stream.Err() + if IsCode(err, ErrorCancelled) || !IsCode(err, ErrorProviderUnavailable) { + t.Fatalf("error = %v, want provider_unavailable", err) + } + var engineErr *Error + if !errors.As(err, &engineErr) || !engineErr.Retryable { + t.Fatalf("error = %#v, want retryable", err) + } + }) + } +} + +func TestStreamCallerCancellationReportsContextCause(t *testing.T) { + tests := []struct { + name string + ctx func() (context.Context, context.CancelFunc) + wantKind string + wantErr error + }{ + {"cancel", func() (context.Context, context.CancelFunc) { + ctx, cancel := context.WithCancel(context.Background()) + cancel() + return ctx, cancel + }, llm.ErrKindCanceled, context.Canceled}, + {"deadline", func() (context.Context, context.CancelFunc) { + return context.WithDeadline(context.Background(), time.Now().Add(-time.Second)) + }, llm.ErrKindTimeout, context.DeadlineExceeded}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + ctx, cancel := tt.ctx() + // The provider reports the consequence of the caller's + // cancellation with a kind that does not match the context. + stream := streamOf(ctx, cancel, core.FluxStreamEvent{ + Type: "error", Error: "stream read error", + ErrorInfo: &llm.StreamErrorInfo{Kind: llm.ErrKindUnavailable, Retryable: true}, + }) + defer stream.Close() + + events := drainEngineStream(t, stream) + if len(events) == 0 { + t.Fatal("stream ended without events") + } + last := events[len(events)-1] + if last.Type != EventCancelled || last.ErrorInfo == nil || last.ErrorInfo.Kind != tt.wantKind { + t.Fatalf("events = %v, last = %+v, want cancelled terminal with kind %s", engineEventTypes(events), last, tt.wantKind) + } + if last.RequestID != "request-1" { + t.Fatalf("request ID = %q, want request-1", last.RequestID) + } + err := stream.Err() + if !IsCode(err, ErrorCancelled) || !errors.Is(err, tt.wantErr) { + t.Fatalf("error = %v, want cancelled caused by %v", err, tt.wantErr) + } + }) + } +} + +func TestStreamErrorKindsMapToEngineCodes(t *testing.T) { + tests := []struct { + kind string + retryable bool + wantCode ErrorCode + wantRetryable bool + }{ + {llm.ErrKindRateLimited, true, ErrorRateLimited, true}, + {llm.ErrKindAuth, false, ErrorAuthentication, false}, + {llm.ErrKindContextExceeded, false, ErrorContextExceeded, false}, + {llm.ErrKindInvalidRequest, false, ErrorInvalidRequest, false}, + {llm.ErrKindContentFiltered, false, ErrorInvalidRequest, false}, + {llm.ErrKindUnavailable, true, ErrorProviderUnavailable, true}, + {llm.ErrKindInternal, true, ErrorProviderUnavailable, true}, + {llm.ErrKindTimeout, false, ErrorProviderUnavailable, true}, + } + for _, tt := range tests { + t.Run(tt.kind, func(t *testing.T) { + ctx, cancel := context.WithCancel(context.Background()) + stream := streamOf(ctx, cancel, core.FluxStreamEvent{ + Type: "error", Error: "provider said no", + ErrorInfo: &llm.StreamErrorInfo{Kind: tt.kind, Retryable: tt.retryable}, + }) + defer stream.Close() + drainEngineStream(t, stream) + + var err *Error + if !errors.As(stream.Err(), &err) { + t.Fatalf("error = %v, want *Error", stream.Err()) + } + if err.Code != tt.wantCode || err.Retryable != tt.wantRetryable { + t.Fatalf("error = {code %s retryable %v}, want {code %s retryable %v}", err.Code, err.Retryable, tt.wantCode, tt.wantRetryable) + } + if err.Provider != lifecycleRoute.Provider || err.Model != lifecycleRoute.Model || err.Message != "provider said no" { + t.Fatalf("error = %+v, want route and provider message", err) + } + }) + } +} + +func TestStreamCancelReleasesProviderBeforeTerminalDelivery(t *testing.T) { + source := make(chan core.FluxStreamEvent, 64) + for i := 0; i < cap(source); i++ { + source <- core.FluxStreamEvent{Type: "content", Content: "x"} + } + released := make(chan struct{}) + var releaseOnce sync.Once + ctx, cancel := context.WithCancel(context.Background()) + stream := newStream(ctx, cancel, llm.NewStreamResult(source, "request-idle", func() { + releaseOnce.Do(func() { close(released) }) + }), lifecycleRoute) + + // The consumer never calls Next: forward fills the event buffer and parks. + deadline := time.Now().Add(time.Second) + for len(stream.events) < cap(stream.events) { + if time.Now().After(deadline) { + t.Fatal("forward did not fill the event buffer") + } + time.Sleep(time.Millisecond) + } + cancel() + + select { + case <-released: + case <-time.After(time.Second): + t.Fatal("provider stream was not released while the terminal waited for the consumer") + } + if err := stream.Err(); !IsCode(err, ErrorCancelled) { + t.Fatalf("error = %v, want cancelled before the terminal is read", err) + } + + // Close unblocks forward, which then closes the event channel. + _ = stream.Close() + _ = stream.Close() + drainEngineStream(t, stream) +} + +func TestStreamNoEventsAfterDone(t *testing.T) { + ctx, cancel := context.WithCancel(context.Background()) + stream := streamOf( + ctx, cancel, + core.FluxStreamEvent{Type: "content", Content: "hi"}, + core.FluxStreamEvent{Type: "done", StopReason: "end_turn"}, + core.FluxStreamEvent{Type: "content", Content: "late"}, + core.FluxStreamEvent{Type: "error", Error: "late failure"}, + ) + defer stream.Close() + + events := drainEngineStream(t, stream) + if got := engineEventTypes(events); len(got) != 3 || got[2] != EventDone { + t.Fatalf("events = %v, want route_selected, content_delta, done", got) + } + if err := stream.Err(); err != nil { + t.Fatalf("error = %v, want nil after done", err) + } +} + +func TestStreamCloseTwiceIsNoop(t *testing.T) { + var closes, cancels int + var mu sync.Mutex + source := llm.NewStreamResult(make(chan core.FluxStreamEvent), "", func() { + mu.Lock() + closes++ + mu.Unlock() + }) + ctx, cancel := context.WithCancel(context.Background()) + stream := newStream(ctx, func() { + mu.Lock() + cancels++ + mu.Unlock() + cancel() + }, source, lifecycleRoute) + + _ = stream.Close() + _ = stream.Close() + drainEngineStream(t, stream) + _ = stream.Close() + + mu.Lock() + defer mu.Unlock() + if closes != 1 || cancels != 1 { + t.Fatalf("source closes = %d, cancels = %d, want 1 each", closes, cancels) + } + if err := stream.Err(); err != nil { + t.Fatalf("error = %v, want nil after Close", err) + } +} diff --git a/engine/types.go b/engine/types.go index 370d31e5..6c25de5e 100644 --- a/engine/types.go +++ b/engine/types.go @@ -76,6 +76,7 @@ const ( EventWarning = "warning" EventProviderBlock = "provider_block" EventTTFT = "ttft" + EventCancelled = "cancelled" EventDone = "done" ) diff --git a/llm/llm_test.go b/llm/llm_test.go index 13b5728c..c51a16b7 100644 --- a/llm/llm_test.go +++ b/llm/llm_test.go @@ -2,6 +2,7 @@ package llm_test import ( "encoding/json" + "sync/atomic" "testing" "github.com/GrayCodeAI/flux/llm" @@ -28,3 +29,17 @@ func TestLlmParity(t *testing.T) { t.Fatalf("schema parity mismatch\n got: %s\nwant: %s", got, want) } } + +func TestStreamResultCloseIsIdempotent(t *testing.T) { + t.Parallel() + + var closes atomic.Int32 + result := llm.NewStreamResult(nil, "", func() { closes.Add(1) }) + copied := *result + result.Close() + copied.Close() + + if got := closes.Load(); got != 1 { + t.Fatalf("close count = %d, want 1", got) + } +} diff --git a/llm/provider.go b/llm/provider.go index f33431fd..34a45b61 100644 --- a/llm/provider.go +++ b/llm/provider.go @@ -34,7 +34,10 @@ type Generator interface { } // EventStreamer is the pull-based host stream contract used by the engine facade. -// Next must not be called concurrently. Close is idempotent. +// Next must not be called concurrently. Close is idempotent. Callers must call +// Close (or read until Next returns false): cancelling the request context ends +// the stream with a terminal "cancelled" event, whose delivery waits for the +// consumer. type EventStreamer interface { Next() bool Event() FluxStreamEvent diff --git a/llm/types.go b/llm/types.go index 981e8aa4..1a685695 100644 --- a/llm/types.go +++ b/llm/types.go @@ -13,6 +13,7 @@ package llm import ( "context" "encoding/json" + "sync" "github.com/GrayCodeAI/flux/tools" ) @@ -84,6 +85,7 @@ const ( ErrKindUnavailable = "unavailable" ErrKindInvalidRequest = "invalid_request" ErrKindCanceled = "canceled" + ErrKindTruncated = "truncated" ErrKindInternal = "internal" ) @@ -272,10 +274,11 @@ type FluxStreamEvent struct { Warning string `json:"warning,omitempty"` // ProviderBlock is set on "provider_block" events: one completed opaque // block (for example a signed thinking block) to replay next turn. - ProviderBlock *ProviderBlock `json:"provider_block,omitempty"` - RequestID string `json:"request_id,omitempty"` - Usage *FluxUsage `json:"usage,omitempty"` - StopReason string `json:"stop_reason,omitempty"` + ProviderBlock *ProviderBlock `json:"provider_block,omitempty"` + RequestID string `json:"request_id,omitempty"` + ErrorInfo *StreamErrorInfo `json:"error_info,omitempty"` + Usage *FluxUsage `json:"usage,omitempty"` + StopReason string `json:"stop_reason,omitempty"` // TTFT and TTFTms both carry time-to-first-token in milliseconds but ride // different events: the dedicated "ttft" event populates TTFT, while the // terminal "done" event populates TTFTms. The engine normalizes the two @@ -286,8 +289,11 @@ type FluxStreamEvent struct { Route *ResolvedRoute `json:"route,omitempty"` } -// StreamResult wraps a streaming response with cleanup. Callers must call Close() -// when done reading events, or cancel the context. +// StreamResult wraps a streaming response with cleanup. Callers must read +// Events until it is closed or call Close. Cancelling the request context is +// not enough on its own: streams coordinated by provider/core then end with a +// terminal "cancelled" event, and the goroutine delivering it waits until that +// event is read or Close is called. Close is idempotent. type StreamResult struct { Events <-chan FluxStreamEvent RequestID string @@ -297,6 +303,11 @@ type StreamResult struct { // NewStreamResult constructs a stream result. The cancel function is optional // and must be idempotent. func NewStreamResult(events <-chan FluxStreamEvent, requestID string, cancel context.CancelFunc) *StreamResult { + if cancel != nil { + cleanup := cancel + var once sync.Once + cancel = func() { once.Do(cleanup) } + } return &StreamResult{Events: events, RequestID: requestID, cancel: cancel} } diff --git a/provider/adapters/anthropic.go b/provider/adapters/anthropic.go index 65b5a745..3deadbd5 100644 --- a/provider/adapters/anthropic.go +++ b/provider/adapters/anthropic.go @@ -603,13 +603,16 @@ func (c *AnthropicClient) Chat(ctx context.Context, messages []core.FluxMessage, // StreamChat sends a streaming message to Anthropic. func (c *AnthropicClient) StreamChat(ctx context.Context, messages []core.FluxMessage, opts core.ChatOptions) (*core.StreamResult, error) { - req, body, err := c.buildAnthropicRequest(ctx, messages, opts, true) + streamCtx, cancel := context.WithCancel(ctx) + req, body, err := c.buildAnthropicRequest(streamCtx, messages, opts, true) if err != nil { + cancel() return nil, err } - resp, err := c.doRequestWithMimoAuthRetry(ctx, req, body) + resp, err := c.doRequestWithMimoAuthRetry(streamCtx, req, body) if err != nil { + cancel() return nil, fmt.Errorf("flux: anthropic stream request failed: %w", err) } @@ -618,14 +621,16 @@ func (c *AnthropicClient) StreamChat(ctx context.Context, messages []core.FluxMe if resp.StatusCode != 200 { detail, readErr := core.ParseProviderError(resp.Body) _ = resp.Body.Close() + cancel() return nil, core.FormatAPIError("anthropic", "stream", resp.StatusCode, requestID, detail, readErr) } - streamCtx, cancel := context.WithCancel(ctx) - sseEvents := core.ParseSSEStream(streamCtx, resp.Body, c.logger) + streamBody, cleanup := core.BindStreamBody(streamCtx, resp.Body, cancel) + sseEvents := core.ParseSSEStream(streamCtx, streamBody, c.logger) events := core.ProcessAnthropicStream(streamCtx, sseEvents, c.logger) + result := llm.NewStreamResult(events, requestID, cleanup) - return llm.NewStreamResult(events, requestID, cancel), nil + return core.CoordinateStreamResult(ctx, result), nil } func (c *AnthropicClient) doRequestWithMimoAuthRetry(ctx context.Context, req *http.Request, body []byte) (*http.Response, error) { diff --git a/provider/adapters/azure.go b/provider/adapters/azure.go index 37ef366b..5144adfc 100644 --- a/provider/adapters/azure.go +++ b/provider/adapters/azure.go @@ -133,14 +133,17 @@ func (c *AzureClient) StreamChat(ctx context.Context, messages []core.FluxMessag return nil, fmt.Errorf("flux: model is required for azure") } + streamCtx, cancel := context.WithCancel(ctx) reqBody := c.buildRequest(messages, opts, true) body, err := json.Marshal(reqBody) if err != nil { + cancel() return nil, fmt.Errorf("flux: azure stream marshal request failed: %w", err) } url := fmt.Sprintf("%s/openai/deployments/%s/chat/completions?api-version=%s", c.endpoint, opts.Model, c.apiVersion) - req, err := http.NewRequestWithContext(ctx, "POST", url, bytes.NewReader(body)) + req, err := http.NewRequestWithContext(streamCtx, "POST", url, bytes.NewReader(body)) if err != nil { + cancel() return nil, fmt.Errorf("flux: azure stream request creation failed: %w", err) } c.setHeaders(req) @@ -149,8 +152,9 @@ func (c *AzureClient) StreamChat(ctx context.Context, messages []core.FluxMessag c.logger.Debug("azure stream", "model", opts.Model, "endpoint", c.endpoint) - resp, err := core.DoWithRetry(ctx, c.httpClient, req, c.retry, c.logger) + resp, err := core.DoWithRetry(streamCtx, c.httpClient, req, c.retry, c.logger) if err != nil { + cancel() return nil, fmt.Errorf("flux: azure stream request failed: %w", err) } @@ -158,14 +162,16 @@ func (c *AzureClient) StreamChat(ctx context.Context, messages []core.FluxMessag if resp.StatusCode != 200 { detail, readErr := core.ParseProviderError(resp.Body) _ = resp.Body.Close() + cancel() return nil, core.FormatAPIError("azure", "stream", resp.StatusCode, requestID, detail, readErr) } - streamCtx, cancel := context.WithCancel(ctx) - sseEvents := core.ParseSSEStream(streamCtx, resp.Body, c.logger) + streamBody, cleanup := core.BindStreamBody(streamCtx, resp.Body, cancel) + sseEvents := core.ParseSSEStream(streamCtx, streamBody, c.logger) events := core.ProcessOpenAIStream(streamCtx, sseEvents, c.logger) + result := llm.NewStreamResult(events, requestID, cleanup) - return llm.NewStreamResult(events, requestID, cancel), nil + return core.CoordinateStreamResult(ctx, result), nil } func (c *AzureClient) Ping(ctx context.Context) error { diff --git a/provider/adapters/bedrock.go b/provider/adapters/bedrock.go index 9ee9f0c6..e402b81b 100644 --- a/provider/adapters/bedrock.go +++ b/provider/adapters/bedrock.go @@ -8,6 +8,7 @@ import ( "encoding/binary" "encoding/hex" "encoding/json" + "errors" "fmt" "hash/crc32" "io" @@ -129,20 +130,24 @@ func (c *BedrockClient) StreamChat(ctx context.Context, messages []core.FluxMess return nil, err } + streamCtx, cancel := context.WithCancel(ctx) streamURL := strings.Replace(c.modelURL(opts.Model), "/invoke", "/invoke-with-response-stream", 1) - req, err := http.NewRequestWithContext(ctx, http.MethodPost, streamURL, bytes.NewReader(streamBody)) + req, err := http.NewRequestWithContext(streamCtx, http.MethodPost, streamURL, bytes.NewReader(streamBody)) if err != nil { + cancel() return nil, fmt.Errorf("flux: bedrock stream request creation failed: %w", err) } req.Header.Set("Accept", "application/vnd.amazon.eventstream") req.Header.Set("Content-Type", "application/json") req.GetBody = func() (io.ReadCloser, error) { return io.NopCloser(bytes.NewReader(streamBody)), nil } if err := c.sign(req, streamBody, time.Now().UTC()); err != nil { + cancel() return nil, err } - resp, err := core.DoWithRetry(ctx, c.httpClient, req, c.retry, c.logger) //nolint:bodyclose // closed in goroutine for streaming + resp, err := core.DoWithRetry(streamCtx, c.httpClient, req, c.retry, c.logger) //nolint:bodyclose // closed in goroutine for streaming if err != nil { + cancel() return nil, fmt.Errorf("flux: bedrock stream request failed: %w", err) } if resp.StatusCode != http.StatusOK { @@ -152,10 +157,12 @@ func (c *BedrockClient) StreamChat(ctx context.Context, messages []core.FluxMess } }() detail, readErr := core.ParseProviderError(resp.Body) + cancel() return nil, core.FormatAPIError("bedrock stream", "stream", resp.StatusCode, resp.Header.Get("X-Amzn-Requestid"), detail, readErr) } - streamCtx, cancel := context.WithCancel(ctx) + boundBody, cleanup := core.BindStreamBody(streamCtx, resp.Body, cancel) + resp.Body = boundBody ch := make(chan core.FluxStreamEvent, 64) go func() { @@ -190,13 +197,13 @@ func (c *BedrockClient) StreamChat(ctx context.Context, messages []core.FluxMess } }() if err != nil { - if err != io.EOF { + if !errors.Is(err, io.EOF) { select { - case ch <- core.FluxStreamEvent{Type: "error", Content: err.Error()}: + case ch <- core.FluxStreamEvent{Type: "error", Error: err.Error()}: case <-streamCtx.Done(): } } - break + return } // Parse the chunk payload @@ -238,25 +245,24 @@ func (c *BedrockClient) StreamChat(ctx context.Context, messages []core.FluxMess if chunk.Delta != nil && chunk.Delta.StopReason != "" { finishReason = chunk.Delta.StopReason } + usage = mergeBedrockUsage(usage, chunk.Usage) case "message_start": - if chunk.Message != nil && chunk.Message.Usage != nil { - usage = chunk.Message.Usage + if chunk.Message != nil { + usage = mergeBedrockUsage(usage, chunk.Message.Usage) } + case "message_stop": + doneEvt := core.FluxStreamEvent{Type: "done", StopReason: finishReason, Usage: usage} + select { + case ch <- doneEvt: + case <-streamCtx.Done(): + } + return } } - - // Send final done event with usage (matching Anthropic/OpenAI pattern) - doneEvt := core.FluxStreamEvent{Type: "done", StopReason: finishReason} - if usage != nil { - doneEvt.Usage = usage - } - select { - case ch <- doneEvt: - case <-streamCtx.Done(): - } }() - return llm.NewStreamResult(ch, resp.Header.Get("X-Amzn-Requestid"), cancel), nil + result := llm.NewStreamResult(ch, resp.Header.Get("X-Amzn-Requestid"), cleanup) + return core.CoordinateStreamResult(ctx, result), nil } func (c *BedrockClient) Ping(ctx context.Context) error { @@ -335,6 +341,11 @@ func (c *BedrockClient) modelURL(model string) string { } // anthropicStreamChunk represents a chunk from Bedrock's streaming response. +type bedrockStreamUsage struct { + InputTokens int `json:"input_tokens,omitempty"` + OutputTokens int `json:"output_tokens,omitempty"` +} + type anthropicStreamChunk struct { Type string `json:"type"` Index int `json:"index,omitempty"` @@ -351,8 +362,26 @@ type anthropicStreamChunk struct { StopReason string `json:"stop_reason,omitempty"` } `json:"delta,omitempty"` Message *struct { - Usage *core.FluxUsage `json:"usage,omitempty"` + Usage *bedrockStreamUsage `json:"usage,omitempty"` } `json:"message,omitempty"` + Usage *bedrockStreamUsage `json:"usage,omitempty"` +} + +func mergeBedrockUsage(current *core.FluxUsage, update *bedrockStreamUsage) *core.FluxUsage { + if update == nil { + return current + } + if current == nil { + current = &core.FluxUsage{} + } + if update.InputTokens > 0 { + current.PromptTokens = update.InputTokens + } + if update.OutputTokens > 0 { + current.CompletionTokens = update.OutputTokens + } + current.TotalTokens = current.PromptTokens + current.CompletionTokens + return current } // eventStreamReader parses Amazon EventStream binary frames from a reader. diff --git a/provider/adapters/bedrock_test.go b/provider/adapters/bedrock_test.go index fb11560c..06e5c144 100644 --- a/provider/adapters/bedrock_test.go +++ b/provider/adapters/bedrock_test.go @@ -12,6 +12,7 @@ import ( "testing" "time" + "github.com/GrayCodeAI/flux/llm" "github.com/GrayCodeAI/flux/provider/core" "github.com/GrayCodeAI/flux/types" ) @@ -385,7 +386,8 @@ func TestBedrockClient_StreamChat_Success(t *testing.T) { `{"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}`, `{"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":" from Bedrock stream!"}}`, `{"type":"content_block_stop","index":0}`, - `{"type":"message_delta","delta":{"stop_reason":"end_turn"}}`, + `{"type":"message_delta","delta":{"stop_reason":"end_turn"},"usage":{"output_tokens":3}}`, + `{"type":"message_stop"}`, } var eventStreamData []byte for _, c := range chunks { @@ -410,14 +412,57 @@ func TestBedrockClient_StreamChat_Success(t *testing.T) { defer result.Close() var content string + var done core.FluxStreamEvent for evt := range result.Events { if evt.Type == "content" { content += evt.Content } + if evt.Type == "done" { + done = evt + } } if content != "Hello from Bedrock stream!" { t.Errorf("content = %q", content) } + if done.StopReason != "end_turn" || done.Usage == nil || done.Usage.TotalTokens != 8 { + t.Fatalf("done = %+v", done) + } +} + +func TestBedrockClient_StreamChat_MissingMessageStopIsTruncated(t *testing.T) { + t.Parallel() + chunks := []string{ + `{"type":"message_start","message":{"usage":{"input_tokens":5}}}`, + `{"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"partial"}}`, + `{"type":"message_delta","delta":{"stop_reason":"end_turn"},"usage":{"output_tokens":1}}`, + } + var eventStreamData []byte + for _, chunk := range chunks { + eventStreamData = append(eventStreamData, buildEventStreamFrame(chunk)...) + } + transport := roundTripFunc(func(*http.Request) (*http.Response, error) { + return &http.Response{ + StatusCode: http.StatusOK, + Header: http.Header{"Content-Type": []string{"application/vnd.amazon.eventstream"}}, + Body: io.NopCloser(bytes.NewReader(eventStreamData)), + }, nil + }) + client := NewBedrockClient("AKID", "secret", "", "us-east-1") + client.retry = core.RetryConfig{RetryConfig: types.RetryConfig{MaxRetries: 0}} + client.httpClient = &http.Client{Transport: transport} + result, err := client.StreamChat(context.Background(), []core.FluxMessage{{Role: "user", Content: "Hi"}}, core.ChatOptions{Model: "anthropic.claude", MaxTokens: 256}) + if err != nil { + t.Fatal(err) + } + defer result.Close() + + var terminal core.FluxStreamEvent + for event := range result.Events { + terminal = event + } + if terminal.Type != "error" || terminal.ErrorInfo == nil || terminal.ErrorInfo.Kind != llm.ErrKindTruncated { + t.Fatalf("terminal = %+v, want truncated error", terminal) + } } func TestBedrockClient_StreamChat_Error(t *testing.T) { diff --git a/provider/adapters/concentrate_responses.go b/provider/adapters/concentrate_responses.go index 60e52faf..b4f24927 100644 --- a/provider/adapters/concentrate_responses.go +++ b/provider/adapters/concentrate_responses.go @@ -5,6 +5,7 @@ import ( "bytes" "context" "encoding/json" + "errors" "fmt" "io" "log/slog" @@ -243,7 +244,10 @@ func (c *ConcentrateResponsesClient) StreamChat(ctx context.Context, messages [] return nil, core.FormatAPIError("concentrate", "stream", resp.StatusCode, requestID, detail, readErr) } - return c.handleStream(streamCtx, cancel, resp, requestID), nil + streamBody, cleanup := core.BindStreamBody(streamCtx, resp.Body, cancel) + resp.Body = streamBody + result := c.handleStream(streamCtx, cleanup, resp, requestID) + return core.CoordinateStreamResult(ctx, result), nil } // Ping checks the health of the Concentrate API. @@ -675,11 +679,20 @@ func newSSEReader(r io.Reader) *sseReader { return &sseReader{reader: bufio.NewReader(r)} } +// errConcentrateEventTooLarge reports an SSE event over core.SSEMaxEventBytes. +var errConcentrateEventTooLarge = fmt.Errorf("concentrate: SSE event exceeds %d bytes", core.SSEMaxEventBytes) + +// Read returns the next event. Lines and the accumulated data of one event are +// bounded by core.SSEMaxEventBytes so a peer cannot grow memory without limit. func (s *sseReader) Read() (streamEvent, error) { var event streamEvent var data []string + size := 0 for { - line, err := s.reader.ReadString('\n') + line, err := s.readLine(core.SSEMaxEventBytes - size) + if errors.Is(err, errConcentrateEventTooLarge) { + return event, err + } if err != nil { if err != io.EOF { return event, err @@ -702,8 +715,28 @@ func (s *sseReader) Read() (streamEvent, error) { return decodeConcentrateSSEData(data) } if strings.HasPrefix(line, "data:") { - data = append(data, strings.TrimSpace(strings.TrimPrefix(line, "data:"))) + field := strings.TrimSpace(strings.TrimPrefix(line, "data:")) + size += len(field) + 1 + data = append(data, field) + } + } +} + +// readLine reads one line including its newline, failing once it would exceed +// limit bytes. At EOF it returns the partial line with io.EOF, like +// bufio.Reader.ReadString. +func (s *sseReader) readLine(limit int) (string, error) { + var line []byte + for { + chunk, err := s.reader.ReadSlice('\n') + if len(line)+len(chunk) > limit { + return "", errConcentrateEventTooLarge + } + line = append(line, chunk...) + if errors.Is(err, bufio.ErrBufferFull) { + continue } + return string(line), err } } diff --git a/provider/adapters/concentrate_responses_test.go b/provider/adapters/concentrate_responses_test.go index 1005af4c..85588d9a 100644 --- a/provider/adapters/concentrate_responses_test.go +++ b/provider/adapters/concentrate_responses_test.go @@ -518,3 +518,59 @@ func TestNormalizeToolParamsDoesNotMutateCallerMap(t *testing.T) { t.Fatalf("explicit additionalProperties = %#v, want preserved true", got["additionalProperties"]) } } + +type endlessSSEReader struct{ line []byte } + +func (r endlessSSEReader) Read(p []byte) (int, error) { + n := 0 + for n < len(p) { + n += copy(p[n:], r.line) + } + return n, nil +} + +func TestConcentrateSSEReaderBoundsEventSize(t *testing.T) { + t.Parallel() + tests := map[string][]byte{ + "many data lines": []byte("data: " + strings.Repeat("x", 64*1024) + "\n"), + "one line without newlines": []byte(strings.Repeat("y", 64*1024)), + } + for name, line := range tests { + t.Run(name, func(t *testing.T) { + t.Parallel() + done := make(chan error, 1) + go func() { + _, err := newSSEReader(endlessSSEReader{line: line}).Read() + done <- err + }() + select { + case err := <-done: + if !errors.Is(err, errConcentrateEventTooLarge) { + t.Fatalf("error = %v, want size limit error", err) + } + case <-time.After(10 * time.Second): + t.Fatal("reader kept accumulating an unbounded event") + } + }) + } +} + +func TestConcentrateSSEReaderParsesEvents(t *testing.T) { + t.Parallel() + body := "event: response.output_text.delta\r\n" + + "data: {\"type\":\"response.output_text.delta\",\n" + + "data: \"delta\":\"hi\"}\n\n" + + "data: [DONE]\n" + reader := newSSEReader(strings.NewReader(body)) + event, err := reader.Read() + if err != nil || event.Type != "response.output_text.delta" || event.Delta != "hi" { + t.Fatalf("event = %+v, err = %v, want the delta event", event, err) + } + event, err = reader.Read() + if err != nil || event.Type != "response.completed" { + t.Fatalf("event = %+v, err = %v, want completion", event, err) + } + if _, err := reader.Read(); !errors.Is(err, io.EOF) { + t.Fatalf("err = %v, want EOF", err) + } +} diff --git a/provider/adapters/gemini.go b/provider/adapters/gemini.go index 238941a4..397a77ee 100644 --- a/provider/adapters/gemini.go +++ b/provider/adapters/gemini.go @@ -4,6 +4,7 @@ import ( "bytes" "context" "encoding/json" + "errors" "fmt" "io" "log/slog" @@ -132,20 +133,24 @@ func (c *GeminiClient) StreamChat(ctx context.Context, messages []core.FluxMessa if opts.Model == "" { return nil, fmt.Errorf("flux: model is required for gemini") } + streamCtx, cancel := context.WithCancel(ctx) body, err := c.buildBody(messages, opts) if err != nil { + cancel() return nil, err } url := fmt.Sprintf("%s/models/%s:streamGenerateContent?alt=sse", c.baseURL, opts.Model) - req, err := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(body)) + req, err := http.NewRequestWithContext(streamCtx, http.MethodPost, url, bytes.NewReader(body)) if err != nil { + cancel() return nil, fmt.Errorf("flux: gemini stream request creation failed: %w", err) } c.setHeaders(req) req.GetBody = func() (io.ReadCloser, error) { return io.NopCloser(bytes.NewReader(body)), nil } - resp, err := core.DoWithRetry(ctx, c.httpClient, req, c.retry, c.logger) + resp, err := core.DoWithRetry(streamCtx, c.httpClient, req, c.retry, c.logger) if err != nil { + cancel() return nil, fmt.Errorf("flux: gemini stream request failed: %w", err) } @@ -154,19 +159,23 @@ func (c *GeminiClient) StreamChat(ctx context.Context, messages []core.FluxMessa if resp.StatusCode != http.StatusOK { detail, readErr := core.ParseProviderError(resp.Body) _ = resp.Body.Close() + cancel() return nil, core.FormatAPIError("gemini", "stream", resp.StatusCode, requestID, detail, readErr) } - streamCtx, cancel := context.WithCancel(ctx) + streamBody, cleanup := core.BindStreamBody(streamCtx, resp.Body, cancel) + var events <-chan core.FluxStreamEvent if geminiSharedParserEnabled() { - sseEvents := core.ParseSSEStream(streamCtx, resp.Body, c.logger) - events := processGeminiStream(streamCtx, sseEvents, c.logger) - return llm.NewStreamResult(events, requestID, cancel), nil + sseEvents := core.ParseSSEStream(streamCtx, streamBody, c.logger) + events = processGeminiStream(streamCtx, sseEvents, c.logger) + } else { + buffered := make(chan core.FluxStreamEvent, 64) + events = buffered + go c.streamLoop(streamCtx, streamBody, buffered) } - // Fallback (opt-out via FLUX_GEMINI_SHARED_PARSER=0): old bespoke parser. - events := make(chan core.FluxStreamEvent, 64) - go c.streamLoop(streamCtx, resp.Body, events) - return llm.NewStreamResult(events, requestID, cancel), nil + result := llm.NewStreamResult(events, requestID, cleanup) + + return core.CoordinateStreamResult(ctx, result), nil } func (c *GeminiClient) Ping(ctx context.Context) error { @@ -553,6 +562,19 @@ func mapGeminiFinishReason(reason string) string { } } +func geminiUsageSnapshot(usage *geminiUsage) *core.FluxUsage { + if usage == nil { + return nil + } + return &core.FluxUsage{ + PromptTokens: usage.PromptTokenCount, + CompletionTokens: usage.CandidatesTokenCount, + TotalTokens: usage.TotalTokenCount, + ThinkingTokens: usage.ThoughtsTokenCount, + CacheReadTokens: usage.CachedContentTokenCount, + } +} + // --- Streaming --- // processGeminiStream converts Gemini SSE events (parsed by the shared @@ -567,26 +589,22 @@ func mapGeminiFinishReason(reason string) string { // original Gemini "done with usage" contract). // - A bare finish reason without usage emits a "done" event with // StopReason but no Usage. -// - If the SSE channel closes without a finish reason, a bare "done" -// is emitted (matches the original "if !doneSent" fallback). func processGeminiStream(ctx context.Context, sseEvents <-chan core.SSEEvent, logger *slog.Logger) <-chan core.FluxStreamEvent { ch := make(chan core.FluxStreamEvent, core.StreamChannelBuffer) go func() { defer close(ch) - doneSent := false + var pendingUsage *core.FluxUsage for { select { case <-ctx.Done(): return case evt, ok := <-sseEvents: if !ok { - if !doneSent { - core.Emit(ctx, ch, core.FluxStreamEvent{Type: "done"}) + if pendingUsage != nil { + core.Emit(ctx, ch, core.FluxStreamEvent{Type: "usage", Usage: pendingUsage}) } return } - // Propagate SSE-level errors (raised by core.ParseSSEStream on - // scanner failure or non-cancel context expiry). if evt.Event == "error" { core.Emit(ctx, ch, core.FluxStreamEvent{Type: "error", Error: evt.Data}) return @@ -600,6 +618,9 @@ func processGeminiStream(ctx context.Context, sseEvents <-chan core.SSEEvent, lo logger.Debug("failed to parse gemini event", "error", err) continue } + if chunk.Usage != nil { + pendingUsage = geminiUsageSnapshot(chunk.Usage) + } if len(chunk.Candidates) == 0 { continue } @@ -622,25 +643,12 @@ func processGeminiStream(ctx context.Context, sseEvents <-chan core.SSEEvent, lo }) } } - // Final-chunk emission: a "done" event with Usage - // (when present) and StopReason (when present). The - // original streamLoop emitted these together in a - // single event when the chunk carried both. - if chunk.Usage != nil || candidate.FinishReason != "" { - evt := core.FluxStreamEvent{Type: "done"} - if chunk.Usage != nil { - evt.Usage = &core.FluxUsage{ - PromptTokens: chunk.Usage.PromptTokenCount, - CompletionTokens: chunk.Usage.CandidatesTokenCount, - TotalTokens: chunk.Usage.TotalTokenCount, - ThinkingTokens: chunk.Usage.ThoughtsTokenCount, - CacheReadTokens: chunk.Usage.CachedContentTokenCount, - } - } - if candidate.FinishReason != "" { - evt.StopReason = mapGeminiFinishReason(candidate.FinishReason) - } - core.Emit(ctx, ch, evt) + if candidate.FinishReason != "" { + core.Emit(ctx, ch, core.FluxStreamEvent{ + Type: "done", + StopReason: mapGeminiFinishReason(candidate.FinishReason), + Usage: pendingUsage, + }) return } } @@ -653,7 +661,16 @@ func (c *GeminiClient) streamLoop(ctx context.Context, body io.ReadCloser, event defer close(events) defer func() { _ = body.Close() }() - doneSent := false + var pendingUsage *core.FluxUsage + processLine := func(line string) bool { + line = strings.TrimSpace(line) + if line == "" || !strings.HasPrefix(line, "data:") { + return false + } + data := strings.TrimSpace(strings.TrimPrefix(line, "data:")) + return c.processStreamChunkWithUsage(ctx, data, events, &pendingUsage) + } + buf := make([]byte, 0, 64*1024) tmp := make([]byte, 4096) for { @@ -667,34 +684,46 @@ func (c *GeminiClient) streamLoop(ctx context.Context, body io.ReadCloser, event } line := string(buf[:idx]) buf = buf[idx+1:] - line = strings.TrimSpace(line) - if line == "" || !strings.HasPrefix(line, "data: ") { - continue - } - data := strings.TrimPrefix(line, "data: ") - if c.processStreamChunk(ctx, data, events) { - doneSent = true + if processLine(line) { + return } } } if readErr != nil { - break - } - } - - if !doneSent { - select { - case events <- core.FluxStreamEvent{Type: "done"}: - case <-ctx.Done(): + if len(buf) > 0 && processLine(string(buf)) { + return + } + if !errors.Is(readErr, io.EOF) { + select { + case events <- core.FluxStreamEvent{Type: "error", Error: readErr.Error()}: + case <-ctx.Done(): + } + return + } + if pendingUsage != nil { + select { + case events <- core.FluxStreamEvent{Type: "usage", Usage: pendingUsage}: + case <-ctx.Done(): + } + } + return } } } func (c *GeminiClient) processStreamChunk(ctx context.Context, data string, events chan<- core.FluxStreamEvent) bool { + var pendingUsage *core.FluxUsage + return c.processStreamChunkWithUsage(ctx, data, events, &pendingUsage) +} + +func (c *GeminiClient) processStreamChunkWithUsage(ctx context.Context, data string, events chan<- core.FluxStreamEvent, pendingUsage **core.FluxUsage) bool { var chunk geminiResponse if err := json.Unmarshal([]byte(data), &chunk); err != nil { return false } + if chunk.Usage != nil { + *pendingUsage = geminiUsageSnapshot(chunk.Usage) + } if len(chunk.Candidates) == 0 { return false } @@ -725,18 +754,9 @@ func (c *GeminiClient) processStreamChunk(ctx context.Context, data string, even } } } - if chunk.Usage != nil { + if candidate.FinishReason != "" { select { - case events <- core.FluxStreamEvent{ - Type: "done", - Usage: &core.FluxUsage{ - PromptTokens: chunk.Usage.PromptTokenCount, - CompletionTokens: chunk.Usage.CandidatesTokenCount, - TotalTokens: chunk.Usage.TotalTokenCount, - ThinkingTokens: chunk.Usage.ThoughtsTokenCount, - CacheReadTokens: chunk.Usage.CachedContentTokenCount, - }, - }: + case events <- core.FluxStreamEvent{Type: "done", StopReason: mapGeminiFinishReason(candidate.FinishReason), Usage: *pendingUsage}: return true case <-ctx.Done(): } diff --git a/provider/adapters/gemini_test.go b/provider/adapters/gemini_test.go index be244a7f..a43e7f82 100644 --- a/provider/adapters/gemini_test.go +++ b/provider/adapters/gemini_test.go @@ -10,6 +10,7 @@ import ( "strings" "testing" + "github.com/GrayCodeAI/flux/llm" "github.com/GrayCodeAI/flux/provider/core" "github.com/GrayCodeAI/flux/types" ) @@ -363,7 +364,7 @@ func TestProcessGeminiStream_ToolCallWithUsage(t *testing.T) { t.Errorf("tool call id = %q", evt.ToolCall.ID) } } - if evt.Type == "done" && evt.Usage != nil { + if evt.Type == "usage" && evt.Usage != nil { usage = evt.Usage } } @@ -668,14 +669,12 @@ func TestGeminiClient_StreamChat_Legacy_InvalidJSON(t *testing.T) { t.Fatalf("StreamChat: %v", err) } defer result.Close() - var gotDone bool + var terminal core.FluxStreamEvent for evt := range result.Events { - if evt.Type == "done" { - gotDone = true - } + terminal = evt } - if !gotDone { - t.Error("expected done event") + if terminal.Type != "error" || terminal.ErrorInfo == nil || terminal.ErrorInfo.Kind != llm.ErrKindTruncated { + t.Fatalf("terminal = %+v, want truncated error", terminal) } } @@ -699,14 +698,12 @@ func TestGeminiClient_StreamChat_Legacy_NoCandidates(t *testing.T) { t.Fatalf("StreamChat: %v", err) } defer result.Close() - var gotDone bool + var terminal core.FluxStreamEvent for evt := range result.Events { - if evt.Type == "done" { - gotDone = true - } + terminal = evt } - if !gotDone { - t.Error("expected done event") + if terminal.Type != "error" || terminal.ErrorInfo == nil || terminal.ErrorInfo.Kind != llm.ErrKindTruncated { + t.Fatalf("terminal = %+v, want truncated error", terminal) } } @@ -764,20 +761,19 @@ func TestGeminiClient_StreamChat_Legacy_NoUsage(t *testing.T) { t.Fatalf("StreamChat: %v", err) } defer result.Close() - var gotContent, gotDone bool + var gotContent bool + var terminal core.FluxStreamEvent for evt := range result.Events { if evt.Type == "content" { gotContent = true } - if evt.Type == "done" { - gotDone = true - } + terminal = evt } if !gotContent { t.Error("expected content event") } - if !gotDone { - t.Error("expected done event (from streamLoop fallback)") + if terminal.Type != "error" || terminal.ErrorInfo == nil || terminal.ErrorInfo.Kind != llm.ErrKindTruncated { + t.Fatalf("terminal = %+v, want truncated error", terminal) } } @@ -1013,7 +1009,7 @@ func TestProcessStreamChunk_UsageDone(t *testing.T) { t.Parallel() c := NewGeminiClient("key", "https://gemini.example") events := make(chan core.FluxStreamEvent, 2) - data := `{"candidates":[{"content":{"parts":[{"text":"Hi"}],"finishReason":"STOP"}}],"usageMetadata":{"promptTokenCount":2,"candidatesTokenCount":5,"totalTokenCount":7}}` + data := `{"candidates":[{"content":{"parts":[{"text":"Hi"}]},"finishReason":"STOP"}],"usageMetadata":{"promptTokenCount":2,"candidatesTokenCount":5,"totalTokenCount":7}}` cont := c.processStreamChunk(context.Background(), data, events) if !cont { t.Fatal("processStreamChunk returned false, expected true (done)") diff --git a/provider/adapters/openai.go b/provider/adapters/openai.go index f4e7f1a3..9dae752e 100644 --- a/provider/adapters/openai.go +++ b/provider/adapters/openai.go @@ -553,13 +553,16 @@ func (c *OpenAIClient) Chat(ctx context.Context, messages []core.FluxMessage, op // StreamChat sends a streaming request. func (c *OpenAIClient) StreamChat(ctx context.Context, messages []core.FluxMessage, opts core.ChatOptions) (*core.StreamResult, error) { - req, body, err := c.buildOpenAIRequest(ctx, messages, opts, true) + streamCtx, cancel := context.WithCancel(ctx) + req, body, err := c.buildOpenAIRequest(streamCtx, messages, opts, true) if err != nil { + cancel() return nil, err } - resp, err := c.doRequestWithMimoAuthRetry(ctx, req, body) + resp, err := c.doRequestWithMimoAuthRetry(streamCtx, req, body) if err != nil { + cancel() return nil, fmt.Errorf("flux: %s stream request failed: %w", c.providerName, err) } @@ -568,14 +571,16 @@ func (c *OpenAIClient) StreamChat(ctx context.Context, messages []core.FluxMessa if resp.StatusCode != 200 { detail, readErr := core.ParseProviderError(resp.Body) _ = resp.Body.Close() + cancel() return nil, core.FormatAPIError(c.providerName, "stream", resp.StatusCode, requestID, detail, readErr) } - streamCtx, cancel := context.WithCancel(ctx) - sseEvents := core.ParseSSEStream(streamCtx, resp.Body, c.logger) + streamBody, cleanup := core.BindStreamBody(streamCtx, resp.Body, cancel) + sseEvents := core.ParseSSEStream(streamCtx, streamBody, c.logger) events := core.ProcessOpenAIStream(streamCtx, sseEvents, c.logger) + result := llm.NewStreamResult(events, requestID, cleanup) - return llm.NewStreamResult(events, requestID, cancel), nil + return core.CoordinateStreamResult(ctx, result), nil } // Ping checks connectivity. diff --git a/provider/adapters/openai_test.go b/provider/adapters/openai_test.go index b85633e4..4ee7487c 100644 --- a/provider/adapters/openai_test.go +++ b/provider/adapters/openai_test.go @@ -6,6 +6,7 @@ import ( "net/http" "strings" "testing" + "time" "github.com/GrayCodeAI/flux/provider/core" "github.com/GrayCodeAI/flux/types" @@ -132,6 +133,48 @@ func TestOpenAIClient_StreamChat_Success(t *testing.T) { } } +func TestOpenAIClient_StreamChat_CloseClosesBody(t *testing.T) { + t.Parallel() + body := newBlockingReadCloser() + transport := roundTripFunc(func(*http.Request) (*http.Response, error) { + return &http.Response{ + StatusCode: http.StatusOK, + Header: http.Header{"Content-Type": []string{"text/event-stream"}}, + Body: body, + }, nil + }) + c := NewOpenAIClient("sk-test", "https://api.openai.com/v1", nil) + c.SetRetry(core.RetryConfig{RetryConfig: types.RetryConfig{MaxRetries: 0}}) + c.httpClient = &http.Client{Transport: transport} + result, err := c.StreamChat(context.Background(), []core.FluxMessage{{Role: "user", Content: "Hi"}}, core.ChatOptions{Model: "gpt-4o", MaxTokens: 256}) + if err != nil { + t.Fatalf("StreamChat: %v", err) + } + + select { + case <-body.started: + case <-time.After(2 * time.Second): + t.Fatal("stream body read did not start") + } + result.Close() + select { + case <-body.closed: + case <-time.After(2 * time.Second): + t.Fatal("stream body did not close") + } + if got := body.closeCount.Load(); got != 1 { + t.Fatalf("body close count = %d, want 1", got) + } + select { + case _, ok := <-result.Events: + if ok { + t.Fatal("unexpected event after close") + } + case <-time.After(2 * time.Second): + t.Fatal("result event channel did not close") + } +} + func TestOpenAIClient_StreamChat_EmptyModel(t *testing.T) { t.Parallel() c := NewOpenAIClient("key", "", nil) diff --git a/provider/adapters/protocol_router.go b/provider/adapters/protocol_router.go index d224a1ed..bce2a29b 100644 --- a/provider/adapters/protocol_router.go +++ b/provider/adapters/protocol_router.go @@ -75,22 +75,27 @@ func (r ProtocolRouter) StreamChat(ctx context.Context, messages []core.FluxMess result, err := primaryClient.StreamChat(ctx, messages, opts) if err != nil { if cfg.FallbackOnError != nil && cfg.FallbackOnError(err) { - return fallbackClient.StreamChat(ctx, messages, opts) + fallbackResult, fallbackErr := fallbackClient.StreamChat(ctx, messages, opts) + if fallbackErr != nil { + return fallbackResult, fallbackErr + } + return core.CoordinateStreamResult(ctx, fallbackResult), nil } return result, err } if cfg.ReasoningOnlyFallback && cfg.Primary == ChatProtocolMessages { fallback := fallbackClient - return newStreamWithReasoningFallback(ctx, messages, opts, result, protocolStreamFallback{ + fallbackStream := newStreamWithReasoningFallback(ctx, messages, opts, result, protocolStreamFallback{ chat: func(ctx context.Context, messages []core.FluxMessage, opts core.ChatOptions) (*core.FluxResponse, error) { return fallback.Chat(ctx, messages, opts) }, stream: func(ctx context.Context, messages []core.FluxMessage, opts core.ChatOptions) (*core.StreamResult, error) { return fallback.StreamChat(ctx, messages, opts) }, - }), nil + }) + return core.CoordinateStreamResult(ctx, fallbackStream), nil } - return result, nil + return core.CoordinateStreamResult(ctx, result), nil } func (r ProtocolRouter) providers(primary ChatProtocol) (core.Provider, core.Provider) { diff --git a/provider/adapters/test_helpers_test.go b/provider/adapters/test_helpers_test.go index dcef8bad..53c36412 100644 --- a/provider/adapters/test_helpers_test.go +++ b/provider/adapters/test_helpers_test.go @@ -5,6 +5,8 @@ import ( "encoding/json" "io" "net/http" + "sync" + "sync/atomic" ) type roundTripFunc func(*http.Request) (*http.Response, error) @@ -35,3 +37,27 @@ func jsonDecodeRequest(req *http.Request, value any) error { } return json.Unmarshal(body, value) } + +type blockingReadCloser struct { + started chan struct{} + closed chan struct{} + startOnce sync.Once + closeCount atomic.Int32 +} + +func newBlockingReadCloser() *blockingReadCloser { + return &blockingReadCloser{started: make(chan struct{}), closed: make(chan struct{})} +} + +func (r *blockingReadCloser) Read([]byte) (int, error) { + r.startOnce.Do(func() { close(r.started) }) + <-r.closed + return 0, io.EOF +} + +func (r *blockingReadCloser) Close() error { + if r.closeCount.Add(1) == 1 { + close(r.closed) + } + return nil +} diff --git a/provider/adapters/vertex.go b/provider/adapters/vertex.go index 108bfd38..402de9da 100644 --- a/provider/adapters/vertex.go +++ b/provider/adapters/vertex.go @@ -104,43 +104,49 @@ func (c *VertexClient) StreamChat(ctx context.Context, messages []core.FluxMessa if opts.Model == "" { return nil, fmt.Errorf("flux: model is required for vertex") } + streamCtx, cancel := context.WithCancel(ctx) body, err := c.buildBody(messages, opts, true) if err != nil { + cancel() return nil, err } // Check request size (30 MB Vertex limit) if len(body) > maxVertexRequestSize { + cancel() return nil, fmt.Errorf("flux: request size %d bytes exceeds Vertex limit of %d bytes", len(body), maxVertexRequestSize) } url := fmt.Sprintf("%s/%s:streamRawPredict", c.baseURL(), opts.Model) - req, err := http.NewRequestWithContext(ctx, "POST", url, bytes.NewReader(body)) + req, err := http.NewRequestWithContext(streamCtx, "POST", url, bytes.NewReader(body)) if err != nil { + cancel() return nil, fmt.Errorf("flux: vertex stream request failed: %w", err) } c.setHeaders(req) req.Header.Set("Accept", "text/event-stream") req.GetBody = func() (io.ReadCloser, error) { return io.NopCloser(bytes.NewReader(body)), nil } - resp, err := core.DoWithRetry(ctx, c.httpClient, req, c.retry, c.logger) + resp, err := core.DoWithRetry(streamCtx, c.httpClient, req, c.retry, c.logger) if err != nil { + cancel() return nil, fmt.Errorf("flux: vertex stream request failed: %w", err) } if resp.StatusCode != 200 { detail, readErr := core.ParseProviderError(resp.Body) _ = resp.Body.Close() + cancel() return nil, core.FormatAPIError("vertex", "stream", resp.StatusCode, resp.Header.Get("X-Goog-Request-Id"), detail, readErr) } requestID := resp.Header.Get("X-Goog-Request-Id") - - streamCtx, cancel := context.WithCancel(ctx) - sseEvents := core.ParseSSEStream(streamCtx, resp.Body, c.logger) + streamBody, cleanup := core.BindStreamBody(streamCtx, resp.Body, cancel) + sseEvents := core.ParseSSEStream(streamCtx, streamBody, c.logger) events := core.ProcessAnthropicStream(streamCtx, sseEvents, c.logger) + result := llm.NewStreamResult(events, requestID, cleanup) - return llm.NewStreamResult(events, requestID, cancel), nil + return core.CoordinateStreamResult(ctx, result), nil } func (c *VertexClient) Ping(ctx context.Context) error { diff --git a/provider/budget_provider_test.go b/provider/budget_provider_test.go index 37c457b8..ece61606 100644 --- a/provider/budget_provider_test.go +++ b/provider/budget_provider_test.go @@ -112,3 +112,93 @@ func TestActualCostUSD(t *testing.T) { t.Errorf("unexpected cost %f", cost) } } + +type doneUsageStreamProvider struct { + includeUsageEvent bool +} + +func (*doneUsageStreamProvider) Name() string { return "done-usage" } +func (*doneUsageStreamProvider) Ping(context.Context) error { return nil } +func (*doneUsageStreamProvider) Chat(context.Context, []FluxMessage, ChatOptions) (*FluxResponse, error) { + return nil, nil +} + +func (p *doneUsageStreamProvider) StreamChat(context.Context, []FluxMessage, ChatOptions) (*StreamResult, error) { + events := make(chan FluxStreamEvent, 2) + usage := &FluxUsage{PromptTokens: 3, CompletionTokens: 5, TotalTokens: 8} + if p.includeUsageEvent { + events <- FluxStreamEvent{Type: "usage", Usage: usage} + } + events <- FluxStreamEvent{Type: "done", Usage: usage} + close(events) + return NewStreamResult(events, func() {}), nil +} + +func TestBudgetProvider_DeduplicatesStreamUsage(t *testing.T) { + t.Parallel() + for _, includeUsageEvent := range []bool{false, true} { + t.Run(map[bool]string{false: "done_only", true: "usage_and_done"}[includeUsageEvent], func(t *testing.T) { + store := observability.NewMemoryBudgetStore() + store.SetBudget("stream", 1) + provider := observability.NewBudgetProvider(&doneUsageStreamProvider{includeUsageEvent: includeUsageEvent}, store) + result, err := provider.StreamChat( + context.Background(), + []FluxMessage{{Role: "user", Content: "hello"}}, + ChatOptions{Model: "gpt-4o", MaxTokens: 100, VirtualKeyID: "stream"}, + ) + if err != nil { + t.Fatal(err) + } + defer result.Close() + for range result.Events { + } + + _, input, output, ok := store.Usage("stream") + if !ok || input != 3 || output != 5 { + t.Fatalf("usage = in:%d out:%d ok:%t, want 3/5", input, output, ok) + } + }) + } +} + +type continuationUsageStreamProvider struct{} + +func (*continuationUsageStreamProvider) Name() string { return "continuation-usage" } +func (*continuationUsageStreamProvider) Ping(context.Context) error { return nil } +func (*continuationUsageStreamProvider) Chat(context.Context, []FluxMessage, ChatOptions) (*FluxResponse, error) { + return nil, nil +} + +func (*continuationUsageStreamProvider) StreamChat(context.Context, []FluxMessage, ChatOptions) (*StreamResult, error) { + events := make(chan FluxStreamEvent, 4) + usage := &FluxUsage{PromptTokens: 3, CompletionTokens: 5, TotalTokens: 8} + events <- FluxStreamEvent{Type: "usage", Usage: usage} + events <- FluxStreamEvent{Type: "continuation"} + events <- FluxStreamEvent{Type: "usage", Usage: usage} + events <- FluxStreamEvent{Type: "done", Usage: usage} + close(events) + return NewStreamResult(events, func() {}), nil +} + +func TestBudgetProvider_ResetsUsageAtContinuation(t *testing.T) { + t.Parallel() + store := observability.NewMemoryBudgetStore() + store.SetBudget("stream", 1) + provider := observability.NewBudgetProvider(&continuationUsageStreamProvider{}, store) + result, err := provider.StreamChat( + context.Background(), + []FluxMessage{{Role: "user", Content: "hello"}}, + ChatOptions{Model: "gpt-4o", MaxTokens: 100, VirtualKeyID: "stream"}, + ) + if err != nil { + t.Fatal(err) + } + defer result.Close() + for range result.Events { + } + + _, input, output, ok := store.Usage("stream") + if !ok || input != 6 || output != 10 { + t.Fatalf("usage = in:%d out:%d ok:%t, want 6/10", input, output, ok) + } +} diff --git a/provider/cache/semantic_cache.go b/provider/cache/semantic_cache.go index b5fa1e0b..e60b0a44 100644 --- a/provider/cache/semantic_cache.go +++ b/provider/cache/semantic_cache.go @@ -5,7 +5,6 @@ import ( "crypto/sha256" "encoding/hex" "encoding/json" - "fmt" "sync" "time" @@ -115,6 +114,9 @@ func (cp *CachedProvider) Chat(ctx context.Context, messages []FluxMessage, opts } key := buildCacheKey(messages, opts) + if key == "" { + return cp.inner.Chat(ctx, messages, opts) + } // Fast path: read lock lookup. if resp, ok := cp.get(key); ok { @@ -293,53 +295,36 @@ func (cp *CachedProvider) evictExpiredLocked() { } } -// buildCacheKey produces a deterministic hash of the request parameters that -// affect the response: system prompt, messages, model, and temperature. +// cacheKeyVersion is part of every key so a change to the key derivation can +// never serve an entry written under the old one. +const cacheKeyVersion = "flux-response-cache/v2" + +// buildCacheKey hashes everything that can change a response: every +// ChatOptions field (tools, tool choice, max tokens, stop sequences, sampling, +// thinking, response format, identity and routing hints) and every message +// field (content parts, images, tool calls and results, provider blocks). A +// spurious miss only costs a provider call; a spurious hit serves a reply +// produced for a different request. +// +// JSON is the request's wire form, so requests that encode identically are +// identical to the provider. It returns "" when the request cannot be +// encoded (for example an unsupported Conversation value or a NaN sampling +// parameter); callers must then bypass the cache. func buildCacheKey(messages []FluxMessage, opts ChatOptions) string { - h := sha256.New() - - // Model. - h.Write([]byte("model:")) - h.Write([]byte(opts.Model)) - h.Write([]byte{0}) - - // System prompt. - h.Write([]byte("system:")) - h.Write([]byte(opts.System)) - h.Write([]byte{0}) - - // Temperature (serialized as string for determinism). - h.Write([]byte("temp:")) - if opts.Temperature != nil { - _, _ = fmt.Fprintf(h, "%.6f", *opts.Temperature) - } else { - h.Write([]byte("nil")) - } - h.Write([]byte{0}) - - // Messages: serialize role + content for each message. - for _, m := range messages { - h.Write([]byte("msg:")) - h.Write([]byte(m.Role)) - h.Write([]byte{0}) - h.Write([]byte(m.Content)) - h.Write([]byte{0}) - // Include tool calls and results if present. - if len(m.ToolUse) > 0 { - b, _ := json.Marshal(m.ToolUse) - h.Write(b) - } - if len(m.ToolResults) > 0 { - b, _ := json.Marshal(m.ToolResults) - h.Write(b) - } - h.Write([]byte{0}) + payload, err := json.Marshal(struct { + Version string `json:"version"` + Options ChatOptions `json:"options"` + Messages []FluxMessage `json:"messages"` + }{cacheKeyVersion, opts, messages}) + if err != nil { + return "" } - - return hex.EncodeToString(h.Sum(nil)) + sum := sha256.Sum256(payload) + return hex.EncodeToString(sum[:]) } -// BuildCacheKey returns the deterministic key used by the response cache. +// BuildCacheKey returns the deterministic key used by the response cache, or +// "" when the request cannot be cached. func BuildCacheKey(messages []FluxMessage, opts ChatOptions) string { return buildCacheKey(messages, opts) } diff --git a/provider/cache/semantic_cache_test.go b/provider/cache/semantic_cache_test.go index 610b7cbb..144b6f24 100644 --- a/provider/cache/semantic_cache_test.go +++ b/provider/cache/semantic_cache_test.go @@ -445,3 +445,82 @@ func TestBuildCacheKeyDeterministic(t *testing.T) { t.Error("different system prompts should produce different keys") } } + +func TestBuildCacheKeyCoversRequestShape(t *testing.T) { + t.Parallel() + msgs := []FluxMessage{{Role: "user", Content: "list files"}} + base := ChatOptions{Model: "gpt-4"} + baseKey := buildCacheKey(msgs, base) + if baseKey == "" { + t.Fatal("base key is empty") + } + + topP := 0.5 + topK := 40 + thinking := true + variants := map[string]ChatOptions{ + "tools": {Model: "gpt-4", Tools: []llm.FluxTool{{Name: "read_file"}}}, + "tool choice": {Model: "gpt-4", ToolChoice: &llm.ToolChoiceOption{Type: "required"}}, + "max tokens": {Model: "gpt-4", MaxTokens: 16}, + "stop sequences": {Model: "gpt-4", StopSequences: []string{"\n"}}, + "top p": {Model: "gpt-4", TopP: &topP}, + "top k": {Model: "gpt-4", TopK: &topK}, + "thinking": {Model: "gpt-4", ThinkingEnabled: &thinking, ThinkingBudgetTokens: 1024}, + "reasoning": {Model: "gpt-4", ReasoningEffort: "high"}, + "response format": {Model: "gpt-4", ResponseFormat: &llm.ResponseFormat{Type: "json_object"}}, + "provider": {Model: "gpt-4", Provider: "azure"}, + "caller identity": {Model: "gpt-4", MetadataUserID: "user-2"}, + } + for name, opts := range variants { + if key := buildCacheKey(msgs, opts); key == baseKey { + t.Errorf("%s: key unchanged, a cached reply would be served for a different request", name) + } + } + + messageVariants := map[string][]FluxMessage{ + "image": {{Role: "user", Content: "list files", Images: []string{"data:image/png;base64,AAAA"}}}, + "content parts": {{Role: "user", Content: "list files", ContentParts: []llm.ContentPart{{Type: "text", Text: "extra"}}}}, + "thinking": {{Role: "user", Content: "list files", Thinking: "plan"}}, + "provider blocks": {{Role: "user", Content: "list files", ProviderBlocks: []llm.ProviderBlock{{Provider: "anthropic", Type: "thinking", Data: []byte(`{}`)}}}}, + } + for name, variant := range messageVariants { + if key := buildCacheKey(variant, base); key == baseKey { + t.Errorf("message %s: key unchanged", name) + } + } +} + +func TestCachedProviderDoesNotShareRepliesAcrossToolSets(t *testing.T) { + t.Parallel() + inner := newCacheMock("fixed") + inner.Response = "tool reply" + cp := NewCachedProvider(inner, CacheConfig{Enabled: true}) + msgs := []FluxMessage{{Role: "user", Content: "list files"}} + + _, _ = cp.Chat(context.Background(), msgs, ChatOptions{Model: "gpt-4", Tools: []llm.FluxTool{{Name: "read_file"}}}) + _, _ = cp.Chat(context.Background(), msgs, ChatOptions{Model: "gpt-4", Tools: []llm.FluxTool{{Name: "delete_file"}}}) + if got := inner.CallCount(); got != 2 { + t.Fatalf("inner calls = %d, want 2: different tool sets must not share a cached reply", got) + } +} + +func TestCachedProviderBypassesUnencodableRequests(t *testing.T) { + t.Parallel() + inner := newCacheMock("fixed") + inner.Response = "fresh" + cp := NewCachedProvider(inner, CacheConfig{Enabled: true}) + msgs := []FluxMessage{{Role: "user", Content: "hi"}} + opts := ChatOptions{Model: "gpt-4", Conversation: func() {}} + + if key := buildCacheKey(msgs, opts); key != "" { + t.Fatalf("key = %q, want empty for an unencodable request", key) + } + _, _ = cp.Chat(context.Background(), msgs, opts) + _, _ = cp.Chat(context.Background(), msgs, opts) + if got := inner.CallCount(); got != 2 { + t.Fatalf("inner calls = %d, want 2: unencodable requests must bypass the cache", got) + } + if got := cp.CacheStats().Size; got != 0 { + t.Fatalf("cache entries = %d, want 0", got) + } +} diff --git a/provider/continuation_test.go b/provider/continuation_test.go index 23e4fc70..125c91fc 100644 --- a/provider/continuation_test.go +++ b/provider/continuation_test.go @@ -2,6 +2,7 @@ package provider import ( "context" + "sync/atomic" "testing" ) @@ -239,6 +240,50 @@ func TestContinuation_StreamContinuesOnMaxTokens(t *testing.T) { } } +type fatalStreamProvider struct { + closes atomic.Int32 +} + +func (*fatalStreamProvider) Name() string { return "fatal-stream" } +func (*fatalStreamProvider) Ping(context.Context) error { return nil } +func (*fatalStreamProvider) Chat(context.Context, []FluxMessage, ChatOptions) (*FluxResponse, error) { + return nil, nil +} + +func (p *fatalStreamProvider) StreamChat(context.Context, []FluxMessage, ChatOptions) (*StreamResult, error) { + events := make(chan FluxStreamEvent, 2) + events <- FluxStreamEvent{Type: "error", Error: "connection reset"} + events <- FluxStreamEvent{Type: "done", StopReason: "stop"} + close(events) + return NewStreamResult(events, func() { p.closes.Add(1) }), nil +} + +func TestContinuation_StreamFatalErrorClosesSourceWithoutDone(t *testing.T) { + t.Parallel() + provider := &fatalStreamProvider{} + result, err := StreamChatWithContinuation( + context.Background(), provider, + []FluxMessage{{Role: "user", Content: "hello"}}, + ChatOptions{Model: "test"}, + ContinuationConfig{MaxContinuations: 2, MaxTotalTokens: 100}, + ) + if err != nil { + t.Fatal(err) + } + defer result.Close() + + var events []FluxStreamEvent + for event := range result.Events { + events = append(events, event) + } + if len(events) != 1 || events[0].Type != "error" || events[0].Error != "connection reset" { + t.Fatalf("events = %+v, want one fatal error", events) + } + if got := provider.closes.Load(); got != 1 { + t.Fatalf("source close count = %d, want 1", got) + } +} + // --- test helpers --- // mockResponse defines a canned response for the sequentialMock. diff --git a/provider/core/stream.go b/provider/core/stream.go index 6f083c7b..4272eef2 100644 --- a/provider/core/stream.go +++ b/provider/core/stream.go @@ -10,7 +10,10 @@ import ( "log/slog" "sort" "strings" + "sync" "time" + + "github.com/GrayCodeAI/flux/llm" ) // SSEEvent represents a single Server-Sent Event. @@ -26,11 +29,435 @@ const ( sseScannerMaxBuf = 2 * 1024 * 1024 // StreamChannelBuffer is the default buffer size for provider event channels. StreamChannelBuffer = 256 + // SSEMaxEventBytes caps the accumulated fields of a single SSE event. + // Each line is already capped at 2 MiB; without this cap a peer could + // stream an unbounded run of data lines into one event and exhaust memory. + SSEMaxEventBytes = 16 * 1024 * 1024 ) +// errSSEEventTooLarge reports an SSE event over SSEMaxEventBytes. +var errSSEEventTooLarge = fmt.Errorf("SSE event exceeds %d bytes", SSEMaxEventBytes) + +type closeOnceReadCloser struct { + io.ReadCloser + once sync.Once + err error +} + +func (r *closeOnceReadCloser) Close() error { + r.once.Do(func() { r.err = r.ReadCloser.Close() }) + return r.err +} + +func BindStreamBody(ctx context.Context, body io.ReadCloser, cancel context.CancelFunc) (io.ReadCloser, context.CancelFunc) { + wrapped := &closeOnceReadCloser{ReadCloser: body} + stop := context.AfterFunc(ctx, func() { _ = wrapped.Close() }) + var once sync.Once + cleanup := func() { + once.Do(func() { + cancel() + stop() + _ = wrapped.Close() + }) + } + return wrapped, cleanup +} + +var ErrStreamTruncated = errors.New("stream ended before terminal event") + +// coordinatedStreams records, for the Events channel of every live stream +// produced by TransformStreamResult, the context and request ID coordinating +// it. CoordinateStreamResult uses it to avoid stacking an identical lifecycle +// wrapper (and its goroutine and buffer) on every layer a stream crosses. +var coordinatedStreams sync.Map // map[<-chan FluxStreamEvent]streamCoordination + +type streamCoordination struct { + done <-chan struct{} + requestID string +} + +type StreamEventHandler func(context.Context, FluxStreamEvent) (FluxStreamEvent, error) + +// TransformStreamResult forwards source through handler (nil forwards events +// unchanged) and enforces the stream lifecycle: exactly one terminal event +// (done, error or cancelled), a terminal cancelled event carrying +// StreamErrorInfo when ctx ends, a truncated error when the source closes +// without a terminal, and StreamErrorInfo on every terminal error. +// +// Cancelling ctx releases the source immediately, but the terminal event is +// delivered like any other: the forwarding goroutine exits once the consumer +// reads it or calls Close. Consumers must therefore read Events until it is +// closed or call Close; Close is idempotent. +func TransformStreamResult(ctx context.Context, source *StreamResult, handler StreamEventHandler) *StreamResult { + if source == nil { + return nil + } + if ctx == nil { + ctx = context.Background() + } + + streamCtx, cancel := context.WithCancel(ctx) + out := make(chan FluxStreamEvent, cap(source.Events)) + closed := make(chan struct{}) + var releaseOnce, closeOnce sync.Once + // releaseSource stops the upstream request and its connection. + releaseSource := func() { + releaseOnce.Do(func() { + cancel() + source.Close() + }) + } + closeSource := func() { + closeOnce.Do(func() { + close(closed) + releaseSource() + }) + } + + // emitCancellation delivers the terminal cancelled event when the + // caller's context ended. Close (closed) suppresses it: the consumer has + // stopped listening. The upstream is released before the terminal is + // handed over, so a consumer that cancels but neither drains nor closes + // the stream strands this goroutine only, not the provider connection. + emitCancellation := func() { + if channelClosed(closed) || ctx.Err() == nil { + return + } + releaseSource() + _ = sendTerminalEvent(out, cancellationEvent(source.RequestID, ctx.Err()), closed) + } + + var events <-chan FluxStreamEvent = out + coordinatedStreams.Store(events, streamCoordination{done: ctx.Done(), requestID: source.RequestID}) + + go func() { + defer close(out) + defer coordinatedStreams.Delete(events) + defer closeSource() + + for { + select { + case <-streamCtx.Done(): + emitCancellation() + return + case event, ok := <-source.Events: + if streamCtx.Err() != nil { + emitCancellation() + return + } + sourceEnded := !ok + if sourceEnded { + if ctx.Err() != nil { + emitCancellation() + return + } + event = FluxStreamEvent{ + Type: "error", + Error: ErrStreamTruncated.Error(), + RequestID: source.RequestID, + ErrorInfo: &llm.StreamErrorInfo{Kind: llm.ErrKindTruncated, Retryable: true}, + } + } + if handler != nil { + var err error + event, err = handler(streamCtx, event) + if err != nil { + if streamCtx.Err() != nil { + emitCancellation() + return + } + event = streamErrorEvent(source.RequestID, err) + } + } + if event.RequestID == "" { + event.RequestID = source.RequestID + } + event = ensureStreamErrorInfo(event) + if !sendLifecycleEvent(streamCtx, out, event) { + // Cancelled while the consumer was behind: the pending + // event is dropped, but the terminal still follows. + emitCancellation() + return + } + if sourceEnded || isTerminalStreamEvent(event) { + return + } + } + } + }() + + return llm.NewStreamResult(out, source.RequestID, closeSource) +} + +// CoordinateStreamResult applies the TransformStreamResult lifecycle to source +// without transforming events. A source that TransformStreamResult already +// coordinates under a context with the same Done channel (the same context, +// or one derived only by adding values) and the same request ID is returned +// unchanged: another wrapper would behave identically. +func CoordinateStreamResult(ctx context.Context, source *StreamResult) *StreamResult { + if ctx == nil { + ctx = context.Background() + } + if source != nil { + if value, ok := coordinatedStreams.Load(source.Events); ok { + if c, _ := value.(streamCoordination); c.done == ctx.Done() && c.requestID == source.RequestID { + return source + } + } + } + return TransformStreamResult(ctx, source, nil) +} + +func cancellationEvent(requestID string, err error) FluxStreamEvent { + kind := llm.ErrKindCanceled + if errors.Is(err, context.DeadlineExceeded) { + kind = llm.ErrKindTimeout + } + message := "context canceled" + if err != nil { + message = err.Error() + } + return FluxStreamEvent{ + Type: "cancelled", + Error: message, + RequestID: requestID, + ErrorInfo: &llm.StreamErrorInfo{Kind: kind}, + } +} + +func streamErrorEvent(requestID string, err error) FluxStreamEvent { + kind := llm.ErrKindInternal + retryable := false + switch { + case errors.Is(err, context.DeadlineExceeded): + kind = llm.ErrKindTimeout + retryable = true + case errors.Is(err, context.Canceled): + kind = llm.ErrKindCanceled + default: + retryable = true + } + message := "stream transform failed" + if err != nil { + message = err.Error() + } + return FluxStreamEvent{ + Type: "error", + Error: message, + RequestID: requestID, + ErrorInfo: &llm.StreamErrorInfo{Kind: kind, Retryable: retryable}, + } +} + +// ensureStreamErrorInfo guarantees that every terminal error-ish event leaving +// this package carries a populated ErrorInfo, so consumers can dispatch on +// Kind/Retryable instead of re-parsing the human-readable Error string. +// +// Events that already carry ErrorInfo are returned untouched — the producer or +// a handler knows more than can be inferred from the message alone. Non-error +// events are returned untouched. +func ensureStreamErrorInfo(event FluxStreamEvent) FluxStreamEvent { + if event.ErrorInfo != nil { + return event + } + switch event.Type { + case "error": + kind, retryable := inferStreamErrorKind(event.Error) + event.ErrorInfo = &llm.StreamErrorInfo{Kind: kind, Retryable: retryable} + case "cancelled", "canceled": + // A cancellation is never an internal fault and never retryable; only + // a deadline distinguishes it from a plain cancel. + kind := llm.ErrKindCanceled + if inferred, _ := inferStreamErrorKind(event.Error); inferred == llm.ErrKindTimeout { + kind = llm.ErrKindTimeout + } + event.ErrorInfo = &llm.StreamErrorInfo{Kind: kind} + } + return event +} + +// inferStreamErrorKind maps a provider error message onto a portable +// StreamErrorInfo kind and retryability. +// +// Like classifyProviderError, it keys on common cross-provider substrings +// rather than a per-provider table, so it does not become a maintenance sink +// across flux's many providers. Matching is deliberately textual only: bare +// HTTP status digits are not matched, because they false-positive against +// ordinary content such as model ids and token counts. Only Go's own +// "context canceled" text maps to the canceled kind: a bare "cancelled" is as +// likely to be provider prose ("subscription cancelled") as a cancellation. +// Whether a canceled or timeout kind is the caller's cancellation is decided +// by the consumer from its own context, never from this inference. +// Unrecognized messages fall back to internal/retryable, matching +// streamErrorEvent. +func inferStreamErrorKind(message string) (kind string, retryable bool) { + msg := strings.ToLower(message) + switch { + case msg == "": + return llm.ErrKindInternal, true + case containsAny(msg, "rate limit", "too many requests", "rate_limit", + "quota exceeded", "exceeded your current quota", "insufficient_quota", + "insufficient credits", "billing"): + return llm.ErrKindRateLimited, true + case containsAny(msg, "api key", "api_key", "unauthorized", "unauthenticated", + "authentication", "invalid_api_key", "forbidden", "permission denied"): + return llm.ErrKindAuth, false + case containsAny(msg, "context length", "context window", "maximum context", + "too many tokens", "context_length_exceeded", "context_exceeded"): + return llm.ErrKindContextExceeded, false + case containsAny(msg, "content filter", "content_filter", "content_policy", + "responsible_ai", "content_filtered"): + return llm.ErrKindContentFiltered, false + case containsAny(msg, "deadline exceeded", "timeout", "timed out", "etimedout"): + return llm.ErrKindTimeout, true + case containsAny(msg, "context canceled", "context cancelled"): + return llm.ErrKindCanceled, false + case containsAny(msg, "service unavailable", "bad gateway", "temporarily unavailable", + "upstream error", "overloaded_error", "overloaded"): + return llm.ErrKindUnavailable, true + case containsAny(msg, "invalid request", "invalid_request", "invalid_request_error", + "unsupported parameter", "unprocessable"): + return llm.ErrKindInvalidRequest, false + default: + return llm.ErrKindInternal, true + } +} + +func containsAny(haystack string, needles ...string) bool { + for _, needle := range needles { + if strings.Contains(haystack, needle) { + return true + } + } + return false +} + +func isTerminalStreamEvent(event FluxStreamEvent) bool { + return event.Type == "done" || event.Type == "cancelled" || event.Type == "canceled" || event.Type == "error" && event.Warning == "" +} + +func sendLifecycleEvent(ctx context.Context, out chan<- FluxStreamEvent, event FluxStreamEvent) bool { + select { + case out <- event: + return true + case <-ctx.Done(): + return false + } +} + +func sendTerminalEvent(out chan<- FluxStreamEvent, event FluxStreamEvent, closed <-chan struct{}) bool { + select { + case out <- event: + return true + case <-closed: + return false + } +} + +func channelClosed(closed <-chan struct{}) bool { + select { + case <-closed: + return true + default: + return false + } +} + +func CloneUsage(usage *FluxUsage) *FluxUsage { + if usage == nil { + return nil + } + cloned := *usage + return &cloned +} + +func MergeUsage(previous, current *FluxUsage) *FluxUsage { + if previous == nil { + merged := CloneUsage(current) + if merged != nil && merged.TotalTokens == 0 { + merged.TotalTokens = merged.PromptTokens + merged.CompletionTokens + } + return merged + } + if current == nil { + return CloneUsage(previous) + } + merged := &FluxUsage{ + PromptTokens: max(previous.PromptTokens, current.PromptTokens), + CompletionTokens: max(previous.CompletionTokens, current.CompletionTokens), + CacheCreationTokens: max(previous.CacheCreationTokens, current.CacheCreationTokens), + CacheReadTokens: max(previous.CacheReadTokens, current.CacheReadTokens), + ThinkingTokens: max(previous.ThinkingTokens, current.ThinkingTokens), + } + merged.TotalTokens = max(usageTotal(previous), usageTotal(current), merged.PromptTokens+merged.CompletionTokens) + return merged +} + +func UsageDelta(previous, current *FluxUsage) *FluxUsage { + if current == nil { + return nil + } + if previous == nil { + delta := CloneUsage(current) + if delta.TotalTokens == 0 { + delta.TotalTokens = delta.PromptTokens + delta.CompletionTokens + } + if usageIsZero(delta) { + return nil + } + return delta + } + + delta := &FluxUsage{ + PromptTokens: positiveUsageDelta(current.PromptTokens, previous.PromptTokens), + CompletionTokens: positiveUsageDelta(current.CompletionTokens, previous.CompletionTokens), + CacheCreationTokens: positiveUsageDelta(current.CacheCreationTokens, previous.CacheCreationTokens), + CacheReadTokens: positiveUsageDelta(current.CacheReadTokens, previous.CacheReadTokens), + ThinkingTokens: positiveUsageDelta(current.ThinkingTokens, previous.ThinkingTokens), + } + currentTotal := usageTotal(current) + previousTotal := usageTotal(previous) + switch { + case current.TotalTokens == 0: + delta.TotalTokens = delta.PromptTokens + delta.CompletionTokens + case currentTotal > previousTotal: + delta.TotalTokens = currentTotal - previousTotal + case delta.PromptTokens > 0 || delta.CompletionTokens > 0: + delta.TotalTokens = delta.PromptTokens + delta.CompletionTokens + } + if usageIsZero(delta) { + return nil + } + return delta +} + +func positiveUsageDelta(current, previous int) int { + if current <= previous { + return 0 + } + return current - previous +} + +func usageTotal(usage *FluxUsage) int { + if usage.TotalTokens > 0 { + return usage.TotalTokens + } + return usage.PromptTokens + usage.CompletionTokens +} + +func usageIsZero(usage *FluxUsage) bool { + return usage.PromptTokens == 0 && + usage.CompletionTokens == 0 && + usage.CacheCreationTokens == 0 && + usage.CacheReadTokens == 0 && + usage.ThinkingTokens == 0 && + usage.TotalTokens == 0 +} + // ParseSSEStream reads an SSE stream and sends events to a channel. // The goroutine closes the channel and body when done or context is cancelled. -// Scanner errors are emitted as SSEEvent with Event="error" so callers can detect truncation. +// Scanner errors, and an event larger than SSEMaxEventBytes, are emitted as +// SSEEvent with Event="error" so callers can detect truncation. func ParseSSEStream(ctx context.Context, body io.ReadCloser, logger *slog.Logger) <-chan SSEEvent { ch := make(chan SSEEvent, sseChannelBuffer) go func() { @@ -41,41 +468,60 @@ func ParseSSEStream(ctx context.Context, body io.ReadCloser, logger *slog.Logger scanner.Buffer(make([]byte, 0, sseScannerInitBuf), sseScannerMaxBuf) var event, data strings.Builder - for scanner.Scan() { + dispatch := func() bool { + if data.Len() == 0 { + event.Reset() + data.Reset() + return true + } select { + case ch <- SSEEvent{Event: strings.TrimSpace(event.String()), Data: strings.TrimSpace(data.String())}: + event.Reset() + data.Reset() + return true case <-ctx.Done(): + return false + } + } + + fits := func(field string) bool { + return event.Len()+data.Len()+len(field)+1 <= SSEMaxEventBytes + } + + var readErr error + for scanner.Scan() { + if ctx.Err() != nil { return - default: } line := scanner.Text() if line == "" { - if data.Len() > 0 { - select { - case ch <- SSEEvent{Event: strings.TrimSpace(event.String()), Data: strings.TrimSpace(data.String())}: - case <-ctx.Done(): - return - } + if !dispatch() { + return } - event.Reset() - data.Reset() continue } - if strings.HasPrefix(line, "event:") { - event.WriteString(strings.TrimPrefix(line, "event:")) - } else if strings.HasPrefix(line, "data:") { + if field, ok := strings.CutPrefix(line, "event:"); ok { + if !fits(field) { + readErr = errSSEEventTooLarge + break + } + event.WriteString(field) + } else if field, ok := strings.CutPrefix(line, "data:"); ok { + if !fits(field) { + readErr = errSSEEventTooLarge + break + } if data.Len() > 0 { data.WriteByte('\n') } - data.WriteString(strings.TrimPrefix(line, "data:")) + data.WriteString(field) } } - if err := scanner.Err(); err != nil { - // Context cancellation produces "context canceled" from the - // scanner when the body is closed; that is the expected - // shutdown path, not a stream error. Skip the warning and - // the synthetic error event so callers (and operators) don't - // see noise for every normal cancel/close. + if readErr == nil { + readErr = scanner.Err() + } + if err := readErr; err != nil { if ctxErr := ctx.Err(); ctxErr != nil || errors.Is(err, context.Canceled) { return } @@ -84,7 +530,9 @@ func ParseSSEStream(ctx context.Context, body io.ReadCloser, logger *slog.Logger case ch <- SSEEvent{Event: "error", Data: fmt.Sprintf("stream read error: %v", err)}: case <-ctx.Done(): } + return } + dispatch() }() return ch } @@ -479,7 +927,6 @@ func ProcessOpenAIStreamWithOpts(ctx context.Context, sseEvents <-chan SSEEvent, return case evt, ok := <-sseEvents: if !ok { - finish("") return } // Propagate SSE-level errors diff --git a/provider/core/stream_test.go b/provider/core/stream_test.go index d35f7e6d..9df7ad90 100644 --- a/provider/core/stream_test.go +++ b/provider/core/stream_test.go @@ -6,8 +6,12 @@ import ( "io" "log/slog" "strings" + "sync" + "sync/atomic" "testing" "time" + + "github.com/GrayCodeAI/flux/llm" ) func testLogger() *slog.Logger { @@ -39,6 +43,20 @@ func TestSSEParseBasicEvents(t *testing.T) { } } +func TestSSEParseFlushesUnterminatedFinalEvent(t *testing.T) { + t.Parallel() + body := io.NopCloser(strings.NewReader("data: [DONE]")) + ch := ParseSSEStream(context.Background(), body, testLogger()) + + var events []SSEEvent + for event := range ch { + events = append(events, event) + } + if len(events) != 1 || events[0].Data != "[DONE]" { + t.Fatalf("events = %+v, want one [DONE] event", events) + } +} + func TestSSEParseMultilineData(t *testing.T) { t.Parallel() sseData := "event:content\ndata:line one\ndata:line two\ndata:line three\n\n" @@ -112,6 +130,387 @@ func TestSSEParseContextCancellation(t *testing.T) { } } +func TestCoordinateStreamResultTerminalContract(t *testing.T) { + t.Parallel() + + tests := []struct { + name string + events []FluxStreamEvent + wantTypes []string + wantErrKind string + }{ + { + name: "truncated", + events: []FluxStreamEvent{{Type: "content", Content: "partial"}}, + wantTypes: []string{"content", "error"}, + wantErrKind: llm.ErrKindTruncated, + }, + { + name: "fatal before done", + events: []FluxStreamEvent{ + {Type: "error", Error: "connection reset"}, + {Type: "done"}, + }, + wantTypes: []string{"error"}, + }, + { + name: "terminal before late usage", + events: []FluxStreamEvent{ + {Type: "done", StopReason: "stop"}, + {Type: "usage", Usage: &FluxUsage{TotalTokens: 1}}, + }, + wantTypes: []string{"done"}, + }, + { + name: "warning before done", + events: []FluxStreamEvent{ + {Type: "error", Error: "empty response", Warning: "empty response"}, + {Type: "done", StopReason: "stop"}, + }, + wantTypes: []string{"error", "done"}, + }, + } + + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + + events := make(chan FluxStreamEvent, len(tt.events)) + for _, event := range tt.events { + events <- event + } + close(events) + + var closes atomic.Int32 + source := llm.NewStreamResult(events, "request-1", func() { closes.Add(1) }) + result := CoordinateStreamResult(context.Background(), source) + var got []FluxStreamEvent + for event := range result.Events { + got = append(got, event) + } + result.Close() + + if len(got) != len(tt.wantTypes) { + t.Fatalf("events = %+v, want types %v", got, tt.wantTypes) + } + for i, wantType := range tt.wantTypes { + if got[i].Type != wantType { + t.Fatalf("event[%d].Type = %q, want %q", i, got[i].Type, wantType) + } + if got[i].RequestID != "request-1" { + t.Fatalf("event[%d].RequestID = %q, want request-1", i, got[i].RequestID) + } + } + if tt.wantErrKind != "" { + last := got[len(got)-1] + if last.ErrorInfo == nil || last.ErrorInfo.Kind != tt.wantErrKind || !last.ErrorInfo.Retryable { + t.Fatalf("terminal error info = %+v, want retryable kind %q", last.ErrorInfo, tt.wantErrKind) + } + } + if count := closes.Load(); count != 1 { + t.Fatalf("source close count = %d, want 1", count) + } + }) + } +} + +func TestTransformStreamResultCloseUnblocksForwarder(t *testing.T) { + t.Parallel() + + events := make(chan FluxStreamEvent) + streamCtx, cancel := context.WithCancel(context.Background()) + producerDone := make(chan struct{}) + go func() { + defer close(events) + defer close(producerDone) + <-streamCtx.Done() + }() + + var closes atomic.Int32 + source := llm.NewStreamResult(events, "request-2", func() { + closes.Add(1) + cancel() + }) + result := TransformStreamResult(context.Background(), source, func(_ context.Context, event FluxStreamEvent) (FluxStreamEvent, error) { + return event, nil + }) + result.Close() + + select { + case <-producerDone: + case <-time.After(2 * time.Second): + t.Fatal("producer did not stop after stream close") + } + select { + case _, ok := <-result.Events: + if ok { + t.Fatal("unexpected event after close") + } + case <-time.After(2 * time.Second): + t.Fatal("transformed event channel did not close") + } + result.Close() + if got := closes.Load(); got != 1 { + t.Fatalf("source close count = %d, want 1", got) + } +} + +func TestCoordinateStreamResultCancellationEmitsTerminal(t *testing.T) { + t.Parallel() + + events := make(chan FluxStreamEvent) + ctx, cancel := context.WithCancel(context.Background()) + result := CoordinateStreamResult(ctx, llm.NewStreamResult(events, "request-cancel", nil)) + cancel() + + select { + case event, ok := <-result.Events: + if !ok { + t.Fatal("stream closed without cancellation terminal") + } + if event.Type != "cancelled" || event.ErrorInfo == nil || event.ErrorInfo.Kind != llm.ErrKindCanceled { + t.Fatalf("event = %+v, want canceled terminal", event) + } + if event.RequestID != "request-cancel" { + t.Fatalf("request ID = %q, want request-cancel", event.RequestID) + } + case <-time.After(time.Second): + t.Fatal("cancellation terminal was not emitted") + } + + select { + case _, ok := <-result.Events: + if ok { + t.Fatal("unexpected event after cancellation terminal") + } + case <-time.After(time.Second): + t.Fatal("stream did not close after cancellation terminal") + } + result.Close() +} + +func TestTransformStreamResultCancelWhileForwardingEmitsTerminal(t *testing.T) { + t.Parallel() + + for run := 0; run < 100; run++ { + source := make(chan FluxStreamEvent) + ctx, cancel := context.WithCancel(context.Background()) + result := CoordinateStreamResult(ctx, llm.NewStreamResult(source, "request-busy", nil)) + + // The unbuffered wrapper takes the event and parks delivering it, + // because nobody is reading yet; then the caller cancels. + source <- FluxStreamEvent{Type: "content", Content: "pending"} + cancel() + + var last FluxStreamEvent + deadline := time.After(time.Second) + drain: + for { + select { + case event, ok := <-result.Events: + if !ok { + break drain + } + last = event + case <-deadline: + t.Fatalf("run %d: stream did not close", run) + } + } + if last.Type != "cancelled" || last.ErrorInfo == nil || last.ErrorInfo.Kind != llm.ErrKindCanceled { + t.Fatalf("run %d: last event = %+v, want cancelled terminal", run, last) + } + result.Close() + } +} + +func TestTransformStreamResultCancelReleasesSourceBeforeTerminalDelivery(t *testing.T) { + t.Parallel() + + source := make(chan FluxStreamEvent, 1) + released := make(chan struct{}) + var releaseOnce sync.Once + ctx, cancel := context.WithCancel(context.Background()) + result := CoordinateStreamResult(ctx, llm.NewStreamResult(source, "request-idle", func() { + releaseOnce.Do(func() { close(released) }) + })) + + // Fill the one-slot output buffer, then park the forwarder on the next + // event: the consumer never reads. + source <- FluxStreamEvent{Type: "content", Content: "one"} + waitFor(t, func() bool { return len(result.Events) == cap(result.Events) }) + source <- FluxStreamEvent{Type: "content", Content: "two"} + cancel() + + select { + case <-released: + case <-time.After(time.Second): + t.Fatal("source was not released while the terminal waited for the consumer") + } + + // Close unblocks the parked forwarder, which then closes the channel. + result.Close() + deadline := time.After(time.Second) + for { + select { + case _, ok := <-result.Events: + if !ok { + return + } + case <-deadline: + t.Fatal("Close did not unblock the forwarder") + } + } +} + +func TestStreamResultCloseTwiceIsNoop(t *testing.T) { + t.Parallel() + + var closes int32 + source := make(chan FluxStreamEvent) + result := CoordinateStreamResult(context.Background(), llm.NewStreamResult(source, "", func() { + atomic.AddInt32(&closes, 1) + })) + result.Close() + result.Close() + if _, ok := <-result.Events; ok { + t.Fatal("expected no events after Close") + } + if got := atomic.LoadInt32(&closes); got != 1 { + t.Fatalf("source closed %d times, want 1", got) + } +} + +func waitFor(t *testing.T, condition func() bool) { + t.Helper() + deadline := time.Now().Add(time.Second) + for !condition() { + if time.Now().After(deadline) { + t.Fatal("condition not reached") + } + time.Sleep(time.Millisecond) + } +} + +func TestCoordinateStreamResultSkipsIdenticalCoordination(t *testing.T) { + t.Parallel() + + type ctxKey struct{} + ctx, cancel := context.WithCancel(context.Background()) + defer cancel() + source := make(chan FluxStreamEvent, 1) + inner := CoordinateStreamResult(ctx, llm.NewStreamResult(source, "request-1", nil)) + defer inner.Close() + + if got := CoordinateStreamResult(ctx, inner); got != inner { + t.Fatal("same context: want the coordinated stream returned unchanged") + } + if got := CoordinateStreamResult(context.WithValue(ctx, ctxKey{}, "v"), inner); got != inner { + t.Fatal("value-only derived context: want the coordinated stream returned unchanged") + } + child, cancelChild := context.WithCancel(ctx) + defer cancelChild() + wrapped := []*StreamResult{ + CoordinateStreamResult(child, inner), + CoordinateStreamResult(ctx, llm.NewStreamResult(inner.Events, "other-request", nil)), + TransformStreamResult(ctx, inner, func(_ context.Context, e FluxStreamEvent) (FluxStreamEvent, error) { return e, nil }), + } + for i, got := range wrapped { + if got.Events == inner.Events { + t.Fatalf("case %d (cancellable child, other request ID, handler): want a new lifecycle wrapper", i) + } + got.Close() + } +} + +func TestCoordinatedStreamRegistryIsReleased(t *testing.T) { + t.Parallel() + + source := make(chan FluxStreamEvent, 1) + result := CoordinateStreamResult(context.Background(), llm.NewStreamResult(source, "", nil)) + source <- FluxStreamEvent{Type: "done"} + if event := <-result.Events; event.Type != "done" { + t.Fatalf("event = %+v, want done", event) + } + if _, ok := <-result.Events; ok { + t.Fatal("stream did not close after done") + } + waitFor(t, func() bool { + _, ok := coordinatedStreams.Load(result.Events) + return !ok + }) +} + +func TestUsageDeltaHandlesCumulativeAndSplitUsage(t *testing.T) { + t.Parallel() + + t.Run("cumulative", func(t *testing.T) { + var state *FluxUsage + var prompt, completion int + for _, usage := range []*FluxUsage{ + {PromptTokens: 1, CompletionTokens: 1, TotalTokens: 2}, + {PromptTokens: 1, CompletionTokens: 2, TotalTokens: 3}, + } { + delta := UsageDelta(state, usage) + state = MergeUsage(state, usage) + if delta != nil { + prompt += delta.PromptTokens + completion += delta.CompletionTokens + } + } + if prompt != 1 || completion != 2 { + t.Fatalf("prompt=%d completion=%d, want 1/2", prompt, completion) + } + }) + + t.Run("split then aggregate", func(t *testing.T) { + var state *FluxUsage + var prompt, completion int + for _, usage := range []*FluxUsage{ + {PromptTokens: 10}, + {CompletionTokens: 5}, + {PromptTokens: 10, CompletionTokens: 5, TotalTokens: 15}, + } { + delta := UsageDelta(state, usage) + state = MergeUsage(state, usage) + if delta != nil { + prompt += delta.PromptTokens + completion += delta.CompletionTokens + } + } + if prompt != 10 || completion != 5 { + t.Fatalf("prompt=%d completion=%d, want 10/5", prompt, completion) + } + }) + + t.Run("reverse split then aggregate", func(t *testing.T) { + var state *FluxUsage + var prompt, completion int + for _, usage := range []*FluxUsage{ + {CompletionTokens: 5}, + {PromptTokens: 10}, + {PromptTokens: 10, CompletionTokens: 5, TotalTokens: 15}, + } { + delta := UsageDelta(state, usage) + state = MergeUsage(state, usage) + if delta != nil { + prompt += delta.PromptTokens + completion += delta.CompletionTokens + } + } + if prompt != 10 || completion != 5 { + t.Fatalf("prompt=%d completion=%d, want 10/5", prompt, completion) + } + }) + + t.Run("duplicate aggregate", func(t *testing.T) { + usage := &FluxUsage{PromptTokens: 3, CompletionTokens: 5, TotalTokens: 8} + state := MergeUsage(nil, usage) + if delta := UsageDelta(state, usage); delta != nil { + t.Fatalf("delta = %+v, want nil", delta) + } + }) +} + // --- ProcessAnthropicStream tests --- func TestSSEAnthropicContentBlockDelta(t *testing.T) { @@ -433,6 +832,26 @@ func TestSSEOpenAIUsage(t *testing.T) { } } +func TestSSEOpenAIChannelCloseDoesNotSynthesizeDone(t *testing.T) { + t.Parallel() + events := make(chan SSEEvent, 1) + events <- SSEEvent{Data: `{"choices":[{"delta":{"content":"partial"},"finish_reason":null}]}`} + close(events) + + var results []FluxStreamEvent + for event := range ProcessOpenAIStream(context.Background(), events, testLogger()) { + results = append(results, event) + } + if len(results) == 0 { + t.Fatal("expected content event") + } + for _, event := range results { + if event.Type == "done" { + t.Fatalf("unexpected done event: %+v", event) + } + } +} + // --- ParseInlineToolCalls tests --- func TestSSEParseInlineToolCallsCanopywave(t *testing.T) { @@ -515,3 +934,183 @@ functions.read_file:1 t.Errorf("second tool args[path] = %q, want /tmp/test.go", path) } } + +// --- ensureStreamErrorInfo / inferStreamErrorKind tests --- + +func TestInferStreamErrorKind(t *testing.T) { + t.Parallel() + tests := []struct { + name string + message string + wantKind string + wantRetries bool + }{ + {"empty message defaults to internal", "", llm.ErrKindInternal, true}, + {"unrecognized defaults to internal", "something went sideways", llm.ErrKindInternal, true}, + {"rate limit", "Rate limit exceeded for gpt-4o", llm.ErrKindRateLimited, true}, + {"too many requests", "429 Too Many Requests", llm.ErrKindRateLimited, true}, + {"quota", "You exceeded your current quota", llm.ErrKindRateLimited, true}, + {"overloaded is unavailable not rate limited", "overloaded_error", llm.ErrKindUnavailable, true}, + {"invalid api key", "invalid_api_key", llm.ErrKindAuth, false}, + {"unauthorized", "Unauthorized", llm.ErrKindAuth, false}, + {"forbidden", "Forbidden", llm.ErrKindAuth, false}, + {"context length", "maximum context length is 8192 tokens", llm.ErrKindContextExceeded, false}, + {"too many tokens", "too many tokens for this model", llm.ErrKindContextExceeded, false}, + {"content filter", "blocked by content_filter", llm.ErrKindContentFiltered, false}, + {"deadline", "context deadline exceeded", llm.ErrKindTimeout, true}, + {"timed out", "request timed out", llm.ErrKindTimeout, true}, + {"canceled", "context canceled", llm.ErrKindCanceled, false}, + {"cancelled spelling", "context cancelled", llm.ErrKindCanceled, false}, + // Provider prose that merely contains "cancelled" is not a + // cancellation; only the consumer's own context can say that. + {"provider prose is not a cancellation", "request was cancelled upstream", llm.ErrKindInternal, true}, + {"client timeout is a timeout", "context deadline exceeded (Client.Timeout exceeded while reading body)", llm.ErrKindTimeout, true}, + {"unavailable", "503 service unavailable", llm.ErrKindUnavailable, true}, + {"bad gateway", "502 bad gateway", llm.ErrKindUnavailable, true}, + {"invalid request", "invalid_request_error: bad param", llm.ErrKindInvalidRequest, false}, + // Status digits alone are intentionally not matched: they false-positive + // against ordinary content such as model ids and token counts. + {"bare digits are not matched", "model gpt-4o-2501 returned 500 tokens", llm.ErrKindInternal, true}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + kind, retryable := inferStreamErrorKind(tt.message) + if kind != tt.wantKind { + t.Errorf("kind = %q, want %q", kind, tt.wantKind) + } + if retryable != tt.wantRetries { + t.Errorf("retryable = %v, want %v", retryable, tt.wantRetries) + } + }) + } +} + +func TestEnsureStreamErrorInfo(t *testing.T) { + t.Parallel() + tests := []struct { + name string + event FluxStreamEvent + wantNil bool + wantKind string + wantRetries bool + }{ + { + name: "content event untouched", + event: FluxStreamEvent{Type: "content", Content: "hi"}, + wantNil: true, + }, + { + name: "done event untouched", + event: FluxStreamEvent{Type: "done"}, + wantNil: true, + }, + { + name: "error gains inferred info", + event: FluxStreamEvent{Type: "error", Error: "rate limit exceeded"}, + wantKind: llm.ErrKindRateLimited, wantRetries: true, + }, + { + name: "bare error gains internal", + event: FluxStreamEvent{Type: "error", Error: "boom"}, + wantKind: llm.ErrKindInternal, wantRetries: true, + }, + { + name: "cancelled with unknown message stays canceled", + event: FluxStreamEvent{Type: "cancelled", Error: "stopped"}, + wantKind: llm.ErrKindCanceled, wantRetries: false, + }, + { + name: "cancelled with deadline is timeout", + event: FluxStreamEvent{Type: "cancelled", Error: "context deadline exceeded"}, + wantKind: llm.ErrKindTimeout, wantRetries: false, + }, + { + name: "cancelled with provider wording stays canceled", + event: FluxStreamEvent{Type: "cancelled", Error: "rate limit"}, + wantKind: llm.ErrKindCanceled, wantRetries: false, + }, + { + name: "existing ErrorInfo is preserved", + event: FluxStreamEvent{ + Type: "error", + Error: "rate limit exceeded", + ErrorInfo: &llm.StreamErrorInfo{Kind: llm.ErrKindAuth, StatusCode: 401}, + }, + wantKind: llm.ErrKindAuth, wantRetries: false, + }, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + got := ensureStreamErrorInfo(tt.event) + if tt.wantNil { + if got.ErrorInfo != nil { + t.Fatalf("ErrorInfo = %+v, want nil", got.ErrorInfo) + } + return + } + if got.ErrorInfo == nil { + t.Fatal("ErrorInfo = nil, want populated") + } + if got.ErrorInfo.Kind != tt.wantKind { + t.Errorf("Kind = %q, want %q", got.ErrorInfo.Kind, tt.wantKind) + } + if got.ErrorInfo.Retryable != tt.wantRetries { + t.Errorf("Retryable = %v, want %v", got.ErrorInfo.Retryable, tt.wantRetries) + } + // A preserved StatusCode must survive the pass-through. + if tt.event.ErrorInfo != nil && got.ErrorInfo.StatusCode != tt.event.ErrorInfo.StatusCode { + t.Errorf("StatusCode = %d, want %d", got.ErrorInfo.StatusCode, tt.event.ErrorInfo.StatusCode) + } + }) + } +} + +// endlessReader repeats line forever, like a hostile peer that never ends an +// SSE event. +type endlessReader struct { + line []byte + offset int +} + +func (r *endlessReader) Read(p []byte) (int, error) { + n := 0 + for n < len(p) { + copied := copy(p[n:], r.line[r.offset:]) + n += copied + r.offset = (r.offset + copied) % len(r.line) + } + return n, nil +} + +func (r *endlessReader) Close() error { return nil } + +func TestParseSSEStreamRejectsOversizedEvent(t *testing.T) { + t.Parallel() + + line := []byte("data: " + strings.Repeat("x", 64*1024) + "\n") + ch := ParseSSEStream(context.Background(), &endlessReader{line: line}, testLogger()) + + deadline := time.After(10 * time.Second) + for { + select { + case event, ok := <-ch: + if !ok { + t.Fatal("stream closed without reporting the oversized event") + } + if event.Event != "error" { + t.Fatalf("event = %+v, want only the size error", event) + } + if !strings.Contains(event.Data, "exceeds") { + t.Fatalf("error = %q, want size limit error", event.Data) + } + if _, ok := <-ch; ok { + t.Fatal("stream kept going after the size error") + } + return + case <-deadline: + t.Fatal("parser kept accumulating an unbounded event") + } + } +} diff --git a/provider/gemini_stream_test.go b/provider/gemini_stream_test.go index 8a5cb2a0..12b531c3 100644 --- a/provider/gemini_stream_test.go +++ b/provider/gemini_stream_test.go @@ -12,6 +12,8 @@ import ( "strings" "testing" "time" + + "github.com/GrayCodeAI/flux/llm" ) // slogDiscard returns a *slog.Logger that discards all output. Used @@ -222,9 +224,6 @@ func TestGemini_Stream_SharedParser_DoneWithUsage(t *testing.T) { } } -// TestGemini_Stream_SharedParser_EmptyStream: a server that returns -// 200 with no body should still emit a "done" event so consumers -// don't hang. func TestGemini_Stream_SharedParser_EmptyStream(t *testing.T) { t.Setenv(geminiSharedParserEnvVar, "1") srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { @@ -243,10 +242,10 @@ func TestGemini_Stream_SharedParser_EmptyStream(t *testing.T) { events := drainGeminiStream(t, sr, 3*time.Second) if len(events) != 1 { - t.Fatalf("expected 1 event (done), got %d: %+v", len(events), events) + t.Fatalf("expected 1 truncated error, got %d: %+v", len(events), events) } - if events[0].Type != "done" { - t.Errorf("event type = %q, want \"done\"", events[0].Type) + if events[0].Type != "error" || events[0].ErrorInfo == nil || events[0].ErrorInfo.Kind != llm.ErrKindTruncated { + t.Fatalf("event = %+v, want truncated error", events[0]) } } @@ -420,7 +419,9 @@ func TestGemini_Stream_SharedParser_FeatureFlag(t *testing.T) { // semantic without going through the HTTP layer. func TestProcessGeminiStream_PreservesDoneWithUsage(t *testing.T) { usageFrame := `{"candidates":[{"content":{"parts":[],"role":"model"},"finishReason":"STOP"}],"usageMetadata":{"promptTokenCount":1,"candidatesTokenCount":2,"totalTokenCount":3}}` + priorUsageFrame := `{"usageMetadata":{"promptTokenCount":1,"candidatesTokenCount":1,"totalTokenCount":2}}` sseCh := make(chan SSEEvent, 2) + sseCh <- SSEEvent{Data: priorUsageFrame} sseCh <- SSEEvent{Data: usageFrame} ctx, cancel := context.WithCancel(context.Background()) @@ -477,10 +478,7 @@ func TestProcessGeminiStream_DoneWithoutUsage(t *testing.T) { } } -// TestProcessGeminiStream_EmptyStream_EmitsDone: when the SSE -// channel closes without a finish reason, a bare "done" event is -// emitted (matches the original streamLoop's if !doneSent fallback). -func TestProcessGeminiStream_EmptyStream_EmitsDone(t *testing.T) { +func TestProcessGeminiStream_EmptyStream_HasNoTerminal(t *testing.T) { ctx, cancel := context.WithCancel(context.Background()) defer cancel() sseCh := make(chan SSEEvent) @@ -488,11 +486,8 @@ func TestProcessGeminiStream_EmptyStream_EmitsDone(t *testing.T) { out := processGeminiStream(ctx, sseCh, slogDiscard()) events := collectFluxStreamEvents(t, out, 2*time.Second) - if len(events) != 1 { - t.Fatalf("expected 1 event, got %d: %+v", len(events), events) - } - if events[0].Type != "done" { - t.Errorf("event type = %q, want done", events[0].Type) + if len(events) != 0 { + t.Fatalf("events = %+v, want none", events) } } diff --git a/provider/guardrails_provider_test.go b/provider/guardrails_provider_test.go index 631616a6..cf86347e 100644 --- a/provider/guardrails_provider_test.go +++ b/provider/guardrails_provider_test.go @@ -4,7 +4,9 @@ import ( "context" "errors" "strings" + "sync/atomic" "testing" + "time" "github.com/GrayCodeAI/flux/provider/resilience" ) @@ -274,6 +276,58 @@ func TestGuardrailProvider_ChatNoGuardrails(t *testing.T) { } } +type guardrailStreamProvider struct { + closes atomic.Int32 +} + +func (*guardrailStreamProvider) Name() string { return "guardrail-stream" } +func (*guardrailStreamProvider) Ping(context.Context) error { return nil } +func (*guardrailStreamProvider) Chat(context.Context, []FluxMessage, ChatOptions) (*FluxResponse, error) { + return nil, nil +} + +func (p *guardrailStreamProvider) StreamChat(ctx context.Context, _ []FluxMessage, _ ChatOptions) (*StreamResult, error) { + streamCtx, cancel := context.WithCancel(ctx) + events := make(chan FluxStreamEvent) + go func() { + <-streamCtx.Done() + close(events) + }() + return NewStreamResult(events, func() { + p.closes.Add(1) + cancel() + }), nil +} + +func TestGuardrailProvider_StreamClosePropagates(t *testing.T) { + t.Parallel() + inner := &guardrailStreamProvider{} + gp := resilience.NewGuardrailProvider(inner, NewGuardrails(GuardrailRule{ + Type: GuardrailCustom, + Name: "block", + Pattern: `blocked`, + Action: GuardrailBlock, + })) + result, err := gp.StreamChat(context.Background(), []FluxMessage{{Role: "user", Content: "hello"}}, ChatOptions{Model: "test"}) + if err != nil { + t.Fatal(err) + } + result.Close() + result.Close() + + select { + case _, ok := <-result.Events: + if ok { + t.Fatal("unexpected event after close") + } + case <-time.After(2 * time.Second): + t.Fatal("guardrail stream did not close") + } + if got := inner.closes.Load(); got != 1 { + t.Fatalf("inner close count = %d, want 1", got) + } +} + // --------------------------------------------------------------------------- // ClientOption tests for WithGuardrails / WithGuardrailType // --------------------------------------------------------------------------- diff --git a/provider/observability/budget_provider.go b/provider/observability/budget_provider.go index 55e9ec77..22e19203 100644 --- a/provider/observability/budget_provider.go +++ b/provider/observability/budget_provider.go @@ -5,6 +5,8 @@ import ( "errors" "fmt" "sync" + + "github.com/GrayCodeAI/flux/provider/core" ) // ErrBudgetExceeded is returned when a virtual key has exhausted its budget. @@ -109,28 +111,22 @@ func (bp *BudgetProvider) StreamChat(ctx context.Context, messages []FluxMessage return nil, err } - // Wrap the events channel to record actual spend from the final usage - // event. Without this, streamed calls under a virtual key never debit the - // budget (unlike the non-streaming Chat path), so streaming-heavy clients - // would underreport spend. Mirrors UsageLimitProvider.StreamChat. - wrappedCh := make(chan FluxStreamEvent, cap(result.Events)) - go func() { - defer close(wrappedCh) - for evt := range result.Events { - if evt.Type == "usage" && evt.Usage != nil { - cost := ActualCostUSD(opts.Model, evt.Usage) - _ = bp.store.RecordUsage(ctx, vk, cost, evt.Usage.PromptTokens, evt.Usage.CompletionTokens) - } - select { - case wrappedCh <- evt: - case <-ctx.Done(): - result.Close() - return + var previousUsage *core.FluxUsage + return core.TransformStreamResult(ctx, result, func(streamCtx context.Context, evt FluxStreamEvent) (FluxStreamEvent, error) { + if evt.Type == "continuation" { + previousUsage = nil + return evt, nil + } + if (evt.Type == "usage" || evt.Type == "done") && evt.Usage != nil { + delta := core.UsageDelta(previousUsage, evt.Usage) + previousUsage = core.MergeUsage(previousUsage, evt.Usage) + if delta != nil { + cost := ActualCostUSD(opts.Model, delta) + _ = bp.store.RecordUsage(streamCtx, vk, cost, delta.PromptTokens, delta.CompletionTokens) } } - }() - - return NewStreamResult(wrappedCh, result.Close), nil + return evt, nil + }), nil } func (bp *BudgetProvider) recordUsage(ctx context.Context, vk, model string, resp *FluxResponse) { diff --git a/provider/observability/callbacks.go b/provider/observability/callbacks.go index 18b8d3dc..3cb214ea 100644 --- a/provider/observability/callbacks.go +++ b/provider/observability/callbacks.go @@ -6,6 +6,8 @@ import ( "log/slog" "sync" "time" + + "github.com/GrayCodeAI/flux/provider/core" ) // ProviderCallback defines hooks that are invoked at various points during @@ -151,30 +153,15 @@ func (cp *CallbackProvider) StreamChat(ctx context.Context, messages []FluxMessa return nil, err } - // Wrap the events channel to invoke OnStreamEvent for each event. cbs := cp.snapshotCallbacks() - origEvents := result.Events - wrappedEvents := make(chan FluxStreamEvent, cap(origEvents)) - - go func() { - defer close(wrappedEvents) - for evt := range origEvents { - // Fire stream event callbacks. - for _, cb := range cbs { - cp.safeCall("OnStreamEvent", func() { - cb.OnStreamEvent(ctx, provider, model, evt) - }) - } - select { - case wrappedEvents <- evt: - case <-ctx.Done(): - result.Close() - return - } + return core.TransformStreamResult(ctx, result, func(_ context.Context, evt FluxStreamEvent) (FluxStreamEvent, error) { + for _, cb := range cbs { + cp.safeCall("OnStreamEvent", func() { + cb.OnStreamEvent(ctx, provider, model, evt) + }) } - }() - - return NewStreamResultWithRequestID(wrappedEvents, result.RequestID, result.Close), nil + return evt, nil + }), nil } // --- internal helpers --- diff --git a/provider/observability/callbacks_test.go b/provider/observability/callbacks_test.go index 3c42c1a8..3402badf 100644 --- a/provider/observability/callbacks_test.go +++ b/provider/observability/callbacks_test.go @@ -9,6 +9,8 @@ import ( "sync/atomic" "testing" "time" + + "github.com/GrayCodeAI/flux/llm" ) // --- test helpers --- @@ -600,6 +602,130 @@ func TestCallbackErrorInNewCallbackProvider(t *testing.T) { } } +type truncatedStreamProvider struct{} + +func (*truncatedStreamProvider) Name() string { return "truncated-stream" } +func (*truncatedStreamProvider) Ping(context.Context) error { return nil } +func (*truncatedStreamProvider) Chat(context.Context, []FluxMessage, ChatOptions) (*FluxResponse, error) { + return nil, nil +} + +func (*truncatedStreamProvider) StreamChat(context.Context, []FluxMessage, ChatOptions) (*StreamResult, error) { + events := make(chan FluxStreamEvent) + close(events) + return NewStreamResult(events, func() {}), nil +} + +func TestCallbackStreamChatObservesTruncation(t *testing.T) { + t.Parallel() + callback := &recordingCallback{} + provider := mustCallbackProvider(t, &truncatedStreamProvider{}) + provider.AddCallback(callback) + result, err := provider.StreamChat(context.Background(), []FluxMessage{{Role: "user", Content: "hello"}}, ChatOptions{Model: "test"}) + if err != nil { + t.Fatal(err) + } + defer result.Close() + for range result.Events { + } + + waitUntil(t, 2*time.Second, func() bool { + callback.mu.Lock() + defer callback.mu.Unlock() + for _, event := range callback.streamEvents { + if event.event.ErrorInfo != nil && event.event.ErrorInfo.Kind == llm.ErrKindTruncated { + return true + } + } + return false + }) +} + +type continuationUsageProvider struct{} + +func (*continuationUsageProvider) Name() string { return "continuation-usage" } +func (*continuationUsageProvider) Ping(context.Context) error { return nil } +func (*continuationUsageProvider) Chat(context.Context, []FluxMessage, ChatOptions) (*FluxResponse, error) { + return nil, nil +} + +func (*continuationUsageProvider) StreamChat(context.Context, []FluxMessage, ChatOptions) (*StreamResult, error) { + events := make(chan FluxStreamEvent, 4) + usage := &FluxUsage{PromptTokens: 3, CompletionTokens: 5, TotalTokens: 8} + events <- FluxStreamEvent{Type: "usage", Usage: usage} + events <- FluxStreamEvent{Type: "continuation"} + events <- FluxStreamEvent{Type: "usage", Usage: usage} + events <- FluxStreamEvent{Type: "done", Usage: usage} + close(events) + return NewStreamResultWithRequestID(events, "continuation", func() {}), nil +} + +func TestUsageLimitResetsUsageAtContinuation(t *testing.T) { + t.Parallel() + tracker := NewUsageTracker() + provider, err := NewUsageLimitProvider(&continuationUsageProvider{}, tracker) + if err != nil { + t.Fatal(err) + } + result, err := provider.StreamChat(context.Background(), []FluxMessage{{Role: "user", Content: "hello"}}, ChatOptions{Model: "test"}) + if err != nil { + t.Fatal(err) + } + defer result.Close() + for range result.Events { + } + if got := tracker.GetUsage().SessionTokens; got != 16 { + t.Fatalf("session tokens = %d, want 16", got) + } +} + +type blockingTraceProvider struct { + closes atomic.Int32 +} + +func (*blockingTraceProvider) Name() string { return "blocking-trace" } +func (*blockingTraceProvider) Ping(context.Context) error { return nil } +func (*blockingTraceProvider) Chat(context.Context, []FluxMessage, ChatOptions) (*FluxResponse, error) { + return nil, nil +} + +func (p *blockingTraceProvider) StreamChat(ctx context.Context, _ []FluxMessage, _ ChatOptions) (*StreamResult, error) { + streamCtx, cancel := context.WithCancel(ctx) + events := make(chan FluxStreamEvent) + go func() { + <-streamCtx.Done() + close(events) + }() + return NewStreamResultWithRequestID(events, "request-trace", func() { + p.closes.Add(1) + cancel() + }), nil +} + +func TestTracingStreamClosePropagates(t *testing.T) { + t.Parallel() + inner := &blockingTraceProvider{} + tracing := NewTracingProvider(inner) + result, err := tracing.StreamChat(context.Background(), []FluxMessage{{Role: "user", Content: "hello"}}, ChatOptions{Model: "test"}) + if err != nil { + t.Fatal(err) + } + result.Close() + result.Close() + + select { + case _, ok := <-result.Events: + if ok { + t.Fatal("unexpected event after close") + } + case <-time.After(2 * time.Second): + t.Fatal("tracing stream did not close") + } + if got := inner.closes.Load(); got != 1 { + t.Fatalf("inner close count = %d, want 1", got) + } +} + // --- helper: countingCallback for thread-safety test --- type countingCallback struct { diff --git a/provider/observability/recorder.go b/provider/observability/recorder.go index 3da419bc..5b040440 100644 --- a/provider/observability/recorder.go +++ b/provider/observability/recorder.go @@ -6,6 +6,8 @@ import ( "os" "sync" "time" + + "github.com/GrayCodeAI/flux/provider/core" ) // RecorderMode controls whether the recorder records new interactions or replays existing ones. @@ -150,7 +152,7 @@ func (r *RecorderProvider) StreamChat(ctx context.Context, messages []FluxMessag return nil, err } if replayResult != nil { - return replayResult, nil + return core.CoordinateStreamResult(ctx, replayResult), nil } // Record mode: call the inner provider's StreamChat @@ -173,7 +175,8 @@ func (r *RecorderProvider) StreamChat(ctx context.Context, messages []FluxMessag } // Drain events from the real stream and accumulate the response - return r.recordStream(ctx, result, messages, opts, hash), nil + recorded := r.recordStream(ctx, result, messages, opts, hash) + return core.CoordinateStreamResult(ctx, recorded), nil } // checkReplay checks if we're in replay mode and returns the replay result. @@ -278,6 +281,10 @@ func (r *RecorderProvider) syntheticStream(ctx context.Context, resp *FluxRespon // recordStream drains the real stream, accumulates data, saves the interaction, and // returns a synthetic stream with the accumulated response. +// +// The interaction is saved before the terminal event is forwarded: the +// caller's stream is coordinated and ends at that terminal, so a caller that +// saves the cassette right after draining must already see the interaction. func (r *RecorderProvider) recordStream(ctx context.Context, result *StreamResult, messages []FluxMessage, opts ChatOptions, hash string) *StreamResult { streamCtx, cancel := context.WithCancel(ctx) ch := make(chan FluxStreamEvent, 64) @@ -290,6 +297,27 @@ func (r *RecorderProvider) recordStream(ctx context.Context, result *StreamResul var toolCalls []ToolCall var usage *FluxUsage var finishReason string + var saveOnce sync.Once + save := func() { + saveOnce.Do(func() { + r.mu.Lock() + defer r.mu.Unlock() + r.cassette.Interactions = append(r.cassette.Interactions, Interaction{ + Request: RecordedRequest{ + Messages: messages, + Model: opts.Model, + System: opts.System, + Hash: hash, + }, + Response: RecordedResponse{ + Content: r.redact(content), + ToolCalls: toolCalls, + Usage: usage, + FinishReason: finishReason, + }, + }) + }) + } // Drain the real stream, forwarding events to the caller for evt := range result.Events { @@ -307,39 +335,37 @@ func (r *RecorderProvider) recordStream(ctx context.Context, result *StreamResul case "done": finishReason = evt.StopReason } + if isTerminalRecordedEvent(evt) { + save() + } // Forward the event to the caller select { case ch <- evt: case <-streamCtx.Done(): + // The caller stopped early: do not record a partial response. result.Close() return } } - - // Save the accumulated interaction - r.mu.Lock() - interaction := Interaction{ - Request: RecordedRequest{ - Messages: messages, - Model: opts.Model, - System: opts.System, - Hash: hash, - }, - Response: RecordedResponse{ - Content: r.redact(content), - ToolCalls: toolCalls, - Usage: usage, - FinishReason: finishReason, - }, - } - r.cassette.Interactions = append(r.cassette.Interactions, interaction) - r.mu.Unlock() + save() }() return NewStreamResult(ch, cancel) } +// isTerminalRecordedEvent reports whether evt ends a stream: done, a +// cancellation, or an error that is not a warning-marked diagnostic. +func isTerminalRecordedEvent(evt FluxStreamEvent) bool { + switch evt.Type { + case "done", "cancelled", "canceled": + return true + case "error": + return evt.Warning == "" + } + return false +} + // redact applies the redactor function if set. func (r *RecorderProvider) redact(s string) string { if r.redactor != nil { diff --git a/provider/observability/recorder_test.go b/provider/observability/recorder_test.go index cf9e0207..910a1fb3 100644 --- a/provider/observability/recorder_test.go +++ b/provider/observability/recorder_test.go @@ -438,3 +438,63 @@ func TestRecorderRedactor(t *testing.T) { t.Errorf("recorded content = %q, want redacted", c.Interactions[0].Response.Content) } } + +// lingeringStreamProvider sends content and done but closes its channel only +// when release is closed, like a provider that finishes reading the body +// after the terminal event. +type lingeringStreamProvider struct { + release chan struct{} +} + +func (p *lingeringStreamProvider) Chat(context.Context, []FluxMessage, ChatOptions) (*FluxResponse, error) { + return &FluxResponse{Content: "unused"}, nil +} + +func (p *lingeringStreamProvider) StreamChat(context.Context, []FluxMessage, ChatOptions) (*StreamResult, error) { + ch := make(chan FluxStreamEvent, 2) + ch <- FluxStreamEvent{Type: "content", Content: "lingering content"} + ch <- FluxStreamEvent{Type: "done", StopReason: "stop"} + go func() { + <-p.release + close(ch) + }() + return NewStreamResult(ch, nil), nil +} + +func (p *lingeringStreamProvider) Ping(context.Context) error { return nil } +func (p *lingeringStreamProvider) Name() string { return "lingering" } + +func TestRecorderStreamSavesInteractionBeforeTerminal(t *testing.T) { + t.Parallel() + inner := &lingeringStreamProvider{release: make(chan struct{})} + defer close(inner.release) + path := filepath.Join(t.TempDir(), "stream.json") + rec, err := NewRecorderProvider(inner, path, RecordModeRecord) + if err != nil { + t.Fatal(err) + } + + result, err := rec.StreamChat(context.Background(), []FluxMessage{{Role: "user", Content: "stream me"}}, ChatOptions{Model: "m"}) + if err != nil { + t.Fatal(err) + } + defer result.Close() + for evt := range result.Events { + if evt.Type == "done" { + break + } + } + + // The consumer has the terminal event; the interaction must already be + // recorded even though the provider has not closed its channel yet. + if err := rec.Save(); err != nil { + t.Fatal(err) + } + c, err := LoadCassette(path) + if err != nil { + t.Fatal(err) + } + if len(c.Interactions) != 1 || c.Interactions[0].Response.Content != "lingering content" || c.Interactions[0].Response.FinishReason != "stop" { + t.Fatalf("interactions = %+v, want the streamed response recorded", c.Interactions) + } +} diff --git a/provider/observability/tracing.go b/provider/observability/tracing.go index 688703c2..3f247fd1 100644 --- a/provider/observability/tracing.go +++ b/provider/observability/tracing.go @@ -3,6 +3,8 @@ package observability import ( "context" + "github.com/GrayCodeAI/flux/llm" + "github.com/GrayCodeAI/flux/provider/core" "go.opentelemetry.io/otel" "go.opentelemetry.io/otel/attribute" "go.opentelemetry.io/otel/codes" @@ -88,25 +90,25 @@ func (tp *TracingProvider) StreamChat(ctx context.Context, messages []FluxMessag span.SetAttributes(attribute.String("request_id", sr.RequestID)) - // Wrap the events channel so the span ends when the stream finishes. - origEvents := sr.Events - wrappedEvents := make(chan FluxStreamEvent, cap(origEvents)) + streamCtx, cancel := context.WithCancel(ctx) + coordinated := core.CoordinateStreamResult(streamCtx, sr) + wrappedEvents := make(chan FluxStreamEvent, cap(coordinated.Events)) go func() { defer span.End() defer close(wrappedEvents) - for evt := range origEvents { + defer coordinated.Close() + terminal := false + for evt := range coordinated.Events { switch evt.Type { case "error": if evt.Warning != "" { - // Non-fatal health diagnostic: record it without - // failing the span (the stream still completes). span.SetAttributes(attribute.String("warning", evt.Warning)) } else { + terminal = true span.SetStatus(codes.Error, evt.Error) span.SetAttributes(attribute.Bool("error", true)) } case "usage": - // Token usage is delivered on the "usage" event, not "done". if evt.Usage != nil { span.SetAttributes( attribute.Int("usage.prompt_tokens", evt.Usage.PromptTokens), @@ -115,22 +117,24 @@ func (tp *TracingProvider) StreamChat(ctx context.Context, messages []FluxMessag ) } case "done": + terminal = true span.SetStatus(codes.Ok, "") } - // Respect cancellation on the send: if the consumer abandons the - // stream, this goroutine must not block forever forwarding events - // (which would leak the goroutine and keep the span open). select { case wrappedEvents <- evt: - case <-ctx.Done(): - sr.Close() + case <-streamCtx.Done(): return } } + if !terminal { + if err := streamCtx.Err(); err != nil { + span.SetStatus(codes.Error, err.Error()) + } + } }() - return &StreamResult{ - Events: wrappedEvents, - RequestID: sr.RequestID, - }, nil + return llm.NewStreamResult(wrappedEvents, sr.RequestID, func() { + cancel() + coordinated.Close() + }), nil } diff --git a/provider/observability/usage_limit.go b/provider/observability/usage_limit.go index 112ca26d..76b520ca 100644 --- a/provider/observability/usage_limit.go +++ b/provider/observability/usage_limit.go @@ -4,6 +4,8 @@ import ( "context" "errors" "fmt" + + "github.com/GrayCodeAI/flux/provider/core" ) // UsageLimitProvider wraps any Provider and enforces token/cost budgets @@ -80,28 +82,25 @@ func (u *UsageLimitProvider) StreamChat(ctx context.Context, messages []FluxMess return nil, err } - // Wrap the events channel to intercept usage events. - wrappedCh := make(chan FluxStreamEvent, cap(result.Events)) - go func() { - defer close(wrappedCh) - for evt := range result.Events { - if evt.Type == "usage" && evt.Usage != nil { - total := evt.Usage.TotalTokens + var previousUsage *core.FluxUsage + return core.TransformStreamResult(ctx, result, func(_ context.Context, evt FluxStreamEvent) (FluxStreamEvent, error) { + if evt.Type == "continuation" { + previousUsage = nil + return evt, nil + } + if (evt.Type == "usage" || evt.Type == "done") && evt.Usage != nil { + delta := core.UsageDelta(previousUsage, evt.Usage) + previousUsage = core.MergeUsage(previousUsage, evt.Usage) + if delta != nil { + total := delta.TotalTokens if total == 0 { - total = evt.Usage.PromptTokens + evt.Usage.CompletionTokens + total = delta.PromptTokens + delta.CompletionTokens } u.tracker.Record(total, 0, opts.Provider, opts.Model) } - select { - case wrappedCh <- evt: - case <-ctx.Done(): - result.Close() - return - } } - }() - - return NewStreamResult(wrappedCh, result.Close), nil + return evt, nil + }), nil } // recordUsage extracts token count from an FluxResponse and records it. diff --git a/provider/resilience/adaptive_ratelimit.go b/provider/resilience/adaptive_ratelimit.go index d1812356..2ff9be6c 100644 --- a/provider/resilience/adaptive_ratelimit.go +++ b/provider/resilience/adaptive_ratelimit.go @@ -8,6 +8,8 @@ import ( "strconv" "sync" "time" + + "github.com/GrayCodeAI/flux/provider/core" ) // RateLimitState holds the current rate limit tracking state for a provider. @@ -276,45 +278,21 @@ func (a *AdaptiveRateLimitProvider) StreamChat(ctx context.Context, messages []F return nil, err } - // Wrap the events channel to intercept usage events for token tracking. - // Also observe ctx cancellation so that: - // - we stop forwarding events promptly (the caller's wrappedCh is no - // longer being drained on the consumer side); - // - we release the inner stream's resources (body, goroutine) by - // closing it; otherwise it hangs on a blocking send. - wrappedCh := make(chan FluxStreamEvent, cap(result.Events)) - go func() { - defer close(wrappedCh) - for { - select { - case <-ctx.Done(): - result.Close() - // Drain remaining events so the inner goroutine can complete - // and close its own body. - for range result.Events { - } - return - case evt, ok := <-result.Events: - if !ok { - return - } - if evt.Type == "usage" && evt.Usage != nil { - tokens := a.extractTokens(evt.Usage) - a.recordTokens(tokens) - } - select { - case wrappedCh <- evt: - case <-ctx.Done(): - result.Close() - for range result.Events { - } - return - } + var previousUsage *core.FluxUsage + return core.TransformStreamResult(ctx, result, func(_ context.Context, evt FluxStreamEvent) (FluxStreamEvent, error) { + if evt.Type == "continuation" { + previousUsage = nil + return evt, nil + } + if (evt.Type == "usage" || evt.Type == "done") && evt.Usage != nil { + delta := core.UsageDelta(previousUsage, evt.Usage) + previousUsage = core.MergeUsage(previousUsage, evt.Usage) + if delta != nil { + a.recordTokens(a.extractTokens(delta)) } } - }() - - return NewStreamResult(wrappedCh, result.RequestID, result.Close), nil + return evt, nil + }), nil } // UpdateFromHeaders updates the rate limit state from HTTP response headers. diff --git a/provider/resilience/adaptive_ratelimit_test.go b/provider/resilience/adaptive_ratelimit_test.go index 9c471c1b..74168ed0 100644 --- a/provider/resilience/adaptive_ratelimit_test.go +++ b/provider/resilience/adaptive_ratelimit_test.go @@ -228,6 +228,31 @@ func TestAdaptiveRateLimitProvider_StreamChat(t *testing.T) { } } +func TestAdaptiveRateLimitProvider_ResetsUsageAtContinuation(t *testing.T) { + t.Parallel() + inner := &mockProvider{name: "test", streamFn: func(context.Context, []FluxMessage, ChatOptions) (*StreamResult, error) { + events := make(chan FluxStreamEvent, 4) + usage := &FluxUsage{PromptTokens: 3, CompletionTokens: 5, TotalTokens: 8} + events <- FluxStreamEvent{Type: "usage", Usage: usage} + events <- FluxStreamEvent{Type: "continuation"} + events <- FluxStreamEvent{Type: "usage", Usage: usage} + events <- FluxStreamEvent{Type: "done", Usage: usage} + close(events) + return NewStreamResult(events, "", func() {}), nil + }} + provider := mustAdaptiveRateLimitProvider(t, inner, AdaptiveRateLimitConfig{}) + result, err := provider.StreamChat(context.Background(), nil, ChatOptions{}) + if err != nil { + t.Fatal(err) + } + defer result.Close() + for range result.Events { + } + if got := provider.Status().TotalTokens; got != 16 { + t.Fatalf("total tokens = %d, want 16", got) + } +} + func TestAdaptiveRateLimitProvider_UpdateFromHeaders(t *testing.T) { t.Parallel() inner := &mockProvider{name: "test"} diff --git a/provider/resilience/continuation.go b/provider/resilience/continuation.go index 324545c5..366fb6b1 100644 --- a/provider/resilience/continuation.go +++ b/provider/resilience/continuation.go @@ -131,8 +131,10 @@ func StreamChatWithContinuation(ctx context.Context, p core.Provider, messages [ core.Emit(cancelCtx, outCh, core.FluxStreamEvent{Type: "error", Error: err.Error()}) return } + stream = core.CoordinateStreamResult(cancelCtx, stream) var stopReason string + sawDone := false for evt := range stream.Events { switch evt.Type { case "content": @@ -148,12 +150,11 @@ func StreamChatWithContinuation(ctx context.Context, p core.Provider, messages [ core.Emit(cancelCtx, outCh, evt) case "done": stopReason = evt.StopReason + sawDone = true case "error": core.Emit(cancelCtx, outCh, evt) - // Warning-marked error events are non-fatal health - // diagnostics emitted just before the terminal done; - // keep consuming so that done event is observed. if evt.Warning == "" { + stream.Close() return } default: @@ -161,6 +162,9 @@ func StreamChatWithContinuation(ctx context.Context, p core.Provider, messages [ } } stream.Close() + if !sawDone { + return + } // Don't continue if: not max_tokens, had tool calls, or hit token cap if stopReason != "max_tokens" && stopReason != "length" { diff --git a/provider/resilience/guardrails.go b/provider/resilience/guardrails.go index 2bd8a85b..d557949f 100644 --- a/provider/resilience/guardrails.go +++ b/provider/resilience/guardrails.go @@ -80,40 +80,17 @@ func (gp *GuardrailProvider) StreamChat(ctx context.Context, messages []FluxMess return result, nil } - origEvents := result.Events - wrappedEvents := make(chan FluxStreamEvent, cap(origEvents)) - - go func() { - defer close(wrappedEvents) - for evt := range origEvents { - if evt.Type == "content" && gp.guardrails != nil { - violations, checkErr := gp.guardrails.Check(ctx, evt.Content) - if checkErr != nil { - select { - case wrappedEvents <- FluxStreamEvent{ - Type: "error", - Error: checkErr.Error(), - }: - case <-ctx.Done(): - } - result.Close() - return - } - if len(violations) > 0 { - evt.Content = core.ApplyRedactions(evt.Content, violations) - } - } - select { - case wrappedEvents <- evt: - case <-ctx.Done(): - result.Close() - return - } + return core.TransformStreamResult(ctx, result, func(streamCtx context.Context, evt FluxStreamEvent) (FluxStreamEvent, error) { + if evt.Type != "content" { + return evt, nil } - }() - - return &StreamResult{ - Events: wrappedEvents, - RequestID: result.RequestID, - }, nil + violations, err := gp.guardrails.Check(streamCtx, evt.Content) + if err != nil { + return evt, err + } + if len(violations) > 0 { + evt.Content = core.ApplyRedactions(evt.Content, violations) + } + return evt, nil + }), nil } diff --git a/reports/Flux OSS landscape roadmap.md b/reports/Flux OSS landscape roadmap.md new file mode 100644 index 00000000..53d1e228 --- /dev/null +++ b/reports/Flux OSS landscape roadmap.md @@ -0,0 +1,287 @@ +# Build Flux around trustworthy provider interoperability + +**Recommendation:** keep Flux a narrow, embeddable Go provider runtime and make correctness, lossless protocol behavior, and evidence-backed conformance its differentiator—not UI, tenancy, agent orchestration, or model execution. The repository has a credible foundation, but several public promises are ahead of the composed runtime: stateful middleware is rebuilt per request, provider and route state is lost, budgets and migration are not transaction-safe, and the advertised first-run path is not independently reliable. Against the selected 20-project OSS landscape, Bifrost, AxonHub, and GoModel are the closest implementation peers; LiteLLM, Agent Router, Portkey, and the official provider SDKs are the strongest semantic donors. The next sequence is therefore: **make current contracts true, unify provider construction, prove conformance, complete OpenAI and Anthropic semantics, then publish a narrow `v0.1.0`**. Model serving, agent loops, and general gateway control planes remain external concerns, while `v1.0` waits for a compatibility baseline, reproducible releases, security ownership, and external-consumer evidence. + +**Research cutoff:** 2026-09-24. **Flux snapshot:** `f9037a82e403001d50868e4100bcfa19d8a387af` on `chore/oss-competitive-roadmap`. **Evidence convention:** “verified” means confirmed in the cited source or recorded local check; “interpretation” means an engineering judgment; “assumption” means a condition that must be validated before implementation. + +## A narrow thesis beats feature-count parity + +Flux already has the right product boundary: hosts import `engine`, `llm`, `graph`, and `tools`, while Flux owns credentials, catalog resolution, provider transports, normalized streams, retry/fallback, usage, and provider telemetry (`README.md:31-65`; `docs/architecture/HOST-ENGINE-BOUNDARY.md:6-36`). Hosts own UX, sessions, permissions, tool execution, and agent loops. That boundary is strategically valuable because the closest peers usually bundle one of the rejected layers: a web console and tenancy, a Kubernetes control plane, a language SDK ecosystem, or model execution. The competitive product is consequently not “another LiteLLM.” It is: + +> **the narrow, auditable Go runtime that makes heterogeneous hosted and self-hosted LLM endpoints behave like one explicit, lossless, conformance-tested contract, without requiring a gateway service, database, control plane, agent runtime, or model server.** + +The research produced two related but distinct views that must not be conflated. `selection_methodology.md` defines a reproducible **evidence-ranked** portfolio using public source, an OSI-approved core, activity within 180 days, Flux-specific evidence, and the score `6F + 4C + 3I + 3M + 2G + O + E`. It caps the evidence set at eight direct peers, six serving runtimes, four provider SDKs, and two adjacent donors. `gap_mapping.md` then reconciles that portfolio with the runtime and SDK findings into the **decision portfolio** below; it explicitly supersedes the scoped lists where they conflict. GitHub counts, release labels, and maintenance signals are 2026-09-24 snapshots, not quality scores, and upstream documentation establishes intent rather than field-by-field wire conformance. ([GitHub repository API](https://docs.github.com/en/rest/repos/repos#get-a-repository)) + +| Decision rank | Project | Portfolio role | Evidence-ranked position | Primary Flux use | Decision | +|---:|---|---|---|---|---| +| 1 | [Bifrost](https://github.com/maximhq/bifrost) | Direct peer; closest Go runtime | 1 (97.0) | Provider abstraction, operation matrices, raw-response inspection, narrow middleware | **Borrow and build evidence surfaces** | +| 2 | [LiteLLM](https://github.com/BerriAI/litellm) | Direct peer; feature benchmark | 2 (95.0) | Per-attempt policy revalidation, route/attempt semantics, spend, privacy-safe OTel | **Borrow semantics; reject platform shape** | +| 3 | [AxonHub](https://github.com/looplj/axonhub) | Direct peer; Go transformer pipeline | 4 (91.5) | Inbound dialect → canonical request → provider → reverse stream/error transforms | **Borrow architecture** | +| 4 | [GoModel](https://github.com/ENTERPILOT/GoModel) | Direct peer; compact Go runtime | 5 (90.0) | Policy composition, aliases, budgets, usage, migration ergonomics | **Borrow selectively** | +| 5 | [Agent Router](https://github.com/theagentrouter/agent-router) | Direct peer; reliability/extension donor | 7 (84.5) | Stream idle deadlines, per-request credentials, provider extensions | **Borrow reliability semantics** | +| 6 | [Portkey Gateway](https://github.com/Portkey-AI/gateway) | Direct peer; maintenance-watch benchmark | 10 (81.5) | Request-attached policy, sticky routing, conditional fallback | **Borrow policy shape only** | +| 7 | [OmniRoute](https://github.com/diegosouzapw/OmniRoute) | Direct peer; high-velocity gateway donor | 3 (94.0) | Provider fallback, aliases, native passthrough, migration diagnostics | **Borrow selectively** | +| 8 | [Manifest (`mnfst/llm-gateway`)](https://github.com/mnfst/llm-gateway) | Direct peer; routing-focused | 8 (83.5) | Route hints, model tiers, affinity, cost-aware selection | **Defer until baselines exist** | +| 9 | [LocalAI](https://github.com/mudler/LocalAI) | Serving runtime; local profile target | 6 (85.5) | Backend-neutral discovery and profile metadata | **Adopt as conformance target** | +| 10 | [Ollama](https://github.com/ollama/ollama) | Serving runtime; local lifecycle profile | 9 (82.0) | Native tags/show/keep-alive and structured-output profile | **Build read-only profile support** | +| 11 | [llama.cpp](https://github.com/ggml-org/llama.cpp) | Serving runtime; portable profile | 15 (69.0) | Native props/model/slot evidence and constrained JSON | **Adopt as endpoint profile** | +| 12 | [vLLM](https://github.com/vllm-project/vllm) | Serving runtime; conformance target | 12 (74.0) | Version-sensitive OpenAI/Anthropic surfaces and serving topology | **Adopt as conformance target** | +| 13 | [SGLang](https://github.com/sgl-project/sglang) | Serving runtime; gateway-boundary donor | 11 (74.5) | Cache/load signals, gateway lifecycle, plane separation | **Adopt as profile and signal donor** | +| 14 | [any-llm-go](https://github.com/mozilla-ai/any-llm-go) | Go optional-interface donor | Secondary scoped donor | Optional operation ports and explicit unsupported states | **Borrow interface shape; validate before dependency use** | +| 15 | [Vercel AI SDK](https://github.com/vercel/ai) | Model-port and middleware donor | Secondary framework donor | Versioned model port, provider metadata, deterministic middleware | **Borrow semantics; reject UI and agent layers** | +| 16 | [OpenAI Python SDK](https://github.com/openai/openai-python) | OpenAI protocol donor | 16 (66.0) | Chat/Responses wire shapes, typed SSE, tool events | **Adopt pinned fixtures** | +| 17 | [Anthropic Python SDK](https://github.com/anthropics/anthropic-sdk-python) | Anthropic protocol/replay donor | 18 (62.0) | Typed Messages events, accumulation, signed/redacted replay | **Adopt pinned fixtures** | +| 18 | [Google GenAI Python](https://github.com/googleapis/python-genai) | Gemini/Vertex protocol donor | 19 (61.5) | Multimodal parts, tools, safety, streaming | **Adopt pinned fixtures** | +| 19 | [OpenTelemetry GenAI conventions](https://github.com/open-telemetry/semantic-conventions-genai) | Interoperability standard | 20 (58.0) | Portable GenAI spans, usage, finish reasons, privacy | **Adopt as standard** | +| 20 | [Gateway API Inference Extension](https://github.com/kubernetes-sigs/gateway-api-inference-extension) | Deployment interoperability donor | Secondary scoped donor | Backend-neutral model/pool identity, endpoint-picker and capability signals | **Borrow concepts through optional adapters** | + +This reconciliation retains all eight direct peers, narrows the serving group to five high-value profiles, and replaces three evidence-ranked entries with donors that close more direct contract or deployment gaps. Specifically, `go-openai` (evidence rank 13), Langfuse (14), and TensorRT-LLM (17) remain valuable secondary donors but leave the decision twenty for any-llm-go, Vercel AI SDK, and Gateway API Inference Extension. “Decision rank” expresses Flux implementation priority, not project quality, popularity, or performance. + +| Layer | Representative projects | What Flux should learn | What Flux must not import | +|---|---|---|---| +| Data plane | Bifrost, AxonHub, GoModel, Agent Router, Portkey, OmniRoute, Manifest | Transform isolation, per-attempt policy, deadline/recovery, declarative routing | UI, tenancy, MCP hosting, cluster management | +| Serving target | LocalAI, Ollama, llama.cpp, vLLM, SGLang | Runtime/version evidence, endpoint profiles, load/capability signals | Kernels, weights, batching, KV cache, GPU scheduling | +| Contract and port | any-llm-go, Vercel AI SDK, OpenAI, Anthropic, Google GenAI | Optional interfaces, deterministic middleware, typed events, partial JSON, provider replay | SDK objects in `engine`, provider-specific host semantics | +| Evidence and deployment | OpenTelemetry GenAI, Gateway API Inference Extension | Portable telemetry and backend-neutral routing signals | Embedded evaluation UI, Kubernetes types in stable `engine` DTOs | + +A wider secondary donor set remains useful: go-openai, Langfuse, TensorRT-LLM, New API’s independent `relaykit` boundary, APISIX and Higress lifecycle ordering, Helicone’s P2C/PeakEWMA strategy, LLM Gateway’s migration/doctor UX, BAML schema descriptors, Pydantic AI test models, and OpenAI Agents stream/usage rules. Each pattern should enter Flux only when it reduces a verified contract or operational gap. + +## Flux has a strong core whose promises outrun composition + +The source has more than leaf scaffolding. Canonical DTOs are centralized in `llm`, and `provider/core` aliases them rather than maintaining a second wire model (`llm/types.go:49-75`; `provider/core/core.go:35-97`). The lower-level `core.Provider` remains only `Chat`, `StreamChat`, `Ping`, and `Name` (`provider/core/core.go:21-33`). Streaming uses a bounded SSE parser and semantic events, routing has real weighted, least-busy, latency, cost, and usage strategies, and signed route manifests validate revisions, signatures, redirects, and last-good state. The test suite is unusually broad, and the recorded audit passed race tests, vet, lint, formatting, and `govulncheck`; CI additionally covers fuzzing, secrets, dead code, duplication, Markdown, and cross-platform builds (`.github/workflows/ci.yml:34-300`). + +The problem is composition. The normal engine path reloads catalog/config and reconstructs adapters, rate limiters, caches, and deployment routers on every request. Selection and execution can therefore observe different state, and the advertised circuit, rate-window, and cache state does not survive a call. This is not a performance footnote: it makes existing reliability features semantically different from their descriptions. + +| Area | Verified strength | Current gap or risk | Evidence | +|---|---|---|---| +| Host boundary | Compile-checked `engine` contract and narrow low-level provider port | Public host `llm.Provider` still bundles several maintenance facets; `runtime` docs call an older entry point “recommended” | `docs/architecture/HOST-ENGINE-BOUNDARY.md:33-36`; `llm/provider.go:16-24`; `runtime/runtime.go:1-22` | +| Canonical DTOs | Multimodal input, tools, reasoning, usage, route, warnings, and provider blocks are modeled | Several fields have no producer or are dropped before the host | `llm/types.go:29-47`; `llm/types.go:49-75`; `llm/types.go:219-307` | +| Provider coverage | 28 registry IDs across Anthropic Messages, Gemini, OpenAI Responses, and OpenAI Chat families | Direct and deployment construction can select different protocols; 28 IDs are not 28 independent conformance levels | `catalog/registry/providers.go:16-341`; `provider/provider_registry.go:74-179`; `setup/deployment.go:183-383` | +| Streaming | Bounded parser, semantic deltas, cancellation-aware provider requests | EOF can look successful; tracing and guardrail wrappers drop the source `Close`; split usage and provider state are lost | `provider/core/stream.go:31-89`; `engine/stream.go:82-165`; `provider/observability/tracing.go:91-135`; `provider/resilience/guardrails.go:83-118` | +| Routing | Deployment validation, fallback, breaker filtering, live atomic snapshots | Fallback ignores caller intent, tools can be stripped, actual deployment/attempt is not returned | `router/deployment_router.go:134-247`; `router/deployment_router.go:320-389`; `router/deployment_router.go:554-576` | +| Resilience | Retry, continuation, rate limits, cache, guardrails, health, tracing | Stateful wrappers are rebuilt; deadline classes and admission limits are absent | `engine/engine.go:303-327`; `provider/core/transport.go:10-11`; `provider/core/transport.go:56-63` | +| Budget/cost | Check, usage records, and analytics exist | Check and record are separate; no reservation, operation ID, or idempotent finalization | `provider/observability/budget_provider.go:73-142`; `storage/budgets.go:157-209` | +| Credential safety | Injected stores, sanitized state, atomic config writes | Migration can remove plaintext after partial failure; endpoint and secret-at-rest policy remain weak | `credentials/migrate.go:31-85`; `storage/budgets.go:44-87` | +| Observability | Public OTel wrapper and metrics exist | No normal engine construction path wires the wrapper; attributes are custom/non-current, and the conversation span ends before asynchronous stream work | `provider/observability/tracing.go:12-136`; `internal/observability/genai_semconv.go:21-58`; `conversation/engine.go:166-175` | +| Delivery | Internal HTTP, optional gRPC, and three SDK drafts exist | These are not coherent public products; OpenAPI and SDKs are incomplete and outside root gates | `internal/api/server.go:101-121`; `api/openapi.yaml:118-340`; `internal/sdk/go/go.mod:1-3` | +| Release | `VERSION` is `0.0.1`; the research snapshot verifies a signed tag; CI has strong static checks | No release PR, API diff, module-proxy gate, SBOM/provenance gate, or repeatable `v0.1` policy | `VERSION:1`; `.github/workflows/release.yml:1-31`; [recorded `v0.0.1` release](https://github.com/GrayCodeAI/flux/releases/tag/v0.0.1) | + +The highest-risk behavior follows directly from the source. `engine.toClientOptions` does not map canonical `ThinkingEnabled`, although adapters consume it (`engine/convert.go:15-34`). Anthropic response parsing skips redacted thinking and the message builder has no provider-block replay branch (`provider/adapters/anthropic.go:233-300`; `provider/adapters/anthropic.go:303-429`). The engine normalizer does not copy provider blocks (`engine/stream.go:124-165`). Conversation persistence reads usage only on `done`, while real Anthropic and OpenAI processors can emit usage earlier or separately (`conversation/engine.go:266-291`; `provider/core/stream.go:236-281`; `provider/core/stream.go:502-512`). The stable engine continuation path has a default continuation and total-token cap, but still synthesizes `done` when a source closes without a terminal event; the separate conversation path has no hard continuation-call bound (`engine/continuation.go:15-23`; `engine/continuation.go:64-70`; `conversation/engine.go:305-313`). + +Routing can silently change semantics. Automatic fallback follows explicit selection, and deployments lacking requested tools remain eligible while `optsForOffering` removes unsupported tools (`router/deployment_router.go:320-389`; `router/deployment_router.go:554-576`). The response route is overwritten with the initially selected route rather than the successful fallback deployment (`engine/convert.go:88-96`; `llm/types.go:235-246`). This undermines billing, audit, debugging, and any future reliability policy that depends on knowing which provider served the request. + +Operational controls are similarly incomplete. Budget checks and later records are not one transaction, and stream usage events are independently recorded while errors are ignored (`provider/observability/budget_provider.go:73-142`). The response cache omits response-affecting fields, is not semantically similar despite its name, and returns copies that still alias provider-block and warning slices (`provider/cache/semantic_cache.go:52-138`; `provider/cache/semantic_cache.go:176-199`; `provider/cache/semantic_cache.go:296-339`; `provider/core/copy.go:3-56`). The coalescer ties shared work to the first caller and does not decrement waiter state on every exit (`provider/resilience/coalesce.go:16-48`; `provider/resilience/coalesce.go:113-189`). + +The onboarding and delivery story compounds those runtime gaps. A normal engine call requires a prepared catalog cache and otherwise points users to a Rho command (`engine/state.go:86-103`; `catalog/v1.go:615-644`). The strict decoder rejected an additive field observed in the then-current [default catalog](https://langdag.com/model-catalog/v1/catalog.json) on 2026-09-24 (`catalog/v1.go:444-460`). All examples use lower-level `provider` packages rather than the accepted host facade; the streaming example claims auto-continuation but calls the non-continuing method; and the multi-provider example performs manual fallback (`examples/basic/main.go:8-20`; `examples/streaming/main.go:1-33`; `examples/multi-provider/main.go:31-44`). OpenAPI documents 11 paths while the internal server registers 15, omitting `/ready`, `/rerank`, and `/v1/chat/completions` (`api/openapi.yaml:118-340`; `internal/api/server.go:101-121`). The Go SDK is a nested module outside root `go test ./...`, while the Python and TypeScript clients live under `internal` (`internal/sdk/go/go.mod:1-3`; `go.mod:1-41`). The Python delete method expects JSON, matching the current server but not OpenAPI’s declared `204` response (`internal/sdk/python/flux.py:68-75`; `internal/api/server.go:266-273`; `api/openapi.yaml:223-237`). + +The honest baseline is therefore **strong primitives, partial end-to-end behavior**. A provider-count-oriented roadmap would hide exactly the defects adopters are most likely to hit. Flux should first reduce semantic edges: options, stream termination, usage, provider replay, fallback requirements, request identity, actual route, and cancellation. + +## Correctness and trust should unlock the roadmap + +### One runtime snapshot and one stream state machine + +The first architectural correction is to make an immutable runtime snapshot the unit of selection and execution. It should contain the compiled catalog, provider configuration, resolved adapters, router, breaker, limiter, cache, and immutable secret references. `Engine` should load once at construction and swap snapshots atomically only after explicit catalog, config, or credential invalidation. Selection and execution must use the same generation, while credential rotation should replace a secret reference without rebuilding unrelated state. This directly fixes the current per-request reconstruction in `engine/engine.go:303-327` and `setup/deployment.go:94-110`. + +The second correction is a shared stream coordinator under `engine` and `provider/core`. Every adapter, wrapper, and facade should produce exactly one terminal state: `done`, `error`, or `cancelled`. A transport EOF without a provider terminal event is truncation, not success. Every wrapper must preserve the upstream `Close` callback, use a request-owned child context, stop forwarding on slow-consumer cancellation, and close/drain the body. Split usage should be accumulated into one final aggregate; provider request IDs, finish reasons, route attempts, warnings, and replay blocks should survive that final event. This is the smallest shared fix for several independent high-severity findings. + +Agent Router’s stream idle timeout and failover behavior demonstrates why Flux should separate response-header, first-event, first-visible-output, inter-event idle, and total-deadline clocks. ([Agent Router v1.1 notes](https://theagentrouter.ai/release-notes/v1.1)) Flux should add bounded admission separately from deadlines: configurable global, provider, tenant, and queue limits; deterministic queue-full errors; cancellation while waiting; and no provider call for a request rejected at admission. The current ten-minute end-to-end HTTP timeout is too coarse to protect memory, connections, or spend under stalled streams (`provider/core/transport.go:10-63`). + +### Provider-native state must be explicit, not best effort + +The canonical DTO should represent two kinds of state deliberately: normalized content that hosts can use and opaque provider blocks required for replay. Anthropic signed/redacted thinking, OpenAI reasoning items, and Gemini thought state should remain namespaced, size-bounded, redacted by policy, and protocol-filtered on the next turn. The official Anthropic SDK’s event accumulation and replay helpers are the primary fixture source, not a runtime dependency. ([Anthropic stream helpers](https://github.com/anthropics/anthropic-sdk-python/blob/main/helpers.md)) + +When a provider does not support a requested field, Flux should transform it only when the transformation is safe and emit a structured warning. If a required capability cannot be preserved, the route should be rejected or the call should fail explicitly. Tools must never be removed merely to make a fallback provider eligible. A universal DTO cannot faithfully represent every future provider extension, so raw passthrough is valuable only as a bounded, versioned, redacted escape hatch after canonical OpenAI and Anthropic behavior is correct. + +### Conformance replaces provider-count confidence + +Flux already has a `verify` package, but it is not wired into adapters or CI. The suite should become the authoritative contract for auth, endpoint construction, blocking and streaming requests, tools, schema output, reasoning/replay, usage, errors, cancellation, close semantics, and unknown additive fields. Each case must return `PASS`, `UNSUPPORTED`, `FAIL`, `TIMEOUT`, or `NOT_TESTED`; unsupported is a valid result only when the capability matrix says so and the route policy rejects semantic loss. + +| Tier | Scope | Deterministic gate | Live evidence | +|---|---|---|---| +| A | Anthropic, OpenAI, Gemini, Azure, Bedrock, Vertex | Pinned official-SDK fixtures plus Flux contract tests | Credential-gated nightly/weekly probes with redacted cassettes | +| B | OpenAI-compatible gateways and specialized adapters | Provider/profile-specific wire fixtures, errors, usage, cancellation | Periodic, budget-capped smoke and last-verified metadata | +| C | User-supplied custom OpenAI endpoints | Generic wire subset only | Host-owned; Flux reports unsupported fields rather than claiming parity | +| Local | Ollama, llama.cpp, vLLM, SGLang, LocalAI | Pinned profile fixtures and discovery parsing | Version-pinned endpoint probes; never provider API calls | + +Operation support should be recorded as an `operation × provider/profile × model × runtime/version × configuration` tuple with source, confidence, observation time, and expiry. A generic “OpenAI compatible” or non-local provider boolean is not sufficient. LocalAI’s discovery document, llama.cpp’s `/props`, Ollama’s tags/show endpoints, and KTransformers’ qualified support matrix all demonstrate why support evidence must preserve model, runtime, hardware, and configuration qualifiers. ([LocalAI discovery](https://localai.io/docs/features/api-discovery/index.html); [llama.cpp server](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md); [Ollama API](https://github.com/ollama/ollama/blob/main/docs/api.md); [KTransformers matrix](https://ktransformers.net/en/docs/support-matrix)) + +### Security must fail closed at host boundaries + +The embedded API is internal, so it should not become a hidden public gateway by accident. `NewServer` does not validate dependencies or safe transport posture, `ServeHTTP` bypasses the listener’s bind check, and plain HTTP is accepted on a non-loopback address when an API key exists (`internal/api/server.go:50-85`; `internal/httputil/httputil.go:88-103`). Flux should require a validated constructor, refuse non-loopback plain HTTP unless an explicit trusted-proxy/TLS mode is configured, attach authorization and tenant scope to every sensitive route, and make readiness test real provider/configuration readiness. + +Credential migration must be transactional: enumerate all recognized values, write and verify every destination, preserve unknown or failed values in a permission-restricted source/backup, and write a success marker only after complete success. SQLite main, WAL, and SHM files must be hardened under a permissive umask. Provider keys should be stored as references or protected secrets rather than plaintext budget rows. Custom endpoint validation must cover scheme, redirect, resolved address, DNS rebinding, port, and private/local-network policy at connection time, with explicit operator opt-in for local endpoints. + +Budgets need reserve → execute → idempotently finalize semantics. A reservation must atomically prevent concurrent overspend; finalization must use a Flux operation ID, aggregate split usage once, reconcile cancellation and partial cost, and surface failed ledger writes as retryable operational events. Cache and coalescer identity must include tenant, provider, deployment, every response-affecting message part, tool/schema definition and choice, response format, reasoning controls, stop conditions, sampling controls, and provider options. Returned metadata must be deeply isolated. + +OpenTelemetry GenAI conventions should be adopted through the OTel API, with content capture off by default and compatibility aliases only during a documented migration. Flux should use an in-memory exporter test to prove provider/model identity, operation, tokens, finish reason, route attempts, error class, and actual stream termination. ([OpenTelemetry GenAI spans](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md)) + +## Turn the landscape into explicit build decisions + +### Build, adopt, borrow, defer, and reject + +| Decision | Flux work or standard | Primary donors | Reason | +|---|---|---|---| +| **Build** | Shared stream terminal/cancellation coordinator | Agent Router, Bifrost, OpenAI/Anthropic SDKs | Fixes cross-layer correctness rather than one adapter | +| **Build** | Lossless canonical state and warnings | Anthropic, OpenAI, Google GenAI, Vercel AI SDK | Makes declared DTO fields real across two turns | +| **Build** | Immutable runtime snapshot and atomic refresh | Bifrost, GoModel, LiteLLM | Makes stateful resilience true and selection deterministic | +| **Build** | Fresh bootstrap and catalog trust pipeline | Flux’s signed-manifest controls | Removes Rho dependency and mutable-source startup failure | +| **Build** | One provider descriptor/factory and 28-ID parity tests | AxonHub, New API `relaykit`, any-llm-go | Eliminates entry-point protocol divergence | +| **Build** | Conformance v1, capability evidence, endpoint profiles | Bifrost matrix, LocalAI discovery, official SDKs | Replaces nominal provider count with reproducible evidence | +| **Build** | Atomic budgets, exact cache identity, safe migration | LiteLLM spend semantics, secure state patterns | Protects money, secrets, and tenant isolation | +| **Adopt** | OpenTelemetry GenAI semantics and privacy defaults | OpenTelemetry, LiteLLM OTel v2 | Portable, current, and less API surface than a custom sink | +| **Adopt** | Official-SDK pinned wire fixtures | OpenAI, Anthropic, Google | Faster and more trustworthy than handwritten assumptions | +| **Adopt** | Go API diff, module-proxy smoke, SBOM/provenance/attestation | Go and GitHub release tooling | Creates a reproducible supply-chain baseline | +| **Adopt** | AIPerf-style methodology and machine-readable results | AIPerf | Measures client overhead without fake GPU/provider ranking | +| **Borrow** | Transformer pipeline layering | AxonHub | Isolates inbound, canonical, outbound, and reverse transforms | +| **Borrow** | Per-attempt auth/budget/capability revalidation | LiteLLM | Prevents fallback policy bypass | +| **Borrow** | Raw-response/debug inspection with redaction | Bifrost | Exposes translation loss without leaking raw secrets | +| **Borrow** | Request-attached policy and sticky affinity | Portkey | Makes route intent explicit after deterministic baselines exist | +| **Borrow** | Idle timeout, per-request credentials, provider extensions | Agent Router | Improves stream recovery and compatibility | +| **Borrow** | Deterministic `transformParams`/wrap generate/wrap stream seam | Vercel AI SDK | Adds extension without agent/plugin sprawl | +| **Borrow** | Optional operation interfaces and explicit unsupported states | any-llm-go | Keeps `core.Provider` narrow | +| **Borrow** | Schema descriptor plus validation/repair metadata | BAML, Pydantic AI | Makes structured output explicit without owning a prompt lab | +| **Borrow** | Backend-neutral endpoint signals | Gateway API Inference Extension | Adds pool/load/capability hints without Kubernetes DTOs | +| **Borrow** | Request-scoped `order`/`only`/`sort` routing hints and OpenRouter provider-preference passthrough | OpenRouter, Vercel AI Gateway (hosted; design donors only) | Normalizes the routing controls hosts already use without widening operator policy; see the 2026-09-27 addendum in `oss_gateways.md` | +| **Adopt** | Models.dev catalog data with a reviewed overlay and per-field provenance | Models.dev (MIT) | Implements the planned WP14 generator from a maintained, openly licensed source | +| **Defer** | Default semantic cache and semantic routing | LiteLLM, Bifrost, Helicone | Exact identity and cost/latency baselines are not yet trustworthy | +| **Defer** | P2C/PeakEWMA and sticky affinity | Helicone, Portkey | Useful but lower priority until persistent snapshots and tenant scope exist | +| **Defer** | Public model-call middleware and custom selectors | Vercel AI SDK, current Flux seams | Lifecycle defects would be multiplied by extension points | +| **Defer** | Responses/raw passthrough and more endpoint breadth | OpenAI SDK, LiteLLM, Agent Router | Add translation paths before canonical losslessness is proven | +| **Defer** | Public HTTP, gRPC, and SDK extraction | Portkey deployment portability | Valuable only as separately versioned, tested modules | +| **Defer** | Evaluation/guardrail ports and model administration | Langfuse, DSPy, Ollama, Triton | Host or admin concerns; Flux should export evidence, not products | +| **Defer** | Wasm or unrestricted plugins | Higress, APISIX | Sandbox, ABI, memory, signing, and credential risks exceed current need | +| **Reject from core** | UI, agent runtime, tools execution, RAG, memory, sessions | Dify, Open WebUI, LangGraph, Pydantic AI | Direct host duplication and dependency-direction violation | +| **Reject from core** | Tenancy, virtual organizations, billing, admin control plane | LiteLLM, Kong, New API | Platform concerns and mandatory state/operations | +| **Reject from core** | Model execution, weights, batching, KV cache, GPU scheduling | vLLM, SGLang, llama.cpp, TensorRT-LLM | Wrong architectural layer | +| **Reject from core** | General API-gateway plugin ecosystem or MCP process host | Kong, APISIX, Higress, Bifrost | Broad attack surface and product identity drift | +| **Reject from core** | Mandatory PostgreSQL, Redis, Kubernetes, or distributed cache | LiteLLM, Agent Router, KServe | Defeats embeddability and host neutrality | + +Bifrost’s 2026 MCP registration vulnerability is a useful boundary lesson: merging management and execution planes expands the consequences of an unauthenticated control surface. ([CVE analysis](https://research.jfrog.com/vulnerabilities/bifrost-is-vulnerable-to-unauthenticated-remote-code-execution-via-mcp-stdio-client-registration-cve-2026-90898/)) Flux should not host command-spawning MCP, unrestricted process plugins, or an ambient management listener merely to match feature lists. + +### Provider and protocol gaps should be closed by importance + +| Provider/protocol concern | Current Flux state | Required direction | +|---|---|---| +| Provider identity and factory | 28 IDs, four declared families, but multiple constructor switches | One descriptor, canonical aliases, one resolver, parity test for every ID | +| Gemini | Native client exists, but direct `gemini` does not select it; Vertex paths diverge | Make `ProviderSpec` and factory authoritative; test native and compatible profiles | +| Vertex | Direct path is Anthropic-on-Vertex while deployment path is native Gemini | Declare distinct profiles explicitly; never select by entry point | +| Concentrate | Deployment uses Responses while direct path uses generic OpenAI | Use the declared Responses factory everywhere or declare two provider IDs | +| Ollama | Thin OpenAI wrapper | Add read-only native tags/show/profile metadata; no model pull/load in `engine` | +| OpenAI Chat | Text/tool builder exists; system/developer and JSON-schema descriptor need conformance | Typed translator with pinned official fixtures and no silently ignored accepted fields | +| Anthropic Messages | Native adapter exists; signed/redacted replay and public facade do not | Preserve blocks, events, usage split, finish reasons, and multi-turn replay | +| Gemini/Vertex | Native multimodal/tool support exists; thinking and URL semantics are partial | Pin current API fixtures and preserve provider-native state | +| Embeddings/batch/rerank/media/moderation | Separate leaf interfaces/clients | Keep optional ports; add independent conformance only when demand warrants it | +| Provider count | Nominal gateway IDs | Publish operation/profile/model/version support tiers and last-verified dates | + +Protocol work should use pinned official fixtures and opt-in live probes. Current documentation is evidence of intended semantics, but no live provider credentials were available in the audit. The OpenAI JSON-schema descriptor, Gemini image URL handling, current provider SDK shapes, and release boundaries must therefore be validated before changing wire behavior. + +### API and OSS maturity need one truth + +| Surface | Current state | Recommendation | +|---|---|---| +| `engine` | Accepted Go host facade with `Generate`/`Stream` | Stabilize the existing minimal generation path before adding breadth | +| Root `flux` | Not yet the primary API | Consider a small root `New/Generate/Stream/Close` facade for `v0.1`, preserving `engine` during migration | +| `llm.Provider` | Broad composition facet | Separate generation from optional controller/maintenance interfaces | +| Internal HTTP | Useful local component; weak secure defaults and lossy conversation path | Keep internal or extract only after typed translation and fail-closed construction | +| OpenAPI | Hand-maintained, incomplete relative to routes | Validate against a public handler or remove from the core promise | +| gRPC | Tagged, skeletal, optional | Keep experimental; do not market until generated clients/protobuf and compatibility exist | +| Go SDK | Nested module, excluded from root CI | Extract as independent module only if delivery is a validated use case | +| Python/TypeScript SDKs | Internal stubs outside root CI | Remove from stable claims or publish tested, separately versioned modules | +| Examples | Lower-level and partly mislabeled | Rewrite against stable API and compile/run them as external-package tests | +| Catalog publication | Remote mutable artifact outside repository governance | Add in-repo generation, schema compatibility, publisher identity, digest/signature, last-good cache | +| Release | Tag-triggered GitHub release notes | Add release PR, API diff, clean consumer, module proxy, SBOM/provenance, verified tag | +| Security operations | CI scanning exists; repository alerts/push protection were disabled in the audit snapshot ([repository security snapshot](https://api.github.com/repos/GrayCodeAI/flux)) | Enable native alerts/secret push protection and own support/security policy in Flux | + +The repository should publish a stability inventory: **stable** for the minimal host generation contract, **advanced** for deliberate extension interfaces, and **experimental/internal** for delivery, distributed routing, and broad package surfaces. No `v1.0` claim should precede an import census, `apidiff` baseline, migration guide, support window, and external-consumer test. + +## Sequence hardening into a credible release + +The backlog, phase gates, versioning targets, and 90-day sequence below are **proposed acceptance criteria and engineering recommendations**, not implemented capabilities, dated commitments, or validated demand estimates. + +### Prioritized backlog + +| Priority | Workstream | Minimum acceptance evidence | +|---|---|---| +| P0 | Central stream terminal/cancellation coordinator | Exactly one terminal event; truncated EOF fails; every `Close` cancels source/body; race tests show no forwarding leak | +| P0 | Lossless canonical options, stream state, provider replay, route, and warnings | Every documented option has conversion tests; signed/redacted/provider state round-trips over two turns | +| P0 | Immutable runtime snapshot and explicit refresh | One state load per generation; limiter/cache/breaker survive calls; selection and execution share a snapshot | +| P0 | Fresh-state bootstrap and catalog trust | Clean temporary home runs mock-backed `New → Generate/Stream`; current artifact validates; last-good recovery works | +| P0 | One provider descriptor and factory | All 28 IDs resolve identically through direct/deployment paths; Gemini, Vertex, and Concentrate are regressions | +| P0 | Transactional credential migration and secret/storage hardening | Failed or partial writes never remove source; sidecar modes and endpoint policy are tested | +| P0 | Atomic budget reserve/finalize | Concurrent calls cannot overspend; duplicate operation IDs are idempotent; cancellation reconciliation is durable | +| P0 | Canonical cache/coalescer identity | Every response-affecting field and scope changes identity; returned state is isolated; “exact” is named truthfully | +| P0 | Admission, lifecycle deadlines, fallback policy, actual route attribution | Bounded queue; distinct clocks; fallback opt-out; required tools reject rather than strip; actual attempt returned | +| P0 | Shared conformance v1 | Every built-in adapter runs shared request/stream/error/usage/cancel cases or explicit unsupported rows | +| P0 | Truthful `v0.1` release and docs | Clean external module consumes the candidate; examples run; provider tables generate; API diff and release gates pass | +| P1 | Full OpenAI Chat Completions translator | Multimodal messages, tools/results, schema, options, usage, stream options, and errors pass official fixtures | +| P1 | Anthropic Messages facade and replay | Thinking, tools, split usage, stop reasons, errors, and two-turn replay pass official fixtures | +| P1 | Operation/profile/model capability evidence | Support is qualified by runtime/version/config with source, confidence, freshness, and expiry | +| P1 | Read-only endpoint profiles | Ollama, llama.cpp, vLLM, SGLang, and LocalAI metadata imports are bounded, redacted, and versioned | +| P1 | Current privacy-safe OTel GenAI contract | In-memory exporter verifies attributes, stream completion, and content-off defaults | +| P1 | Performance, soak, and live-provider evidence | Versioned local results and nightly `PASS/UNSUPPORTED/FAIL/TIMEOUT/NOT_TESTED` reports | +| P1 | Public selector and narrow model-call middleware | Deterministic order; generate/stream wrappers preserve terminal, cancel, usage, and route state | +| P1 | Signed-manifest production proof | TLS, separate process, peer outage, stale revision, restart, persisted last-good, and drain; otherwise remain experimental | +| P2 | Optional operation ports | Embeddings, batch, rerank, media, and moderation have independent interfaces/suites; `core.Provider` stays four methods | +| P2 | Migration/compatibility diagnostics | Report effective route and every transformed, dropped, unsupported, or required-host field | +| P2 | Responses/raw passthrough | Explicit, versioned, size-bounded, redacted; canonical mode warns on every lossy conversion | +| P2 | Genuinely adoptable delivery modules | Any retained HTTP/gRPC/SDK is public, independently versioned, route/schema tested, and consumer-tested | +| P3 | Evaluation and semantic-guardrail ports | Privacy-safe fixture replay and typed callbacks; no dataset, optimizer, policy catalog, or UI | +| P3 | Separate model-server administration | Explicit admin credentials and operations; never imported by stable `engine` | +| P3 | Wasm/edge adapter only after a proven gap | Documented use case, ABI, sandbox, limits, signing, rollback, and no ambient credentials | + +### Phased delivery gates + +| Phase | Objective | Exit gate | +|---|---|---| +| 1. Hardening and truth | Make current claims true and establish one canonical contract | Fresh quickstart; P0 race tests; 28-ID factory parity; conformance v1; docs/API/release truth; clean candidate consumer | +| 2. Protocol completeness | Make OpenAI Chat and Anthropic Messages trustworthy | Official-SDK fixtures pass; no flattened history or silently ignored accepted fields; replay and unknown-event semantics explicit | +| 3. Production reliability | Prove behavior under load, failure, cancellation, and restart | Admission/deadlines, budget/cache concurrency, route/health semantics, benchmarks, soak, and signed-manifest proof or experimental label | +| 4. Interoperability | Support hosted and self-hosted profiles with portable evidence | Versioned capability reports, local endpoint profiles, OTel conformance, optional backend-neutral signals, migration diagnostics | +| 5. Ecosystem and adoption | Publish one coherent library and validate external demand | Truthful `v0.1`, stability labels, automated release/security gates, extracted optional modules, at least one independent Go host | + +### First 90 days: sequence, not a staffing promise + +The repository evidence does not establish maintainer capacity, so this is an outcome-oriented 90-day sequence rather than a delivery commitment. Scope should be reduced before dates move. + +| Window | Primary outcome | Exit evidence | +|---|---|---| +| Days 0-30 | Reproduce and freeze the P0 contract: stream terminal model, option/state matrix, fresh bootstrap, factory table, migration failure cases, budget concurrency tests | Failing tests become explicit acceptance gates; no new provider or endpoint breadth enters the branch | +| Days 31-60 | Implement runtime snapshot, stream coordinator, provider blocks/usage aggregation, fallback policy/actual route, cache identity, budgets, and conformance v1 | Two-request state tests, blocked-provider cancellation, concurrent budget/cache tests, and all built-in adapter fixture results pass under race | +| Days 61-90 | Complete OpenAI translator, Anthropic replay, current OTel semantics, truthful docs/examples/OpenAPI scope, release automation, and clean consumer test | Candidate can be installed from a clean external module; API diff approved; release/security evidence generated; unresolved features explicitly experimental | + +If the P0 exit criteria are not met by day 90, the correct action is to delay `v0.1.0`, not pull P1 work into the release. A narrow pre-1.0 release with explicit limitations is more valuable than a feature-rich release whose claims are not reproducible. + +### Versioning and `v1.0` gates + +`v0.1.0` should be a **library-foundation release**: fresh onboarding, stable core generation contract, persistent runtime state, canonical stream behavior, verified provider construction, and truthful package/API scope. It should not imply a complete gateway, SDK suite, distributed control plane, or full endpoint catalog. + +`v0.2.x` should add evidence-backed extension: provider plugin descriptors, public selector/middleware seams where justified, Tier-A conformance, privacy-safe OTel, endpoint profiles, reproducible benchmarks, and optional separately versioned delivery modules. Security alerts, dependency automation, API diffs, release PRs, module-proxy checks, SBOM/provenance, and verified tags should be established by this stage. + +`v1.0` requires at least one compatibility baseline and external-consumer history; a documented Flux-owned support/security policy; reproducible release from a clean clone; no unresolved incompatibility in the supported contract; a migration guide from Eyrie and the current pre-1.0 API; provider support tiers with last-verified dates; and a credible maintenance/succession path. The current `v0.0.1` history does not establish those conditions. + +### Principal risks and guardrails + +| Risk | Early warning | Guardrail | +|---|---|---| +| Universal DTO remains inherently lossy | More provider-specific fields enter the stable DTO or raw vendor objects leak into `engine` | Canonical normalized state plus bounded, namespaced replay blocks and explicit warnings | +| Correctness fixes become an uncontrolled API break | `engine.ContractVersion` changes without migration notes | Inventory public imports, add `apidiff`, publish migration guide, keep pre-1.0 scope narrow | +| Provider-count pressure resumes | New registry IDs land before factory/conformance changes | Freeze nominal additions until every registered ID has one factory and a conformance row | +| Middleware amplifies lifecycle bugs | Public hooks appear before terminal/cancel coordinator | No public middleware until wrappers have shared conformance and ownership | +| Exact semantics are mislabeled as semantic AI | Similarity or model-name heuristics become default | Ship exact cache/routing first; require measured opt-in semantic features | +| Serving scope expands into model infrastructure | Import of vLLM/SGLang/Triton internals or model pull/load appears in `engine` | Read-only profiles only; administration in a separate module | +| Catalog publication becomes a supply-chain dependency | Mutable external artifact lacks identity, digest, or compatibility test | Generator, signed/digested release artifact, tolerant additive schema, last-good cache | +| Distribution/gateway scope expands the project | HTTP/gRPC/SDKs appear in the stable promise without independent ownership | Extract and version separately or remove from core claims | +| One-maintainer capacity invalidates roadmap dates | P0 work grows or bypasses race/conformance gates | Sequence gates, reduced scope, design partners, explicit defer decisions | +| Benchmark claims distort positioning | Vendor throughput or stars become Flux quality evidence | Publish client-only methodology and distinguish upstream claims from Flux measurements | + +### Non-goals and unresolved evidence + +The following are explicit non-goals for Flux core: UI or playground; chat UX; agent execution; tool execution; sessions and memory; RAG; workflow/durable orchestration; tenancy, organizations, virtual teams, billing, or admin control planes; model weights, pull/load/delete, compilation, batching, KV cache, or GPU scheduling; Kubernetes CRDs; a general API gateway; MCP hosting; a mandatory database/cache; unrestricted plugins or Wasm; and a feature-count race with LiteLLM. + +Several evidence gaps must remain visible. No live provider calls or independent common-harness benchmark were run, so current hosted compatibility, model availability, regional behavior, and latency remain unverified. The mutable default catalog may have changed since 2026-09-24 and must be revalidated. The official provider SDKs establish wire references, not proof that every Flux field is implemented. Production call volume, slow-consumer behavior, restart evidence, TLS/process manifest operation, partition behavior, secret-store latency, and downstream adoption are not established. Upstream release labels and component boundaries can move quickly, especially for Portkey, Bifrost, rolling llama.cpp/TensorRT tags, and newer GenAI conventions. The public API proposal, roadmap sequence, and support tiers are engineering interpretations until validated with independent Go hosts and design partners. + +## Conclusion + +The landscape does not show that Flux needs more breadth; it shows why a narrower promise is stronger. Flux can own the difficult seam between heterogeneous providers and host applications, but only if the same contract governs options, provider replay, streaming, usage, fallback, route attribution, and cancellation. The winning release is not a miniature gateway product—it is a small runtime whose claims survive fresh installation, concurrent calls, provider failure, and two-turn protocol replay. + +The next strategic move is to make the repository’s strongest abstraction real end to end. A persisted runtime snapshot, one stream state machine, one provider factory, one conformance corpus, and one truthful release contract would convert Flux from a broad collection of provider capabilities into a trustworthy interoperability layer. Everything beyond that—agents, serving, tenancy, control planes, and product UI—should remain outside the core until external demand proves that a separate module is justified. diff --git a/research_notes/Flux OSS landscape roadmap/code_architecture.md b/research_notes/Flux OSS landscape roadmap/code_architecture.md new file mode 100644 index 00000000..2b644097 --- /dev/null +++ b/research_notes/Flux OSS landscape roadmap/code_architecture.md @@ -0,0 +1,304 @@ +# Flux OSS landscape roadmap: code architecture audit + +**Audit date:** 2026-09-24 +**Checkout:** `github.com/GrayCodeAI/flux` at `f9037a8` on `chore/oss-competitive-roadmap` +**Scope:** local source, tests, CI, build tags, SDKs, examples, and architecture documentation. No competitor benchmark was performed. +**Evidence labels:** **Confirmed defect** means the behavior follows directly from current source; **Design opportunity** means the code is intentional or plausible but leaves a contract/operational gap; **Uncertainty** means the source establishes a risk but product intent needs confirmation. + +## Cited Findings + +### Takeaway + +Flux has a strong host-neutral foundation: canonical DTOs, a narrow engine facade, feature-oriented provider packages, validated deployment routing, atomic live snapshots, signed control-plane manifests, and a broad test/CI baseline. The largest risks are not missing packages; they are **contract features that are declared ahead of their implementations** and **middleware that changes stream/usage behavior after the provider boundary**. + +The highest-value work is to close the canonical request/stream contract first: restore `ThinkingEnabled`, define one usage event policy, preserve provider blocks and route attempts, and make every stream decorator preserve cancellation. Then harden migration, budget, cache, and proxy semantics. The existing green test suite does not exercise these cross-layer cases. + +### Architectural map + +| Area | Responsibility | Evidence | +|---|---|---| +| Host contract | `engine`, `llm`, `graph`, and `tools` are the accepted host-facing packages; providers own credentials, catalog, routing, transports, resilience, usage, and provider telemetry. | `README.md:46-65`; `docs/architecture/HOST-ENGINE-BOUNDARY.md:6-36`; `docs/architecture/FEATURE-MONOREPO.md:11-35` | +| Engine composition | `engine.New` accepts host-owned secret store and paths, snapshots custom gateways, and composes state, catalog, deployment transport, rate limiting, and caching. | `engine/engine.go:22-117`; `engine/engine.go:303-326` | +| Canonical DTOs | `provider/core` aliases `llm` types rather than maintaining parallel wire definitions. | `provider/core/core.go:35-97`; `llm/types.go:49-75`; `engine/types.go:155-164` | +| Provider runtime | `provider` composes adapters and feature packages; `provider/core` owns neutral stream and wire primitives. | `provider/client.go`; `provider/core/core.go:1-12`; `provider/core/stream.go:31-89` | +| Catalog and control plane | Catalog compilation, deployment policy, live replacement, and signed credential-free manifests are separated. | `catalog/v1.go:584-613`; `router/live_deployment_router.go:12-59`; `router/controlplane/snapshot.go:68-140`; `router/controlplane/peers.go:65-110` | +| Delivery | HTTP and optional gRPC are delivery adapters over the conversation/provider runtime. | `internal/api/server.go:101-121`; `internal/grpc/grpc.go:1-8`; `internal/grpc/server_grpc.go:1-78` | +| Persistence and usage | SQLite conversation DAG, budget store, provider observability, and a separate internal telemetry package coexist. | `storage/sqlite.go:39-66`; `storage/budgets.go:35-68`; `provider/observability/usage_limit.go:9-35`; `internal/observability/observability.go:1-8` | + +### Strengths confirmed + +- **Host boundary is explicit and compile-checked.** `engine/contract_assert.go` asserts the engine and stream implement the `llm` ports; `engine/types.go:142-146` re-exports the provider/stream contract. The accepted boundary document explicitly says Rho owns UX, sessions, tools, permissions, and agent loops (`docs/architecture/HOST-ENGINE-BOUNDARY.md:8-12`, `:137-154`). +- **DTO ownership is centralized.** `llm/types.go:1-10` calls itself the canonical contract, and `provider/core/core.go:35-86` aliases those types. This avoids the older provider/client DTO duplication described in the migration plan. +- **Provider layering is mechanically checked.** `scripts/check-provider-layering.sh:10-24` prevents `provider/core` and feature packages from importing siblings or the provider facade. The ecosystem script rejects Rho imports (`scripts/check-ecosystem-boundaries.sh:7-23`). +- **Deployment routing validates more than the legacy router.** `NewDeploymentRouter` rejects missing catalog, empty deployments, nil providers, and inconsistent deployment IDs (`router/deployment_router.go:74-112`). It filters by tools and circuit-breaker state (`:365-389`). +- **Live routing is atomic and last-good-safe.** `LiveDeploymentRouter.Replace` validates a complete replacement before publishing an atomic snapshot (`router/live_deployment_router.go:41-59`); control-plane replicas reject stale/conflicting revisions and retain the prior route on failure (`router/controlplane/snapshot.go:68-140`). +- **Control-plane trust is materially stronger than ordinary catalog refresh.** Remote peers require HTTPS and pinned Ed25519 keys, cap manifests at 16 MiB, reject redirects, and reject disagreement at the newest revision (`router/controlplane/peers.go:65-110`, `:113-169`). +- **Provider state persistence is defensive.** `SaveProviderConfig` rejects credential fields, writes a restrictive temporary file, fsyncs, and atomically renames (`config/provider_env.go:490-548`); credential migration support has a broad table of recognized fields (`config/provider_secrets.go:22-146`). +- **Transport streams are genuinely streaming-first.** The SSE parser uses a bounded scanner and cancellation-aware forwarding (`provider/core/stream.go:31-89`); Anthropic and OpenAI processors emit content, reasoning, tool, usage, TTFT, and terminal events rather than buffering complete responses (`:110-291`, `:396-588`). +- **Testing and CI are broad and currently green.** See the verification record under Gaps. The repository has extensive adapter, routing, storage, API, and contract tests; `go test -race`, vet, lint, formatting, vulnerability scanning, and boundary checks all pass. + +### Confirmed defects and contract gaps + +#### F1. Host `ThinkingEnabled` is silently dropped + +**Classification:** Confirmed defect, P1. + +`engine.toClientOptions` maps `GLMThinkingEnabled` but never maps `advanced.ThinkingEnabled` (`engine/convert.go:15-34`). The field is part of both the host request contract and client options (`llm/provider.go:104-113`; `llm/types.go:135-155`). Provider adapters do consume it: Anthropic gives it precedence in `resolveThinking` (`provider/adapters/anthropic.go:177-181`), and OpenAI-compatible translation uses it as the canonical toggle (`provider/adapters/openai.go:409-411`). + +The result is host/provider behavior divergence: a host can request reasoning enabled through the new field, the engine drops it, and the adapter falls back to deprecated or budget-based behavior. Existing `engine/convert_test.go` cases do not assert `ThinkingEnabled`. + +**Remediation:** map `opts.ThinkingEnabled = advanced.ThinkingEnabled`, retain the deprecated alias only as an explicit fallback policy, and add adapter-to-engine contract tests for enabled, disabled, and nil cases. + +#### F2. The stream contract declares provider metadata that the engine drops + +**Classification:** Confirmed defect, P1. + +`llm.FluxStreamEvent` carries `ProviderBlock` and `Route` (`llm/types.go:265-287`), and the stable event vocabulary includes `provider_block` (`engine/types.go:65-84`). However, `engine.normalizeEvent` copies content, thinking, request ID, usage, stop reason, and TTFT, then maps only selected event types; it never copies `event.ProviderBlock` or `event.Route` (`engine/stream.go:124-165`). The initial synthetic `route_selected` event is present (`engine/stream.go:82-86`), but provider event route metadata is lost. + +`StreamErrorInfo` is declared (`llm/types.go:90-98`) but is not a field on `FluxStreamEvent`; parser errors are emitted as only `Type:"error", Error:string` (`provider/core/stream.go:283-285`, `:485-488`). `FluxResponse.Warnings` is also declared (`llm/types.go:248-263`) but no provider adapter populates it. + +**Remediation:** either remove the unimplemented fields from the stable contract or make them mandatory end-to-end: parse and preserve provider blocks, attach route/attempt metadata to terminal events, and add a structured error field or a documented error-classification path. Add black-box tests that feed a provider block through adapter → core → engine. + +#### F3. Provider-native reasoning blocks cannot round-trip + +**Classification:** Confirmed defect, P1. + +The canonical message/response contract explicitly promises opaque provider state for Anthropic thinking signatures, redacted thinking, Gemini thought state, and OpenAI reasoning items (`llm/types.go:58-75`, `:248-260`). The Anthropic wire response type captures `thinking` and `signature`, but the parser concatenates only thinking text and silently skips redacted blocks (`provider/adapters/anthropic.go:233-300`). `buildAnthropicMessages` handles tool use, tool results, content parts, and images, but has no branch that emits `m.Thinking` or `m.ProviderBlocks` (`:303-429`). Repository-wide adapter search found no `ProviderBlock` or `ProviderBlocks` construction. + +This means a host cannot safely continue a signed-thinking conversation even though the DTO comments instruct it to store and replay those blocks. The API currently loses protocol state rather than exposing a clear unsupported warning. + +**Remediation:** define the provider-block wire envelope per adapter, preserve the raw block exactly, filter by protocol on replay, and test Anthropic signed/redacted thinking and Gemini/OpenAI reasoning metadata across two turns. + +#### F4. Real streaming usage is ignored by conversation persistence and the OpenAI proxy + +**Classification:** Confirmed defect, P1. + +Anthropic emits input usage on `message_start` and output usage on `message_delta` as separate `usage` events (`provider/core/stream.go:236-281`). OpenAI emits a `usage` event before the terminal event when `stream_options.include_usage` is present (`provider/core/stream.go:502-512`). Gemini attaches usage to its terminal `done` event (`provider/adapters/gemini.go:625-644`). + +`conversation.streamAndSave` records usage only in the `case "done"` branch (`conversation/engine.go:272-290`). It therefore persists zero tokens for Anthropic and OpenAI streaming calls. The OpenAI proxy reads usage back from the persisted node (`internal/api/openai_proxy.go:267-283`), so its compatibility response can also report zero usage. + +The current mocks hide the defect: `conversation/engine_test.go:23-28` and `internal/api/server_test.go:26-31` put usage directly on `done`, unlike the real Anthropic/OpenAI processors. The Anthropic adapter test also asserts content only and does not assert usage (`provider/adapters/anthropic_test.go:98-125`). + +**Remediation:** define one canonical usage event contract (prefer a final aggregate event), have every adapter conform, aggregate split usage events in the conversation layer, and add per-adapter tests that assert persisted tokens and proxy usage. Do not infer “usage on done” from mocks. + +#### F5. Stream decorators drop the upstream cancellation function + +**Classification:** Confirmed defect, P1. + +`llm.StreamResult.Close` is the contract cleanup mechanism (`llm/types.go:289-307`). `TracingProvider` wraps the event channel in a goroutine but returns a struct literal with no cancellation callback (`provider/observability/tracing.go:91-135`). `GuardrailProvider` does the same (`provider/resilience/guardrails.go:83-118`). A caller that closes the returned result cannot cancel the underlying provider stream; the forwarding goroutine can remain blocked until the provider completes. + +The same code paths do forward on context cancellation (`tracing.go:120-128`, `guardrails.go:106-111`), but the public `Close` path cannot trigger that context. `UsageLimitProvider` and `BudgetProvider` demonstrate the safer pattern by returning `NewStreamResult(wrappedCh, result.Close)` (`provider/observability/usage_limit.go:83-104`; `provider/observability/budget_provider.go:112-133`). + +**Remediation:** return a constructor that preserves the original `Close` callback, and add a test that closes before the producer finishes and asserts the provider context/body is canceled. + +#### F6. Streaming guardrails do not provide the cross-chunk behavior already implemented elsewhere + +**Classification:** Confirmed defect/contract gap, P1/P2. + +The streaming decorator checks each content chunk and applies only redactions (`provider/resilience/guardrails.go:70-105`). A block rule does return an error from `Guardrails.Check` and is converted to a terminal error (`provider/core/guardrails.go:152-201`), so blocking itself is not wholly absent. However: + +- PII patterns split across chunks are not accumulated. +- `Warn` violations are not logged or emitted despite the documented behavior (`guardrails.go:14-18`). +- The richer `core.StreamGuardrails` implementation explicitly supports buffering, injection blocking, and final PII checks (`provider/core/stream_guardrails.go:62-166`) but is not wired into the provider decorator. + +**Remediation:** make `GuardrailProvider.StreamChat` use `StreamGuardrails`, define whether a final retrospective block is possible after bytes have already been delivered, and add chunk-boundary tests for PII, secrets, and warning events. + +#### F7. The response cache is exact, incomplete-key, and not safely deep-copied + +**Classification:** Confirmed defect, P1/P2. + +The engine describes caching as semantic and deterministic by default (`engine/engine.go:45-51`; `README.md:136-140`), but the implementation is a deterministic hash/LRU cache (`provider/cache/semantic_cache.go:52-75`, `:105-138`). The key includes model, system, temperature, role/content, tool calls, and tool results (`:296-339`), but omits response-affecting fields such as `ContentParts`, images, thinking, provider blocks, `MaxTokens`, `TopP`, `TopK`, stop sequences, tool choice, response format, reasoning controls, provider identity, and metadata. Two materially different requests can therefore share a cached response. + +There are two additional correctness/configuration gaps: + +- `CacheConfig.Enabled` is documented as defaulting to true (`:15-28`, `:31-38`), but `NewCachedProvider` copies the zero value `false` and does not apply that default (`:73-92`). `Engine.Options.EnableCaching` with a zero `CacheConfig` therefore silently disables the cache. +- `core.CopyResponse` deep-copies usage and tool arguments but leaves response slices such as `ProviderBlocks` and `Warnings` aliased (`provider/core/copy.go:3-23`). A caller can mutate cached metadata. + +A separate unused `internal/cache` package implements another in-memory/Redis backend and a hard-coded Anthropic cache warmer (`internal/cache/backend.go:1-9`, `:97-121`; `internal/cache/cache_warmer.go:10-31`), increasing architectural ambiguity. + +**Remediation:** rename the feature to exact response cache unless true similarity is implemented; version and hash all response-affecting canonical fields; normalize zero config through `DefaultCacheConfig`; deep-copy every response slice/raw block; and either integrate or remove the internal cache subsystem. + +#### F8. Deployment response provenance is not populated + +**Classification:** Confirmed defect, P1/P2. + +`ResolvedRoute` documents `DeploymentID` as the backend that actually served the request and `Attempts` as the number of attempts (`llm/types.go:235-246`). `DeploymentRouter` tracks the selected deployment internally and returns the adapter response unchanged (`router/deployment_router.go:134-180`, `:432-439`); it does not assign `resp.Route`. The engine then overwrites any response route with the route selected before the call (`engine/convert.go:88-96`). A repository-wide search found no production assignment to `DeploymentID` or `Attempts` outside tests. + +The result is correct route selection but incomplete operational attribution, especially for failover, usage pricing, and the operations graph projection. + +**Remediation:** attach a resolved route at the deployment boundary after success, including actual deployment and attempt count; preserve it through engine conversion and stream terminal events; add a failover test asserting provider `A` attempt 1 and provider `B` attempt 2 are reported distinctly. + +#### F9. Budget enforcement is not atomic across check and record + +**Classification:** Confirmed defect, P1. + +`BudgetProvider.Chat` and `StreamChat` call `CheckBudget`, invoke the provider, and later call `RecordUsage` (`provider/observability/budget_provider.go:73-91`, `:94-133`). The SQLite store’s `CheckBudget` is a read and `RecordUsage` is a later transaction (`storage/budgets.go:157-208`). Concurrent requests can all pass the same pre-charge check and then collectively exceed the limit. The single connection serializes individual statements, not the business-level check/use/record sequence. + +The streaming wrapper also records each usage event independently (`:119-123`), while Anthropic supplies split input/output events; the accounting policy should explicitly aggregate or debit deltas. + +**Remediation:** reserve estimated cost atomically before the provider call, reconcile with actual usage in a transaction, and define behavior for over-reservation, cancellation, mid-stream failure, and split usage events. + +#### F10. Credential migration can delete plaintext that was not successfully migrated + +**Classification:** Confirmed security defect, P1. + +`migrateEnvFileAt` reads all key/value pairs, attempts only recognized discovery keys, and removes the file when `len(secrets) > 0` even if `migrated == 0` (`credentials/migrate.go:50-85`). Thus a file containing only unknown keys, or recognized keys whose keyring writes all fail, is deleted without a durable replacement. `MigrateEnvFileCredentials` then writes a completion marker (`:31-47`). + +The existing migration tests cover empty files and successful/key-existing cases (`credentials/migrate_test.go:142-270`) but do not cover partial or total write failure. + +**Remediation:** remove a plaintext file only after every recognized secret is confirmed stored (or explicitly acknowledged as unmapped), preserve the original file on failure, and make the migration marker conditional on successful completion. + +#### F11. Conversation SQLite sidecars are not permission-hardened + +**Classification:** Confirmed security defect, P1/P2. + +`storage.Open` chmods only the main database path after opening WAL mode (`storage/sqlite.go:39-66`). SQLite creates `-wal` and `-shm` sidecars, and those files can contain conversation pages. The budget store demonstrates the missing hardening explicitly by chmodding all three paths (`storage/budgets.go:58-67`). The security tests cover SQL escaping but not filesystem sidecar modes (`storage/security_test.go:9-220`). + +**Remediation:** create/verify the parent directory mode, chmod the main file and existing sidecars, and add platform-aware tests that inspect sidecar permissions after writes. + +#### F12. The OpenAI-compatible proxy is a convenience adapter, not semantic compatibility + +**Classification:** Confirmed compatibility defect, P1/P2. + +The endpoint claims that existing OpenAI clients can talk to Flux unchanged (`internal/api/openai_proxy.go:16-20`). Its request message type contains only `Role` and string `Content` (`:22-26`), and `splitOpenAIMessages` flattens the entire conversation into one prompt string (`:300-331`). Tool-call messages, tool results, multimodal content, tool choice, response format, `stop`, sampling controls, and several other accepted fields are either ignored or unavailable (`internal/api/openai_proxy.go:28-44`; `handleOpenAIChatCompletions` explicitly documents lenient partial decoding at `:109-140`). + +This is acceptable only if the endpoint is clearly documented as a lossy single-turn compatibility facade. It is not a faithful OpenAI conversation adapter. + +**Remediation:** either implement a typed translation into the engine request/message contract, including tool lifecycle and structured output, or narrow the endpoint contract and add explicit capability warnings/headers for unsupported fields. + +#### F13. The deployment catalog has two protocol sources of truth + +**Classification:** Confirmed design drift, P1/P2. + +The registry declares Anthropic and Gemini as native protocols (`catalog/registry/providers.go:20-29`, `:42-50`, `:183-194`), and its protocol-matrix tests explicitly require `anthropic-messages` for Anthropic/Bedrock (`catalog/registry/protocol_matrix_test.go:76-87`). The v1 bootstrap catalog, however, assigns `openai-chat-completions` to `anthropic-direct` and `anthropic-bedrock` (`catalog/v1_defaults.go:43-49`) and only defines OpenAI protocols (`:35-40`). `EnsureCredentialRegistryInCatalog` fills missing deployments but does not correct existing ones (`catalog/credential_registry.go:20-35`). + +This is a data-model consistency defect even if current adapter selection works: route metadata, credential sync, protocol diagnostics, and future generated catalogs can disagree. + +**Remediation:** make one registry authoritative, generate the other catalog table from it, and add a test that every deployment protocol equals the registry protocol and is present in the catalog protocol map. + +### Other confirmed lower-priority defects + +- **Adaptive rate-limit zero mode is contradictory.** The configuration says `MaxDelay == 0` returns an error instead of delaying (`provider/resilience/adaptive_ratelimit.go:153-164`), but the constructor replaces every non-positive value with 10 seconds (`:214-226`). The test labels zero as “don't delay” but does not assert the outcome (`provider/resilience/adaptive_ratelimit_test.go:149-177`). +- **Usage tracker zero limits block all calls.** `CanProceed` uses `>=` without a disabled sentinel (`:90-115`), so a zero hourly/daily/session/cost limit is immediately exhausted. `UsageLimitProvider` records cost as zero on both blocking and streaming paths (`:56-67`, `:83-116`), so its cost limit never advances. Either document `<=0` as unlimited or normalize it to a safe default. +- **Batch retry reuses a consumed request body.** `Submit` creates one request and repeatedly calls `Do` without rewinding or invoking `GetBody` (`provider/batch/batch.go:97-121`). The adapter request path explicitly sets `GetBody` for retries (`provider/adapters/anthropic.go:549-557`), but batch submission does not. `WaitUntilDone` also leaves 429/5xx response bodies open before sleeping (`provider/batch/batch_async.go:103-131`), and `Poll` decodes non-200 bodies without checking status (`provider/batch/batch.go:142-166`). +- **Structured output option is a no-op.** `provider/media.WithStructuredOutput` returns an empty `core.ClientOption` and ignores both arguments (`provider/media/structured.go:195-198`). The separate validator supports only a small subset of JSON Schema (`:26-145`), and the array implementation reads a `minimum` keyword where JSON Schema normally uses `minItems` (`:120-145`). +- **Same-role merging is lossy for modern message parts.** `MergeConsecutiveRoles` merges only text and legacy images (`:28-38`) even though canonical messages also carry `Thinking`, `ContentParts`, and `ProviderBlocks` (`llm/types.go:49-64`). It is currently mainly exercised through compatibility tests/aliases, so the immediate blast radius is smaller than the contract suggests, but the exported helper is not safe for multimodal/provider-state history. +- **Graph validation is intentionally minimum-only, not structural.** `GraphSpec.Validate` checks IDs/types and edge fields but not edge endpoint existence, duplicate node IDs, or cycles (`graph/graph.go:254-273`). This is acceptable for a data vocabulary, but should be documented as such or paired with a structural validator before runtime consumers assume more. +- **Small contract inconsistencies remain.** `types.ToAgentId` returns `(nil, nil)` for an invalid ID (`types/types.go:19-25`); `BehaviorPresetFrom` accepts the zero value despite its comment saying empty is invalid (`tools/versioning.go:12-31`); `runtime.ModelIDs` promises sorted output but never sorts (`runtime/runtime.go:164-174`); `operationsgraph/operations_graph.go:1` has the wrong package comment name. +- **Legacy router construction is under-validated.** `router.New` dereferences `e.Provider.Name()` while collecting stats and accepts nil providers, zero/negative weights, and fallback-only configurations (`router/router.go:55-85`). The newer `DeploymentRouter` validates much more carefully (`router/deployment_router.go:74-112`). + +### Design drift and simplification opportunities + +- **Old multi-package public surface remains.** `runtime/runtime.go:1-22` calls itself the recommended host entry point and lists `provider`, `catalog`, `config`, `credentials`, `setup`, and `storage` as public API, while the accepted boundary says hosts should use `engine` (`README.md:46-65`; `docs/architecture/HOST-ENGINE-BOUNDARY.md:33-36`). Make `runtime` a deprecated compatibility package or update its documentation and ownership. +- **Provider registry composition is coupled.** `provider/adapters/provider_registry.go:5-10` imports catalog, config, and credentials to derive maps and detect providers. The target architecture says adapters translate protocols and do not own routing/credentials (`docs/architecture/FEATURE-MONOREPO.md:48-56`). Move registry composition into `provider/` or `setup`, and make the layering guard cover the full direction rule rather than only provider subpackage imports. +- **Retry models are duplicated.** `types.RetryConfig` (`types/retry.go:5-12`), `provider/core.RetryConfig` (`provider/core/retry.go:15-21`), and `router.RetryConfig` (`router/retry.go:10-16`) represent overlapping policy. Keep one canonical policy plus transport/router-specific behavior, or add compile-time conversion tests and a deprecation plan. +- **Observability has two stacks.** OTel wrappers use `provider/observability/tracing.go:12-136`; the separate zero-dependency package describes its own OTel-compatible model (`internal/observability/observability.go:1-8`, `:23-42`) but has no production import hits. Choose OTel as the execution path and retain only a deliberately isolated test/export implementation. +- **Configuration duplicates provider metadata.** `config/profiles.go:39-64` and `catalog/registry/providers.go:16-29` both define provider modes, endpoints, credentials, and model behavior. `config/provider_env.go:16` also describes the config as mirroring Rho. Generate config profiles from the registry or explicitly scope config as legacy host state. +- **Rho-specific paths remain in supposedly host-neutral code.** `credentials/migrate.go:10-30` hard-codes `~/.rho` and `~/.hawk`; `runtime` documents Rho-specific assembly (`runtime/runtime.go:1-22`). This is acceptable only as an explicit legacy migration package, not as the default engine surface. +- **The optional gRPC surface is intentionally skeletal.** The untagged contract and default service return `ErrUnimplemented` (`internal/grpc/grpc.go:42-67`); the tagged server uses a hand-written JSON codec and a single unary method (`internal/grpc/server_grpc.go:15-78`). The tagged test proves the loopback round trip (`internal/grpc/server_grpc_test.go:21-52`), but there are no generated protobuf clients or SDK gRPC clients. Keep it clearly experimental or make the compatibility contract explicit. +- **The internal cache warmer contains hard-coded pricing and Anthropic assumptions.** `internal/cache/cache_warmer.go:10-31` embeds a model-specific price and cache multipliers. If retained, source pricing/capability data from the catalog and make the warmer provider/profile-specific. + +### Cross-layer test gaps + +The passing suite is broad, but the most important cross-layer cases are absent: + +- No engine conversion test asserts `GenerationOptions.ThinkingEnabled` reaches `ChatOptions.ThinkingEnabled`. +- No adapter-to-engine test carries `ProviderBlock`, `ProviderMetadata`, or structured stream errors. +- Anthropic/OpenAI stream tests do not assert the split `usage` events are persisted by `conversation.Engine`; mocks currently put usage on `done`. +- No test closes a tracing/guardrail-wrapped stream early and verifies upstream cancellation. +- Cache tests do not vary all response-affecting fields or mutate returned provider blocks/warnings. +- Migration tests do not simulate all keyring writes failing or leave unknown secrets in the plaintext file. +- Budget tests are sequential and do not exercise concurrent check/use/record interleavings. +- Storage security tests do not inspect `-wal`/`-shm` permissions. +- Batch tests do not verify body replay after a transient submit response, non-2xx polling, or response-body closure on retry. +- No failover test asserts actual `DeploymentID` and `Attempts` in returned route metadata. +- No black-box conformance matrix runs the same request/stream contract against every adapter. + +## Inferences + +### Priority roadmap + +1. **Close the canonical generation contract before adding more providers.** + - Restore `ThinkingEnabled` mapping. + - Standardize aggregate usage events and update conversation/API tests. + - Decide whether provider blocks, warnings, and structured stream errors are supported or remove them from the stable DTO. + - Preserve actual deployment route/attempt provenance. + - Add an adapter conformance harness that treats provider-native metadata as a required contract, not a per-adapter extra. + +2. **Fix stream lifecycle and safety wrappers.** + - Preserve `Close` in every wrapper. + - Wire `core.StreamGuardrails` or remove the duplicate implementation. + - Define cancellation behavior for provider errors, pre-first-byte errors, and mid-stream failures. + - Add fault-injection tests for abandoned streams and slow consumers. + +3. **Make budgets and credential migration fail closed.** + - Replace check-then-record with reservation/reconciliation. + - Preserve plaintext files until all recognized secrets are durably stored. + - Define cost/token aggregation for split stream events. + - Add race tests and migration failure injection. + +4. **Consolidate source-of-truth surfaces.** + - Make the provider registry generate catalog/config profiles. + - Mark `runtime` and legacy `router.Router` as compatibility APIs or remove them from host guidance. + - Choose one observability implementation. + - Rename the exact response cache and remove/ integrate the unused internal cache stack. + +5. **Make compatibility boundaries explicit.** + - Either implement full OpenAI request/message translation or label the proxy lossy and expose unsupported-field diagnostics. + - Keep gRPC experimental until protobuf/client parity exists. + - Add source-level tests for host import boundaries, not just provider-internal layering. + +6. **Harden operational state.** + - Secure SQLite sidecars and test actual filesystem modes. + - Use per-cache unique temporary files and verify atomic rename under concurrent refresh. + - Enforce HTTPS/host/signature policy for remote catalog refreshes, or document explicit operator trust. + - Avoid mutating compiled catalog input/owned slices. + +### Recommended acceptance criteria + +A future “contract v2 hardened” milestone should require: + +- every documented `GenerationOptions` field has a conversion test; +- every documented stream event has a producer and host-preservation test; +- usage, route, request ID, and finish reason survive adapter, router, conversation, and proxy boundaries; +- stream `Close` cancels every wrapper and provider body; +- budget tests demonstrate no overspend under concurrent requests; +- migration tests demonstrate no plaintext loss on keyring failure; +- failover responses identify the actual deployment and attempt count; +- every registered provider has a catalog/protocol/config fixture generated from one registry; +- all adapters pass the same local `httptest` streaming/error/usage contract suite. + +## Gaps + +### Verification record + +All commands below passed on the audited checkout: + +- `go test ./...` +- `go test -race ./...` +- `go test ./... -count=1 -timeout=120s` +- `go test -race ./... -count=1 -timeout=180s` +- `go vet ./...` +- `go test -tags grpc ./...` +- `golangci-lint run ./... --timeout=5m` — 0 issues +- `gofumpt -l .` — no output +- `goimports -l .` — no output +- `govulncheck ./...` — no vulnerabilities found +- `scripts/check-provider-layering.sh` — passed +- `scripts/check-ecosystem-boundaries.sh` — passed +- `scripts/check-no-replace-directives.sh` — passed + +`make ci` was not run because its `tidy` and `fmt` phases mutate files; the read-only equivalents above were run instead. No source files were changed during this audit. + +### Remaining uncertainties + +- The default catalog’s OpenAI protocol for Anthropic deployments may be a deliberate legacy normalization rather than an accidental mismatch; the registry matrix and current comments disagree, so the owning team must choose the canonical meaning. +- “Semantic caching” may be a product roadmap label for a future similarity implementation; current code is unequivocally exact hashing. The public name and behavior should be reconciled. +- `runtime` may need to remain temporarily for backward compatibility. The boundary is clear, but the deprecation/migration path is not. +- Deployment route fields may be intended to be populated by a host-facing wrapper rather than the lower-level router. Current source has no such wrapper in this repository. +- SQLite sidecar exposure depends on process umask and SQLite creation timing; the source still lacks an explicit guarantee and regression test. +- No live provider credentials, production workloads, or multi-process control-plane deployment were exercised. The signed-manifest and budget findings are static contract findings, not production traffic measurements. + +### Audit boundary + +This report is a local architectural audit, not a competitor benchmark, security certification, provider conformance certification, or production readiness sign-off. External landscape research is kept in the companion files under `research_notes/Flux OSS landscape roadmap/`; this document records only repository architecture, implementation behavior, and test evidence. diff --git a/research_notes/Flux OSS landscape roadmap/gap_mapping.md b/research_notes/Flux OSS landscape roadmap/gap_mapping.md new file mode 100644 index 00000000..9cf1623c --- /dev/null +++ b/research_notes/Flux OSS landscape roadmap/gap_mapping.md @@ -0,0 +1,377 @@ +# Flux OSS Competitive Gap Map and Decision-Ready Roadmap + +- **Research cutoff:** 2026-09-24 +- **Flux snapshot reviewed:** commit `f9037a82e403001d50868e4100bcfa19d8a387af`, local branch `chore/oss-competitive-roadmap` +- **Method:** synthesized all eight notes in this directory; current Flux source was consulted read-only to validate high-impact claims and package touchpoints. No live provider calls, benchmarks, or new upstream research were performed. +- **Decision rule:** optimize for a small, trustworthy, host-neutral Go provider runtime. Do not turn Flux into a UI, agent framework, tenancy/control plane, gateway product, or model server. + +## 1. Decision frame, canonical landscape, and baseline + +### Executive decision + +1. **Make correctness and trust the next release, not breadth.** Flux already has substantial leaf functionality, but the host path currently drops declared state, reconstructs stateful middleware per call, requires a Rho-prepared catalog, and markets compatibility and caching more strongly than the implementation supports. These are release blockers, not polish. +2. **Build one versioned provider-call contract around streams, usage, route attempts, and provider replay state.** The same contract must survive adapter → router → engine/conversation → transport wrapper → host. +3. **Turn provider breadth into evidence-backed conformance.** A provider count is not a capability. Each provider/deployment/profile needs versioned support, unknown, limitations, and last-verified evidence. +4. **Differentiate on trustworthy interoperability, not feature-count parity with LiteLLM.** Flux's smallest thesis is: *the embeddable Go runtime that makes heterogeneous hosted and self-hosted LLM endpoints behave like one explicit, lossless, conformance-tested contract—without requiring a gateway service, database, control plane, agent loop, or model server.* +5. **Keep model execution and product semantics external.** Model weights, batching, KV cache, GPU scheduling, autoscaling, UI, sessions, tools, RAG, agent execution, and tenancy remain external. Any future model-management client must live in a separate admin-only module, never the stable chat/provider contract. + +### Canonical top-20 OSS portfolio + +This is a **Flux decision-value ranking**, not a project quality, popularity, or performance ranking. It supersedes the three scoped top-20 lists where they conflict. Serving rows are interoperability targets and conformance donors; **do not build their execution/control-plane layers in Flux core**. + +| Rank | Project | Role | Why it belongs | Flux decision | +|---:|---|---|---|---| +| 1 | [Bifrost](https://github.com/maximhq/bifrost) | Closest Go runtime peer | Closest current combination of Go SDK/gateway modes, provider abstraction, fallback, load balancing, semantic caching, telemetry, operation matrices, raw-response inspection, and narrow middleware. | Borrow its transformation isolation, capability matrix, raw-fidelity test hook, and Go seam; do not copy UI, MCP hosting, or clustering. | +| 2 | [LiteLLM](https://github.com/BerriAI/litellm) | Feature and policy benchmark | Broadest verified public reference for provider translation, retry versus fallback, cache-key discipline, per-attempt authorization/budget revalidation, spend, and privacy-safe GenAI OTel. | Use as a semantic benchmark; do not adopt its Python/Rust gateway, UI, tenancy, PostgreSQL, or Redis-centered product shape. | +| 3 | [AxonHub](https://github.com/looplj/axonhub) | Go transformer-pipeline peer | Best direct reference for inbound dialect → unified request → outbound provider → reverse stream/error transformation in Go. | Borrow transformer layering and testability; not a dependency. | +| 4 | [GoModel](https://github.com/ENTERPILOT/GoModel) | Compact Go peer | Useful for aliases, passthrough, scoped policy composition, caching/budgets, usage, and an embeddable operational DX smaller than LiteLLM. | Borrow plugin/policy composition and migration ergonomics. | +| 5 | [Agent Router](https://github.com/theagentrouter/agent-router) | Reliability and extension donor | Its v1.x stream idle timeout, per-request credentials, provider-extension fields, compatibility-preserving rename, and staged cross-provider translation address real Flux gaps. | Borrow deadline/recovery and extension semantics; reject its Kubernetes control plane and MCP scope for core. | +| 6 | [Portkey Gateway](https://github.com/Portkey-AI/gateway) | Declarative routing benchmark | Strong reference for request-attached JSON policy, conditional routing, sticky affinity, timeouts, and fallback, but its current 1.x/2.0 OSS boundary and maintenance require caution. | Borrow policy shape and affinity; validate source before using any implementation detail. | +| 7 | [OmniRoute](https://github.com/diegosouzapw/OmniRoute) | High-velocity gateway donor | Current source and docs expose local/cloud mixing, aliases, retries/fallbacks, circuit breakers, caching, tools, streaming, and native passthrough. | Borrow migration and passthrough semantics; avoid its broader CLI/dashboard/product surface. | +| 8 | [Manifest (`mnfst/llm-gateway`)](https://github.com/mnfst/llm-gateway) | Routing-focused peer | Distinct donor for prompt-complexity scoring, model tiers, local/cloud routing, sticky/session behavior, and cost-aware fallback. | Borrow route-hint and affinity concepts only after deterministic cost/latency baselines. | +| 9 | [LocalAI](https://github.com/mudler/LocalAI) | Local runtime profile | Best Go donor for multi-backend discovery, machine-readable capability metadata, model configuration, and the limits of a single “OpenAI-compatible” claim. | Compatibility/profile target; do not absorb model lifecycle or backend control into core. | +| 10 | [Ollama](https://github.com/ollama/ollama) | Local lifecycle profile | Rich native tags/show/running/version surface and native structured-output behavior are currently hidden behind Flux's thin OpenAI wrapper. | Add native metadata/profile support; keep pull/load/delete and weight lifecycle in an optional admin package. | +| 11 | [llama.cpp / llama-server](https://github.com/ggml-org/llama.cpp) | Portable runtime profile | Exposes richer runtime evidence than `/v1/models`: `/props`, load/unload, slots, metrics, version, multimodal, tools, JSON schema, and OpenAI/Anthropic routes. | Consume bounded metadata and conformance-test the endpoint; never embed kernels or scheduler logic. | +| 12 | [vLLM](https://github.com/vllm-project/vllm) | High-throughput conformance target | Canonical source for feature-by-hardware capability evidence, OpenAI/Anthropic compatibility surfaces, async streaming, and version-sensitive behavior. | Profile and conformance-test; Flux does not own batching, KV memory, parallelism, or execution. | +| 13 | [SGLang](https://github.com/sgl-project/sglang) | Serving/gateway boundary donor | Contributes model-gateway lifecycle, heterogeneous protocol routing, cache/load signals, health, and the distinction between request/control/state planes. | Borrow signals and plane separation; exclude SGLang-specific history, MCP execution, workers, and infrastructure dependencies. | +| 14 | [any-llm-go](https://github.com/mozilla-ai/any-llm-go) | Go optional-interface donor | Small provider port plus optional embedding, listing, capability, and error-conversion interfaces is the right way to avoid widening the four-method provider contract. Evidence is newer than Flux's mature implementation, so parity claims are low-confidence. | Borrow interface shape and explicit unsupported states; do not adopt as a dependency before validating license, maintenance, and wire parity. | +| 15 | [Vercel AI SDK](https://github.com/vercel/ai) | Model-port and middleware donor | Versioned language-model specification, explicit provider metadata, deterministic middleware order, and separate generate/stream wrapping closely match Flux's needed seam. | Borrow versioned port/middleware semantics; reject UI, RSC, harnesses, and agent-loop abstractions. | +| 16 | [OpenAI Python SDK](https://github.com/openai/openai-python) | Canonical OpenAI protocol donor | Primary reference for Chat Completions versus Responses, typed SSE, response iteration, tool events, and structured request evolution. | Use pinned wire fixtures and generated/checked expectations; do not expose SDK types in `engine`. | +| 17 | [Anthropic Python SDK](https://github.com/anthropics/anthropic-sdk-python) | Canonical Messages/replay donor | Primary reference for typed stream events, accumulation, partial tool JSON, signed/redacted thinking state, and multi-turn replay. | Use to close provider-block and stream lifecycle correctness; no dependency needed. | +| 18 | [Google GenAI Python](https://github.com/googleapis/python-genai) | Gemini/Vertex protocol donor | Distinct source for typed multimodal parts, tool calling, streaming, safety settings, and Gemini/Vertex request modeling. | Use for conformance fixtures; preserve provider-native Gemini state in adapters. | +| 19 | [OpenTelemetry GenAI conventions](https://github.com/open-telemetry/semantic-conventions-genai) | Interoperability standard | Canonical current vocabulary for provider/model identity, operation, token usage, finish reasons, streaming, privacy, and extensions. Flux currently emits older/custom attributes and does not wire them coherently through `engine`. | Adopt the standard through the OTel API, with content capture off by default and compatibility aliases only during migration. | +| 20 | [Gateway API Inference Extension](https://github.com/kubernetes-sigs/gateway-api-inference-extension) | Deployment interoperability seam | Supplies backend-neutral public-model, pool/deployment, endpoint-picker, fallback, rollout, and capability concepts without requiring Flux to implement Kubernetes CRDs. | Borrow identity/signals into optional deployment metadata; do not add Kubernetes types to the stable engine DTO. | + +The selection is based primarily on [`selection_methodology.md`](selection_methodology.md), then reconciled with the direct-peer evidence in [`oss_gateways.md`](oss_gateways.md), the downstream runtime boundary in [`model_runtimes.md`](model_runtimes.md), and the contract donors in [`sdk_frameworks.md`](sdk_frameworks.md). + +### How the inconsistent project lists were reconciled + +| Scoped list | Conflicting shape | Canonical resolution | +|---|---|---| +| [`selection_methodology.md`](selection_methodology.md) | 8 direct peers + 6 runtimes + 4 SDKs + 2 adjacent donors | Retain the eight direct peers. Retain five highest-value serving targets. Replace the broad SDK/adjacent tail with the five protocol donors plus OTel GenAI and GAIE. | +| [`model_runtimes.md`](model_runtimes.md) | 11 engines/runtimes plus packaging, deployment, and control-plane projects | Treat engines as external profiles and control planes as signal donors. Do not turn model-server selection into a top-20 Flux product category. | +| [`sdk_frameworks.md`](sdk_frameworks.md) | Runtime contracts mixed with LangGraph, Pydantic AI, Genkit, ADK, Dify, Open WebUI, and other agent/UI products | Promote only provider/model-port donors. Use agent frameworks as negative boundary evidence, not dependencies or Flux feature targets. | +| [`oss_gateways.md`](oss_gateways.md) | Direct peers plus general gateways, console products, and infrastructure control planes | Retain Bifrost, LiteLLM, Agent Router, Portkey, OmniRoute, and the Go transformer peers. Demote general gateways and console products to subsystem-specific donors. | +| [`provider_coverage.md`](provider_coverage.md) | 28 registry IDs but multiple construction paths and four protocol families | Count registry entries as configured gateways, not independent conformance levels. Unify construction before adding more IDs. | + +Important projects intentionally outside the canonical 20 remain useful **secondary** donors: + +- [Helicone AI Gateway](https://github.com/Helicone/ai-gateway) for P2C + PeakEWMA, but its gateway activity and license are contradictory; use the algorithm as a reference, not code or a dependency. +- [New API](https://github.com/QuantumNous/new-api) for its independently testable Go `relaykit` translation boundary, but not its console, tenancy, billing, or JavaScript task runtime. +- [Apache APISIX](https://github.com/apache/apisix) and [Higress](https://github.com/higress-group/higress) for lifecycle ordering and incremental streaming transforms, not for etcd/Istio/plugin sprawl. +- [LLM Gateway (`theopenco/llmgateway`)](https://github.com/theopenco/llmgateway) for coding-agent migration/doctor UX, not its database-backed product shell. +- [Langfuse](https://github.com/langfuse/langfuse) for observability/evaluation product requirements, not an embedded platform. +- [llm-d](https://github.com/llm-d/llm-d), [Dynamo](https://github.com/ai-dynamo/dynamo), and [KServe](https://github.com/kserve/kserve) for the explicit boundary between request routing and model-serving control planes. +- [LangGraph](https://github.com/langchain-ai/langgraph), [Pydantic AI](https://github.com/pydantic/pydantic-ai), [Google Genkit](https://github.com/genkit-ai/genkit), and similar projects for typed DTOs and lifecycle vocabulary only. Their runners, sessions, workflows, memory, UI, and agent loops are explicit non-goals. +- [Hugging Face TGI](https://github.com/huggingface/text-generation-inference) is a historical donor only because its repository is archived/maintenance-only at the research cutoff. + +### Flux baseline strengths to protect + +| Strength | Verified baseline | What to protect | Current caveat | +|---|---|---|---| +| Host-neutral boundary | `engine`, `llm`, `graph`, and `tools` are the stated host surface; UX, agent loops, tools, permissions, sessions, and product semantics belong to the host. See [`code_architecture.md`](code_architecture.md) and [`product_engineering.md`](product_engineering.md). | One-way dependency, compile-checked host imports, and a genuinely small stable facade. | The current `llm.Provider` still bundles too many maintenance/configuration facets, and internal delivery surfaces are marketed too broadly. | +| Canonical DTO direction | Provider/core aliases the central `llm` types instead of maintaining a second wire model; the DTOs already model content parts, tools, reasoning, usage, route, warnings, and provider replay blocks. | Additive, versioned DTO evolution with one owner. | Several declared fields are not populated or preserved end to end; the direction is right, the contract is ahead of implementation. | +| Small lower-level provider port | `provider/core.Provider` remains `Chat`, `StreamChat`, `Ping`, and `Name`. | Keep optional operations behind narrow interfaces rather than widening this contract. | Public package stability and optional-interface discovery are not yet consistently documented. | +| Streaming-first transport | SSE parsing is bounded, context-aware, and adapters emit semantic deltas rather than buffering full responses. | Pull-based streaming, backpressure, cancellation, and incremental middleware. | Terminal state, EOF/truncation, cancellation-preserving wrappers, deadlines, and usage aggregation are inconsistent. | +| Routing and resilience inventory | Weighted, least-busy, latency, cost, and usage strategies; circuit breaking; deployment fallback; continuation; rate limiting; cache; health; OTel wrappers exist as real code. | The breadth of primitives and the deployment-router validation model. | Stateful middleware is rebuilt per engine call, fallback can violate requirements, and route/attempt provenance is incomplete. | +| Last-good live routing | Atomic live snapshots, revision conflict handling, last-good retention, HTTPS/pinned Ed25519 peer manifests, redirect rejection, and size caps are strong controls. | Signed, immutable, fail-safe route distribution. | It is an in-process foundation, not a production-distributed control plane; TLS/process/restart/partition evidence is absent. | +| Defensive secret/state foundations | Injected credential stores, sanitized persisted state, atomic config writes, strict body limits, and no normal provider-secret serialization are useful. | Host-owned secret injection and safe status projection. | Legacy migration can delete plaintext after failed writes; SQLite sidecars, plaintext budget keys, URL/redirect policy, and namespace isolation remain gaps. | +| Provider/catalog breadth | The registry contains 28 gateway IDs, live catalog machinery exists, and native Anthropic/Gemini plus OpenAI-compatible/Responses/cloud paths are implemented in adapters. | Provider-qualified identity, catalog provenance, and multi-protocol support. | Registry count overstates independent protocol implementations; direct and deployment construction diverge. | +| Test and CI investment | The notes record broad deterministic tests, race runs, vet, lint, formatting, vulnerability scanning, and boundary checks passing. | Mock-first contract tests, race coverage, and package-level security tests. | Green leaf tests mask missing cross-layer, fresh-install, provider-parity, live-provider, process, and soak coverage. | +| Useful extension seeds | Custom OpenAI-compatible gateways, injected secret stores, lower-level HTTP client injection, catalog live metadata, and lower-level decorators exist. | Small explicit ports and default-free composition. | Public selectors, transports, cache backends, telemetry sinks, and native provider factories are fragmented or inaccessible. | + +These strengths are corroborated by [`code_architecture.md`](code_architecture.md), [`reliability_security.md`](reliability_security.md), [`provider_coverage.md`](provider_coverage.md), and [`product_engineering.md`](product_engineering.md). They are implementation facts unless a caveat explicitly says the feature is only declared, internal, leaf-level, or miswired. + +### Smallest differentiating thesis + +> **Flux is the narrow, auditable Go provider runtime for heterogeneous LLM endpoints: one explicit contract for lossless provider state, evidence-backed capabilities, deterministic stream/fallback behavior, and reproducible conformance, with no required gateway service, control plane, agent runtime, or model server.** + +This thesis is intentionally narrower than: + +- **LiteLLM:** Flux does not win by matching endpoint count, virtual keys, teams, spend databases, UI, or multi-instance gateway operations. +- **vLLM/SGLang/llama.cpp/Ollama/LocalAI:** Flux does not execute models, own weights, batch tokens, allocate KV, compile artifacts, or schedule GPUs. +- **LangGraph/Pydantic AI/Genkit/OpenAI Agents SDK:** Flux does not own the agent loop, tool execution, memory, sessions, checkpoints, durable workflow execution, or UI. + +The competitive product is therefore **trustworthy interoperability**, expressed as a small API plus conformance artifacts and migration diagnostics—not feature-count parity. + +## 2. Gap map, ranked backlog, and phased roadmap + +### Matrix legend and decision semantics + +**Flux status** + +- **Implemented:** verified on the intended path in current source/tests. +- **Partial:** real leaf implementation, but incomplete, lossy, internal-only, or not wired through the normal host path. +- **Declared:** public DTO/docs exist without a complete producer/consumer implementation. +- **Missing:** no adequate implementation in the reviewed checkout. +- **External/host:** deliberately not a core capability; may be a separate optional adapter or host concern. + +**Gap severity:** **S0** release/security/correctness blocker; **S1** high-value trust or production gap; **S2** meaningful interoperability/operations gap; **S3** low-return or strategically out of boundary. + +**Strategic value / complexity:** H/M/L; **VH** means very high. Complexity is comparative, not an effort estimate. + +**Recommendation:** **Build** in Flux-owned source; **Adopt** an external standard/tool without making it a runtime platform; **Borrow** a design pattern without dependency/code copying; **Defer** to a later phase or separate optional module; **Reject from core** when it violates the runtime boundary. + +### Consolidated capability matrix + +| Capability | Flux status | Best OSS donor(s) | Evidence | Gap severity | Strategic value | Implementation complexity | Dependency | Principal risk | Recommendation | +|---|---|---|---|---|---|---|---|---|---| +| Fresh-state library onboarding | **Partial:** `engine.New` can succeed, but normal generation requires a valid cache prepared through Rho/Flux-specific steps; README quickstart is not independently executable. | Bifrost SDK/gateway split; Portkey quickstart | [`product_engineering.md`](product_engineering.md), [`oss_gateways.md`](oss_gateways.md), [Bifrost](https://docs.getbifrost.ai/overview) | S0 | H | M | Catalog bootstrap, external consumer test | First adopter cannot run the advertised path | **Build** a minimal embedded/operator-explicit bootstrap and compile-tested quickstart | +| Catalog schema and publication trust | **Partial/confirmed break at audit:** strict decoding conflicts with the then-current remote artifact; provenance metadata is descriptive, and remote publication is not an in-repo trust workflow. | Agent Router compatibility discipline; current signed-manifest controls | [`product_engineering.md`](product_engineering.md), [current default catalog](https://langdag.com/model-catalog/v1/catalog.json), [Flux parser](https://github.com/GrayCodeAI/flux/blob/main/catalog/v1.go) | S0 | H | M | Schema evolution, CI artifact validation, publisher identity/digest | Mutable or additive catalog changes can break startup or silently change routes | **Build** tolerant/versioned schema plus CI validation and signed/digested last-good artifacts | +| Public API, release, and supply-chain truth | **Partial:** one signed `v0.0.1`, 29+ later commits on main, no API-diff baseline, module-proxy consumer gate, SBOM/attestation, or release PR; internal HTTP/gRPC/SDKs are presented inconsistently. | Product audit; Go module compatibility tooling | [`product_engineering.md`](product_engineering.md), [`reliability_security.md`](reliability_security.md), [v0.0.1](https://github.com/GrayCodeAI/flux/releases/tag/v0.0.1) | S0 | H | H | Stable package inventory, migration policy, release workflow | Accidental breaking changes and false production-readiness claims | **Build** a truthful pre-1.0 release process; **adopt** API diff and module-proxy gates | +| One provider descriptor and constructor path | **Partial:** 28 registry entries, but direct and deployment switches can choose different adapters; Gemini, Vertex, and Concentrate already diverge. | any-llm-go, AxonHub, New API `relaykit`, Bifrost | [`provider_coverage.md`](provider_coverage.md), [any-llm-go](https://pkg.go.dev/github.com/mozilla-ai/any-llm-go/providers), [New API](https://github.com/QuantumNous/new-api) | S0 | H | H | `ProviderSpec`, factory ownership, compatibility data | Same provider ID can silently use the wrong wire protocol or auth mode | **Build** one registry-driven resolver and 28-ID parity tests | +| Canonical generation option preservation | **Partial:** most options are modeled, but `ThinkingEnabled` is dropped and system/developer, JSON-schema descriptor, tool-choice, and warning semantics are uneven. | any-llm-go, Bifrost, Pydantic AI | [`code_architecture.md`](code_architecture.md) F1/F12, [`provider_coverage.md`](provider_coverage.md), [Bifrost matrix](https://docs.getbifrost.ai/providers/supported-providers/overview) | S0 | H | M | Engine conversion, adapter tests, capability profile | Valid-looking response with materially different behavior | **Build** one versioned option mapping and warning contract | +| Stream terminal state and cancellation | **Partial/defective:** channel EOF can become success; wrappers can drop `Close`; terminal `done/error/cancelled` invariants are not centralized. | OpenAI Agents SDK, Bifrost, Haystack, Agent Router | [`reliability_security.md`](reliability_security.md) F-03, [`code_architecture.md`](code_architecture.md) F5, [OpenAI streaming](https://openai.github.io/openai-agents-python/streaming) | S0 | H | H | One stream coordinator under `engine` and `provider/core` | Goroutine/body leaks and truncated streams reported as successful | **Build** a shared terminal/cancellation coordinator | +| Usage, cost, and continuation accounting | **Partial/defective:** real Anthropic/OpenAI streams emit split/earlier usage, while conversation persistence only reads usage on `done`; no hard continuation-call cap. | OpenAI Agents SDK, LiteLLM | [`reliability_security.md`](reliability_security.md) F-02/F-05, [`code_architecture.md`](code_architecture.md) F4/F9, [OpenAI usage](https://openai.github.io/openai-agents-python/usage) | S0 | H | H | Canonical usage event, accumulator, idempotent accounting | Uncharged spend, zero usage, or unbounded continuation | **Build** one final aggregate usage contract and hard continuation limits | +| Provider-native replay and raw state | **Declared/partial:** `ProviderBlock`, `ProviderBlock` events, raw tool arguments, and metadata exist, but adapters do not construct/preserve/replay them end to end. | Anthropic SDK, OpenAI SDK, LangChain | [`code_architecture.md`](code_architecture.md) F2/F3, [`provider_coverage.md`](provider_coverage.md), [Anthropic helpers](https://github.com/anthropics/anthropic-sdk-python/blob/main/helpers.md) | S0 | H | H | Adapter envelopes, persistence, privacy policy | Follow-up turns fail or silently lose signed reasoning state | **Build** required, protocol-filtered replay with explicit unsupported warnings | +| Tool and structured-output contract | **Partial:** tools and schemas are modeled, but the OpenAI schema shape is uncertain, `WithStructuredOutput` is a no-op, validator coverage is narrow, and proxy tool calls are lost. | Pydantic AI, BAML, Bifrost, OpenAI SDK | [`provider_coverage.md`](provider_coverage.md), [`code_architecture.md`](code_architecture.md) F6/F12, [BAML](https://docs.boundaryml.com/home) | S1 | H | H | Schema descriptor, capability negotiation, raw output/validation result | Silent schema/tool incompatibility and nondeterministic fallback | **Build** native modes plus explicit validation/repair metadata; do not execute tools | +| Per-attempt route policy and provenance | **Partial:** selected route exists; actual deployment, attempts, reason, policy decision, and externally-visible-output state are incomplete. | LiteLLM, Portkey, Manifest, Bifrost | [`oss_gateways.md`](oss_gateways.md), [`reliability_security.md`](reliability_security.md) F-08, [LiteLLM reliability](https://docs.litellm.ai/docs/proxy/reliability) | S0 | H | H | Route attempt object, policy port, wrappers | Cost, audit, and debugging point to the wrong backend | **Build** first-class attempt trace and actual-serving provenance | +| Fallback and capability preservation | **Partial/defective:** automatic fallback may occur when not allowed; unsupported tools can be stripped; non-transient errors can affect circuits; pre-output buffering is unbounded. | LiteLLM, Agent Router, Manifest | [`reliability_security.md`](reliability_security.md) F-07, [LiteLLM fallback enforcement](https://docs.litellm.ai/docs/proxy/reliability#enforce-key-model-access-on-fallbacks) | S0 | H | M | Explicit fallback policy, hard requirements, route events | Semantically different request or policy bypass | **Build** per-attempt revalidation and reject—not strip—unsupported requirements | +| Stateful runtime snapshot | **Partial/defective:** normal `Engine` selection and transport resolution reload state and recreate rate/cache/circuit wrappers per request. | Bifrost, GoModel, LiteLLM immutable/middleware separation | [`reliability_security.md`](reliability_security.md) F-01, [`product_engineering.md`](product_engineering.md), [Bifrost](https://docs.getbifrost.ai/overview) | S0 | H | H | Engine lifecycle, generation-based invalidation, credential rotation | Headline resilience features reset on every request and selection/execution can disagree | **Build** an immutable runtime snapshot with explicit refresh | +| Response cache and coalescer identity | **Partial/defective:** implementation is exact hash/LRU, zero config disables it despite docs, keys omit many request/deployment/tenant fields, and coalescer lifetime/key handling is unsafe. | LiteLLM caching, Bifrost | [`reliability_security.md`](reliability_security.md) F-09/F-10, [`code_architecture.md`](code_architecture.md) F7, [LiteLLM caching](https://docs.litellm.ai/docs/proxy/caching) | S0 | H | H | Canonical request identity, tenant scope, deep copy | Cross-tenant/deployment response collision and mutable metadata races | **Build** exact-cache correctness first; defer semantic similarity by default | +| Atomic budget reservation/finalization | **Partial/defective:** check and record are separate; streaming records every usage event and ignores errors; no idempotency key. | LiteLLM spend/rate-limit semantics | [`reliability_security.md`](reliability_security.md) F-05, [`code_architecture.md`](code_architecture.md) F9, [LiteLLM architecture](https://docs.litellm.ai/docs/proxy/architecture) | S0 | H | H | Storage transaction API, operation/request ID, usage finalization | Concurrent overspend and unreconciled spend | **Build** reserve → execute → idempotent finalize with durable failure path | +| Admission control and stream deadlines | **Partial:** broad ten-minute HTTP timeout exists; no distinct connect/header, first-event, first-output, idle, total deadlines or bounded global/provider queue. | Agent Router, llm-d, Dynamo | [`reliability_security.md`](reliability_security.md) F-04, [`provider_coverage.md`](provider_coverage.md), [Agent Router v1.1](https://theagentrouter.ai/release-notes/v1.1) | S0 | H | H | Transport settings, router admission, stream coordinator | Resource exhaustion and noisy retry amplification | **Build** bounded admission plus independent lifecycle clocks | +| Credential migration and secret lifecycle | **Partial/defective:** migration can remove a source after zero/partial successful writes and write a success marker regardless; default paths remain Rho/Hawk-specific. | No direct donor; apply transaction principles from provider state/config code | [`reliability_security.md`](reliability_security.md) F-12, [`code_architecture.md`](code_architecture.md) F10, [`product_engineering.md`](product_engineering.md) | S0 | H | M | Secret store interface, migration transaction | Permanent credential loss | **Build** transactional migration with verification and fail-closed marker | +| Secret-at-rest and endpoint isolation | **Partial:** strong local config controls, but budget DB can contain plaintext provider keys, SQLite sidecars are not uniformly hardened, and custom URL/redirect/private-address policy is permissive. | Go security patterns; current signed peer URL checks | [`reliability_security.md`](reliability_security.md) F-13/F-14, [`code_architecture.md`](code_architecture.md) F11 | S1 | H | H | Secret references, filesystem mode, URL policy, connection-time checks | SSRF, credential disclosure, cross-tenant leakage | **Build** explicit threat-model controls; keep strict host mode non-ambient | +| Shared provider conformance suite | **Partial:** `verify` exists but is not wired into adapters; core protocols have tests, while many compatibility providers are smoke-only. | Pydantic `TestModel`, OpenAI verification guide, Bifrost matrix | [`product_engineering.md`](product_engineering.md), [`provider_coverage.md`](provider_coverage.md), [OpenAI verification](https://developers.openai.com/cookbook/articles/gpt-oss/verifying-implementations) | S0 | H | H | Canonical contract, fixtures, capability skip model | Provider count mistaken for compatibility | **Build** on the existing `verify` package and run it for every adapter | +| Operation-level capability evidence | **Partial:** live catalog preserves richer model metadata, but provider registry uses broad booleans and lacks profile/model/runtime/version evidence. | Bifrost, LocalAI, KTransformers, LangChain profiles | [`provider_coverage.md`](provider_coverage.md), [Bifrost matrix](https://docs.getbifrost.ai/providers/supported-providers/overview), [KTransformers support matrix](https://ktransformers.net/en/docs/support-matrix) | S1 | H | H | Catalog schema, provider/profile identity, freshness | False claims from model names or generic OpenAI shape | **Build** an operation × profile × model × version evidence matrix | +| Endpoint profiles and native runtime metadata | **Partial:** Ollama is a thin OpenAI wrapper; llama.cpp/vLLM/SGLang/LocalAI native metadata is not consumed; local “compatible” profiles are not distinct. | Ollama, llama.cpp, LocalAI, vLLM, SGLang | [`model_runtimes.md`](model_runtimes.md), [Ollama API](https://github.com/ollama/ollama/blob/main/docs/api.md), [llama.cpp server](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md), [LocalAI discovery](https://localai.io/docs/features/api-discovery/index.html) | S1 | H | M | Optional profile package, bounded metadata importer | Generic compatibility masks semantic and lifecycle differences | **Build** read-only endpoint profiles; do not add model execution | +| OpenAI Chat Completions interoperability | **Partial:** internal proxy flattens history, accepts string content, ignores multiple fields, loses tool-call results, and can emit normal finish after an error. | OpenAI Python SDK, Bifrost, New API `relaykit` | [`code_architecture.md`](code_architecture.md) F12, [`reliability_security.md`](reliability_security.md) F-16, [OpenAI repository](https://github.com/openai/openai-python) | S1 | H | H | Isolated wire translator, conversation accumulator | Clients receive successful-looking but lossy generations | **Build** full typed translator in an isolated compatibility layer; delivery route remains optional | +| Anthropic Messages interoperability | **Missing as a public facade/partial internally:** native adapter exists, but signed blocks, native pass-through, and public `/v1/messages` do not. | Anthropic SDK, Agent Router, LiteLLM native endpoint | [`provider_coverage.md`](provider_coverage.md), [`code_architecture.md`](code_architecture.md) F3, [Anthropic SDK](https://github.com/anthropics/anthropic-sdk-python), [Agent Router capabilities](https://theagentrouter.ai/docs/capabilities/llm-integrations/supported-providers) | S1 | H | H | Lossless provider-block contract, isolated wire translator | Two-turn thinking/tool workflows fail outside direct adapter use | **Build** after canonical stream/state gates pass | +| Responses and raw native passthrough | **Declared/partial:** Concentrate has a Responses adapter; canonical public passthrough, lossless native facade, and warning contract are absent. | OpenAI SDK, LiteLLM pass-through, Agent Router extensions | [`provider_coverage.md`](provider_coverage.md), [LiteLLM pass-through](https://docs.litellm.ai/docs/pass_through/intro), [Agent Router release notes](https://theagentrouter.ai/release-notes/) | S2 | H | H | Chat/Anthropic compatibility, namespaced options, raw redaction | Universal DTO becomes lossy or untyped vendor objects leak into core | **Defer**, then **build** an explicit opt-in passthrough interface if canonical losslessness is proven | +| Optional operation ports | **Implemented but separate:** embeddings, Anthropic batch, rerank, media, and moderation exist outside the four-method provider contract. | any-llm-go, LiteLLM endpoint breadth | [`provider_coverage.md`](provider_coverage.md), [`model_runtimes.md`](model_runtimes.md), [any-llm-go](https://pkg.go.dev/github.com/mozilla-ai/any-llm-go/providers) | S2 | M | M | Capability matrix and host demand | Widening `Provider` or cloning LiteLLM's endpoint zoo | **Borrow** optional-interface pattern; keep operations separate until demonstrated demand | +| Public selector and model-call middleware | **Missing/partial:** routing uses a string switch; decorators exist but ordering, stream preservation, and public registration are inconsistent. | Vercel AI SDK, LangChain model middleware, APISIX phases | [`sdk_frameworks.md`](sdk_frameworks.md), [`oss_gateways.md`](oss_gateways.md), [Vercel middleware](https://ai-sdk.dev/docs/ai-sdk-core/middleware), [APISIX terminology](https://apisix.apache.org/docs/apisix/terminology/plugin/) | S2 | M | M | Stable core contracts, runtime snapshot | Plugin sprawl, cancellation loss, non-deterministic transforms | **Borrow** a narrow deterministic seam after P0; no agent/process plugins | +| Sticky affinity and P2C/PeakEWMA | **Missing:** current strategies are weighted, shuffle, least-busy, EWMA latency, cost, and usage. | Helicone, Portkey | [`oss_gateways.md`](oss_gateways.md), [Helicone routing](https://github.com/Helicone/ai-gateway#-smart-provider-selection) | S2 | M | M | Persistent strategy state, host-supplied scope/key | State resets, cross-tenant affinity, latency herding | **Borrow** as optional strategies; benchmark before enabling by default | +| Endpoint signal and deployment interoperability | **Missing:** route DTO lacks queue/KV/token-load/SLO/role/adapter signals; current manifest controls do not prove multi-process operation. | GAIE, llm-d, Dynamo, SGLang | [`model_runtimes.md`](model_runtimes.md), [GAIE API](https://gateway-api-inference-extension.sigs.k8s.io/concepts/api-overview/), [llm-d scheduler](https://github.com/llm-d/llm-d/blob/main/docs/architecture/core/router/epp/scheduling.md) | S2 | M | H | Capability/profile schema, optional adapters | Kubernetes-shaped stable API or pretending unknown signals exist | **Borrow** backend-neutral concepts; optional adapters only | +| Current GenAI OpenTelemetry semantics | **Partial:** an OTel wrapper exists, but engine does not wire it, wrapper attributes are custom/older, spans can end before async stream work, and content/audit policy is not coherent. | OTel GenAI, LiteLLM OTel v2, Vercel telemetry | [`reliability_security.md`](reliability_security.md) F-18, [`oss_gateways.md`](oss_gateways.md), [OTel GenAI](https://github.com/open-telemetry/semantic-conventions-genai), [LiteLLM OTel](https://docs.litellm.ai/docs/observability/opentelemetry_v2) | S1 | H | M | Stream coordinator, usage/route metadata, injected tracer | Nonportable telemetry and privacy leakage | **Adopt** OTel GenAI conventions and privacy defaults | +| Reproducible benchmark, soak, and failure evidence | **Missing:** benchmarks are not a CI baseline; no live provider/common-harness benchmark, process restart, TLS, partition, or soak evidence exists. | AIPerf/MLPerf methodology; Agent Router compatibility discipline | [`model_runtimes.md`](model_runtimes.md), [`product_engineering.md`](product_engineering.md), [AIPerf](https://github.com/ai-dynamo/aiperf) | S1 | H | M | Stable runtime and local deterministic harness | “Production ready” remains an assertion | **Adopt** AIPerf methodology; build a Flux client-only benchmark/conformance result schema | +| Signed distributed route manifests | **Partial/experimental foundation:** atomic snapshots, revisions, key pinning, size limits, and last-good logic are implemented; multi-process/TLS/restart/partition evidence is absent. | Agent Router, llm-d, current Flux controls | [`code_architecture.md`](code_architecture.md), [`model_runtimes.md`](model_runtimes.md), [Agent Router](https://github.com/theagentrouter/agent-router) | S1 | M | H | TLS/process test harness, persisted last-good state | Scope expands into a control plane without evidence | **Build** production proof while keeping the feature experimental; do not add consensus/autoscaling | +| HTTP/gRPC/SDK delivery surfaces | **Declared/internal/partial:** HTTP is internal and lossy; gRPC is tagged/skeletal; Python/TypeScript SDKs are internal stubs and Go SDK is a nested untested module. | Portkey deployment portability; OpenAI/Anthropic SDK semantics | [`product_engineering.md`](product_engineering.md), [`code_architecture.md`](code_architecture.md), [Portkey repository](https://github.com/Portkey-AI/gateway) | S2 | M | H | Stable library truth and explicit ownership | Core promise expands into half-maintained products | **Defer** extraction to separately versioned/tested modules; do not make delivery the identity | +| Evaluation and semantic guardrail ports | **Partial:** deterministic guardrails/cross-chunk core exist but streaming wrapper wiring is incomplete; no host-neutral evaluation platform. | DSPy, Pydantic Evals, Langfuse, Bifrost/Portkey guardrails | [`sdk_frameworks.md`](sdk_frameworks.md), [`code_architecture.md`](code_architecture.md) F6, [DSPy](https://dspy.ai/) | S2 | M | M | Privacy-safe event/trace export | Vendor latency, data leakage, product-layer creep | **Defer** external hooks; fix deterministic local behavior first | +| Model list/load/unload control | **Missing from core; leaf providers expose partial generation paths.** | Ollama, llama.cpp, Triton, Xinference | [`model_runtimes.md`](model_runtimes.md), [Triton model repository](https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/model_repository.html) | S3 | M | H | Separate admin credentials, explicit intent, audit | SSRF/RCE-like trust boundary, accidental downloads, control-plane creep | **Defer** to a separate admin-only module; **do not build in core** | +| Wasm and unrestricted plugin execution | **Missing:** no stable Go hook ABI or sandbox. | Higress, Bifrost security lesson | [`oss_gateways.md`](oss_gateways.md), [Higress](https://higress.ai/en/ai-gateway), [Bifrost CVE analysis](https://research.jfrog.com/vulnerabilities/bifrost-is-vulnerable-to-unauthenticated-remote-code-execution-via-mcp-stdio-client-registration-cve-2026-90898/) | S3 | L | VH | Stable Go middleware, explicit ABI and trust model | Large sandbox, supply-chain, memory/fuel, credential-access surface | **Defer**; never ship process/management plugins | +| UI, agents, orchestration, tenancy/control plane, RAG, memory, and model execution | **Missing by design** | Dify, Open WebUI, LangGraph, Pydantic AI, vLLM/SGLang/llama.cpp | [`sdk_frameworks.md`](sdk_frameworks.md), [`model_runtimes.md`](model_runtimes.md), [Dify](https://github.com/langgenius/dify), [Open WebUI](https://github.com/open-webui/open-webui) | S3 | Boundary-negative | VH | Host product or external model/control plane | Flux clone, security/maintenance explosion, dependency direction reversal | **Reject from core**; retain only contracts/observability needed by hosts | + +### Ranked backlog + +Priority meanings: **P0** must be resolved before a truthful `v0.1.0`; **P1** establishes protocol and production depth; **P2** is selective interoperability/ecosystem expansion; **P3** is a demand-gated experiment or a boundary guardrail. No calendar or effort estimate is asserted because maintainer capacity and design-partner availability are not established. + +| Rank | Priority | Work item and rationale | Likely Flux file/package touchpoints | Dependencies | Measurable acceptance criteria | +|---:|:---:|---|---|---|---| +| 1 | P0 | **Central stream terminal/cancellation coordinator.** Fixes truncation, EOF, `Close`, goroutine, body, and terminal-event defects across every wrapper. | `engine/stream.go`, `engine/continuation.go`, `llm/types.go`, `provider/core/stream.go`, tracing/guardrail/cache/budget/usage wrappers | None | Every adapter emits exactly one `done`, `error`, or `cancelled`; truncated EOF never emits `done`; closing every wrapped stream cancels the blocking fake/provider request; race tests show no leaked forwarding goroutine. | +| 2 | P0 | **Make the canonical contract lossless.** Map `ThinkingEnabled`; preserve provider blocks, raw tool arguments/metadata, warnings, usage source, request ID, actual route, finish reason, and structured errors. | `engine/convert.go`, `engine/stream.go`, `llm/types.go`, `provider/core/stream.go`, `provider/adapters/*`, `conversation/engine.go` | 1 | Every documented generation option has an engine conversion test; Anthropic signed/redacted thinking and OpenAI/Gemini provider state round-trip across two turns; request ID/usage/route/finish survive blocking and streaming paths. | +| 3 | P0 | **Persist one immutable runtime snapshot.** Make breaker, rate, cache, router, and selected catalog/config state survive requests and use one snapshot for selection and execution. | `engine/engine.go`, `engine/state.go`, `setup/deployment.go`, `router/*`, provider feature wrappers | 1 for stream-safe wrappers | Two sequential `Engine` calls load state once each at construction/refresh, not twice per request; the same limiter/cache/breaker instance is observed across calls; explicit catalog/config/credential invalidation swaps the snapshot atomically. | +| 4 | P0 | **Fresh-state bootstrap and catalog trust.** Make the library usable without Rho and prevent the current schema/trust failure. | `catalog/v1.go`, `catalog/v1_defaults.go`, `catalog/refresh.go`, `engine/state.go`, catalog generation/CI scripts | 3 | A clean temporary `HOME`/state dir runs a mock-backed `engine.New → Generate/Stream` without pre-existing files or Rho commands; CI parses the published default artifact; additive unknown fields are retained/ignored safely; invalid current artifact keeps a verified last-good state. | +| 5 | P0 | **One provider registry/factory.** Eliminate direct/deployment protocol divergence before adding providers. | `catalog/registry/*`, `provider/provider_registry.go`, `provider/adapters/provider_registry.go`, `setup/deployment.go`, `config/profiles.go` | 2 for capability metadata | A table-driven test covers all 28 registry IDs and proves direct/deployment construction resolves the same declared adapter/protocol; Gemini, Vertex, and Concentrate are explicit regression cases; aliases canonicalize once at the public boundary. | +| 6 | P0 | **Transactional credential migration and secret-at-rest hardening.** Prevent credential loss and secure local persistence. | `credentials/migrate.go`, `credentials/*_test.go`, `storage/sqlite.go`, `storage/budgets.go`, `engine/host_runtime.go`, `internal/probehttp` | None | Total/partial keychain write failure never removes source data and never writes success marker; SQLite main/WAL/SHM permissions are checked under a permissive umask; non-loopback plain HTTP/custom endpoint policy fails closed; no secret appears in cache keys/logs. | +| 7 | P0 | **Atomic budget reserve/finalize.** Close concurrent overspend and duplicate/missing charges. | `provider/observability/budget_provider.go`, `provider/observability/usage_limit.go`, `storage/budgets.go`, `llm/types.go` | 1, 2 | Concurrent requests cannot exceed a configured budget; duplicate finalization with the same operation ID is idempotent; split usage is finalized once; failed finalization is observable/retryable rather than ignored; cancellation reconciliation is tested. | +| 8 | P0 | **Canonical cache/coalescer identity.** Remove response collisions and rename the current exact cache truthfully. | `provider/cache/semantic_cache.go`, `provider/resilience/coalesce.go`, `provider/core/copy.go`, cache analytics | 2, 3 | Cache/coalescer keys differ when any response-affecting message part, tool/schema, generation option, provider/deployment, tenant scope, or provider option differs; returned metadata is deeply isolated; docs never call the default exact cache “semantic”; zero config behavior matches documentation. | +| 9 | P0 | **Admission, deadlines, fallback policy, and route attribution.** Make resilience policy-correct under load and failure. | `provider/core/transport.go`, `router/deployment_router.go`, `router/strategy.go`, `engine/*`, `internal/api/server.go` | 1, 3 | Configured queue limit `N=1` rejects/cancels a second wait without a second provider call; distinct connect/header/first-event/first-output/idle/total clocks are tested; `AllowFallback=false` never falls back; missing tool capability rejects rather than strips; actual deployment/attempt is reported after failover. | +| 10 | P0 | **Shared conformance v1.** Turn existing `verify` into an authoritative adapter gate. | `verify/*`, `provider/*_test.go`, `provider/adapters/*_test.go`, `testdata/conformance/*` | 1, 2, 5 | Every adapter runs the same sanitized request/response/SSE/error/cancellation suite; unsupported cases report explicit `UNSUPPORTED`; Tier-A Anthropic/OpenAI/Gemini/Azure/Bedrock/Vertex pass identical semantic assertions; unknown fields/events remain safely additive. | +| 11 | P0 | **Truthful v0.1 release and docs.** Remove false quickstart, provider-count, cache, proxy, SDK, security, and release claims. | README, `examples/*`, credential guide, `api/openapi.yaml`, `docs/ARCHITECTURE.md`, `.github/workflows/*`, VERSION/changelog | 4, 10 | A clean external module installs the candidate tag and compiles/runs `New/Generate/Stream` against a mock; all README examples are external-package tests; provider/env/count tables are generated from the registry; a documented stable/advanced/experimental inventory decides whether root `flux` or `engine` is primary without an unreviewed break; API compatibility diff is approved; module proxy, tag signature, and release metadata gates pass. | +| 12 | P1 | **Full OpenAI Chat Completions translator.** Replace the lossy internal proxy subset with typed translation. | `internal/api/openai_proxy.go`, `conversation/engine.go`, proposed isolated wire-translation package, OpenAI fixtures | 2, 9, 10 | Array/object multimodal messages, tools/results/choice, tool choice, schema descriptor, stop, seed, logprobs, stream options, usage detail, and error envelopes pass pinned OpenAI fixtures; no accepted field is silently ignored; an error never emits normal finish plus `[DONE]`. | +| 13 | P1 | **Native Anthropic Messages facade and replay.** Complete the second major SDK contract. | Anthropic adapter/core stream, isolated wire translator, conversation accumulator | 2, 10, 12 | Official Anthropic SDK blocking/streaming fixtures pass; signed/redacted thinking, tool lifecycle, usage split/timing, stop reasons, and errors round-trip; history is not flattened to transcript text. | +| 14 | P1 | **Operation × profile × model capability evidence.** Replace coarse provider booleans. | `catalog/live/*`, `catalog/v1.go`, `provider/adapters/provider_registry.go`, new profile schema | 5 | Every built-in deployment has explicit `supported/unsupported/unknown` rows for chat, stream, tools, schema, reasoning, image/audio input, output media, embeddings, rerank, batch, and passthrough, qualified by model/profile/runtime version/evidence/freshness. | +| 15 | P1 | **Endpoint profiles and bounded native metadata.** Distinguish generic OpenAI, Ollama, llama.cpp, vLLM, SGLang, and LocalAI. | new read-only profile package; Ollama/llama.cpp/LocalAI adapters; catalog live importer | 10, 14 | Ollama native `/api/tags` and `/api/show` are read without weight lifecycle; llama.cpp `/props` and model/slot metadata are bounded and namespaced; vLLM/SGLang/LocalAI profiles record version/config evidence; unknown remains unknown. | +| 16 | P1 | **Current privacy-safe OTel GenAI contract.** Replace custom/deprecated semantics and end spans at stream completion. | `provider/observability/tracing.go`, `internal/observability/*`, `engine/*`, GrayCodeAI conventions | 1, 2, 9 | In-memory exporter test sees canonical provider/model/operation/token/reason/TTFT/route/error attributes; spans end at actual terminal state; content capture is off by default; no model/request/tenant IDs become metric labels. | +| 17 | P1 | **Performance, soak, and live-provider evidence.** Replace assertion with reproducible measurements. | `benchmarks_test.go`, `verify`, new benchmark result schema, nightly workflow, process tests | 3, 10, 14 | Versioned local results report construction/catalog/keychain/cache/stream overhead, allocations, concurrency, TTFT overhead, success/error behavior; nightly tests report `PASS/UNSUPPORTED/FAIL/TIMEOUT/NOT_TESTED`; no vendor claim is relabeled as Flux evidence. | +| 18 | P1 | **Public selector and narrow model-call middleware.** Add deterministic extension without agent/plugin sprawl. | `router/strategy.go`, provider decorator layer, proposed stable interface, boundary tests | 1, 3 | Custom selector has compile-time tests; middleware order is deterministic; `wrapGenerate` and `wrapStream` preserve terminal/cancel/usage/route state; middleware cannot launch processes, mutate management state, or receive raw ambient credentials. | +| 19 | P1 | **Sticky affinity and P2C/PeakEWMA.** Improve local/fleet routing without shared infrastructure. | `router/strategy.go`, `router/router.go`, runtime snapshot state | 3, 9, 17 | Host supplies bounded scope/key; opt-out works; P2C selection is deterministic under seeded fixtures and avoids fastest-instance herding; state persists across engine calls and never crosses tenant scope. | +| 20 | P1 | **Prove signed manifest operation or keep experimental.** Add TLS/process/restart/partition evidence, not new control-plane features. | `router/controlplane/*`, persisted last-good state, process/TLS test harness | 3, 11 | A second process fetches a signed manifest over TLS, survives peer outage, rejects stale/conflicting revision, restarts from persisted last-good state, and is drained on shutdown; no consensus/autoscaling/UI scope is added. | +| 21 | P2 | **Optional operation ports.** Formalize embeddings, batch, rerank, media, and moderation only as separate capabilities. | `provider/core/embedding.go`, `provider/batch/*`, `provider/media/*`, rerank/moderation, capability registry | 10, 14 | Each operation has an independent interface and conformance suite; unsupported state is explicit; `core.Provider` remains four methods; asynchronous polling remains host-driven and bounded. | +| 22 | P2 | **Migration and compatibility diagnostics.** Provide a library/CLI report, not a migration control plane. | preflight/capability APIs, catalog/registry, optional CLI outside core | 11, 14, 15 | Given a target gateway/request, Flux reports effective provider/model/deployment, supported/transformed/dropped fields, loss warnings, and required host action; provider-qualified aliases import/export reproducibly. | +| 23 | P2 | **Responses and raw native passthrough.** Add only after canonical chat/Anthropic losslessness. | optional Responses adapter/facade, namespaced provider options, raw redaction hook | 2, 10, 12, 13 | A pinned Responses suite passes; raw mode is explicit, versioned, redacted, size-bounded, and never leaks untyped SDK objects into stable DTOs; canonical mode warns on every lossy conversion. | +| 24 | P2 | **Extract genuinely adoptable delivery modules.** Make optional surfaces public or remove them from the promise. | `internal/api`, `internal/grpc`, `internal/sdk/*`, new separate modules/workflows | 11, 12, 13 | Any retained HTTP/gRPC/SDK is independently versioned, externally importable, route/schema tested, and covered by root/consumer CI; unsupported stubs are removed or clearly experimental. | +| 25 | P3 | **Evaluation and external semantic-guardrail ports.** Export stable evidence without shipping products. | telemetry/event export, deterministic guardrail wrapper, optional host callback | 1, 2, 16 | A host can replay a privacy-safe completed fixture and receive a typed score/callback result; fail-open/closed and latency budget are explicit; Flux ships no datasets, prompt lab, policy catalog, optimizer, or UI. | +| 26 | P3 | **Separate model-server administration module.** Support explicit host workflows without contaminating generation. | new admin-only module; optional Ollama/llama.cpp/Triton/Xinference clients | 11, 14, 15 | Generation never implicitly downloads/loads/deletes; admin calls require separate credentials, explicit operation, idempotency key, timeout, audit event, and safe status; module is not imported by stable `engine`. | +| 27 | P3 | **Wasm or edge adapter only after a proven gap.** Re-evaluate trusted Go hooks before adding a sandbox. | separate edge/Wasm module, no stable DTO dependency | 18 | A documented target use case cannot be solved by the Go hook; ABI/version negotiation, deterministic streaming, memory/fuel limits, no ambient credentials, signing policy, and rollback are proven; otherwise the proposal is rejected/deferred. | + +### Phased roadmap and exit criteria + +The phases are **sequence gates, not schedules**. No date or effort estimate is justified from the notes because Flux currently has one visible human contributor and unknown design-partner capacity. + +#### Phase 1 — Hardening and truth + +**Objective:** make current claims true, remove release/security blockers, and establish one canonical contract. + +**Sequence:** backlog 1 → 2 → 3; in parallel where independent, 4–9; then 10 and 11. Provider factory work precedes broad conformance because otherwise fixtures can bless inconsistent constructor paths. + +**Exit criteria** + +- Fresh temporary-state mock quickstart passes without Rho or pre-existing state. +- Current published catalog is validated or clearly non-default; invalid updates preserve last-good state. +- All P0 stream, replay, usage, route, cache, budget, credential, admission, and storage tests pass under `-race`. +- All 28 provider IDs resolve consistently through direct and deployment paths. +- Every built-in adapter runs conformance v1 or reports an explicit unsupported case. +- Documentation, examples, provider counts, OpenAPI, SDK/server claims, and release gates match actual externally testable behavior. +- A release candidate can be consumed from a clean external module with no unreviewed incompatible API diff. + +#### Phase 2 — Provider protocol completeness + +**Objective:** make OpenAI Chat and Anthropic Messages behavior lossless enough to trust, before adding endpoint breadth. + +**Sequence:** canonical provider state and terminal semantics → OpenAI translator → Anthropic facade/replay → operation capability matrix → endpoint profiles. + +**Exit criteria** + +- Pinned official-SDK fixtures pass for blocking and streaming OpenAI Chat and Anthropic Messages. +- Tools, structured output, reasoning/replay state, multimodal input, usage, errors, cancellation, and unknown additive events have explicit tested semantics. +- The proxy no longer flattens multi-turn history or silently ignores accepted request fields. +- No universal DTO is declared lossless without a warning or error for every known lossy edge. +- Native profile evidence is versioned, bounded, redacted, and preserves `unknown`. + +#### Phase 3 — Production reliability + +**Objective:** prove behavior under concurrency, failure, restart, and sustained load rather than asserting it. + +**Sequence:** runtime snapshot and stream coordinator are prerequisites; then admission/deadlines → budget/cache concurrency → route/breaker health → TLS/process manifest proof → benchmark/soak/live tiers. + +**Exit criteria** + +- Configured concurrency/queue limits hold under cancellation and slow-consumer tests. +- First-event/idle/total deadlines and safe pre-output failover are tested for every core protocol. +- Budget reservation and finalization remain correct under duplicate, partial, failed, and concurrent events. +- Cache/coalescer tests show no cross-scope collision or mutable-state leak. +- Ping/health semantics, circuit transitions, route attempts, and request IDs survive wrappers. +- Versioned local benchmark and soak reports include success rate, latency/TTFT overhead, allocations, concurrency, stream behavior, and failure classes. +- Signed manifests either pass the multi-process/TLS/restart/outage gate or remain explicitly experimental. + +#### Phase 4 — Interoperability + +**Objective:** make Flux portable across hosted and self-hosted endpoints and observable without coupling to model-serving control planes. + +**Sequence:** capability evidence → endpoint profiles → OTel GenAI → GAIE/llm-d-style optional signals → migration diagnostics → selective operation ports. + +**Exit criteria** + +- Capability reports are versioned by provider/profile/model/runtime/config and preserve source, confidence, observation time, and expiry. +- Ollama, llama.cpp, vLLM, SGLang, and LocalAI profiles are reproducible against pinned versions or fixtures. +- OTel output passes a shared attribute/privacy conformance test and external collectors can interpret it without Flux-specific joins. +- Optional routing signals use backend-neutral DTOs; no Kubernetes dependency enters stable `engine` types. +- Migration diagnostics identify every transformed/dropped/unsupported field before a host changes a base URL or model ID. +- Optional operation interfaces pass their own suites without changing `core.Provider`. + +#### Phase 5 — Ecosystem and adoption + +**Objective:** publish a coherent library, make optional surfaces genuinely consumable, and collect evidence that external Go hosts value the thesis. + +**Sequence:** truthful `v0.1.0` → public API/stability labels and migration guide → module/release/security automation → optional delivery-module extraction → design-partner validation. + +**Exit criteria** + +- `v0.1.0` is a narrow library release with supported/advanced/experimental package tiers and a compatibility policy owned by Flux. +- CI produces API diffs, module-proxy smoke results, SBOM/provenance/attestation, signed tags, vulnerability/secret alerts, and release-PR automation. +- README quickstart, examples, provider matrix, conformance badges/tiers, and last-verified dates are generated or continuously checked. +- HTTP/gRPC/SDK surfaces are either independently maintained public modules or removed from the public promise. +- At least one independent Go host has integrated the stable API without importing internal packages; unresolved support exceptions are documented. +- P2/P3 experiments advance only when an explicit use case, dependency budget, security model, and measurable adoption benefit exist. + +### Sequencing guardrails + +- **No new provider-count work before P0 contract/factory/conformance work.** Another thin wrapper increases nominal breadth without reducing competitive risk. +- **No `/v1/responses` or raw passthrough before OpenAI/Anthropic canonical losslessness.** It would multiply translation paths while core semantics are unstable. +- **No public middleware before stream terminal/cancellation is centralized.** Middleware amplifies every lifecycle defect. +- **No default semantic cache, semantic routing, or LLM classifier before deterministic exact behavior and cost/latency baselines.** +- **No model-server control plane, UI, agent runtime, tenancy, RAG, or memory in core.** These are hard product-boundary constraints, not merely lower priorities. +- **No v1.0 claim until at least one compatibility baseline, release cadence, security/support owner, and external-consumer history exist.** The notes do not establish that history. + +## 3. Contradictions, confidence gaps, assumptions, and sources + +### Contradictions and low-confidence evidence requiring validation + +| Topic | Contradiction or uncertainty | Decision / validation required | +|---|---|---| +| Flux source snapshot | [`code_architecture.md`](code_architecture.md) labels commit `f9037a8` on `chore/oss-competitive-roadmap`; [`product_engineering.md`](product_engineering.md) calls the same commit `main`/`origin/main`. The synthesis checkout is currently on the feature branch. | Treat `f9037a8` as authoritative for this document; re-run implementation checks against the release branch before roadmap execution. | +| Provider count | README/source say 28 registry gateways; runtime docs say 16; credential guide says 15; AGENTS says 75+. | Use **28 registered gateway IDs** only. Generate docs from the registry. Do not call them 28 independent protocols or 75+ providers. | +| Scoped top-20 lists | Gateway, model-runtime, SDK/framework, and methodology notes each selected 20 for different universes. | Use the canonical portfolio above. Secondary donors remain reference material, not competing “top-20” claims. | +| LiteLLM release | One note observed `v1.102.1`; another metadata table observed `v1.99.3` on the same date. | Revalidate exact release before publication; this does not affect the architectural recommendation. | +| Bifrost release meaning | Latest repository tag may be a Helm chart while transport has a separate release line. | Do not use a repository “latest release” as proof of gateway/transport version; pin components. | +| Portkey 2.0 | Public announcement calls the production gateway open source, while the default branch labels 2.0 pre-release and stable OSS 1.x differs in feature boundary. | Treat 1.x source and 2.0 claims separately; current maintenance is a validation item. | +| Helicone license/activity | Gateway metadata/license indicate GPL-3.0 while README says Apache; gateway source is older than the separate observability repository. | Use P2C/PeakEWMA only as a cited algorithm idea. Do not copy or depend until license/activity is resolved. | +| Agent Router identity | Former Envoy AI Gateway canonical name changed in 2026. | Use `theagentrouter/agent-router`; preserve old CRD/CLI/module compatibility only when discussing upstream migration. | +| TGI status | Historically influential, but the repository is archived/maintenance-only. | Keep as historical donor, not active dependency or canonical portfolio member. | +| Default catalog | At audit time, the mutable remote artifact contained a field rejected by the strict decoder. It may change after the snapshot. | Re-run the artifact test on every publication/CI run. Keep a pinned fallback and report artifact digest/time. | +| OpenAI JSON-schema shape | Flux appears to emit `json_schema` directly where the current reference describes a descriptor object; no live credential validation was performed. | Validate with pinned official fixtures and an opt-in live probe before changing wire shape. | +| “Semantic cache” | README/engine call it semantic; normal implementation is deterministic exact hash/LRU. A separate embedding cache is not wired into the engine. | Rename the implemented path to exact response cache. Treat semantic similarity as a separate, opt-in future capability. | +| Streaming usage | Mocks often put usage on `done`, while real Anthropic/OpenAI processors emit separate/earlier usage events. | Replace mock-defined behavior with protocol fixtures and one final aggregate usage contract. | +| Provider replay state | DTO comments promise lossless replay, but no production adapter constructs or engine preserves the blocks. | Treat the fields as **declared, not implemented** until two-turn tests pass. | +| Structured-output helper | `WithStructuredOutput` is a no-op and the validator supports only a subset. | Do not claim provider-neutral structured-output repair; fix or remove the option. | +| OpenAI-compatible proxy | Docs say clients work “unchanged”; source accepts a lossy subset and can emit normal finish after an error. | Either complete the translator or explicitly publish a capability-limited facade; do not market drop-in parity before fixtures pass. | +| gRPC/SDKs | README presents optional gRPC/SDK surfaces, but gRPC is tagged/skeletal and SDKs are internal or outside root CI. | Defer external claims until modules are public, tested, documented, and independently versioned—or remove them from the core story. | +| OTel attributes | Flux contains old/custom and newer `gen_ai.*` paths; current upstream conventions moved to a dedicated repository. | Treat current OTel output as partial/nonconformant. Add a versioned migration and exporter test. | +| Optional cache/audit claims | README advertises distributed cache/audit sinks, but implementations are internal, incomplete, or unwired. | Remove from the stable promise until public, production-wired, and tested. | +| Provider/model support | No live API credentials or common conformance/performance benchmark was available. | Registry/test evidence proves construction and fixtures, not current hosted availability, regional behavior, or performance. | +| Vendor benchmarks/model claims | No independent common harness exists; many model/server support matrices are hardware-, version-, or config-dependent. | Preserve upstream claims as claims. Require pinned Flux conformance/benchmark evidence for competitive claims. | +| Public API demand | The proposed root `flux.New/Generate/Stream` API and v1 gates are engineering judgments, not a validated demand study. | Test the shape in at least two independent Go hosts before freezing v1; publish v0.1 only as pre-1.0. | +| Effort and dates | Notes contain a few illustrative week ranges but maintainer capacity, catalog publisher ownership, and design-partner availability are unknown. | This roadmap gives sequence and exit criteria only. Do not convert them into dates or staffing commitments. | + +### Evidence confidence and assumptions + +- **High confidence:** current Flux source behavior at `f9037a8`, cited line-level static findings, recorded test/race/vet results, exact public repository identity/licensing where inspected, and upstream architecture docs directly cited by the notes. +- **Medium confidence:** official project documentation as intent/implementation description; it may lag deployed versions or omit provider/model combinations. +- **Low confidence:** vendor provider/model counts, throughput/latency/adoption claims, rolling “latest” docs, pre-release feature boundaries, mutable remote artifacts, and any feature inferred from a generic “OpenAI-compatible” label. +- **Not established:** production SLOs, real call volume, downstream adoption, maintainer bus factor beyond the note’s Git history observation, legal advice, or independent provider/model performance. +- **Assumption for prioritization:** a P0 finding is release-blocking because it can make a public claim false, lose money/secrets, corrupt state, or break a core stream. A host that never enables an internal proxy, cache, budget store, or API may lower that feature’s immediate operational risk, but it does not remove the trust/documentation problem. +- **Assumption for boundary:** “optional” means a separately versioned package/module with no mandatory dependency from stable `engine`; an internal helper wired into normal `engine` is not optional. + +### Local note index + +All eight notes were read and used: + +1. [`selection_methodology.md`](selection_methodology.md) — comparison-universe criteria, weighted Flux-specific selection, direct-peer taxonomy, evidence hierarchy, and exclusions. +2. [`oss_gateways.md`](oss_gateways.md) — direct gateway peers, Bifrost/Agent Router/Helicone/Portkey/New API patterns, stream/policy/fidelity gaps, selective roadmap, and explicit non-goals. +3. [`model_runtimes.md`](model_runtimes.md) — model-server/runtime/control-plane boundary, local runtime profiles, capability evidence, endpoint metadata, conformance, benchmarking, and “do not execute models” guidance. +4. [`sdk_frameworks.md`](sdk_frameworks.md) — provider-port, stream, usage, middleware, schema, telemetry, and agent/host boundary donors. +5. [`code_architecture.md`](code_architecture.md) — line-level implementation audit, confirmed defects, design drift, source-of-truth conflicts, test gaps, and contract-v2 acceptance criteria. +6. [`provider_coverage.md`](provider_coverage.md) — 28-ID inventory, four protocol families, direct/deployment constructor divergence, capability/test coverage, replay and operation-surface gaps. +7. [`product_engineering.md`](product_engineering.md) — fresh-install failure, documentation/API/release drift, test maturity, extension seams, stable API proposal, and `v0.1`/`v1.0` gates. +8. [`reliability_security.md`](reliability_security.md) — production-significant stream, budget, routing, admission, credential, storage, API, batch, and telemetry findings with prioritized remediation. + +### Curated upstream references used by the decision + +#### Direct runtime peers + +- [Bifrost repository](https://github.com/maximhq/bifrost) and [provider/operation matrix](https://docs.getbifrost.ai/providers/supported-providers/overview) +- [LiteLLM repository](https://github.com/BerriAI/litellm), [request architecture](https://docs.litellm.ai/docs/proxy/architecture), [reliability/fallback](https://docs.litellm.ai/docs/proxy/reliability), [caching](https://docs.litellm.ai/docs/proxy/caching), and [OTel v2](https://docs.litellm.ai/docs/observability/opentelemetry_v2) +- [AxonHub](https://github.com/looplj/axonhub) +- [GoModel](https://github.com/ENTERPILOT/GoModel) and [documentation](https://gomodel.enterpilot.io/) +- [Agent Router](https://github.com/theagentrouter/agent-router), [capabilities](https://theagentrouter.ai/docs/capabilities/), and [v1.1 release notes](https://theagentrouter.ai/release-notes/v1.1) +- [Portkey Gateway](https://github.com/Portkey-AI/gateway) and [conditional routing](https://portkey.ai/docs/product/ai-gateway/conditional-routing.md) +- [OmniRoute](https://github.com/diegosouzapw/OmniRoute) +- [Manifest routing documentation](https://mnfst-manifest.mintlify.app/concepts/routing) + +#### Runtime/interoperability targets + +- [Ollama API](https://github.com/ollama/ollama/blob/main/docs/api.md) +- [llama.cpp server](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) +- [LocalAI discovery](https://localai.io/docs/features/api-discovery/index.html) +- [vLLM documentation](https://docs.vllm.ai/) and [feature matrix](https://docs.vllm.ai/en/stable/features/) +- [SGLang repository](https://github.com/sgl-project/sglang) and [model gateway design](https://github.com/sgl-project/sglang/blob/main/docs/advanced_features/sgl_model_gateway.md) +- [GAIE API overview](https://gateway-api-inference-extension.sigs.k8s.io/concepts/api-overview/) +- [llm-d scheduler](https://github.com/llm-d/llm-d/blob/main/docs/architecture/core/router/epp/scheduling.md) +- [Dynamo architecture](https://docs.nvidia.com/dynamo/v1.4.0/knowledge-base/overview) +- [AIPerf](https://github.com/ai-dynamo/aiperf) + +#### Protocol and contract donors + +- [any-llm-go provider interfaces](https://pkg.go.dev/github.com/mozilla-ai/any-llm-go/providers) and [provider matrix](https://github.com/mozilla-ai/any-llm-go/blob/main/docs/providers.md) +- [Vercel Language Model V4 specification](https://github.com/vercel/ai/blob/main/packages/provider/src/language-model/v4/language-model-v4.ts) and [middleware](https://ai-sdk.dev/docs/ai-sdk-core/middleware) +- [OpenAI Python SDK](https://github.com/openai/openai-python), [streaming guide](https://developers.openai.com/api/docs/guides/streaming-responses), and [implementation verification](https://developers.openai.com/cookbook/articles/gpt-oss/verifying-implementations) +- [Anthropic Python SDK](https://github.com/anthropics/anthropic-sdk-python) and [streaming helpers](https://github.com/anthropics/anthropic-sdk-python/blob/main/helpers.md) +- [Google GenAI Python](https://github.com/googleapis/python-genai) +- [BAML](https://github.com/BoundaryML/baml) +- [Pydantic AI](https://github.com/pydantic/pydantic-ai) +- [OpenAI Agents SDK streaming](https://openai.github.io/openai-agents-python/streaming) and [usage](https://openai.github.io/openai-agents-python/usage) +- [OpenTelemetry GenAI semantic conventions](https://github.com/open-telemetry/semantic-conventions-genai) + +#### Boundary and security references + +- [Flux host-engine boundary](https://github.com/GrayCodeAI/flux/blob/main/docs/architecture/HOST-ENGINE-BOUNDARY.md) +- [APISIX plugin lifecycle](https://apisix.apache.org/docs/apisix/terminology/plugin/) +- [Higress AI Gateway](https://higress.ai/en/ai-gateway) +- [Bifrost CVE-2026-90898 analysis](https://research.jfrog.com/vulnerabilities/bifrost-is-vulnerable-to-unauthenticated-remote-code-execution-via-mcp-stdio-client-registration-cve-2026-90898/) +- [Dify](https://github.com/langgenius/dify) and [Open WebUI](https://github.com/open-webui/open-webui) as explicit product-boundary comparators, not core dependencies + +### Final decision rule + +Flux should proceed only in this order: **make current contracts true → prove provider conformance → finish core protocol compatibility → demonstrate production reliability → add evidence-backed interoperability → publish and validate a narrow library release**. Any proposal that introduces UI, agent orchestration, a tenancy/control plane, mandatory external infrastructure, a general gateway/plugin platform, or model execution into core should be rejected regardless of competitor feature parity. diff --git a/research_notes/Flux OSS landscape roadmap/model_runtimes.md b/research_notes/Flux OSS landscape roadmap/model_runtimes.md new file mode 100644 index 00000000..852a226e --- /dev/null +++ b/research_notes/Flux OSS landscape roadmap/model_runtimes.md @@ -0,0 +1,473 @@ +# Flux OSS model-serving and local-runtime landscape (2026-09-24) + +## Landscape, scope, and final top-20 + +### Takeaway + +Flux overlaps most directly with provider SDKs and gateways such as LiteLLM, not with the execution engines underneath vLLM, SGLang, TensorRT-LLM, llama.cpp, and Ollama. The useful comparison set therefore needs both execution engines and the newer orchestration/routing layers. The final active top-20 below includes 11 engine/runtime projects, 5 packaging or deployment platforms, and 4 distributed routing/control-plane projects. Hugging Face TGI remains important historically but is not in the active top-20 because its repository is archived and explicitly in maintenance mode as of the research date ([TGI repository](https://github.com/huggingface/text-generation-inference), accessed 2026-09-24). + +### Method and evidence rules + +- **Cutoff:** 2026-09-24. Repository metadata, archived status, and activity below are a live GitHub API snapshot from that date; values will change. +- **Primary-source preference:** official repositories, architecture documents, API specifications, model manifests, release notes, and official issue/PR evidence. Third-party summaries were not used for core claims. +- **Source-confirmed:** a feature or boundary is stated in current upstream source/docs. **Project claim:** an upstream benchmark, adoption, or performance statement that was not independently reproduced. **Inference:** a recommendation or architectural interpretation derived from the cited evidence. +- **Feature notation:** “supports” means the cited project explicitly documents the capability; it does not imply that every model, backend, quantization, or hardware combination supports it. Project feature matrices are especially important where combinations are conditional. +- **Flux baseline:** the local Flux worktree was audited on 2026-09-24. Local line references are included after the corresponding repository links because the checkout is on `chore/oss-competitive-roadmap`, not necessarily identical to public `main`. + +### Final top-20 selection + +| # | Project | Architectural role | Why it belongs in the final landscape | +|---:|---|---|---| +| 1 | **Ollama** | Local model manager and inference daemon | Best reference for consumer-grade model manifests, pull/list/show lifecycle, keep-alive, and an OpenAI-compatible local endpoint. Flux currently has a thin Ollama adapter, so this is the highest-value local-runtime profile. Sources: [repository](https://github.com/ollama/ollama) and [API reference](https://github.com/ollama/ollama/blob/main/docs/api.md) (accessed 2026-09-24). | +| 2 | **llama.cpp / llama-server** | Portable CPU/GPU execution engine and HTTP server | Broadest practical local hardware/format reference. Its native `/props`, `/models`, slot, metrics, and OpenAI/Anthropic surfaces expose much richer runtime state than `/v1/models` alone. Sources: [repository](https://github.com/ggml-org/llama.cpp) and [server documentation](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) (accessed 2026-09-24). | +| 3 | **vLLM** | High-throughput serving engine | The de facto reference for continuous batching, paged KV memory, prefix caching, speculative decoding, disaggregation, broad hardware/model coverage, and OpenAI compatibility. It is a target Flux should profile, not a layer Flux should absorb. Sources: [documentation](https://docs.vllm.ai/) and [feature matrix](https://docs.vllm.ai/en/stable/features/) (accessed 2026-09-24). | +| 4 | **SGLang** | High-throughput engine and language/runtime | Important source of RadixAttention, HiCache, multi-LoRA batching, prefill/decode disaggregation, and a new model gateway. It shows which advanced cache/routing signals should be visible to a provider runtime without moving them into Flux. Sources: [repository](https://github.com/sgl-project/sglang), [documentation](https://docs.sglang.ai/), and [model gateway design](https://github.com/sgl-project/sglang/blob/main/docs/advanced_features/sgl_model_gateway.md) (accessed 2026-09-24). | +| 5 | **LocalAI** | Multi-backend, multi-modal local inference server | Strong reference for backend discovery, model configuration metadata, P2P/distributed operation, authentication, and a very broad OpenAI-compatible surface. Its surface area also demonstrates why generic “OpenAI-compatible” is insufficient. Sources: [repository](https://github.com/mudler/LocalAI), [quickstart](https://localai.io/docs/basics/getting_started/index.html), and [API discovery](https://localai.io/docs/features/api-discovery/index.html) (accessed 2026-09-24). | +| 6 | **llamafile** | Single-file packaging and distribution | Unique operational pattern: model/runtime/dependencies become a portable executable with no installation. It is the clearest reference for asset packaging, local execution, and inherited llama.cpp APIs. Sources: [repository](https://github.com/mozilla-ai/llamafile) and [documentation](https://docs.mozilla.ai/llamafile) (accessed 2026-09-24). | +| 7 | **TensorRT-LLM** | NVIDIA-optimized inference engine | Essential for advanced GPU scheduling, in-flight batching, paged KV, quantization, wide expert/context parallelism, prefill/decode disaggregation, and hardware-specific operational tradeoffs. Sources: [repository](https://github.com/NVIDIA/TensorRT-LLM), [architecture](https://nvidia.github.io/TensorRT-LLM/architecture/overview.html), and [parallelism](https://nvidia.github.io/TensorRT-LLM/features/parallel-strategy.html) (accessed 2026-09-24). | +| 8 | **MLC LLM** | Compiler plus cross-platform deployment engine | Reference for model conversion, compilation caching, precompiled libraries, device-specific manifests, and native Web/iOS/Android/desktop deployment. Its documentation/version lag is itself an important operational warning. Sources: [repository](https://github.com/mlc-ai/mlc-llm) and [documentation](https://llm.mlc.ai/docs/) (accessed 2026-09-24). | +| 9 | **KTransformers** | CPU/GPU heterogeneous MoE execution | Narrow but strategically important for large MoE models on workstations. Its explicit support matrix distinguishes current, narrow, needs-smoke, legacy, and unsupported combinations—better capability honesty than broad model-family claims. Sources: [repository](https://github.com/kvcache-ai/ktransformers), [inference guide](https://ktransformers.net/en/docs/inference), and [support matrix](https://ktransformers.net/en/docs/support-matrix) (accessed 2026-09-24). | +| 10 | **Qualcomm GenieX** | On-device NPU/GPU/CPU runtime | Current successor/redirect target for the Nexa SDK lineage and a strong edge-runtime reference: one C ABI, pluggable backends, model-format-specific capabilities, model pulls, multimodal input, and an OpenAI-compatible server. Sources: [repository](https://github.com/qualcomm/GenieX), [documentation](https://geniex.aihub.qualcomm.com/), and [SDK architecture](https://github.com/qualcomm/geniex/blob/main/sdk/README.md) (accessed 2026-09-24). | +| 11 | **Xinference** | Multi-engine model supervisor and worker platform | Strongest compact reference for “Flux-like” local operational concerns over heterogeneous engines: supervisor/workers, dynamic model launch/termination, model registries, custom manifests, replicas, UI, and OpenAI-compatible APIs. Sources: [repository](https://github.com/xorbitsai/inference) and [documentation](https://inference.readthedocs.io/) (accessed 2026-09-24). | +| 12 | **BentoML** | Model packaging, service runtime, OCI deployment | Useful for immutable service artifacts, model references, OCI packaging, local/Kubernetes deployment, concurrency, autoscaling, and optional managed cloud control planes. It is a deployment platform, not an inference kernel. Sources: [repository](https://github.com/bentoml/BentoML), [services](https://docs.bentoml.com/en/latest/build-with-bentoml/services.html), and [vLLM example](https://docs.bentoml.com/en/latest/examples/vllm.html) (accessed 2026-09-24). | +| 13 | **KServe** | Kubernetes inference control plane | Defines standards and operational behavior for model initialization, revisions, canaries, autoscaling, storage, and LLM-specific serving. Its `LLMInferenceService` and llm-d integration show what should remain declarative infrastructure around Flux. Sources: [repository](https://github.com/kserve/kserve), [LLMInferenceService overview](https://kserve.github.io/website/docs/model-serving/generative-inference/llmisvc/llmisvc-overview), and [configuration](https://kserve.github.io/website/docs/model-serving/generative-inference/llmisvc/llmisvc-configuration) (accessed 2026-09-24). | +| 14 | **Ray Serve LLM** | Distributed application/runtime orchestration | Strong reference for engine-independent deployment graphs, multi-model ingress, placement groups, autoscaling, custom routing, prefill/decode separation, and observability. Its cost is a large distributed-runtime dependency, so Flux should borrow interfaces and semantics rather than embed it. Sources: [repository](https://github.com/ray-project/ray), [Serve LLM documentation](https://docs.ray.io/en/latest/serve/llm/), and [architecture](https://docs.ray.io/en/latest/serve/llm/architecture/overview.html) (accessed 2026-09-24). | +| 15 | **NVIDIA Dynamo** | Datacenter-scale inference orchestrator | The clearest statement of separate request, control, and state planes: SLO planner, KV-aware routing, multi-tier cache, worker lifecycle, load shedding, and multiple inference backends. It directly identifies missing layers around Flux without asking Flux to execute models. Sources: [repository](https://github.com/ai-dynamo/dynamo) and [architecture](https://docs.nvidia.com/dynamo/v1.4.0/knowledge-base/overview) (accessed 2026-09-24). | +| 16 | **llm-d** | Kubernetes-native distributed inference stack | Best current reference for composable Router/InferencePool/Model Server boundaries, filter-score-pick scheduling, prefix/cache/session/LoRA affinity, fairness, SLO-aware autoscaling, and vendor-neutral model-server contracts. Sources: [repository](https://github.com/llm-d/llm-d), [architecture](https://llm-d.ai/docs/architecture), [router](https://llm-d.ai/docs/dev/architecture/core/router), and [scheduler](https://github.com/llm-d/llm-d/blob/main/docs/architecture/core/router/epp/scheduling.md) (accessed 2026-09-24). | +| 17 | **Gateway API Inference Extension** | Standard inference-routing protocol | Not an inference server, but a critical seam: standard model/pool identity, endpoint-picker protocol, model-aware routing, LoRA/rollout metadata, criticality, and observability. Flux should be able to consume equivalent signals without implementing Kubernetes CRDs. Sources: [documentation](https://gateway-api-inference-extension.sigs.k8s.io/), [repository](https://github.com/kubernetes-sigs/gateway-api-inference-extension), and [API overview](https://gateway-api-inference-extension.sigs.k8s.io/concepts/api-overview/) (accessed 2026-09-24). | +| 18 | **AIBrix** | Kubernetes control plane for GenAI serving | Adds the missing system layer around vLLM/SGLang/TensorRT-LLM: gateway policies, model-aware routing, inference autoscaling, unified runtime sidecars, cold-start artifact management, heterogeneous placement, distributed KV cache, and GPU diagnostics. Sources: [repository](https://github.com/vllm-project/aibrix), [documentation](https://aibrix.readthedocs.io/), and [KV-cache design](https://aibrix.readthedocs.io/latest/designs/aibrix-kvcache-offloading-framework.html) (accessed 2026-09-24). | +| 19 | **Triton Inference Server** | General multi-backend model server | Retains value as the mature generic model-repository/server reference: versioned model assets, load/unload/index APIs, per-model schedulers and batching, ensembles, multiple protocols/backends, and explicit operational control. Sources: [repository](https://github.com/triton-inference-server/server), [architecture](https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/introduction/index.html), [model repository](https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/model_repository.html), and [model-management extension](https://docs.nvidia.com/deeplearning/triton-inference-server/archives/triton-inference-server-2650/user-guide/docs/protocol/extension_model_repository.html) (accessed 2026-09-24). | +| 20 | **LiteLLM** | Provider SDK and multi-tenant AI gateway | The closest operational peer to Flux: OpenAI-format translation, routing, retries/fallbacks, virtual keys, budgets, rate limits, spend logging, Redis/PostgreSQL, multi-instance state, and production deployment. Flux’s differentiator should be a pure-Go embedded engine facade, first-class local runtimes, normalized streams, and less mandatory external state—not a feature checklist clone. Sources: [repository](https://github.com/BerriAI/litellm), [architecture](https://docs.litellm.ai/docs/proxy/architecture), [routing](https://docs.litellm.ai/docs/proxy/load_balancing), and [production deployment](https://docs.litellm.ai/docs/proxy/deploy) (accessed 2026-09-24). | + +### Adoption, activity, language, and license snapshot + +GitHub stars are an adoption signal, not a quality measure. “Last push” is the repository’s `pushed_at` value in the API snapshot. All values were accessed 2026-09-24. + +| Project | Stars | Primary GitHub language | Repository license signal | Last push | Interpretation | +|---|---:|---|---|---|---| +| Ollama | 181,528 | Go | MIT | 2026-09-23 | Very high adoption and current activity. [GitHub API](https://api.github.com/repos/ollama/ollama) | +| LocalAI | 49,242 | Go | MIT | 2026-09-23 | Large ecosystem and current activity. [GitHub API](https://api.github.com/repos/mudler/LocalAI) | +| llama.cpp | 129,334 | C/C++ | MIT | 2026-09-23 | Largest portable-runtime ecosystem in this set. [GitHub API](https://api.github.com/repos/ggml-org/llama.cpp) | +| vLLM | 92,533 | Python, CUDA/C++ components | Apache-2.0 | 2026-09-23 | Very high adoption and current activity. [GitHub API](https://api.github.com/repos/vllm-project/vllm) | +| SGLang | 36,379 | Python, CUDA/C++/Rust components | Apache-2.0 | 2026-09-23 | Fast-growing server/runtime with a foundation-oriented public identity. [GitHub API](https://api.github.com/repos/sgl-project/sglang) | +| MLC LLM | 23,183 | Python/C++ | Apache-2.0 | 2026-08-17 | Active but less recently pushed than the top server projects; docs still expose an old `0.1.0` version label. [GitHub API](https://api.github.com/repos/mlc-ai/mlc-llm), [docs](https://llm.mlc.ai/docs/) | +| KServe | 5,991 | Go | Apache-2.0 | 2026-09-23 | Active Kubernetes serving standard. [GitHub API](https://api.github.com/repos/kserve/kserve) | +| Ray Serve LLM | 43,910 (Ray) | Python/C++ | Apache-2.0 | 2026-09-23 | Large, active distributed platform; Serve LLM is one product surface. [GitHub API](https://api.github.com/repos/ray-project/ray) | +| TensorRT-LLM | 14,703 | Python/C++/CUDA | Apache-2.0 root, with listed component exceptions | 2026-09-23 | Active NVIDIA project. [GitHub API](https://api.github.com/repos/NVIDIA/TensorRT-LLM), [license](https://github.com/NVIDIA/TensorRT-LLM/blob/main/LICENSE) | +| llamafile | 26,038 | C/C++ | Apache-2.0 | 2026-09-23 | High adoption for a distribution format rather than an engine. [GitHub API](https://api.github.com/repos/mozilla-ai/llamafile), [license](https://github.com/mozilla-ai/llamafile/blob/main/LICENSE) | +| Dynamo | 8,153 | Rust, Python/C++ components | Apache-2.0 with separately licensed test data | 2026-09-23 | Young but highly active orchestration layer. [GitHub API](https://api.github.com/repos/ai-dynamo/dynamo), [license](https://github.com/ai-dynamo/dynamo/blob/main/LICENSE) | +| llm-d | 4,640 | Shell/Go/Kubernetes manifests | Apache-2.0 | 2026-09-23 | Young, active CNCF sandbox project. [GitHub API](https://api.github.com/repos/llm-d/llm-d), [repository overview](https://github.com/llm-d/llm-d/blob/main/docs/getting-started/README.md) | +| Gateway API Inference Extension | 770 | Go | Apache-2.0 | 2026-09-21 | Early but active Kubernetes SIG project; influence is greater than star count suggests. [GitHub API](https://api.github.com/repos/kubernetes-sigs/gateway-api-inference-extension) | +| AIBrix | 5,107 | Go/Python | Apache-2.0 | 2026-09-23 | Active vLLM-centered control plane. [GitHub API](https://api.github.com/repos/vllm-project/aibrix) | +| Triton Server | 11,003 | C++/Python | BSD-3-Clause | 2026-09-23 | Mature and active generic inference server. [GitHub API](https://api.github.com/repos/triton-inference-server/server) | +| BentoML | 8,856 | Python | Apache-2.0 | 2026-09-07 | Active packaging/deployment framework; managed cloud is optional and not equivalent to the OSS core. [GitHub API](https://api.github.com/repos/bentoml/BentoML) | +| Xinference | 9,592 | Python | Apache-2.0 | 2026-09-23 | Active local/cluster model platform. [GitHub API](https://api.github.com/repos/xorbitsai/inference) | +| KTransformers | 19,535 | Python/C++ | Apache-2.0 | 2026-09-23 | Active, specialized heterogeneous-MoE project. [GitHub API](https://api.github.com/repos/kvcache-ai/ktransformers) | +| Qualcomm GenieX | 8,395 | Rust/C/C++ | BSD-3-Clause | 2026-09-23 | Active edge runtime and current Nexa repository redirect target. [GitHub API](https://api.github.com/repos/qualcomm/GenieX), [license](https://github.com/qualcomm/GenieX/blob/main/LICENSE) | +| LiteLLM | 59,489 | Python with Rust core | MIT outside `enterprise/`; enterprise subtree separately licensed | 2026-09-23 | Very active direct gateway peer. [GitHub API](https://api.github.com/repos/BerriAI/litellm), [license](https://github.com/BerriAI/litellm/blob/main/LICENSE) | + +### Stale names, qualification, and exclusions + +- **Hugging Face TGI — important legacy reference, not final active top-20.** The repository is archived, says it is in maintenance mode, and limits contributions to minor fixes/docs while recommending vLLM, SGLang, llama.cpp, and MLX interoperability ([repository warning](https://github.com/huggingface/text-generation-inference), accessed 2026-09-24). Its Rust-router/Python-model-server split, zero-config sizing, OpenTelemetry, Prometheus, continuous batching, and long-prompt prefix cache remain useful design evidence ([architecture](https://huggingface.co/docs/text-generation-inference/architecture), [TGI v3 overview](https://huggingface.co/docs/text-generation-inference/main/en/conceptual/chunking), accessed 2026-09-24). TGI should appear in a historical/lessons subsection, not as a new Flux dependency. +- **Aethon — exclude from the model-server top-20.** The suggested `NVIDIA/Aethon` GitHub endpoint returned 404 on 2026-09-24 ([GitHub API](https://api.github.com/repos/NVIDIA/Aethon), accessed 2026-09-24). Search results for “Aethon” resolve to unrelated agent/UI projects such as [utensils/aethon](https://github.com/utensils/aethon) and [mertozbas/aethon](https://github.com/mertozbas/aethon), neither of which is an OSS model-serving engine (accessed 2026-09-24). No reliable current primary source for an NVIDIA model-serving Aethon was found. +- **Nexa SDK — follow the current redirect/rename to GenieX.** The former repository URL resolves to Qualcomm’s current GenieX project, whose docs and SDK are materially different from the earlier Nexa architecture ([Nexa SDK URL](https://github.com/NexaAI/Nexa-SDK), [GenieX](https://github.com/qualcomm/GenieX), accessed 2026-09-24). Cite GenieX, not stale Nexa release notes. +- **Red Hat vLLM Production Stack — do not list as a separate current project.** Its former GitHub endpoint did not resolve, while the official vLLM launch material says AIBrix was the new vLLM control-plane stack and the production stack was starting from a separate implementation ([former endpoint](https://api.github.com/repos/Red-Hat-AI-Innovation-Team/vLLM-production-stack), [vLLM launch post](https://vllm.ai/blog/2025-02-21-aibrix-release), accessed 2026-09-24). AIBrix is the current comparison target. +- **SGLang Model Gateway is a component, not a 21st project.** It is still important because it exposes worker lifecycle, heterogeneous protocol routing, history, MCP, privacy controls, and service discovery, but those features should be compared under SGLang ([design](https://github.com/sgl-project/sglang/blob/main/docs/advanced_features/sgl_model_gateway.md), accessed 2026-09-24). +- **Open WebUI, text-generation-webui, FastChat, and MCP servers are not selected as model servers.** They are valuable clients, evaluation harnesses, or tool protocols but do not replace the execution/control-plane coverage in the selected set. No exhaustive negative claim is made; they remain integration targets for Flux. + +### Inferences + +- The landscape has separated into four layers: model/engine execution, gateway/provider normalization, endpoint scheduling, and deployment control. A single “model server” comparison hides these boundaries. +- Flux should learn most from LiteLLM’s operational semantics, llm-d/Dynamo/GAIE’s model-aware routing, and llama.cpp/Ollama/LocalAI/Xinference’s local lifecycle—not from copying GPU kernels or autoscaler code. +- A final report should avoid a single performance ranking. Hardware, quantization, input/output distributions, prefix reuse, software versions, and SLO definitions dominate the results. + +### Gaps + +- GitHub stars and push timestamps do not establish production quality, contributor diversity, or support guarantees. +- No common live benchmark was run across the 20 projects. All upstream performance numbers below are explicitly treated as project claims. +- Several projects publish rolling “latest” docs rather than immutable versioned docs. Feature statements should be rechecked against a pinned Flux compatibility-test fixture before implementation. + +## Project evidence and comparison with Flux + +### Flux baseline and boundary + +- Flux describes itself as a provider runtime between an application and model APIs, owning credentials, catalog/route resolution, provider transports, normalized streams, resilience, usage, and telemetry; hosts own UX, agents, tools, permissions, sessions, and product semantics ([Flux README](https://github.com/GrayCodeAI/flux/blob/main/README.md), local `README.md:31-44`, accessed 2026-09-24). +- The stable host contract is deliberately small: `Generate`, `Stream`, model/catalog/credential/gateway/health/preflight operations, and a four-method lower-level `Provider` contract (`Chat`, `StreamChat`, `Ping`, `Name`) ([host boundary](https://github.com/GrayCodeAI/flux/blob/main/docs/architecture/HOST-ENGINE-BOUNDARY.md), local `docs/architecture/HOST-ENGINE-BOUNDARY.md:58-78`; [core contract](https://github.com/GrayCodeAI/flux/blob/main/provider/core/core.go), local `provider/core/core.go:21-33`, accessed 2026-09-24). +- The current DTOs already cover multimodal image/audio parts, tools, tool choice, structured output, reasoning controls, normalized usage, route attribution, provider replay blocks, warnings, typed stream errors, and pull-based streams ([LLM types](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go), local `llm/types.go:29-191,219-307`, accessed 2026-09-24). +- Flux already has weighted, least-busy, latency, cost, and usage routing; circuit breakers; deployment failover; live signed-manifest replicas; response and semantic caching; rate limiting; OpenTelemetry tracing/metrics; health/readiness; and an optional OpenAI-compatible proxy ([README feature sections](https://github.com/GrayCodeAI/flux/blob/main/README.md), local `README.md:112-176`; [router](https://github.com/GrayCodeAI/flux/blob/main/router/router.go), local `router/router.go:32-165`; [decentralized routing](https://github.com/GrayCodeAI/flux/blob/main/docs/architecture/DECENTRALIZED-FLUX.md), local `docs/architecture/DECENTRALIZED-FLUX.md:1-88`, accessed 2026-09-24). +- Flux’s own distributed routing document correctly labels the implementation as a foundation, not a finished control plane: it lacks persisted last-good state, global service discovery, fleet health, resource-reuse continuity, rollout/rollback, and measured scaling/soak limits ([distributed-routing limits](https://github.com/GrayCodeAI/flux/blob/main/docs/architecture/DECENTRALIZED-FLUX.md), local `docs/architecture/DECENTRALIZED-FLUX.md:64-88`, accessed 2026-09-24). + +### Cross-project capability map + +This is a synthesis of the official sources below, not a claim that every feature combination is available. + +| Concern | Projects with strong source-confirmed coverage | Flux implication | +|---|---|---| +| **Model discovery/loading** | Ollama (`tags/show/pull/copy/delete/running`), llama.cpp (`/models`, explicit load/unload, `/props`), Triton (repository index/load/unload/version policy), Xinference (registries/launch/terminate/replicas), KServe (storage initializer and CRDs), MLC/GenieX/Bento (pull, package, or artifact references) ([Ollama API](https://github.com/ollama/ollama/blob/main/docs/api.md), [llama.cpp server](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md), [Triton repository](https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/model_repository.html), [Xinference](https://inference.readthedocs.io/en/latest/models/index.html), [KServe LLMISVC](https://kserve.github.io/website/docs/model-serving/generative-inference/llmisvc/llmisvc-overview), [MLC packaging](https://llm.mlc.ai/docs/compilation/package_libraries_and_weights.html), [GenieX](https://github.com/qualcomm/GenieX), [BentoML vLLM](https://docs.bentoml.com/en/latest/examples/vllm.html), accessed 2026-09-24) | Flux should discover and describe models, not pull, compile, or own their lifecycle. Add optional, separate control adapters only where a host explicitly requests it. | +| **Hardware backends** | llama.cpp/LocalAI/Ollama for portable local hardware; vLLM/SGLang/TensorRT-LLM/AIBrix/KServe for accelerator fleets; MLC/GenieX for compiled or edge targets ([llama.cpp server](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md), [LocalAI quickstart](https://localai.io/docs/basics/getting_started/index.html), [vLLM](https://docs.vllm.ai/), [SGLang](https://docs.sglang.ai/), [TensorRT-LLM parallelism](https://nvidia.github.io/TensorRT-LLM/features/parallel-strategy.html), [GenieX platforms](https://geniex.aihub.qualcomm.com/en/get-started/platforms), accessed 2026-09-24) | Backend selection belongs to the model server/control plane. Flux can store backend family and health as deployment metadata, but should not link kernels or accelerators. | +| **Scheduling and batching** | vLLM continuous/chunked batching and scheduler policies; SGLang scheduler/RadixAttention; TensorRT-LLM in-flight/overlap scheduler; llama.cpp continuous batching/slots; TGI router continuous batching; llm-d/Dynamo/AIBrix endpoint scheduling ([vLLM V1](https://docs.vllm.ai/en/stable/usage/v1_guide.html), [SGLang index](https://github.com/sgl-project/sglang/blob/efee62ef/docs/index.rst), [TensorRT architecture](https://nvidia.github.io/TensorRT-LLM/architecture/overview.html), [llama.cpp server](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md), [TGI architecture](https://huggingface.co/docs/text-generation-inference/architecture), [llm-d scheduler](https://github.com/llm-d/llm-d/blob/main/docs/architecture/core/router/epp/scheduling.md), [Dynamo architecture](https://docs.nvidia.com/dynamo/v1.4.0/knowledge-base/overview), accessed 2026-09-24) | Flux should route among deployments and expose load/SLO hints. It should not batch model tokens or own a GPU scheduler. | +| **Streaming** | All selected serving projects support streaming to varying degrees; vLLM, SGLang, llama.cpp, Ollama, and LocalAI are the most relevant Flux targets ([vLLM](https://docs.vllm.ai/), [SGLang](https://docs.sglang.ai/), [llama.cpp](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md), [Ollama](https://github.com/ollama/ollama/blob/main/docs/api.md), [LocalAI](https://localai.io/docs/basics/getting_started/index.html), accessed 2026-09-24) | Flux should preserve normalized pull-based streams and make endpoint-specific SSE behavior conformance-tested. | +| **Multimodal input** | vLLM/SGLang/TensorRT-LLM; llama.cpp multimodal projector; Ollama images; LocalAI vision/audio/video; GenieX VLM/audio; Xinference multimodal ([vLLM](https://docs.vllm.ai/), [SGLang](https://docs.sglang.ai/), [llama.cpp](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md), [Ollama generate API](https://docs.ollama.com/api/generate), [LocalAI](https://localai.io/docs/basics/getting_started/index.html), [GenieX quickstart](https://geniex.aihub.qualcomm.com/en/run/cli/quickstart), [Xinference models](https://inference.readthedocs.io/en/latest/models/index.html), accessed 2026-09-24) | Flux already has image/audio DTOs. Missing work is per-deployment evidence and media normalization, not expanding the host product surface. | +| **Structured output and tools** | vLLM/SGLang structured generation and parsers; llama.cpp grammar-constrained JSON and tools; Ollama JSON-schema `format`; LocalAI multiple compatible APIs; LiteLLM normalizes tools/structured APIs; orchestration layers pass them through ([vLLM](https://docs.vllm.ai/), [SGLang](https://docs.sglang.ai/), [llama.cpp](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md), [Ollama generate API](https://docs.ollama.com/api/generate), [LocalAI quickstart](https://localai.io/docs/basics/getting_started/index.html), [LiteLLM architecture](https://docs.litellm.ai/docs/proxy/architecture), accessed 2026-09-24) | Flux should report tool/JSON support as capability evidence and warnings, not infer it solely from an OpenAI-looking endpoint. | +| **Speculative decoding** | llama.cpp draft model; vLLM EAGLE/MTP/draft/n-gram/suffix/custom/dynamic methods; SGLang; TensorRT-LLM EAGLE/MTP/n-gram; Xinference delegates by engine ([llama.cpp README](https://github.com/ggml-org/llama.cpp), [vLLM speculative decoding](https://docs.vllm.ai/en/stable/features/speculative%5Fdecoding/), [SGLang](https://docs.sglang.ai/), [TensorRT-LLM](https://nvidia.github.io/TensorRT-LLM/), [Xinference backends](https://inference.readthedocs.io/en/stable/user_guide/backends.html), accessed 2026-09-24) | This is server configuration. Flux may expose it as a deployment capability and benchmark dimension, but must not configure kernels in its core. | +| **KV/prefix cache** | vLLM automatic prefix caching and KV connectors; SGLang RadixAttention/HiCache; llama.cpp prompt reuse/slots; TGI long-prompt cache; Dynamo/llm-d/AIBrix multi-tier/event-indexed cache and affinity routing ([vLLM](https://docs.vllm.ai/), [SGLang](https://docs.sglang.ai/), [llama.cpp server](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md), [TGI v3](https://huggingface.co/docs/text-generation-inference/main/en/conceptual/chunking), [Dynamo](https://docs.nvidia.com/dynamo/v1.4.0/knowledge-base/overview), [llm-d KV management](https://llm-d.ai/docs/dev/architecture/advanced/kv-management), [AIBrix KV design](https://aibrix.readthedocs.io/latest/designs/aibrix-kvcache-offloading-framework.html), accessed 2026-09-24) | Flux must not own KV blocks. It should consume optional cache-affinity/load signals and expose response-cache semantics separately. | +| **Distributed serving** | TensorRT/vLLM/SGLang engine parallelism; llama.cpp router mode; Ray, KServe, llm-d, Dynamo, AIBrix, Bento, Xinference; Triton model instances/ensembles but not the same tensor-parallel model execution ([sources linked in project profiles](https://github.com/ai-dynamo/dynamo), accessed 2026-09-24) | Keep Flux’s independent replicas and route manifests; consume endpoint-picker and model-server health rather than becoming a scheduler. | +| **OpenAI compatibility** | All selected projects expose some subset; vLLM, llama.cpp, LocalAI, LiteLLM, KServe, Ray, Dynamo, and GenieX are especially relevant. “Compatible” is endpoint-specific and not full behavioral parity ([llama.cpp server](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md), [LocalAI](https://localai.io/docs/basics/getting_started/index.html), [LiteLLM](https://docs.litellm.ai/docs/simple_proxy), [KServe SDK integration](https://kserve.github.io/website/docs/model-serving/generative-inference/sdk-integration), accessed 2026-09-24) | Flux needs conformance tiers and endpoint profiles, not one boolean `openai_compatible` capability. | + +### Ollama + +- **Source-confirmed execution and model lifecycle:** Ollama exposes generate/chat, create, list, show, copy, delete, pull, push, embeddings, list-running, and version endpoints. Generate supports images, streaming, JSON or JSON-schema structured format, thinking output, log probabilities, and model keep-alive ([API reference](https://github.com/ollama/ollama/blob/main/docs/api.md), [generate schema](https://docs.ollama.com/api/generate), accessed 2026-09-24). +- **Operational model:** a local daemon and CLI manage model manifests, blobs, model processes, and a REST API. The project provides OpenAI compatibility as an additional surface; the native API remains richer for model inspection and lifecycle ([repository](https://github.com/ollama/ollama), [docs index](https://github.com/ollama/ollama/blob/main/docs/README.md), accessed 2026-09-24). +- **Flux comparison:** Flux’s Ollama adapter currently delegates through the OpenAI client at an Ollama base URL, so it does not consume native tags/show/keep-alive or native structured-output semantics directly ([adapter](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/ollama.go), local `provider/adapters/ollama.go:10-38`, accessed 2026-09-24). +- **Inference:** add an `ollama` endpoint profile that uses `/api/tags` and `/api/show` for model evidence, preserves native structured-output semantics, and optionally distinguishes “installed” from “currently loaded.” Do not add weight pulling or model-runner control to the mandatory Provider contract. +- **Gap:** the cited API confirms lifecycle and request features but does not provide the same detailed token-bucket scheduling, prefix-cache, distributed serving, or SLO contract documented by vLLM/llm-d. Flux should not invent those capabilities for Ollama. + +### llama.cpp and llama-server + +- **Source-confirmed execution:** llama-server supports CPU/GPU inference, parallel decoding, continuous batching, multimodal input, streaming, OpenAI chat/completions/responses/embeddings, Anthropic Messages, reranking, JSON-schema-constrained output, function/tool use, and speculative decoding ([server README](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md), accessed 2026-09-24). +- **Source-confirmed model/server control:** the current server exposes `/models`, `/models/load`, `/models/unload`, `/props`, `/slots`, `/metrics`, LoRA adapter listing/scaling, health, and a model router mode. The router selects by request `model` and proxies OpenAI, Anthropic, native, embeddings, rerank, and health routes ([server API](https://llama.app/docs/api), [server source](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/server.cpp), accessed 2026-09-24). +- **Operational model:** C/C++ executable, GGUF assets, local files or model directories, optional API key, hardware-specific backends, and no mandatory external database. llamafile adds a single-file packaging layer rather than changing the execution contract ([llama.cpp repository](https://github.com/ggml-org/llama.cpp), [llamafile](https://github.com/mozilla-ai/llamafile), accessed 2026-09-24). +- **Flux comparison:** llama.cpp can provide almost every model-server optimization Flux lacks, but doing so would violate the host-neutral engine boundary. The clean seam is a server-profile adapter and an optional model-server-control client. +- **Inference:** Flux should ingest `/props` and `/v1/models` as **runtime evidence**, including context, modalities, loaded status, slots, and server version, while keeping raw provider-specific JSON namespaced and non-semantic. +- **Gap:** the source does not establish a multi-node inference control plane equivalent to llm-d or Dynamo. Router mode is a local multi-model HTTP router, not tensor/pipeline parallelism. + +### vLLM + +- **Source-confirmed execution:** vLLM documents PagedAttention, continuous batching, chunked prefill, prefix caching, structured output, tool/reasoning parsers, streaming, multiple quantization formats, speculative methods, multimodal models, and tensor/pipeline/data/expert/context parallelism ([documentation](https://docs.vllm.ai/), accessed 2026-09-24). +- **Source-confirmed API and distributed model:** its server implements OpenAI-compatible APIs plus Anthropic Messages and gRPC. Current V1 defaults chunked prefill where possible, uses a unified scheduler, and publishes a feature-by-hardware compatibility matrix rather than universal support claims ([quickstart](https://docs.vllm.ai/en/latest/getting_started/quickstart/), [V1 guide](https://docs.vllm.ai/en/stable/usage/v1_guide.html), [compatibility matrix](https://docs.vllm.ai/en/stable/features/), accessed 2026-09-24). +- **Operational model:** Python package or container, model weights from Hugging Face/model storage, accelerator-specific images, and server flags. The official feature matrix is essential because e.g. prefix caching, speculative decoding, LoRA, multimodal, and parallelism have hardware/feature incompatibilities ([compatibility matrix](https://docs.vllm.ai/en/stable/features/), accessed 2026-09-24). +- **Flux comparison:** vLLM is a target server profile and a source of capability probes, not a dependency. Flux already handles retries/fallbacks; vLLM owns batching, paged memory, kernels, model execution, and disaggregated worker coordination. +- **Inference:** create a capability-evidence importer for `/v1/models` plus a pinned conformance probe set for tools, JSON schema, multimodal, reasoning blocks, usage, and streaming. Preserve server version because compatibility changes rapidly. +- **Gap:** project-level benchmark numbers cannot be transferred to Flux. Flux has no execution engine and therefore cannot be ranked on tokens/GPU, TTFT under continuous batching, or KV-cache efficiency. + +### SGLang + +- **Source-confirmed execution:** SGLang documents RadixAttention prefix reuse, a low-overhead scheduler, continuous batching, paged attention, chunked prefill, prefill/decode and encode/prefill/decode disaggregation, speculative decoding, structured output, quantization, multi-LoRA batching, and tensor/pipeline/expert/data parallelism ([project index](https://github.com/sgl-project/sglang/blob/efee62ef/docs/index.rst), accessed 2026-09-24). +- **Source-confirmed API and hardware:** it supports OpenAI APIs and Hugging Face model formats, with NVIDIA, AMD, Intel CPU, TPU, and Ascend backends. The project claims use across more than 400,000 GPUs; this is an upstream adoption claim, not an independently audited result ([documentation](https://docs.sglang.ai/), accessed 2026-09-24). +- **Operational model:** Python/C++/CUDA server with pip/container installation, model weights, Prometheus/OTel-oriented observability, and advanced object-storage/KV-cache options. The model gateway adds worker lifecycle, heterogeneous HTTP/gRPC/OpenAI routing, policy, history, MCP, and Kubernetes service discovery ([model gateway](https://github.com/sgl-project/sglang/blob/main/docs/advanced_features/sgl_model_gateway.md), accessed 2026-09-24). +- **Flux comparison:** SGLang Model Gateway overlaps conceptually with LiteLLM/Flux, but SGLang couples it to the SGLang worker runtime. Flux should adopt protocol-neutral routing, lifecycle, and observability ideas without importing agent history, MCP execution, or SGLang dependencies. +- **Inference:** cache locality, role (`prefill`/`decode`), worker load, and route hints belong in a server/control-plane telemetry contract. Flux can choose an endpoint using external signals but should not treat raw prompt prefixes as routing state. +- **Gap:** the cited docs provide a broad feature list but not one universally stable compatibility surface; exact parser and model support remain version/feature dependent. + +### LocalAI + +- **Source-confirmed execution and lifecycle:** LocalAI is a Go server for LLMs plus image, audio, video, and other models, using multiple backends and no mandatory GPU. It supports OpenAI-compatible APIs, Anthropic Messages, OpenAI Responses, gallery/model assets, model configuration, and backend control ([quickstart](https://localai.io/docs/basics/getting_started/index.html), [repository](https://github.com/mudler/LocalAI), accessed 2026-09-24). +- **Source-confirmed discovery:** `/.well-known/localai.json` reports available endpoints and runtime capabilities; `/api/instructions` returns machine/LLM-readable capability guides; model configuration metadata and VRAM estimation expose backend/config fields ([API discovery](https://localai.io/docs/features/api-discovery/index.html), accessed 2026-09-24). +- **Operational model:** local/on-prem process with API key gating, optional multi-user auth/RBAC/usage tracking, configurable backends, and documented P2P/distributed features. A 2026 PR additionally documents backend monitor/shutdown and system/version endpoints, showing a broad and changing management API ([quickstart](https://localai.io/docs/basics/getting_started/index.html), [PR #8852](https://github.com/mudler/LocalAI/pull/8852), accessed 2026-09-24). +- **Flux comparison:** LocalAI validates Flux’s need for server discovery metadata but goes much further into asset/backend control. `/api/instructions` and config metadata are useful patterns for capability evidence; server control belongs behind an optional administrative interface. +- **Inference:** LocalAI’s multiple API dialects are a warning against a global “supports OpenAI” bit. Flux should bind a capability to provider profile, endpoint, model, runtime version, and evidence timestamp. +- **Gap:** P2P/distributed operation is upstream-documented, but the research did not establish parity with production tensor parallelism, SLO-aware routing, or Kubernetes control planes. + +### llamafile + +- **Source-confirmed packaging:** llamafile combines llama.cpp and Cosmopolitan Libc into a portable single-file executable. It can run across common CPU architectures and operating systems with no installation and also packages Whisper transcription in the same style ([repository](https://github.com/mozilla-ai/llamafile), accessed 2026-09-24). +- **Source-confirmed API:** server mode inherits llama.cpp’s OpenAI and Anthropic-compatible APIs; combined mode starts a terminal chat and HTTP server together. Built-in tools are explicitly dangerous in untrusted environments and restricted to localhost origins by default ([quickstart](https://github.com/mozilla-ai/llamafile/blob/main/docs/quickstart.md), [running guide](https://github.com/mozilla-ai/llamafile/blob/main/docs/running_llamafile.md), accessed 2026-09-24). +- **Operational model:** the model/runtime/dependency bundle is the deployable artifact. This solves distribution and cold-start packaging but not fleet scheduling, model-weight governance, or fleet-wide observability. +- **Flux comparison:** Flux should catalog llamafile artifacts and endpoint profiles, not generate or sign single-file binaries. A future packaging adapter could validate an artifact’s model ID, runtime version, license, digest, and supported API profile before treating it as a deployment. +- **Inference:** Flux’s existing model identity fields and catalog provenance are a good base; add artifact digest/runtime metadata rather than inventing a new manifest ecosystem. +- **Gap:** llamafile’s value is primarily distribution; its serving architecture is llama.cpp’s. It should not be counted as a separate scheduler or model-execution layer. + +### TensorRT-LLM + +- **Source-confirmed execution:** TensorRT-LLM provides a modular Python `LLM` API, C++/Python executor components, in-flight batching, overlap scheduling, CUDA graphs, paged KV, chunked context, quantizations, guided decoding, LoRA, speculative methods, and multimodal/VG models ([repository](https://github.com/NVIDIA/TensorRT-LLM), [architecture](https://nvidia.github.io/TensorRT-LLM/architecture/overview.html), [overview](https://nvidia.github.io/TensorRT-LLM/overview.html), accessed 2026-09-24). +- **Source-confirmed distributed execution:** tensor, pipeline, data, expert, context, and wide expert parallelism are supported; wide expert parallelism includes expert slots, dynamic placement, and load balancing ([parallelism](https://nvidia.github.io/TensorRT-LLM/features/parallel-strategy.html), accessed 2026-09-24). +- **Operational model:** NVIDIA GPU/CUDA/ PyTorch stack, high-level Python API, lower-level runtimes, and integration with Dynamo and Triton. Build/runtime versions and hardware compatibility are major operational constraints. +- **Flux comparison:** TensorRT-LLM is entirely below Flux’s intended layer. The useful additions are deployment capabilities: hardware class, parallelism mode, disaggregated role, KV-transfer support, quantization, and health. +- **Inference:** a `BackendProfile` should represent these as optional facts with explicit `supported`, `unsupported`, or `unknown` state. Flux should not attempt to auto-select TensorRT engines or embed NVIDIA-specific options in the stable host DTO. +- **Gap:** the root license is Apache-2.0 with explicitly listed component exceptions; the current license file includes a non-Apache LTX-2 subtree, so a blanket “Apache-2.0” claim would be inaccurate for all bundled model code ([license](https://github.com/NVIDIA/TensorRT-LLM/blob/main/LICENSE), accessed 2026-09-24). + +### MLC LLM + +- **Source-confirmed compilation/deployment:** MLC downloads prequantized weights, compiles a target-specific model library through Apache TVM, caches both, then runs a native chat engine. It exposes Python, REST/OpenAI, JavaScript/WebGPU/WebAssembly, iOS, and Android paths and supports tensor-parallel shards ([introduction](https://llm.mlc.ai/docs/get_started/introduction), [quickstart](https://llm.mlc.ai/docs/get_started/quick_start), accessed 2026-09-24). +- **Source-confirmed packaging metadata:** the package manifest carries model repository/path, model ID, estimated VRAM, bundle-weight choice, context-window override, and target device—useful evidence that runtime capability is target-specific rather than model-name-specific ([package libraries and weights](https://llm.mlc.ai/docs/compilation/package_libraries_and_weights.html), accessed 2026-09-24). +- **Operational model:** pip/conda, compiler cache, platform-specific generated libraries, and compiled application artifacts. Deployment dependencies are much heavier than a pure client adapter. +- **Flux comparison:** MLC is evidence for artifact metadata and cross-hardware deployment, not for Flux to own compilation. A model library should enter the Flux catalog as an opaque deployment artifact referenced by digest and target. +- **Inference:** Flux capability rows should support hardware/precision/runtime qualifiers rather than one row per model. +- **Gap:** the current documentation title still says `mlc-llm 0.1.0`, while the repository was pushed on 2026-08-17. This mismatch is a documentation freshness risk and a reason not to infer current support solely from prose. + +### KTransformers + +- **Source-confirmed scope:** KTransformers is a specialized CPU/GPU heterogeneous system for large MoE inference and LoRA fine-tuning. Its current inference path uses `kt-kernel`/SGLang integration, with hot GPU experts and cold/deferred CPU experts, AMX/AVX kernels, NUMA-aware memory, and multiple quantization methods ([repository](https://github.com/kvcache-ai/ktransformers), [inference](https://ktransformers.net/en/docs/inference), accessed 2026-09-24). +- **Source-confirmed honesty pattern:** the support matrix labels combinations as Current, Current/narrow, Needs smoke, Needs reconciliation, Legacy, or Not current, and requires model family + checkpoint + method + hardware + entry point to match before a result is transferable ([support matrix](https://ktransformers.net/en/docs/support-matrix), accessed 2026-09-24). +- **Operational model:** Python/C++ plus SGLang and model-specific prepared weights. The project is a research/engineering platform, not a general model catalog or fleet gateway. +- **Flux comparison:** KTransformers should be represented as a specialized server profile/deployment, with model/hardware/method qualifiers. Generic tools/vision support should be unknown unless the exact SGLang/model path confirms it. +- **Inference:** this is one of the strongest external precedents for Flux’s capability evidence model: validity is a tuple of model, artifact revision, runtime version, backend, hardware, and configuration. +- **Gap:** the project itself labels several combinations as legacy, narrow, or needing smoke tests. A Flux catalog must preserve those qualifiers rather than flatten them to a boolean. + +### Qualcomm GenieX + +- **Source-confirmed architecture:** GenieX has one C ABI under thin Go CLI, Python, Android, Docker, and server bindings. `llama_cpp` and `qairt` are dynamic plugins, allowing a build to link only the engines it needs ([SDK README](https://github.com/qualcomm/geniex/blob/main/sdk/README.md), [repository](https://github.com/qualcomm/GenieX), accessed 2026-09-24). +- **Source-confirmed model/hardware model:** community GGUF runs through llama.cpp on CPU, Adreno GPU, or Hexagon NPU; Qualcomm AI Hub bundles run through QAIRT on NPU. Bundle precision, context length, and KV-cache size are compile-time constraints, so the same logical model can have materially different runtime capabilities by artifact ([platforms](https://geniex.aihub.qualcomm.com/en/get-started/platforms), [supported models](https://geniex.aihub.qualcomm.com/en/models/supported), accessed 2026-09-24). +- **Operational model:** model pull/cache, CLI/SDK/server, signed release variants for HTP, Docker, and OpenAI-compatible local serving. It targets Snapdragon devices and explicitly rejects CPU/GPU placement for QAIRT bundles rather than silently pretending support ([server quickstart](https://github.com/qualcomm/GenieX), [platforms](https://geniex.aihub.qualcomm.com/en/get-started/platforms), accessed 2026-09-24). +- **Flux comparison:** GenieX is a concrete model for a pluggable backend and capability resolver. Flux can mirror the logical separation—model source, runtime plugin, compute unit, artifact constraints—without implementing the C ABI. +- **Inference:** Flux catalog keys should be deployment/offering-specific, not only model/provider-specific. A Qualcomm bundle and a GGUF model with the same family name may have different context, precision, and audio capabilities. +- **Gap:** the official pages do not establish fleet-scale serving, prefix routing, or broad structured/tool conformance; those should remain unknown rather than inherited from llama.cpp. + +### Xinference + +- **Source-confirmed model operations:** Xinference has model registries plus launch, list, get, terminate, replicas, supervisor/workers, UI, and cluster deployment. Current backends include vLLM, SGLang, llama.cpp, Transformers, and MLX, and launch options are engine-specific ([using Xinference](https://inference.readthedocs.io/en/latest/getting_started/using_xinference.html), [launching](https://inference.readthedocs.io/en/latest/user_guide/launch.html), accessed 2026-09-24). +- **Source-confirmed metadata:** custom model definitions include path/URI, family, context length, dimensions, max tokens, language, abilities, formats, size, quantizations, revision, hub/source, chat template, stop tokens/tags, reasoning tags, cache config, and optional isolated virtual environment ([custom models](https://inference.readthedocs.io/en/latest/models/custom.html), accessed 2026-09-24). +- **Operational model:** pip package, local or Docker launch, supervisor plus workers for clusters, OpenAI-compatible APIs, model lifecycle, and backend-specific extras including continuous batching and speculative decoding where the selected engine supports them ([backends](https://inference.readthedocs.io/en/stable/user_guide/backends.html), accessed 2026-09-24). +- **Flux comparison:** Xinference is close to the desired separation between provider runtime and model control plane but combines them. Flux should import its rich manifest as deployment metadata; it should not absorb supervisor/worker placement. +- **Inference:** Xinference’s engine-selection metadata supports a Flux concept of `model offering → engine profile → deployment instance`, which avoids assuming one backend for a model. +- **Gap:** no source-confirmed inference was found for SLO-aware cache routing, federated control, or production-grade multi-region failover. Treat those as absent/unknown, not unsupported in principle. + +### BentoML + +- **Source-confirmed packaging:** a Bento packages source, Python dependencies, model references, image/system packages, and runtime configuration. Models can be excluded from the image and pulled from Hugging Face, improving image build and cold-start behavior ([vLLM example](https://docs.bentoml.com/en/latest/examples/vllm.html), accessed 2026-09-24). +- **Source-confirmed service model:** class-based services expose typed APIs, lifecycle hooks, sync/async workers, custom commands, tasks, and custom inference code. The LLM example can run vLLM’s built-in OpenAI server while retaining Bento resource and deployment configuration ([services](https://docs.bentoml.com/en/latest/build-with-bentoml/services.html), [vLLM example](https://docs.bentoml.com/en/latest/examples/vllm.html), accessed 2026-09-24). +- **Operational model:** local serve, OCI image, Kubernetes deployment, and optional BentoCloud. OSS core capabilities and managed-cloud control-plane capabilities must be distinguished. +- **Flux comparison:** Bento is a model/service deployment platform, not a provider protocol runtime. Flux should borrow artifact identity, reproducibility, resource declarations, lifecycle hooks, and health/readiness conventions—not Bento’s Python service model. +- **Inference:** a clean Flux integration is a preflight validator for a deployment artifact, not a dependency on Bento. +- **Gap:** BentoML’s own LLM serving examples delegate optimization to vLLM, so it is not evidence for kernels, batching, or KV management. + +### KServe + +- **Source-confirmed control plane:** KServe supports the traditional `InferenceService` path and a GenAI-specific `LLMInferenceService` built around llm-d concepts. It provisions model workloads, storage initialization, Gateway API routes, Endpoint Picker schedulers, multi-node workers, prefill/decode pools, TP/DP/EP, monitoring, and authentication/RBAC ([LLMInferenceService overview](https://kserve.github.io/website/docs/model-serving/generative-inference/llmisvc/llmisvc-overview), accessed 2026-09-24). +- **Source-confirmed operations:** configuration composition layers well-known presets, user `baseRefs`, and service specs; supports model URIs, LoRA adapters, HPA/KEDA or Workload Variant Autoscaler, canary traffic splitting, managed/custom HTTPRoutes, and prefill/decode independent scaling ([configuration](https://kserve.github.io/website/docs/model-serving/generative-inference/llmisvc/llmisvc-configuration), [composition](https://kserve.github.io/website/docs/model-serving/generative-inference/llmisvc/llmisvc-config-composition), accessed 2026-09-24). +- **Operational model:** Kubernetes CRDs/controllers, object storage/PVC/Hugging Face initialization, Gateway API/GAIE, KServe components, and external vLLM/SGLang/other engines. It does not execute model tokens itself. +- **Flux comparison:** KServe is an external control plane Flux can target through OpenAI-compatible endpoints and metadata. Flux should not add CRDs or own Kubernetes reconciliation. +- **Inference:** KServe’s `spec.model` versus workload/router/parallelism separation is a useful model for Flux’s stable host DTO versus optional deployment metadata. +- **Gap:** advanced scheduling, distributed serving, and autoscaling are only as available as the chosen model server and infrastructure; KServe itself is not an inference kernel. + +### Ray Serve LLM + +- **Source-confirmed execution/deployment:** Ray Serve LLM provides OpenAI-compatible chat/completions/embeddings, multi-model and multi-node deployment, autoscaling/load balancing, TP/PP/EP/data-parallel attention, prefill/decode disaggregation, prefix-aware routing, LoRA, vLLM/SGLang backends, metrics, and Grafana dashboards ([Serve LLM](https://docs.ray.io/en/latest/serve/llm/), accessed 2026-09-24). +- **Source-confirmed architecture:** `LLMServer` owns one engine and physical placement groups; `OpenAiIngress` owns API compatibility, model multiplexing, and custom routing. The design explicitly separates application concerns from infrastructure and exposes protocol-based engine/server extension points ([architecture](https://docs.ray.io/en/latest/serve/llm/architecture/overview.html), accessed 2026-09-24). +- **Operational model:** Ray cluster, Python runtime, model source from Hugging Face or S3/GCS/Azure, placement groups, autoscaling, and a distributed control plane. Startup includes node provisioning, image/library startup, model loading, compile, memory profiling, CUDA graph capture, and warmup ([deployment initialization](https://docs.ray.io/en/latest/serve/llm/user-guides/deployment-initialization.html), accessed 2026-09-24). +- **Flux comparison:** Ray’s separation of ingress and engine is correct for Flux, but Ray itself is too large and stateful to embed. Flux should define equivalent narrow interfaces for endpoint metadata, routing signals, and health. +- **Inference:** a Flux deployment adapter can read Ray-style model placement/resource metadata, but should not make Ray a core dependency. +- **Gap:** the overview advertises engine-agnostic SGLang support while the current configuration reference says vLLM is the supported `llm_engine` value. Treat advertised extensibility and currently configurable backends as different claims ([overview](https://docs.ray.io/en/latest/serve/llm/architecture/overview.html), [configuration](https://docs.ray.io/en/latest/serve/llm/user-guides/configuration.html), accessed 2026-09-24). + +### NVIDIA Dynamo + +- **Source-confirmed planes:** Dynamo explicitly separates a fast request plane, responsive control plane, and state/event plane. The request path includes frontend normalization, load/KV-overlap routing, prefill, decode, and streaming; control includes planner, operator, discovery, topology-aware placement, and optional model management; state includes KV events, block management, and NIXL transfer ([architecture](https://docs.nvidia.com/dynamo/v1.4.0/knowledge-base/overview), accessed 2026-09-24). +- **Source-confirmed operations:** health checks, stale endpoint removal, graceful draining, request migration/cancellation, and load shedding are described as normal operating events. Standalone and Gateway API/GAIE topologies both expose OpenAI-compatible APIs ([architecture](https://docs.nvidia.com/dynamo/v1.4.0/knowledge-base/overview), [repository](https://github.com/ai-dynamo/dynamo), accessed 2026-09-24). +- **Source-confirmed model-aware optimization:** disaggregated prefill/decode, KV-aware routing, KV block offload/recall, SLO planner, cold-start model streaming, and cache/event state are major components. Upstream 2x TTFT/7x cold-start claims are project claims tied to particular systems and should not be generalized ([repository](https://github.com/ai-dynamo/dynamo), accessed 2026-09-24). +- **Operational model:** Rust performance components, Python backend integrations, integrated or experimental sidecar engines, Kubernetes CRDs/EndpointSlices, optional event infrastructure, and multi-backend support. +- **Flux comparison:** Dynamo provides the cleanest external reference for Flux’s current “routing foundation, not control plane” limitation. Flux should expose request/deployment health and route metadata, while Dynamo/llm-d/Kubernetes own the serving control plane. +- **Inference:** Flux’s normalized stream should add enough stage information to correlate prefill/decode or retry/fallback events without depending on Dynamo-specific fields. +- **Gap:** the sidecar mode is documented as experimental and has less feature coverage than integrated backends. Backend support should therefore be version-qualified, not merely vendor-qualified. + +### llm-d + +- **Source-confirmed architecture:** llm-d defines Router, InferencePool, and Model Server. The Router separates a standard L7 proxy from an Endpoint Picker; InferencePool groups equivalent model servers and Variants express role/cost/performance; Model Server is vLLM, SGLang, TensorRT-LLM, or another engine ([architecture](https://llm-d.ai/docs/architecture), [model servers](https://github.com/llm-d/llm-d/blob/3d74f4a6/docs/architecture/core/model-servers.md), accessed 2026-09-24). +- **Source-confirmed scheduling:** EPP uses a pluggable Filter → Score → Pick pipeline. Current scorers include prefix, KV utilization, queue depth, running requests, token load, latency, LoRA affinity, and session affinity; P/D profiles choose separate prefill/decode endpoints ([scheduler](https://github.com/llm-d/llm-d/blob/main/docs/architecture/core/router/epp/scheduling.md), accessed 2026-09-24). +- **Source-confirmed operations:** central flow control supports admission, throttling, fairness and ordering; HPA/KEDA and Workload Variant Autoscaler are supported; OpenAI-compatible batch gateway and async processor are separate optional components ([architecture](https://llm-d.ai/docs/architecture), [getting started](https://github.com/llm-d/llm-d/blob/main/docs/getting-started/README.md), accessed 2026-09-24). +- **Source-confirmed cache:** event-driven KV indexing, CPU/SSD offload, peer sharing, and a three-layer model of routing intelligence, observability, and capacity are documented ([KV management](https://llm-d.ai/docs/dev/architecture/advanced/kv-management), accessed 2026-09-24). +- **Operational model:** Kubernetes-native, CNCF sandbox, Gateway API/GAIE, LeaderWorkerSet, EPP sidecars, Prometheus model-server metrics, and vLLM/SGLang/TensorRT-LLM workers. +- **Flux comparison:** llm-d is the strongest architectural peer for Flux’s routing layer but is deliberately a fleet-serving stack. Flux should implement the same **signals and filter/score concepts**, not the Kubernetes resources. +- **Inference:** Flux can expose a small `EndpointSignals` contract with queue depth, in-flight/token load, TTFT/ITL, KV utilization, loaded LoRA IDs, supported roles, health, and SLO headroom. A local process can compute these directly; a Kubernetes deployment can bridge them from EPP/Prometheus. +- **Gap:** the EPP configuration is read at startup and changes require restart, according to upstream docs. This is an operational limitation of llm-d, not a Flux recommendation ([EPP source overview](https://github.com/llm-d/llm-d/tree/main/docs/architecture/core/router/epp), accessed 2026-09-24). + +### Gateway API Inference Extension + +- **Source-confirmed purpose:** GAIE is an official Kubernetes project that turns an ext-proc/Gateway API proxy into an inference gateway. It provides a reference Endpoint Picker for conformance and defines routing around model servers using metrics and capabilities such as prefix cache and loaded LoRAs ([documentation](https://gateway-api-inference-extension.sigs.k8s.io/), [repository](https://github.com/kubernetes-sigs/gateway-api-inference-extension), accessed 2026-09-24). +- **Source-confirmed APIs:** `InferencePool` groups model-server pods and references an Endpoint Picker; `InferenceModel` maps a public model name to backing models/adapters and rollout policy. The EPP protocol communicates selected and optional fallback endpoint through ext-proc metadata, supports request payload/header stages, and defines fail-open behavior ([API overview](https://gateway-api-inference-extension.sigs.k8s.io/concepts/api-overview/), [implementer guide](https://gateway-api-inference-extension.sigs.k8s.io/guides/implementers/), [InferencePool](https://gateway-api-inference-extension.sigs.k8s.io/api-types/inferencepool/), accessed 2026-09-24). +- **Operational model:** Kubernetes Gateway API plus a conformant L7 gateway, EPP extension, and model-server metrics. It is a protocol/control seam, not token execution, weight loading, or a queue engine. +- **Flux comparison:** this is the best interoperability target for Flux’s deployment identities and request/response metadata. Flux need not implement Kubernetes CRDs, but its route DTO should map cleanly to public model, deployment/pool, serving role, variant, adapter, and fallback concepts. +- **Inference:** adopt names and semantics in a backend-neutral schema, but avoid Kubernetes types in the stable `engine` DTO. Any Kubernetes adapter belongs outside the host contract. +- **Gap:** the project is still building toward GA. Flux should target stable concepts and versioned schemas, not assume every gateway implementation supports every planned feature. + +### AIBrix + +- **Source-confirmed scope:** AIBrix is a Kubernetes control plane and serving stack with an LLM gateway, model-aware routing, app-tailored autoscaling, unified AI runtime sidecars, distributed inference, distributed KV cache, heterogeneous cost/SLO serving, and GPU failure detection ([repository](https://github.com/vllm-project/aibrix), accessed 2026-09-24). +- **Source-confirmed cache design:** the current framework supports L1 DRAM and optional L2 remote cache, pluggable connectors and eviction policies, and cross-engine reuse; the documented Connector interface expresses capabilities such as batch operations, prefetch, RDMA, and GPUDirect transfer ([KV-cache design](https://aibrix.readthedocs.io/latest/designs/aibrix-kvcache-offloading-framework.html), accessed 2026-09-24). +- **Source-confirmed routing dependency:** precise prefix-state synchronization currently uses vLLM KV events, ZMQ, a remote tokenizer, and a central index feeding gateway routing. This is useful evidence, not a portable Flux dependency ([KV event sync](https://aibrix.readthedocs.io/latest/features/kv-event-sync.html), accessed 2026-09-24). +- **Operational model:** Kubernetes control/data planes, enhanced vLLM/SGLang images, NIXL/UCX/RDMA for advanced cache/disaggregation, Redis/ZMQ, and GPU diagnostic components. +- **Flux comparison:** AIBrix validates a runtime-sidecar/control-plane boundary. Flux should be able to consume standardized engine signals and model lifecycle metadata, but not become the sidecar. +- **Inference:** Flux’s adapter layer should tolerate partial implementations: some deployments expose full model-server management, some only OpenAI APIs, and some only a handful of Prometheus metrics. +- **Gap:** AIBrix’s strongest claims and current docs are vLLM-centric. Cross-engine claims should be version-qualified until the exact backend integration is conformance-tested. + +### Triton Inference Server + +- **Source-confirmed architecture:** Triton routes HTTP/gRPC/C API requests to per-model schedulers, performs configurable batching, and invokes backend-specific inference. The model repository is the organizational boundary for model files, versions, config, labels, and custom configuration ([architecture](https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/introduction/index.html), [model repository](https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/user_guide/model_repository.html), accessed 2026-09-24). +- **Source-confirmed model management:** repository index/load/unload APIs expose installed-but-unloaded models, configuration overrides, and dependent unload behavior. Models can live on local filesystems, GCS, S3, or Azure storage and use explicit version policies ([model repository extension](https://docs.nvidia.com/deeplearning/triton-inference-server/archives/triton-inference-server-2650/user-guide/docs/protocol/extension_model_repository.html), accessed 2026-09-24). +- **Operational model:** C++ server with framework backends and Python bindings, containers, explicit repositories, model ensembles/BLS, per-model scheduling/batching, and operational profiling/model-analyzer tooling ([repository](https://github.com/triton-inference-server/server), accessed 2026-09-24). +- **Flux comparison:** Triton’s model repository API is the best source for a separate optional **model-server control interface**. It is much broader than a provider `Ping`, but it should not be added to `core.Provider`. +- **Inference:** Flux may ship an optional Triton control adapter that lists/loads/unloads models for host setup workflows. The data-plane provider should only call the model once loaded and ready. +- **Gap:** Triton is not LLM-specialized. Its generic model lifecycle is excellent, but token streaming semantics, LLM parser support, prefix routing, and speculative decoding depend on backend and model integration. NVIDIA’s developer page calls Dynamo the successor to Triton; that is an upstream positioning claim, not a reason to remove Triton from the landscape ([Dynamo developer page](https://developer.nvidia.com/dYNAMO), accessed 2026-09-24). + +### LiteLLM + +- **Source-confirmed architecture:** LiteLLM’s gateway authenticates virtual keys, checks budgets, applies global/key/user/team rate limits, routes with retries/fallbacks, translates providers through its SDK, and performs spend/logging work asynchronously after responses ([life of a request](https://docs.litellm.ai/docs/proxy/architecture), [architecture source](https://github.com/BerriAI/litellm/blob/main/ARCHITECTURE.md), accessed 2026-09-24). +- **Source-confirmed routing:** strategies include simple shuffle, least busy, usage, latency, and cost; deployment `order` provides staged failover; Redis shares rate-limit and router state across gateway replicas; encrypted-content affinity can pin Responses follow-ups to the deployment that created the encrypted state ([load balancing](https://docs.litellm.ai/docs/proxy/load_balancing), accessed 2026-09-24). +- **Source-confirmed operations:** LiteLLM supports monolithic or independently scaled gateway/backend/UI services. PostgreSQL is required for auth/tracking, Redis for multi-instance rate limits/cache, and Helm/Terraform provide HPA/KEDA, Prometheus, disruption budgets, migrations, and cloud infrastructure ([production deployment](https://docs.litellm.ai/docs/proxy/deploy), accessed 2026-09-24). +- **Flux comparison:** LiteLLM overlaps Flux most directly and is more operationally complete in multi-tenancy, virtual keys, budgets, database-backed spend, and UI. Flux is stronger as an embeddable pure-Go provider runtime with a small stable host facade, lower deployment friction, local-runtime evidence, and a normalized pull-based stream vocabulary. +- **Inference:** Flux should not copy LiteLLM’s mandatory Postgres/Redis topology. Instead it should offer a smaller embedded default and optional sink/control adapters for teams that need enterprise gateway state. +- **Governance/license qualification:** most of the repository is MIT, but the `enterprise/` subtree has separate terms. Any Flux comparative or integration recommendation must distinguish OSS core from enterprise features ([license](https://github.com/BerriAI/litellm/blob/main/LICENSE), [multi-tenant architecture](https://docs.litellm.ai/docs/proxy/multi_tenant_architecture), accessed 2026-09-24). +- **Gap:** the gateway does not itself provide model kernels, KV-cache implementation, or hardware-aware model-server scheduling. It is a peer for Flux’s data/control routing, not a substitute for vLLM/SGLang/TensorRT-LLM. + +### Hugging Face TGI: legacy comparison, outside active top-20 + +- **Source-confirmed architecture:** a Rust router/webserver batches and schedules requests and calls one or more Python model servers over gRPC. Model servers load/shard models and perform inference; routers and model servers may run on different machines ([TGI architecture](https://huggingface.co/docs/text-generation-inference/architecture), accessed 2026-09-24). +- **Source-confirmed features:** continuous batching, SSE streaming, OpenAI-compatible APIs, tensor parallelism, Flash/Paged Attention, KV caching, quantization, guided decoding, Prometheus, and OpenTelemetry ([repository](https://github.com/huggingface/text-generation-inference), accessed 2026-09-24). +- **Project claim:** TGI v3 reports a microsecond-scale prefix lookup and large long-prompt gains over vLLM in selected hardware/model tests. These are upstream comparative claims, not independently verified and not a reason to prefer an archived project ([TGI v3 overview](https://huggingface.co/docs/text-generation-inference/main/en/conceptual/chunking), accessed 2026-09-24). +- **Flux comparison:** preserve the router/engine split as a design lesson and its zero-config hardware-aware sizing idea. Do not build a Flux adapter around TGI as a strategic dependency while it is archived/maintenance-only. +- **Inference:** TGI belongs in a final report’s “important but now legacy” callout and can still be a compatibility target for existing deployments. +- **Gap:** no current feature roadmap or active engine development should be inferred from old TGI feature documentation. + +## Flux seams, additive roadmap, and interoperability practices + +### Takeaway + +Flux should remain a provider runtime and become the best normalized client/router for heterogeneous APIs and model-server deployments. It should not own weights, compilation, token batching, KV blocks, GPU placement, or autoscaling. The highest-value additions are evidence-backed model metadata, endpoint profiles, conformance tests, standardized route/load/SLO signals, and benchmark-compatible observability. + +### Architectural boundary: what stays in Flux versus external servers + +| Concern | Flux owns | External model server/control plane owns | +|---|---|---| +| Provider credentials | Secret resolution, provider-scoped environment, safe status, rotation orchestration boundary | Model-server credentials remain external; Flux may reference a credential alias without reading secret material. | +| Model discovery | Query public/native list and metadata APIs, compile a catalog, retain provenance and staleness | Asset pull, blob store, compilation, conversion, signing, and deletion | +| Model loading | Preflight/readiness and optional explicit control calls | Weights-to-device load, warmup, model registry, keep-alive, KV allocation, LoRA load/unload | +| Request normalization | Provider-neutral DTOs, system/tool/content normalization, warnings, stream normalization | Tokenization, prompt template implementation, grammar enforcement, execution | +| Routing | Provider/deployment selection, retries before output, fallback, circuit breaking, cost/latency policy | Token scheduling, continuous batching, prefill/decode placement, KV transfer, load shedding inside a pool | +| Cache | Exact/semantic response cache and provider prompt-cache controls | Prefix/KV caches, block allocation, offload/recall, cache events, peer transfer | +| Speculation | Optional capability/benchmark metadata | Draft model, proposer, acceptance, target/draft execution | +| Distribution | Deployment catalog and optional route manifests | Node placement, tensor/pipeline/expert parallelism, autoscalers, rollouts, multi-region placement | +| Observability | End-to-end normalized request/stream spans, cost, route, retry, fallback, token usage, optional endpoint signals | TTFT/ITL, queue/KV/slot/batch/server metrics, accelerator telemetry | +| API serving | Existing optional OpenAI-compatible proxy as a convenience | A model server’s own API remains the execution boundary; Flux should not acquire its lifecycle responsibilities | + +This split follows Flux’s accepted host-engine boundary and the explicit model-server/router/control-plane boundaries in llm-d, Dynamo, GAIE, KServe, and Triton ([Flux boundary](https://github.com/GrayCodeAI/flux/blob/main/docs/architecture/HOST-ENGINE-BOUNDARY.md), [llm-d architecture](https://llm-d.ai/docs/architecture), [Dynamo architecture](https://docs.nvidia.com/dynamo/v1.4.0/knowledge-base/overview), [GAIE](https://gateway-api-inference-extension.sigs.k8s.io/), [KServe LLMISVC](https://kserve.github.io/website/docs/model-serving/generative-inference/llmisvc/llmisvc-overview), [Triton](https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/introduction/index.html), accessed 2026-09-24). + +### Recommended additive interfaces + +These are concepts for roadmap discussion, not code changes requested in this research task. + +1. **Keep `core.Provider` unchanged.** Preserve `Chat`, `StreamChat`, `Ping`, and `Name`; Flux’s engine contract is a deliberate host boundary ([core contract](https://github.com/GrayCodeAI/flux/blob/main/provider/core/core.go), local `provider/core/core.go:21-33`, accessed 2026-09-24). Add optional narrow interfaces detected by type assertion: + - `ModelDiscoverer`: list public models/load status. + - `ModelDescriber`: return runtime-specific metadata and native raw JSON. + - `HealthReporter`: structured readiness/liveness/capacity snapshot. + - `MetricsProvider`: per-deployment counters/gauges/histograms. + - `ModelServerController`: explicit list/load/unload/warmup; never part of generation. + - `CapabilityProbe`: run bounded, side-effect-free compatibility checks. +2. **Add an endpoint profile separate from provider ID.** Profiles should encode known server families—OpenAI generic, Ollama native/OpenAI, llama.cpp, vLLM, SGLang, LocalAI, TensorRT-LLM, Triton—plus auth, endpoint path, streaming dialect, model listing, usage semantics, and native metadata routes. LiteLLM and LocalAI both show why a single “OpenAI compatible” label loses important behavior ([LiteLLM architecture](https://docs.litellm.ai/docs/proxy/architecture), [LocalAI discovery](https://localai.io/docs/features/api-discovery/index.html), accessed 2026-09-24). +3. **Add capability evidence, not just booleans.** A capability row should include capability name, state, model/offering/deployment, provider/profile, runtime version, evidence source, config fingerprint, observed-at, expires-at, and raw evidence reference. This generalizes KTransformers’ support tuple and LocalAI/Xinference metadata ([KTransformers matrix](https://ktransformers.net/en/docs/support-matrix), [Xinference custom models](https://inference.readthedocs.io/en/latest/models/custom.html), accessed 2026-09-24). +4. **Add route hints to the normalized request without making them provider wire fields.** Useful optional hints are priority/class, deadline/SLO, expected input/output tokens, session/conversation ID, cache-affinity token, requested adapter/role, and idempotency/correlation ID. The host contract already permits additive metadata and event fields ([Flux DTO compatibility](https://github.com/GrayCodeAI/flux/blob/main/docs/architecture/HOST-ENGINE-BOUNDARY.md), local `docs/architecture/HOST-ENGINE-BOUNDARY.md:188-193`, accessed 2026-09-24). +5. **Add a pluggable routing-signal interface.** Normalize queue depth, running/waiting requests, input/output token load, TTFT/ITL distributions, KV utilization, loaded LoRA IDs, serving role, active models, GPU/backend class, and SLO headroom. This follows llm-d’s filter/score plugins, GAIE’s metrics/capabilities model, and Dynamo’s load/KV-overlap routing ([llm-d scheduler](https://github.com/llm-d/llm-d/blob/main/docs/architecture/core/router/epp/scheduling.md), [GAIE](https://gateway-api-inference-extension.sigs.k8s.io/), [Dynamo](https://docs.nvidia.com/dynamo/v1.4.0/knowledge-base/overview), accessed 2026-09-24). +6. **Add an optional `ModelServerController` admin package.** A safe design can support Ollama, llama.cpp, Triton, and Xinference model list/load/unload. It must use separate credentials, explicit user intent, idempotency, timeouts, and audit events; generation must never trigger an implicit download/load. +7. **Add a stream-stage and recovery contract.** Retrying before the first output is safe; after emitted content/thinking/tool deltas, transparent failover is generally not. Emit typed `retryable_stage` or error context and support optional server resume tokens only when an endpoint explicitly advertises them. This matches llm-d’s mid-stream error categories and Flux’s current behavior, which retries stream setup but not mid-stream errors ([llm-d EPP flow](https://github.com/llm-d/llm-d/tree/main/docs/architecture/core/router/epp), [Flux router](https://github.com/GrayCodeAI/flux/blob/main/router/router.go), local `router/router.go:226-255`, accessed 2026-09-24). +8. **Keep management UI/database state optional.** LiteLLM demonstrates the value and cost of Postgres/Redis-backed multi-tenant gateway state; Flux should offer interfaces and pluggable sinks rather than make them core dependencies ([LiteLLM production deployment](https://docs.litellm.ai/docs/proxy/deploy), accessed 2026-09-24). + +### Capability metadata technique to adopt + +Use a layered model rather than one row per provider/model: + +```text +Model identity + owner / canonical family / revision / digest / license / source +Offering + provider / gateway / deployment / model ID / format / quantization / context +Capability evidence + capability / supported|unsupported|unknown / source / version / observed-at +Runtime profile + server family / version / hardware / backend / parallelism / role / limits +``` + +Specific rules: + +- **Never infer support from model naming alone.** “vision model,” “tool model,” or “thinking model” labels are not evidence. Ollama, llama.cpp, LocalAI, GenieX, and Xinference all show that artifact/backend/template details change capabilities ([Ollama API](https://docs.ollama.com/api/generate), [llama.cpp server](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md), [GenieX platforms](https://geniex.aihub.qualcomm.com/en/get-started/platforms), [Xinference custom models](https://inference.readthedocs.io/en/latest/models/custom.html), accessed 2026-09-24). +- **Preserve unknowns.** Flux already has supported/unsupported/unknown capability states ([capability compiler](https://github.com/GrayCodeAI/flux/blob/main/catalog/v1_defaults.go), local `catalog/v1_defaults.go:187-224,296-304`, accessed 2026-09-24). Extend that pattern to server/runtime capabilities and retain source/confidence. +- **Use native sources first.** Merge embedded catalog, provider live listing, server-native `/props`/show/model metadata, config/manifest, and active probe results. LocalAI’s discovery endpoint, llama.cpp’s props/models, Ollama’s show/tags, Triton’s repository index, and Xinference’s custom manifests are concrete patterns ([LocalAI discovery](https://localai.io/docs/features/api-discovery/index.html), [llama.cpp API](https://llama.app/docs/api), [Ollama API](https://github.com/ollama/ollama/blob/main/docs/api.md), [Triton extension](https://docs.nvidia.com/deeplearning/triton-inference-server/archives/triton-inference-server-2650/user-guide/docs/protocol/extension_model_repository.html), accessed 2026-09-24). +- **Version and expire evidence.** Tool/parser/JSON support changes by runtime release and flags. Store runtime version, test suite version, and configuration fingerprint; do not cache a one-time probe forever. +- **Treat model licenses independently from runtime licenses.** GenieX’s BSD runtime does not grant rights to arbitrary downloaded model weights; TensorRT-LLM’s root Apache license has component exceptions; LiteLLM’s OSS/enterprise split is separate from model licenses ([GenieX license](https://github.com/qualcomm/GenieX/blob/main/LICENSE), [TensorRT-LLM license](https://github.com/NVIDIA/TensorRT-LLM/blob/main/LICENSE), [LiteLLM license](https://github.com/BerriAI/litellm/blob/main/LICENSE), accessed 2026-09-24). + +### Interoperability and conformance technique to adopt + +Build three test layers: + +1. **Wire-level contract tests against mocks:** exact JSON/SSE shapes, error envelopes, cancellation, malformed chunks, unknown events, tool argument deltas, structured-output refusal/parse errors, usage timing, and provider replay blocks. This is deterministic and suitable for every CI run. +2. **Behavioral endpoint profiles:** run the same scenarios through real adapters/mocks representing Ollama, llama.cpp, vLLM, SGLang, LocalAI, LiteLLM, and generic OpenAI. OpenAI’s gpt-oss verification guide demonstrates two-turn tool-call tests plus model evals, while its own warning says smoke tests do not guarantee full compatibility ([OpenAI verification guide](https://developers.openai.com/cookbook/articles/gpt-oss/verifying-implementations), accessed 2026-09-24). +3. **Live conformance suites:** opt-in tests against configured endpoints using capability discovery and small deterministic prompts. Report `PASS`, `UNSUPPORTED`, `FAIL`, `TIMEOUT`, and `NOT_TESTED`; never turn unsupported optional behavior into a hard failure. + +Minimum suite: + +- `/v1/models` discovery and model-ID consistency. +- Chat blocking and SSE; termination event; cancellation; backpressure. +- Stream and non-stream usage semantics, including cache/reasoning token fields. +- Tool call start/delta/done, custom/free-form tools if supported, and a second turn carrying the tool result. +- JSON mode and JSON-schema-constrained output, including schema-invalid provider behavior. +- Vision/audio where advertised, with bounded fixtures. +- Reasoning content and provider replay blocks where advertised. +- Retriable/non-retriable errors, `Retry-After`, rate limits, and auth failures. +- Mid-stream failure semantics and explicit “cannot resume” behavior. +- Unknown additive fields/events are ignored. +- Client profiles for OpenAI SDKs, Agents SDK/Codex-style flows, and the actual Flux host. + +These cases can also be informed by upstream gateway test suites, but Flux should make its own suite versioned, privacy-safe, and authoritative for its own contract rather than claiming universal OpenAI conformance ([LiteLLM repository](https://github.com/BerriAI/litellm), accessed 2026-09-24). + +### Benchmarking technique to adopt + +Flux does not need its own GPU benchmark loop. It needs a reproducible client benchmark and result schema that can target any OpenAI-compatible server. + +Adopt these patterns from AIPerf and MLPerf: + +- Control input/output length distributions, concurrency and request-rate modes, warmup, duration, grace period, and streaming mode. AIPerf supports synthetic distributions, trace replay, soak testing, and raw JSONL exports ([AIPerf repository](https://github.com/ai-dynamo/aiperf), [time-based benchmarking](https://github.com/ai-dynamo/aiperf/blob/5ad08166/docs/tutorials/time-based-benchmarking.md), accessed 2026-09-24). +- Report TTFT, time to first non-reasoning output, inter-token latency/TPOT distributions, end-to-end latency, output/system throughput, success rate, and **goodput** under explicit SLO thresholds. AIPerf’s reasoning-token guidance shows why TTFT and first-output-token must be distinct ([AIPerf migration guide](https://docs.nvidia.com/aiperf/getting-started/migrating-from-gen-ai-perf), accessed 2026-09-24). +- Measure prefix/KV reuse with privacy-preserving trace synthesis and user-centric multi-turn timing; optionally collect server metrics/GPU telemetry. AIPerf documents KV efficiency, TTL, Prometheus, and GPU telemetry tests ([AIPerf comprehensive guide](https://docs.nvidia.com/aiperf/getting-started/ai-perf-comprehensive-llm-benchmarking), accessed 2026-09-24). +- Preserve quality constraints. MLPerf ties performance to a dataset, quality metric, and threshold; its server/interactive scenarios measure TTFT and TPOT under different load patterns ([MLPerf Llama3.1-8B](https://mlcommons.org/2025/09/small-llm-inference-5-1/), [MLPerf v6.1 analysis](https://mlcommons.org/2026/09/chairs-mlperf-inference-v6-1/), accessed 2026-09-24). +- Report the full system boundary: client, Flux version, adapter/profile, server/runtime version, model digest/quantization, hardware, prompt/temperature/seed, batch/scheduling flags, prefix-cache setting, and dataset/trace version. Otherwise results are not comparable. +- Make project benchmarks explicitly “upstream claim” until Flux reproduces them on a named system. + +### Observability technique to adopt + +Flux has real OpenTelemetry integration in `provider/observability/tracing.go` and a separate older in-memory telemetry implementation with custom `llm.*` attributes ([tracing source](https://github.com/GrayCodeAI/flux/blob/main/provider/observability/tracing.go), local `provider/observability/tracing.go`; [legacy telemetry](https://github.com/GrayCodeAI/flux/blob/main/internal/observability/observability.go), local `internal/observability/observability.go:1-42`, accessed 2026-09-24). The next layer should standardize output and correlate with endpoint metrics rather than replace tracing. + +Recommended normalized span/metric vocabulary: + +- Request spans: `gen_ai.operation.name`, provider/deployment/model identifiers, requested vs served model, attempt number, route strategy, retry/fallback reason, and normalized error class. +- Token/cost attributes: input, output, total, cached input, cache-write, reasoning tokens, and cost with currency/unit/source. +- Latency: queue, connection, provider TTFT, first output token, inter-token distribution, total latency, and continuation latency. +- Route events: route selected/changed, prefill/decode or variant selected, cache-affinity outcome, fallback, load shedding, and mid-stream failure. +- Runtime resource signals, when available: queue depth, running/waiting requests, tokens queued, KV utilization, loaded adapters, batch size, model load state, and server version. +- Privacy: do not put prompts, responses, tool arguments, or raw secret-bearing headers into telemetry. Use digests/opaque correlation IDs, consistent with Flux’s operations-graph privacy model ([operations graph](https://github.com/GrayCodeAI/flux/blob/main/README.md), local `README.md:329-332`, accessed 2026-09-24). + +Adopt the GrayCodeAI shared OTEL conventions referenced by Flux rather than inventing another `llm.*` vocabulary ([Flux contributor guidance](https://github.com/GrayCodeAI/flux/blob/main/AGENTS.md), accessed 2026-09-24). Generate low-cardinality Prometheus metrics separately; never put model IDs, request IDs, tenant IDs, or session IDs into unbounded metric labels. + +### Prioritized additive roadmap for Flux + +#### P0: correctness and interoperability + +1. **Versioned capability evidence and endpoint profiles** for generic OpenAI, Ollama, llama.cpp, vLLM, SGLang, and LocalAI. +2. **Mock-first conformance suite** for streaming, tools, second turns, structured output, usage, errors, cancellation, and unknown fields. +3. **Native runtime metadata importers** (`/api/show`, `/props`, `/v1/models`, server manifests) while preserving raw namespaced metadata. +4. **OTel semantic-convention alignment** and request/route/error correlation. +5. **Explicit stream recovery semantics** and regression tests for pre-first-byte versus mid-stream failures. + +#### P1: operations and routing + +1. **Endpoint signal snapshot** and pluggable routing scorer inputs. +2. **Route hints** for priority, SLO, expected tokens, session, and cache affinity. +3. **Per-deployment health/readiness detail** with model-loaded state and safe native error classes. +4. **Benchmark command or external AIPerf profile** producing versioned JSON results and conformance reports. +5. **Optional server-profile documentation** covering capability differences, known quirks, and minimum versions. + +#### P2: optional external control-plane adapters + +1. Separate admin-only clients for model list/load/unload on llama.cpp, Ollama, Triton, and Xinference. +2. Bridge/metrics exporter for GAIE, llm-d EPP, AIBrix, or Dynamo signals without importing their APIs into the stable engine DTO. +3. Resume-token support only where a server contract explicitly guarantees safe continuation. +4. Signed remote runtime manifests with last-good persistence, continuing the current signed-manifest direction but not adding consensus or autoscaling. + +### Explicit non-goals + +- Do not add model weight pull/delete/compile/sign operations to the generation path. +- Do not add token batching, continuous batching, paged KV, speculative decoding, CUDA/ROCm kernels, tensor/pipeline/expert parallelism, or prefill/decode workers. +- Do not make Postgres, Redis, Kubernetes, Ray, KServe, llm-d, Dynamo, or Triton mandatory dependencies. +- Do not make the optional HTTP proxy Flux’s primary identity; the library/engine contract should remain sufficient without serving a daemon. +- Do not copy SGLang Model Gateway’s agent history, MCP execution, or product UX into the provider runtime. +- Do not report a server feature as supported until the exact model/artifact/runtime/hardware/configuration evidence is recorded. + +### Final inferences + +- Flux already has a defensible provider-runtime core: stable host contract, normalized streams, tools/multimodal/structured output, credentials, catalog/discovery, route/fallback, health, and telemetry. The gap is not “more provider features”; it is evidence and interoperability around external server profiles. +- LiteLLM is the main feature peer, llama.cpp/Ollama/LocalAI/Xinference are the main local-runtime peers, and llm-d/Dynamo/GAIE/AIBrix/KServe define the external control-plane boundary. +- The best additive Flux differentiator is a small, pure-Go, host-neutral runtime that can describe and route to heterogeneous servers with evidence-backed capabilities and reproducible conformance—not owning model execution. +- A small number of optional interfaces can cover the useful missing layers while keeping the stable `Provider` contract unchanged and preserving Flux’s one-way dependency boundary. + +### Gaps and risks + +- Exact support matrices change quickly. Every adapter should pin the server/runtime version and fixture revision used for capability evidence. +- Upstream benchmark claims are not directly comparable; no common live run was performed here. +- Some official docs are internally inconsistent or ahead of repository configuration, notably MLC’s version label, Ray’s advertised versus configured engine backends, and Dynamo’s experimental sidecar coverage. +- A future report should verify latest releases again on the publication date and preserve this 2026-09-24 snapshot for reproducibility. diff --git a/research_notes/Flux OSS landscape roadmap/oss_gateways.md b/research_notes/Flux OSS landscape roadmap/oss_gateways.md new file mode 100644 index 00000000..32fe69a6 --- /dev/null +++ b/research_notes/Flux OSS landscape roadmap/oss_gateways.md @@ -0,0 +1,371 @@ +# OSS LLM/API gateway landscape and selective Flux roadmap — 2026-09-24 + +## 1. Verified landscape, boundaries, and exclusions + +### Takeaway + +The active OSS field is larger than the proposed candidate list. The closest product-level peers to Flux are **LiteLLM, Portkey Gateway, Bifrost, Helicone AI Gateway, New API, and LLM Gateway (`theopenco/llmgateway`)**. The strongest infrastructure/control-plane comparators are **Agent Router (formerly Envoy AI Gateway), Kong AI Gateway, Apache APISIX, and Higress**. **OpenZiti LLM Gateway** is a useful but much smaller Go implementation comparator. None combines Flux's current combination of an embeddable Go runtime, a narrow host facade, normalized streaming events, provider-specific replay state, and no mandatory UI/control plane. + +### Cited Findings + +**Evidence labels used below** + +- **Repo/API-confirmed:** public source tree, GitHub metadata, code, tests, or release artifact inspected directly. +- **Official-docs:** current first-party documentation; treated as implementation intent, not proof of every deployed configuration. +- **Project claim:** provider/model counts, throughput, latency, or availability published by the project; not independently reproduced. +- **Independent evidence:** vulnerability research, foundation announcements, or acquisition announcements from a source other than the project vendor. +- All web sources below were accessed on **2026-09-24**. Counts are a same-day snapshot and are not a quality ranking. + +#### Current project metadata and OSS boundary + +| Project | Current repository state | Language / license | Latest published release captured | Architectural boundary | +|---|---|---|---|---| +| LiteLLM | 59,489 stars; not archived; source pushed 2026-09-23. **Repo/API-confirmed.** [GitHub API](https://api.github.com/repos/BerriAI/litellm) | Python today, with a staged Rust core; core advertised as MIT, with commercial enterprise features. [OSS page](https://www.litellm.ai/oss) · [feature comparison](https://www.litellm.ai/features) | `v1.102.1`, 2026-09-23. [Release](https://github.com/BerriAI/litellm/releases/tag/v1.102.1) | Python SDK plus self-hosted proxy; optional PostgreSQL/Redis and a separately deployable UI/backend. | +| Portkey Gateway | 13,069 stars; not archived; default-branch source last pushed 2026-05-25. **Repo/API-confirmed.** [GitHub API](https://api.github.com/repos/Portkey-AI/gateway) | TypeScript, MIT. [Repository](https://github.com/Portkey-AI/gateway) | Stable OSS `v1.15.2`, 2026-01-12; `2.0.0` remains labeled pre-release and its branch was last committed 2026-03-14. [Stable release](https://github.com/Portkey-AI/gateway/releases/tag/v1.15.2) · [issue asking for 2.0 status](https://github.com/Portkey-AI/gateway/issues/1606) | Edge-friendly gateway process, with optional connection to Portkey's control plane and separate enterprise private deployment. | +| Bifrost | 8,286 stars; not archived; source pushed 2026-09-23. **Repo/API-confirmed.** [GitHub API](https://api.github.com/repos/maximhq/bifrost) | Go, Apache-2.0. [Repository](https://github.com/maximhq/bifrost) | GitHub's latest tag is `helm-chart-v2.1.43`; HTTP transport `2.1.0` contains the September security fix. [Releases](https://github.com/maximhq/bifrost/releases) · [security fix PR](https://github.com/maximhq/bifrost/pull/6757). **Re-snapshot 2026-09-27 (GitHub releases API):** HTTP transport `transports/v2.2.3` (2026-09-24; adds provider keys pinned per routing fallback) and Core `core/v1.10.4` (2026-09-25). [HTTP v2.2.3](https://github.com/maximhq/bifrost/releases/tag/transports/v2.2.3) | Go HTTP gateway or in-process SDK, with built-in UI; clustering and several governance/guardrail features are enterprise additions. | +| Helicone AI Gateway | 631 stars; not archived; gateway source last pushed 2025-11-21, while the main observability repository was pushed 2026-09-16. **Repo/API-confirmed.** [Gateway API](https://api.github.com/repos/Helicone/ai-gateway) · [observability API](https://api.github.com/repos/Helicone/helicone) | Rust gateway. The actual `LICENSE` and GitHub metadata say GPL-3.0, while the README says Apache. [License](https://github.com/Helicone/ai-gateway/blob/main/LICENSE) · [contradictory README](https://github.com/Helicone/ai-gateway#-license) | No GitHub “latest release”; public beta. [Releases](https://github.com/Helicone/ai-gateway/releases) | Lightweight Rust data plane, optionally joined to Redis/S3 and the separately licensed Apache-2.0 Helicone observability platform. | +| Agent Router, formerly Envoy AI Gateway | 2,131 stars; not archived; source pushed 2026-09-21. **Repo/API-confirmed.** [GitHub API](https://api.github.com/repos/theagentrouter/agent-router) | Go control plane/ext-proc, Apache-2.0. [Repository](https://github.com/theagentrouter/agent-router) | `v1.1.0`, 2026-08-21. [Release](https://github.com/theagentrouter/agent-router/releases/tag/v1.1.0) | Agent Router controls; Envoy/Envoy Gateway carries. Kubernetes CRDs are the normal deployment, with a standalone `aigw run` mode added in 2026. | +| Kong AI Gateway | Kong repository: 44,184 stars; not archived; source pushed 2026-09-23. **Repo/API-confirmed.** [GitHub API](https://api.github.com/repos/Kong/kong) | Kong Gateway is Lua, Apache-2.0. [Repository](https://github.com/Kong/kong) | Kong repository `3.9.3`; the separate AI Gateway runtime changelog reached `2.0.3` on 2026-08-31. [Kong release](https://github.com/Kong/kong/releases/tag/3.9.3) · [AI Gateway changelog](https://developer.konghq.com/ai-gateway/changelog/) | Open-source Kong Gateway/plugins versus AI Gateway 2.x's dedicated runtime, entity model, Konnect control plane, and analytics pipeline. The public `Kong/kong` source does not establish that the entire 2.x product is OSS. | +| Apache APISIX AI Gateway | 17,161 stars; not archived; source pushed 2026-09-22. **Repo/API-confirmed.** [GitHub API](https://api.github.com/repos/apache/apisix) | Lua/LuaJIT/OpenResty, Apache-2.0, ASF governance. [Repository](https://github.com/apache/apisix) | `3.18.0`, 2026-08-20. [Release](https://github.com/apache/apisix/releases/tag/3.18.0) | Generic API gateway with AI plugins; etcd-backed dynamic configuration in clustered mode, with a standalone mode. | +| Higress | 9,448 stars; not archived; source pushed 2026-09-23. **Repo/API-confirmed.** [GitHub API](https://api.github.com/repos/higress-group/higress) | Go plus Istio/Envoy data plane; Apache-2.0; CNCF Sandbox. [Repository](https://github.com/higress-group/higress) · [CNCF project page](https://www.cncf.io/projects/higress) | `v2.2.4`, 2026-08-13. [Release](https://github.com/higress-group/higress/releases/tag/v2.2.4) | Cloud-native data plane plus console and Wasm plugin system; accepted into CNCF Sandbox on 2026-03-15. [CNCF](https://www.cncf.io/projects/higress) | +| New API | 48,782 stars; not archived; source pushed 2026-09-23. **Repo/API-confirmed.** [GitHub API](https://api.github.com/repos/QuantumNous/new-api) | Go/Gin backend, React/Bun frontend, AGPL-3.0 plus Section 7 UI attribution/link terms. [Repository](https://github.com/QuantumNous/new-api) · [license](https://github.com/QuantumNous/new-api/blob/main/LICENSE) | `v1.0.0-rc.40`, 2026-09-21. [Release](https://github.com/QuantumNous/new-api/releases/tag/v1.0.0-rc.40) | Unified gateway, web console, user/group/tenant management, billing/quota state, and task-plugin runtime in one deployable product. | +| LLM Gateway (`theopenco/llmgateway`) | 1,658 stars; not archived; source pushed 2026-09-23. **Repo/API-confirmed.** [GitHub API](https://api.github.com/repos/theopenco/llmgateway) | TypeScript; core AGPL-3.0, `ee/` commercial. [Repository](https://github.com/theopenco/llmgateway) | `v1.18.0`, 2026-09-21. [Release](https://github.com/theopenco/llmgateway/releases/tag/v1.18.0) | Unified container with API gateway, UI/playground, PostgreSQL, Redis, and commercial enterprise administration. | +| OpenZiti LLM Gateway | 96 stars; not archived; source pushed 2026-09-15. **Repo/API-confirmed.** [GitHub API](https://api.github.com/repos/openziti/llm-gateway) | Go, Apache-2.0. [Repository](https://github.com/openziti/llm-gateway) | `v0.1.7`, 2026-08-12. [Release](https://github.com/openziti/llm-gateway/releases/tag/v0.1.7) | Single Go binary; optional OpenZiti/zrok overlay for private transport. | + +#### Candidate verification and classification + +- **Direct product peers:** LiteLLM, Portkey, Bifrost, and Helicone most directly compete on normalized provider access, reliability, caching, and observability. New API and LLM Gateway are also OSS product peers, but their center of gravity is a web console, tenancy, quotas, and hosted-product migration rather than an embeddable runtime. Their repositories confirm those boundaries. [LiteLLM](https://github.com/BerriAI/litellm) · [Portkey](https://github.com/Portkey-AI/gateway) · [Bifrost](https://github.com/maximhq/bifrost) · [Helicone](https://github.com/Helicone/ai-gateway) · [New API](https://github.com/QuantumNous/new-api) · [LLM Gateway](https://github.com/theopenco/llmgateway) +- **Infrastructure peers:** Agent Router, Kong, APISIX, and Higress are built around an API gateway/control-plane architecture. They are excellent sources of routing, policy, protocol, and operational patterns, but adopting their control planes would conflict with Flux's stated host-neutral runtime identity. [Agent Router](https://github.com/theagentrouter/agent-router) · [Kong](https://github.com/Kong/kong) · [APISIX](https://github.com/apache/apisix) · [Higress](https://github.com/higress-group/higress) +- **Niche implementation comparator:** OpenZiti is too small to be a market leader, but its single-binary Go design, OpenZiti transport, weighted endpoint balancing, and optional three-stage semantic router are concrete patterns worth studying. [Repository](https://github.com/openziti/llm-gateway) +- **Material 2026 governance change:** Envoy AI Gateway was renamed **Agent Router** and joined the Agentic AI Foundation on 2026-09-09. The announcement says the 1.x compatibility commitment, CRDs, images, module path, maintainers, and Apache-2.0 license remain unchanged. [AAIF announcement](https://aaif.io/blog/agent-router-joins-aaif) · [migration-preserving README](https://github.com/theagentrouter/agent-router#formerly-envoy-ai-gateway) +- **Portkey governance change:** Palo Alto Networks completed its Portkey acquisition on 2026-05-29 and is positioning it inside Prisma AIRS. This may improve enterprise security integration but makes corporate roadmap direction more important to self-hosted users. **Independent evidence.** [Palo Alto Networks announcement](https://www.paloaltonetworks.com/company/press/2026/palo-alto-networks-completes-acquisition-of-portkey-to-secure-ai-agents) + +#### Current-state and OSS-boundary warnings + +- **Portkey's “fully open source” claim is not matched by a current stable 2.0 artifact.** The March 2026 announcement says the production gateway was open-sourced, but the default README still labels 2.0 pre-release; stable `v1.15.2` dates to January, and the 2.0 branch's last commit found was 2026-03-14. Treat the 1.x OSS gateway and the 2.0 control-plane feature set separately. **Project claim versus repo-confirmed state.** [Announcement](https://portkey.ai/blog/gateway-2-0) · [repository](https://github.com/Portkey-AI/gateway) · [2.0 status issue](https://github.com/Portkey-AI/gateway/issues/1606) +- **Helicone's gateway license is internally contradictory.** GitHub metadata, the license file, and the sidebar identify GPL-3.0; the README says Apache. Current docs also call the old cloud gateway “legacy” and say it is being phased out in favor of a future cloud offering based on the self-hosted gateway. [License](https://github.com/Helicone/ai-gateway/blob/main/LICENSE) · [README](https://github.com/Helicone/ai-gateway#-license) · [current introduction](https://docs.helicone.ai/ai-gateway/introduction) +- **Bifrost demonstrates the cost of merging a management plane with an execution plane.** CVE-2026-90898 allowed unauthenticated stdio MCP client registration, and therefore command execution, when management authentication was disabled; transport `2.1.0` rejects unauthenticated stdio registration. **Independent evidence, corroborated by the project fix.** [JFrog analysis](https://research.jfrog.com/vulnerabilities/bifrost-is-vulnerable-to-unauthenticated-remote-code-execution-via-mcp-stdio-client-registration-cve-2026-90898/) · [fix PR](https://github.com/maximhq/bifrost/pull/6757) +- **Kong is open-core at the product boundary.** `Kong/kong` and the AI plugins are Apache-2.0, while AI Gateway 2.x documentation describes a dedicated runtime, Konnect control plane, first-class entities, and Konnect analytics. The public Kong repository confirms the gateway/plugin code, not the full 2.x management product. [Open-source repository](https://github.com/Kong/kong) · [AI Gateway 2.x documentation](https://developer.konghq.com/ai-gateway/) + +#### Notable exclusions from the core OSS set + +- **OpenRouter:** a major hosted marketplace/gateway, but its own comparison says “Open source: No” and “Self-hostable: No”; its public GitHub organization does not expose the gateway server. [Official comparison](https://openrouter.ai/blog/insights/llm-gateway/) · [GitHub organization](https://github.com/openrouter) +- **Cloudflare AI Gateway:** the service is managed through Cloudflare accounts and APIs. Cloudflare's `ai-gateway-provider` package is OSS client integration code, not the gateway data plane/control plane. [Service docs](https://developers.cloudflare.com/ai-gateway/) · [OSS client package](https://github.com/cloudflare/ai/tree/main/packages/ai-gateway-provider) +- **Vercel AI Gateway:** the AI SDK and `@ai-sdk/gateway` provider are OSS, but the routed service, billing, budgets, and management API are delivered through Vercel. [Gateway docs](https://vercel.com/docs/ai-gateway) · [OSS client package](https://github.com/vercel/ai/tree/main/packages/gateway) +- **Martian Gateway:** official docs expose a managed 200+ model gateway, while the `withmartian` GitHub organization publishes ARES, TensorZero, and research tools but not the Gateway server implementation. [Gateway docs](https://gateway-docs.withmartian.com/gateway) · [GitHub organization](https://github.com/withmartian) +- **TrueFoundry AI Gateway:** its organization publishes infrastructure charts, SDKs, model metadata, and the separate TrueForge agent runtime, but not the LLM/MCP/Agent Gateway server. [GitHub organization](https://github.com/TrueFoundry) · [product description](https://www.truefoundry.com/ai-gateway) +- **One API:** the MIT project is not archived and received source pushes through 2026-01-09, but New API identifies it as its upstream base and is the materially more active successor. [One API metadata](https://api.github.com/repos/songquanpeng/one-api) · [New API repository](https://github.com/QuantumNous/new-api) + +### Inferences + +- The market has converged around a common data-plane core—provider adapters, normalized APIs, retries/fallbacks, caching, telemetry, and policy—but differentiates through either a **developer control plane** (LiteLLM, Portkey, Bifrost, New API) or an **infrastructure control plane** (Agent Router, Kong, Higress, APISIX). +- Flux should compete on **runtime correctness, protocol fidelity, embeddability, and migration quality**, not raw provider counts, UI breadth, tenant administration, or a general-purpose API gateway. +- Repository health and OSS boundaries are first-class competitive dimensions in 2026: Portkey 2.0, Helicone's license/activity split, Bifrost's management-plane RCE, and Kong's 2.x boundary are concrete reasons to avoid assuming that a feature page equals a stable, auditable runtime. +- “Open-source gateway” is not one category. Flux can selectively borrow data-plane techniques while explicitly rejecting the product layers that would turn it into a SaaS or enterprise control plane. + +### Gaps + +- GitHub stars, model counts, and provider counts are volatile and inconsistent. Portkey's own repository alternates between “250+ LLMs” and “1,600+ language, vision, audio, and image models”; this analysis does not normalize those marketing counts into an adapter count. [Repository](https://github.com/Portkey-AI/gateway) +- No independent, common-harness performance or conformance benchmark was found across the projects. Vendor throughput/latency claims are therefore not used to rank them. +- The public source boundary of Kong AI Gateway 2.x's dedicated runtime/control plane could not be established from `Kong/kong`; it should be treated as not verified rather than definitively closed source. +- This is a source/documentation landscape review, not a line-by-line security audit of every project. + +## 2. Project-by-project evidence and Flux-relevant implementation patterns + +### Takeaway + +The most valuable patterns for Flux are narrow and technical: **operation-level capability matrices, raw-response inspection, independently testable protocol translation, policy revalidation on every fallback, explicit stream deadlines, sticky/P2C routing, staged transformation parity checks, and phase-based middleware**. The least suitable patterns are organization administration, web consoles, billing, mandatory databases/Kubernetes control planes, MCP/A2A registries, vector stores, memory, skills marketplaces, and a general plugin ecosystem. + +### Cited Findings + +#### Cross-project capability map + +| Project | Provider/model and normalized API | Streaming, tools, structured output, multimodal, embeddings, batches | Routing and policy | Runtime/deployment | Flux reading | +|---|---|---|---|---|---| +| LiteLLM | Project claims 100+ provider APIs. OpenAI format, native `/v1/messages`, pass-through routes, Responses, completions, embeddings, audio, images, video, rerank, realtime, MCP, and more. **Project claim + official docs.** [Supported endpoints](https://docs.litellm.ai/docs/supported_endpoints) | Streaming and provider tools are core SDK behavior; structured-output fidelity remains provider-dependent. Native Anthropic, Responses, batch, embeddings, image/audio/video, realtime, and pass-through routes are explicitly listed. | Weighted routing, retries, cooldowns, health checks, ordered fallback, context/content-policy fallback, request/key budgets, and spend tracking. [Reliability docs](https://docs.litellm.ai/docs/proxy/reliability) | Python/FastAPI today; staged Rust translation/router/server; Docker/Kubernetes; PostgreSQL and Redis for full governance/state. [Rust migration](https://docs.litellm.ai/blog/litellm-rust-launch) · [deployment](https://docs.litellm.ai/docs/proxy/deploy) | Best reference for API breadth and policy-correct fallback. Avoid its product/UI/DB center of gravity. | +| Portkey | Project claims 45+ providers and 1,600+ models, with OpenAI-compatible chat plus vision/audio/image and realtime. Counts conflict with the README's “250+ LLMs,” so only the provider surface—not the count—is reliable. [Repository](https://github.com/Portkey-AI/gateway) | Streaming is marked for listed providers. Tools and structured output are not given a complete cross-provider fidelity matrix. Batching and fine-tuning are marketed capabilities; their exact OSS 1.x/2.0 implementation boundary is unclear. | JSON configs for retry, fallback, load balancing, conditional routing, timeouts, sticky routing, guardrails, and caching; policy metadata is attached per request. [Conditional routing](https://portkey.ai/docs/product/ai-gateway/conditional-routing.md) | TypeScript/Node; `npx`, Docker, Cloudflare Workers, Kubernetes/Helm; optional control-plane connection. | Borrow declarative route configs, sticky affinity, and error-specific fallback. Avoid the guardrail catalog and UI/control-plane expansion. | +| Bifrost | Official docs say 20+ providers; repository description says 23+. OpenAI-compatible and provider-compatible façades, with an operation-by-provider matrix and raw-response mode. [Overview](https://docs.getbifrost.ai/overview) · [matrix](https://docs.getbifrost.ai/providers/supported-providers/overview) | Matrix covers streaming chat/responses, images, embeddings, TTS/STT, files, batch, rerank, OCR, video, containers, and passthrough. Tool and structured-output support vary by provider. | Automatic fallback, weighted key distribution, virtual keys, hierarchical budgets/rate limits, semantic caching, Prometheus/OTLP, and Go/WASM middleware. | Go gateway, built-in web UI, or Go SDK; enterprise P2P clustering, service discovery, and gossip state. [Clustering](https://docs.getbifrost.ai/enterprise/clustering) | Closest Go technical peer. Borrow operation matrices, raw fidelity, and narrow middleware—not MCP process hosting or clustering. | +| Helicone | Rust gateway, OpenAI syntax, project claim of 100+ models and 20+ providers; provider list is embedded YAML. [Repository](https://github.com/Helicone/ai-gateway) | Streaming is core. The cited gateway README does not establish complete tools, structured output, embeddings, or batch parity. | P2C + PeakEWMA latency routing, model-latency, weighted and cost strategies; request/token/dollar rate limits; Redis/S3 response cache. | Lightweight binary; optional Redis/S3 and Helicone platform. Main gateway repository is stale relative to the active observability repository. | Borrow low-coordination P2C/PeakEWMA and sidecar separation. Do not couple Flux runtime health to a vendor analytics service. | +| Agent Router | 16 providers, one OpenAI-compatible API, native Anthropic Messages, cross-provider translation, Responses, embeddings, image and audio endpoints, and MCP. [Provider docs](https://theagentrouter.ai/docs/capabilities/llm-integrations/supported-providers) · [1.0 announcement](https://tetrate.io/blog/envoy-ai-gateway-v1-0-release) | Streaming, multimodal image/audio/video, reasoning, tool/MCP authorization, guided output, prompt caching, Responses, and per-provider reasoning-token accounting are documented. | Cross-backend retry/failover, dynamic endpoint picking, model virtualization, quota-aware rate limiting, prompt caching, per-request credentials, and stream idle timeout with failover. [Release notes](https://theagentrouter.ai/release-notes/) | Go controller/ext-proc configures Envoy; Kubernetes Gateway API/CRDs, Helm, or standalone `aigw run`; two-tier hosted/self-hosted inference pattern. | Borrow stream idle timeout, explicit provider extension fields, stable compatibility policy, and a one-command local mode. Avoid the Kubernetes control plane and MCP scope. | +| Kong | Provider-agnostic API over OpenAI, Anthropic, Azure, Bedrock, Gemini, Vercel, and others; exact operation parity is plugin/provider specific. [AI Gateway docs](https://developer.konghq.com/ai-gateway/) | Streaming, MCP, A2A, semantic cache, prompt transforms, semantic routing, guardrails, token accounting, and GenAI OTel are documented; no single authoritative cross-provider operation matrix was found. | Consistent hash, lowest latency, usage, round robin, semantic matching; retries/fallback, token quotas, auth/secrets, transforms, guardrails, metering/billing. | Lua/OpenResty plugins; Konnect, self-hosted, hybrid, DB-less, and Kubernetes. AI Gateway 2.x adds entities, policies, dedicated runtime, and control plane. | Borrow phase ordering and migration tooling. Avoid becoming a general API gateway or Konnect-style entity/control plane. | +| APISIX | `ai-proxy` targets OpenAI, DeepSeek, Azure, Anthropic, OpenRouter, Gemini, Vertex, Bedrock, and compatible APIs; `ai-proxy-multi` adds multi-provider routing. [AI proxy](https://apisix.apache.org/docs/apisix/plugins/ai-proxy) | LLM and embedding proxying; streaming, tools, and structured output are primarily passthrough/provider-dependent. MCP stdio-to-HTTP/SSE bridge exists. [AI Gateway](https://apisix.apache.org/ai-gateway/) | Weighted/consistent-hash/semantic routing, retries/fallback, health checks, token rate limits, exact/semantic cache, RAG, moderation, prompt transforms, and request/response logging. | OpenResty/LuaJIT data plane; etcd dynamic config, Admin API, hot plugin reload, standalone YAML; 100+ plugins. [Repository](https://github.com/apache/apisix) | Borrow lifecycle phases and local-first behavior. Avoid etcd and generic gateway plugin sprawl. | +| Higress | Project claims 100+ models through a unified OpenAI-compatible protocol, plus native MCP hosting and Gateway API Inference Extension support. [AI Gateway](https://higress.ai/en/ai-gateway) | Streaming SSE is first-class in Wasm plugins; tools and structured output are provider/pass-through dependent. MCP can be hosted or generated from OpenAPI. | Multi-model load balancing/fallback, token limits, semantic cache, auth, content protection, observability, and Wasm plugins. | Go control/data-plane components on Istio/Envoy; Docker all-in-one or Helm; Wasm plugins in Go/Rust/JS; console and container registry. [Repository](https://github.com/higress-group/higress) | Borrow genuinely streaming plugin transforms and protocol adapters. Do not adopt Wasm or a console before core parity. | +| New API | OpenAI Chat/Responses, native Anthropic Messages, Gemini generate/stream, realtime, image/audio, embeddings, rerank, and task plugins. Protocol conversion is isolated in an independently buildable Go `relaykit` module. [Repository](https://github.com/QuantumNous/new-api) | Streaming, tools, reasoning, and multimodal are documented where channel/model support permits; wide endpoint coverage is source-confirmed. | Channel priority/weight, retries, affinity, multiple keys, quotas/subscriptions, expression pricing, cache accounting, users/groups/permissions, OAuth/OIDC/passkeys/2FA. | Go/Gin + React, SQLite/MySQL/PostgreSQL, optional Redis and ClickHouse, JS task plugins, Docker/Electron, built-in console. | Borrow `RelayKit` separation and task-plugin host boundaries. Avoid the full tenancy/billing/UI product. | +| LLM Gateway | Project docs state 200+ models/40+ providers and now also implement the AI SDK gateway protocol. OpenAI, Anthropic, embeddings, rerank, image/audio/video, realtime, and AI SDK surfaces are documented. [Docs](https://docs.llmgateway.io/) · [AI SDK protocol note](https://llmgateway.io/changelog/ai-sdk-gateway-protocol) | Streaming, tools, image/video, audio, realtime, embeddings, and rerank are documented; support varies by provider. | Automatic provider failover, caching, per-key/project/team spend, TTL keys, provider compliance filters, usage analytics, and coding-agent launcher. [Changelog](https://llmgateway.io/changelog) | TypeScript/Hono/Next unified image; PostgreSQL and Redis; AGPL core plus commercial `ee/`. | Borrow migration and coding-agent launch UX, not the hosted platform/control plane. | +| OpenZiti | OpenAI Chat Completions only, with transparent OpenAI↔Anthropic translation and any OpenAI-compatible local backend. [Repository](https://github.com/openziti/llm-gateway) | SSE streaming; no cited tools, structured output, embeddings, or batch surface. | Weighted round robin, active/passive health, VM sleep detection, passive failover, and optional heuristic→embedding→LLM classifier semantic routing. | One Go binary, Prometheus metrics, optional zrok/OpenZiti overlay for private ingress/egress. | Borrow semantic routing only as an explicit, host-selected strategy. Avoid network-overlay scope and premature model auto-selection. | + +#### LiteLLM + +- **API and model surface — official docs:** LiteLLM exposes a very broad endpoint catalog: Chat/Completions, Responses, Anthropic Messages and pass-through, embeddings, images/edits, audio, video, files, batches, fine-tuning, rerank, realtime/WebRTC, evals, MCP, vector stores, memory, search, and agent-oriented endpoints. This is the most complete public endpoint index in the verified set. [Supported endpoints](https://docs.litellm.ai/docs/supported_endpoints) +- **Provider extensions and fidelity:** Native Anthropic and provider pass-through routes let clients keep provider-specific request shapes. That is strategically important because a universal DTO cannot represent every provider-native search, reasoning, cache, or tool feature without loss. [Anthropic unified endpoint](https://docs.litellm.ai/docs/anthropic_unified/) · [pass-through docs](https://docs.litellm.ai/docs/pass_through/intro) +- **Routing and policy — official docs/source:** Fallback happens after retries and can be ordered across model groups. LiteLLM distinguishes content-policy and context-window failures, supports cooldowns and health routing, and records original model group plus attempted fallbacks in spend logs. [Reliability docs](https://docs.litellm.ai/docs/proxy/reliability) +- **Security/policy correctness worth copying:** Current LiteLLM docs explicitly recheck fallback targets against the caller's model access groups and budget, preventing a free primary model from failing over to an unauthorized or unaffordable paid target. This is a concrete implementation rule, not a marketing claim. [Fallback access enforcement](https://docs.litellm.ai/docs/proxy/reliability#enforce-key-model-access-on-fallbacks) · [fallback budget enforcement](https://docs.litellm.ai/docs/proxy/reliability#enforce-budget-on-fallbacks) +- **Implementation migration — official docs:** The Rust migration deliberately moves one route at a time, uses flag-gated PyO3 transforms, and requires parity checks before enabling a Rust path. The Python shell retains I/O, auth, callbacks, database, and spend while the Rust core performs pure transforms. [Rust migration](https://docs.litellm.ai/blog/litellm-rust-launch) +- **Operational boundary — official docs:** Monolithic and microservices deployments are supported; PostgreSQL is needed for auth/tracking and Redis becomes necessary for multi-instance rate limits/cache/router state. The full product is materially more than a data-plane library. [Production deployment](https://docs.litellm.ai/docs/proxy/deploy) +- **Observability/evaluation/governance:** LiteLLM has callbacks/integrations, Prometheus metrics, spend logs, virtual keys, budgets, guardrail hooks, and an Evals endpoint. SSO, SCIM, audit logs, and some advanced governance are commercial. [Feature comparison](https://www.litellm.ai/features) · [Evals endpoint](https://docs.litellm.ai/docs/evals_api) +- **Flux implication:** Copy the endpoint breadth, provider pass-through escape hatch, fallback policy revalidation, attempt metadata, and parity-test discipline. Do not copy the Python/control-plane/UI product shape or its long endpoint tail before Flux's core chat compatibility is lossless. + +#### Portkey Gateway + +- **API and model surface — repo-confirmed:** Portkey's stable OSS gateway is OpenAI-oriented and supports chat, vision, audio/image generation, and OpenAI realtime. Its README says 45+ providers and 1,600+ models, while another heading says 250+ LLMs; those counts are project claims and internally inconsistent. [Repository](https://github.com/Portkey-AI/gateway) +- **Routing/policy — official docs/source:** JSON `Config` objects can be attached per request and compose retries, fallback, load balancing, conditional routing, parameter overrides, guardrails, and caching. Conditional routing can match model, temperature/max tokens, metadata, or URL path. [Conditional routing](https://portkey.ai/docs/product/ai-gateway/conditional-routing.md) · [repository config example](https://github.com/Portkey-AI/gateway#3-routing--guardrails) +- **Useful routing details:** Automatic retries use exponential backoff; load balancing can distribute across API keys/providers; timeouts are configurable; sticky load balancing can hash a caller identifier to preserve conversation affinity and stable A/B assignment. [Repository](https://github.com/Portkey-AI/gateway#core-features) · [February 2026 announcement](https://new.portkey.ai/announcements) +- **Guardrails/cost:** The stable repo exposes many deterministic/custom guardrails and simple caching. The README marks some semantic caching, provider optimization, analytics, and prompt management as hosted/enterprise features, so Flux should not assume all marketed policy features are in the stable OSS branch. [Repository feature caveats](https://github.com/Portkey-AI/gateway#collaboration--workflows) +- **Runtime/deployment — repo-confirmed:** TypeScript runs via `npx`, Node, Docker, Cloudflare Workers, Replit, and Kubernetes/Helm. This is one of the clearest OSS gateway implementations of edge deployment. [Deployment docs in repository](https://github.com/Portkey-AI/gateway/blob/main/docs/installation-deployments.md) +- **Migration UX:** Adoption is base-URL plus provider-qualified model change; a local console and request-attached config make routing visible. Its later 2.0 product adds a separately connected management/control plane. [Repository quickstart](https://github.com/Portkey-AI/gateway#quickstart-2-mins) · [Gateway 2.0 announcement](https://portkey.ai/blog/gateway-2-0) +- **Flux implication:** Copy declarative per-attempt route configuration, sticky session affinity, and edge-friendly handler boundaries. Avoid a built-in catalog of dozens of guardrails and do not make Flux route execution depend on a SaaS control plane. + +#### Bifrost + +- **API and model surface — official docs:** Bifrost's provider matrix distinguishes models, chat/text completions, streaming, Responses, images/edits, embeddings, TTS/STT, files, batch, token counting, rerank, OCR, video, containers, and passthrough. It also offers provider-compatible façades such as Anthropic, Gemini, Cohere, and Bedrock. [Provider matrix](https://docs.getbifrost.ai/providers/supported-providers/overview) +- **Transformation fidelity:** `send_back_raw_response` lets callers inspect the provider's original response and the docs explicitly position it for cost analysis and integration testing of provider-specific transformations. This is one of the most directly reusable ideas for Flux. [Provider matrix response format](https://docs.getbifrost.ai/providers/supported-providers/overview#response-format) +- **Tools/structured output:** Tool calling and reasoning are documented for multiple providers, but the matrix demonstrates why Flux should expose operation- and capability-level support rather than a single boolean “OpenAI compatible.” [Overview](https://docs.getbifrost.ai/overview) +- **Routing/policy:** Open-source features include fallback, weighted key distribution, virtual keys with access/budget/rate policy, semantic cache, real-time monitoring, Prometheus, OTLP, and Go/WASM plugins. Adaptive load balancing, clustering, identity providers, RBAC, audit logs, and several guardrail integrations are enterprise. [Overview](https://docs.getbifrost.ai/overview) +- **Runtime:** A Go HTTP gateway with UI or an in-process Go SDK; no database is required for the basic path. Enterprise clustering introduces peer discovery, gossip state, and a broker mode. [Overview](https://docs.getbifrost.ai/overview) · [Clustering](https://docs.getbifrost.ai/enterprise/clustering) +- **Observability/evaluation:** The gateway exposes request telemetry, Prometheus and OTLP; its vendor also operates Maxim for evaluation/observability, but evaluation is a separate product rather than a necessary gateway subsystem. [Overview](https://docs.getbifrost.ai/overview) +- **Security lesson:** The September 2026 MCP stdio RCE shows why a gateway's management listener, plugin/process-launch surface, and provider credentials need a deny-by-default trust boundary. **Independent evidence.** [JFrog](https://research.jfrog.com/vulnerabilities/bifrost-is-vulnerable-to-unauthenticated-remote-code-execution-via-mcp-stdio-client-registration-cve-2026-90898/) · [fix](https://github.com/maximhq/bifrost/pull/6757) +- **Flux implication:** This is the most useful implementation peer for a Go runtime. Borrow the operation matrix, raw-response/debug mode, Go middleware boundary, and provider-compatible façades. Avoid MCP hosting, command-spawning integrations, default-open management APIs, UI, and distributed cluster state until Flux has a clear host-owned use case. + +#### Helicone AI Gateway + +- **API and model surface — repo-confirmed:** The gateway is a Rust binary with one OpenAI-SDK surface, an embedded provider YAML, and a project claim of 100+ models. It is intentionally narrower than LiteLLM/Bifrost/New API. [Repository](https://github.com/Helicone/ai-gateway) +- **Routing:** Documented strategies include model-latency, provider latency using **P2C + PeakEWMA**, weighted distribution, and lowest cost. P2C avoids repeatedly selecting one currently fastest backend, while PeakEWMA gives a low-state latency estimate. [Repository](https://github.com/Helicone/ai-gateway#-smart-provider-selection) +- **Policy/caching:** Rate limits can count requests, tokens, or dollars per user/team/global scope. Response caching supports in-memory, Redis, and S3 with TTL/invalidation. The gateway README does not establish semantic guardrails, budgets, or a full operation matrix. [Repository](https://github.com/Helicone/ai-gateway#-control-your-spending) · [cache section](https://github.com/Helicone/ai-gateway#-improve-performance) +- **Observability:** OpenTelemetry logs/metrics/traces and optional Helicone integration are built in. The larger Apache-2.0 Helicone repository is active and includes evaluation/experiment UI, but that is a separate observability platform. [Gateway repository](https://github.com/Helicone/ai-gateway) · [Helicone platform](https://github.com/Helicone/helicone) +- **Deployment/migration:** Standalone binary or Docker, optionally with Redis/S3/Helicone; migration is a base-URL and `provider/model` change. [Repository](https://github.com/Helicone/ai-gateway#-self-host-the-ai-gateway) +- **Current-state caution:** The dedicated gateway repository has had no source push since 2025-11-21, has no latest GitHub release, and conflicts with itself on GPL-3.0 versus Apache licensing. Current docs call the old cloud gateway legacy. [Repository metadata](https://api.github.com/repos/Helicone/ai-gateway) · [license](https://github.com/Helicone/ai-gateway/blob/main/LICENSE) · [docs](https://docs.helicone.ai/ai-gateway/introduction) +- **Flux implication:** Copy P2C/PeakEWMA as an optional strategy behind Flux's existing `router.Strategy`, and preserve the separation between a low-latency data plane and optional observability. Do not couple routing health to Helicone or copy the stale/license-ambiguous state. + +#### Agent Router (formerly Envoy AI Gateway) + +- **API and provider surface — official docs:** Version 1.x offers one OpenAI-compatible interface over 16 providers, native Anthropic Messages, OpenAI Responses, chat/completions, embeddings, image generation, audio, multimodal inputs, and MCP. Cross-provider translation is implemented in the Go ext-proc/control plane and enforced by Envoy. [1.0 capabilities](https://tetrate.io/blog/envoy-ai-gateway-v1-0-release) · [providers](https://theagentrouter.ai/docs/capabilities/llm-integrations/supported-providers) +- **Provider extensions:** Release notes document provider-specific fields, reasoning controls, image/audio/video translation, prompt caching, and body redaction. This confirms that Agent Router retains a provider-extension path instead of forcing every feature into one lossy schema. [Release notes](https://theagentrouter.ai/release-notes/) +- **Reliability:** It supports retry/failover, InferencePool endpoint selection, model virtualization, quota-aware rate limiting, hostname routing, and AWS/GCP workload identity. [Release notes](https://theagentrouter.ai/release-notes/) +- **Streaming-specific technique:** v1.1 adds a stream idle timeout with failover and per-request upstream credentials. This is directly relevant to Flux: distinguish initial response timeout, first-visible-event timeout, inter-chunk idle timeout, and total request deadline. [v1.1 release notes](https://theagentrouter.ai/release-notes/v1.1) +- **Observability:** OpenTelemetry GenAI tracing, OpenInference compatibility, token and separate reasoning-token metrics, usage attribution, and request/response redaction are documented. [1.0 capabilities](https://tetrate.io/blog/envoy-ai-gateway-v1-0-release) · [release notes](https://theagentrouter.ai/release-notes/) +- **Deployment/boundary:** A Go controller turns CRDs into Envoy configuration; Envoy Gateway carries traffic. The same project now offers a one-command standalone router, but Kubernetes remains the richer operating model. [Repository](https://github.com/theagentrouter/agent-router) +- **Migration/governance:** The September rebrand preserved CRDs, API group, images, Helm namespace, CLI, module path, release cadence, maintainers, and license. The AAIF structure has nine maintainer seats shared across Bloomberg, Nutanix, AMD, Tetrate, and Netflix, with no majority company. **Independent foundation evidence.** [AAIF announcement](https://aaif.io/blog/agent-router-joins-aaif) +- **Flux implication:** Copy stream idle timeout/failover semantics, per-attempt credentials, provider extension fields, privacy-safe redaction configuration, and an explicit compatibility policy. Avoid adopting Kubernetes CRDs, an MCP gateway, or a controller merely to operate a Go library. + +#### Kong AI Gateway + +- **API and policy surface — official docs:** Kong's AI plugins expose a provider-agnostic API, streaming, semantic caching/routing, prompt compression/decorating, PII sanitization, prompt/response guards, token budgets, model pricing, audit logs, OTel, MCP, and A2A. Provider and operation fidelity is plugin-specific rather than a single published matrix. [AI Gateway documentation](https://developer.konghq.com/ai-gateway/) +- **Routing:** Documented algorithms include consistent hashing, lowest latency, usage, round robin, and semantic matching, with retry/fallback. This is a useful catalog of policies but also shows how quickly a provider runtime becomes a general API gateway. [AI Gateway docs](https://developer.konghq.com/ai-gateway/) +- **Runtime/control plane:** Traditional Kong AI plugins run in Kong Gateway's Lua/OpenResty plugin pipeline. AI Gateway 2.x has a dedicated data plane, entity model, control plane, admin API, analytics, and Konnect UX. [2.0 architecture announcement](https://konghq.com/blog/product-releases/kong-ai-gateway-2-0-agentic-ai) · [current docs](https://developer.konghq.com/ai-gateway/) +- **Migration UX:** Kong provides a V1-plugin to V2-entity mapping, a migration guide, and `kongctl` conversion from decK configuration. This is a strong pattern for capability-driven migration diagnostics. [Migration guide](https://developer.konghq.com/ai-gateway/v2-migration-guide/) +- **OSS boundary:** The public Kong repository and AI plugins are Apache-2.0, but the 2.x runtime/control plane was not verified in that source tree. Kong's docs also position Konnect as the management path. [Repository](https://github.com/Kong/kong) · [FAQ/deployment](https://developer.konghq.com/ai-gateway/) +- **Flux implication:** Copy explicit migration tooling and ordered policy phases. Avoid entity-oriented administration, a general API gateway, Konnect analytics, metering/billing, and a catalog of interchangeable enterprise policies. + +#### Apache APISIX AI Gateway + +- **API surface — official docs:** `ai-proxy` builds provider-specific requests for OpenAI, DeepSeek, Azure, Anthropic, OpenRouter, Gemini, Vertex AI, Bedrock, and compatible endpoints; `ai-proxy-multi` adds load balancing, retries, fallbacks, and health checks. [AI proxy](https://apisix.apache.org/docs/apisix/plugins/ai-proxy) · [AI Gateway](https://apisix.apache.org/ai-gateway/) +- **Policy:** The AI plugin set includes token rate limiting, prompt decorators/templates, request rewrite, moderation/Lakera, RAG, semantic cache/routing, and token observability. [Plugin hub](https://apisix.apache.org/plugins) +- **Implementation pattern:** APISIX executes plugins in lifecycle phases (`rewrite`, `access`, `before_proxy`, `header_filter`, `body_filter`, `log`, and custom balancer hooks). That explicit ordering is valuable for a provider runtime because request transforms, routing, upstream auth, stream transforms, and telemetry have different safety boundaries. [Plugin terminology](https://apisix.apache.org/docs/apisix/terminology/plugin/) +- **Extensibility:** Plugins can be Lua, Java, Go, Python, Node, or experimental Proxy Wasm through local RPC/Wasm. Hot reload and Admin API configuration are first-class. [Plugin development](https://apisix.apache.org/docs/apisix/plugin-develop/) · [repository](https://github.com/apache/apisix) +- **Operations:** The clustered architecture uses stateless APISIX nodes and etcd configuration; standalone YAML mode avoids the control-plane dependency. [Repository architecture](https://github.com/apache/apisix) +- **Flux implication:** Copy lifecycle ordering and hot-reloadable Go-native extension points behind a narrow interface. Do not adopt etcd, a general plugin runner, or generic API management merely to serve model calls. + +#### Higress + +- **API/provider surface — official docs:** Higress claims unified protocol conversion for 100+ models, model-level fallback, MCP hosting/proxying, and conformant Gateway API Inference Extension support. [AI Gateway](https://higress.ai/en/ai-gateway) · [repository](https://github.com/higress-group/higress) +- **Streaming/plugin pattern:** Higress emphasizes complete streaming request/response processing so Wasm plugins can transform SSE without buffering the whole response. This is the most relevant implementation detail for Flux's streaming-first identity. [Higress overview](https://higress.cn/en/docs/latest/overview/what-is-higress) +- **Policy/operations:** It provides load balancing/fallback, token quota/rate limiting, semantic cache, authentication, content protection, observability, API-key pools, and Wasm plugins. [AI Gateway](https://higress.ai/en/ai-gateway) +- **Runtime:** Go components and an Istio/Envoy data plane, with Docker all-in-one or Kubernetes/Helm; plugins compile from Go, Rust, or JavaScript to Wasm. The project includes a console and publishes regional container images. [Repository](https://github.com/higress-group/higress) +- **Governance/migration:** Higress entered CNCF Sandbox in March 2026 and provides an Ingress NGINX-to-Gateway API migration path plus OpenAPI-to-MCP conversion. [CNCF](https://www.cncf.io/projects/higress) · [repository](https://github.com/higress-group/higress) +- **Flux implication:** Copy the principle that streaming middleware must be incremental and cancellable. Do not adopt Wasm, a console, Istio, or a full ingress gateway before Flux has trusted Go hooks and complete protocol parity. + +#### New API + +- **API surface — repo-confirmed:** New API exposes OpenAI Chat/Responses, Anthropic Messages, Gemini generate/stream, OpenAI realtime/Responses WebSocket, images/audio, embeddings/rerank, and task-plugin routes. It explicitly warns that protocol-specific tools and fields may not map exactly. [Repository](https://github.com/QuantumNous/new-api) +- **Implementation pattern:** `relaykit` is an independently buildable Go module for request, response, and streaming conversion among four text protocols. This separation is directly applicable to Flux: protocol translators should be testable independently from routing, credentials, storage, and HTTP serving. [Repository architecture](https://github.com/QuantumNous/new-api#development-and-extensions) +- **Routing/governance:** Source documents model mappings, channel priority/weight, retries, affinity, multiple upstream keys, quotas, subscriptions, usage logs, cache accounting, expression pricing, users/groups, fine-grained permissions, OAuth/OIDC, passkeys, and 2FA. [Repository](https://github.com/QuantumNous/new-api) +- **Runtime:** A single Go/Gin service embeds a React/Bun UI. SQLite is sufficient for a single instance; production examples use PostgreSQL/MySQL and Redis, with optional ClickHouse logs. It also has a JavaScript task-plugin runtime and Electron packaging. [Repository](https://github.com/QuantumNous/new-api#deployment) +- **Observability/evaluation:** The gateway provides usage/audit logs and playground analytics, but its identity is a management console and distribution system rather than a vendor-neutral telemetry runtime. No first-party evaluation engine was established from the reviewed repository. [Repository](https://github.com/QuantumNous/new-api) +- **License:** AGPL-3.0 with Section 7 obligations to preserve attribution and a visible link in modified UIs. [License](https://github.com/QuantumNous/new-api/blob/main/LICENSE) +- **Flux implication:** Copy the independent translation-module boundary and host-safe plugin contract. Avoid the user/group/subscription/billing/UI product, JavaScript task runtime, database requirements, and AGPL implications unless Flux intentionally changes product and licensing direction. + +#### LLM Gateway (`theopenco/llmgateway`) + +- **API surface — official docs:** Project documentation advertises 200+ models/40+ providers, OpenAI and Anthropic compatibility, AI SDK integration, embeddings, rerank, audio/image/video, realtime, and provider-native web search. The project also implemented the AI SDK gateway protocol in August 2026 to preserve model resolution, model lists, and provider-native search during migration. [Docs](https://docs.llmgateway.io/) · [AI SDK protocol announcement](https://llmgateway.io/changelog/ai-sdk-gateway-protocol) +- **Routing/governance:** Current changelog material documents provider failover, response caching, per-project/key/member spend, key TTL, provider compliance filters, and organization analytics. [Changelog](https://llmgateway.io/changelog) +- **Migration UX:** The CLI can configure and launch Claude Code, OpenCode, Codex CLI, and other coding agents against the gateway, verify keys, and preserve each tool's expected environment/config format. This is a notably concrete migration pattern. [CLI launch announcement](https://llmgateway.io/changelog/cli-launch-coding-agents) +- **Runtime:** The unified self-hosted image combines Hono/Next services, PostgreSQL, Redis, UI, playground, and gateway. The core is AGPL-3.0; advanced billing, retention, team/org management, and related features in `ee/` are commercial. [Repository](https://github.com/theopenco/llmgateway) +- **Flux implication:** Borrow compatibility-protocol awareness and tool-specific configuration/doctor commands. Avoid the hosted marketplace, database-backed product console, subscription model, and commercial/core licensing split. + +#### OpenZiti LLM Gateway + +- **API surface — repo-confirmed:** OpenZiti targets `POST /v1/chat/completions`, translates OpenAI requests to Anthropic where needed, supports SSE, and routes to Ollama, llama.cpp, vLLM, SGLang, or any OpenAI-compatible server. Tools, structured output, embeddings, and batches are not documented. [Repository](https://github.com/openziti/llm-gateway) +- **Routing:** Weighted round robin, active health checks, passive failover, and VM sleep detection are aimed at real inference-server pools. Its optional semantic router cascades keyword heuristics, embedding similarity, and an LLM classifier when no model is supplied. [Repository](https://github.com/openziti/llm-gateway) +- **Runtime:** One Go binary, YAML, Prometheus metrics, and no database/message queue/sidecar. OpenZiti/zrok is optional and provides identity-based, encrypted ingress/egress through overlay networks. [Package documentation](https://pkg.go.dev/github.com/openziti/llm-gateway) +- **Flux implication:** Its single-binary scope is highly aligned, but the project is much smaller. Copy only the explicit semantic-routing cascade as an opt-in strategy after measuring its latency/cost; do not make an extra LLM call or overlay network a default dependency. + +### Inferences + +- **The strongest shared architecture is a pure data-plane core plus optional policy/storage adapters.** LiteLLM's pure Rust transform stage, Bifrost's Go SDK/gateway split, and New API's `relaykit` all separate protocol transformation from management concerns. Flux already has `provider/core` and `provider/adapters`; the opportunity is to make conformance and extension boundaries equally explicit. +- **Compatibility is broader than “OpenAI-shaped JSON.”** Agent Router, Bifrost, New API, and Vercel all retain or add native Anthropic/Responses/provider-specific surfaces. Flux's current OpenAI HTTP adapter is text-centric and should not be marketed as full drop-in compatibility until tools, content parts, structured output, tool-call streams, and error semantics are covered. +- **Policy must be evaluated per route attempt, not once per client request.** LiteLLM's fallback authorization/budget recheck and Portkey's per-config overrides are the clearest implementation precedents. A fallback that changes provider/model can change price, data residency, tool semantics, and policy. +- **Streaming reliability needs more lifecycle clocks than a single HTTP timeout.** Agent Router's idle timeout is a useful pattern, but Flux should go further by distinguishing connect/header timeout, first-event timeout, first-content/tool delta, inter-chunk idle timeout, and total deadline. Transparent failover is safe only before an externally visible event unless the stream protocol explicitly permits replay. +- **Raw fidelity and conformance tests are more valuable than another ten shallow adapters.** Bifrost's raw-response mode and LiteLLM's staged parity checks show how to prevent normalization regressions as provider APIs change. +- **Go remains a credible runtime choice.** Bifrost, Agent Router's controller, Higress, New API, and OpenZiti all demonstrate substantial Go gateway/control-plane systems. Flux does not need Rust or a service mesh merely to match their language. +- **Observability should remain a port, not a platform.** Helicone demonstrates the value of separating a fast router from analytics, while LiteLLM/Bifrost/New API demonstrate the operational cost of merging the two. Flux's hashed audit projection and OTel instrumentation are directionally aligned. + +### Gaps + +- Several projects publish feature lists but not a machine-readable operation × provider × streaming × tool × modality matrix. Exact parity for tools, structured output, and provider-native extensions therefore remains unverified. +- LiteLLM's Rust migration is staged through late 2026; its exact production cutover status for every endpoint was not established from the reviewed migration page. +- Portkey's 2.0 source/feature boundary and Helicone's cloud roadmap are moving targets; conclusions are limited to the public state captured on 2026-09-24. +- Evaluations are generally separate products (Maxim, Helicone, LangSmith/LiteLLM Evals) rather than gateway-runtime requirements. No common evaluation interoperability standard was found. + +## 3. Flux baseline, selective adoption roadmap, and explicit non-goals + +### Takeaway + +Flux's highest-value work is **not** copying LiteLLM's or Kong's breadth. It is making Flux the most trustworthy embeddable Go provider runtime: full OpenAI/Anthropic compatibility, operation-level capability truth, per-attempt policy enforcement, correct streaming deadlines/failover, lossless provider extensions, and current OpenTelemetry semantics. The project should selectively add hooks, route algorithms, and migration tooling while leaving organization control, UI, billing, tenancy, MCP/A2A, vector stores, memory, and mandatory infrastructure to hosts or external systems. + +### Cited Findings + +#### Flux baseline verified from the current checkout + +- **Identity and boundary:** Flux is explicitly a host-neutral provider runtime. Rho owns UX, agents, tools, permissions, sessions, and product semantics; Flux owns credentials, catalog/route resolution, transports, normalized streams, retry/fallback, usage, and telemetry. Hosts may import only `engine`, `llm`, `graph`, and `tools`. `README.md:31-62` +- **Runtime, not service monolith:** The stable `engine` facade exposes generation, catalog, credential setup/status, preflight, and normalized events. `engine/types.go:50-146` +- **Provider breadth:** The current registry contains 28 provider gateways, including OpenAI, Anthropic, Gemini, Azure, Bedrock, Vertex, OpenRouter, Groq, Fireworks, Ollama, and multiple hosted/coding-plan providers. `README.md:200-234` +- **Core normalized model:** Flux already models multimodal content parts, tools/tool choice, JSON-schema response format, reasoning/thinking controls, prompt caches, provider-native opaque replay blocks, warnings, route/deployment/attempt data, usage including cache and reasoning tokens, and streaming tool-call deltas/TTFT. `llm/types.go:29-106` · `llm/types.go:122-191` · `llm/types.go:219-287` +- **Reliability:** Flux has weighted/simple-shuffle/least-busy/EWMA-latency/cost/usage routing, retries, cooldowns, and fallback. It deliberately retries only stream setup errors, not errors after the event channel is handed to the caller. `router/router.go:106-165` · `router/router.go:198-255` · `router/strategy.go:16-30` +- **Policy primitives:** Flux has deterministic guardrails, streaming cross-chunk PII accumulation, budget/usage providers, cache and audit ports, embeddings, rerank, and an Anthropic batch client. `provider/core/guardrails.go:12-78` · `provider/core/stream_guardrails.go:9-21` · `provider/observability/budget_provider.go:34-53` · `internal/cache/backend.go:25-33` · `internal/observability/audit.go:21-40` · `provider/batch/batch.go:38-58` +- **Current HTTP compatibility gap:** `/v1/chat/completions` accepts only string message content and a small option subset; it ignores several common parameters, flattens multi-turn history into transcript text, converts tool declarations but does not return OpenAI tool calls, and only streams content deltas. It is useful, but it is not full OpenAI drop-in parity. `internal/api/openai_proxy.go:22-44` · `internal/api/openai_proxy.go:109-188` · `internal/api/openai_proxy.go:190-258` · `internal/api/openai_proxy.go:300-348` +- **Current public operation gap:** The HTTP server exposes `/v1/chat/completions`, `/rerank`, prompt/conversation routes, analytics, and health, but no `/v1/messages`, `/v1/responses`, general batch routes, or multimodal OpenAI facade. `internal/api/server.go:101-121` +- **Batch is provider-specific and outside the stable facade:** the current batch client targets Anthropic Message Batches and reconstructs a limited params map; it does not provide a provider-neutral batch interface or OpenAI batch translation. `provider/batch/batch.go:38-106` +- **Telemetry drift:** Flux's constants use `gen_ai.system` and other attributes that OpenTelemetry now marks moved/deprecated; current guidance is `gen_ai.provider.name` and the dedicated `open-telemetry/semantic-conventions-genai` repository. `internal/observability/genai_semconv.go:21-47` · [OpenTelemetry registry](https://opentelemetry.io/docs/specs/semconv/registry/attributes/gen-ai/) · [new GenAI conventions repository](https://github.com/open-telemetry/semantic-conventions-genai) +- **Control-plane drift to watch:** Flux's optional SQLite budget store owns virtual-key records and stores provider API keys in plaintext in a `0600` database and WAL sidecars. That is locally protected but is still a credential store/tenant-state implementation that pulls the runtime toward control-plane responsibilities. `storage/budgets.go:22-68` · `storage/budgets.go:74-108` + +#### External evidence supporting the recommendations + +- LiteLLM's current docs recheck authorization and budget on fallback targets and record fallback attempt metadata. [Reliability docs](https://docs.litellm.ai/docs/proxy/reliability) +- Bifrost publishes an operation-level provider matrix and an optional raw provider response for transformation testing. [Provider matrix](https://docs.getbifrost.ai/providers/supported-providers/overview) +- New API isolates multi-protocol request/response/stream conversion in an independently buildable Go module. [Repository](https://github.com/QuantumNous/new-api) +- Agent Router v1.1 adds stream idle timeout with failover and preserved per-request upstream credentials. [Release notes](https://theagentrouter.ai/release-notes/v1.1) +- Helicone documents P2C + PeakEWMA provider routing. [Repository](https://github.com/Helicone/ai-gateway#-smart-provider-selection) +- Portkey documents conditional routing and request-attached JSON policy configuration. [Conditional routing](https://portkey.ai/docs/product/ai-gateway/conditional-routing.md) +- Higress and APISIX both make lifecycle/plugin ordering explicit; Higress specifically supports transforming complete SSE streams in Wasm. [Higress overview](https://higress.cn/en/docs/latest/overview/what-is-higress) · [APISIX plugin lifecycle](https://apisix.apache.org/docs/apisix/terminology/plugin/) +- OpenTelemetry moved GenAI conventions to a dedicated repository in 2026 and marks `gen_ai.system` deprecated in favor of `gen_ai.provider.name`. [OTel registry](https://opentelemetry.io/docs/specs/semconv/registry/attributes/gen-ai/) · [GenAI repository](https://github.com/open-telemetry/semantic-conventions-genai) + +### Inferences + +#### Priority 0 — make the existing runtime contract trustworthy + +1. **Build a dedicated wire-compatibility layer and conformance suite.** Keep HTTP/OpenAI/Anthropic translation outside the engine and provider decorators, following New API's `relaykit` separation. [New API architecture](https://github.com/QuantumNous/new-api) + + Complete `/v1/chat/completions` with: + + - array/content-part messages, image and audio inputs; + - `response_format`/JSON schema, `tool_choice`, parallel tool controls, tool calls and streamed tool-call deltas; + - `stop`, `seed`, `logprobs`, `stream_options`, service tier, and other already modeled fields; + - provider/model error bodies and headers rather than converting every upstream failure to HTTP 500; + - usage detail for cache/reasoning tokens where known; + - preservation of provider-native state through existing `ProviderBlock`; + - explicit warnings for ignored or lossy fields using `CallWarning`. + + Add a native `/v1/messages` façade for Anthropic SDKs and an optional `/v1/responses` façade only after the core model can represent the required state. Do **not** flatten conversation history into one prompt. Acceptance should be fixture-based: recorded provider request/response/SSE pairs, cross-provider semantic assertions, and SDK integration tests. + +2. **Make route attempts first-class policy boundaries.** Introduce an additive route trace/attempt event containing requested alias, candidate provider/model/deployment, attempt number, start/end, status/error class, retry/fallback reason, policy decision, actual usage, and whether the response was already externally visible. Revalidate host-provided authorization, model access, data-region policy, and budget for every candidate—not only the originally requested model. This directly adopts LiteLLM's fallback recheck and corrects the ambiguity that arises when a budget wrapper surrounds a router. [LiteLLM reliability](https://docs.litellm.ai/docs/proxy/reliability) + +3. **Define streaming deadlines and safe failover state.** Add independently configurable: + + - connection/response-header timeout; + - first-event timeout; + - first externally meaningful delta timeout; + - inter-event idle timeout; + - total request deadline; + - cancellation propagation and close guarantees. + + Retry/fail over only while no content, tool-call delta, reasoning delta, provider block, or other externally visible event has escaped. After that point emit a typed terminal error with an explicit `partial_output=true`; never silently switch models and concatenate two answers. Agent Router's v1.1 idle-timeout work is the direct peer evidence. [Agent Router release notes](https://theagentrouter.ai/release-notes/v1.1) + +4. **Adopt Bifrost-style capability and fidelity truth.** Generate an operation × provider × model capability matrix for chat, stream, tools, structured output, images/audio, embeddings, rerank, batch, and native passthrough. Add an opt-in raw request/response capture hook for tests and debugging, with redaction and size limits. Continue using `CallWarning` to report lossy translation. This is more valuable than chasing headline provider counts. [Bifrost matrix](https://docs.getbifrost.ai/providers/supported-providers/overview) + +5. **Update GenAI observability to the 2026 conventions repository.** Emit `gen_ai.provider.name`, operation/model/usage/finish-reason/TTFT attributes, separate reasoning/cache token counts, route-attempt events, and privacy-safe cost. Keep content recording opt-in and aligned with the existing hash-only audit design. Maintain compatibility aliases during migration. [OTel GenAI repository](https://github.com/open-telemetry/semantic-conventions-genai) · [Flux audit projection](internal/observability/audit.go) + +#### Priority 1 — selective runtime capabilities that remain host-neutral + +6. **Add sticky affinity and P2C/PeakEWMA behind the existing strategy interface.** Affinity should use a host-supplied stable key, bounded TTL, model/deployment scope, and explicit opt-out. P2C+PeakEWMA improves tail routing under concurrency without requiring a shared coordination service. Helicone documents the algorithm; Flux already has a strategy abstraction and EWMA state, so this is a contained addition. [Helicone repository](https://github.com/Helicone/ai-gateway) · `router/strategy.go` + +7. **Generalize operation ports selectively.** Define provider-neutral `BatchProvider`, `EmbeddingProvider`, and `RerankProvider` contracts with capability discovery and provider-native result metadata. Implement Anthropic and OpenAI batch adapters before exposing HTTP batch routes. Keep asynchronous polling host-driven; do not add a queue/control plane. The existing Anthropic-only client is a useful starting point. `provider/batch/batch.go` + +8. **Introduce narrow, trusted middleware phases.** A minimal Go interface should cover `BeforeRoute`, `AfterRoute`, `BeforeAttempt`, `OnStreamEvent`, and `OnError`, with immutable/context-carried data and deterministic ordering. Keep existing decorators for built-in behavior. Borrow the phase model from APISIX/Higress, but do not expose arbitrary process launch, filesystem, network, or management mutation to plugins. [APISIX lifecycle](https://apisix.apache.org/docs/apisix/terminology/plugin/) · [Higress overview](https://higress.cn/en/docs/latest/overview/what-is-higress) + +9. **Turn guardrails into deterministic primitives plus external policy hooks.** Retain fast local redaction/block rules and streaming accumulation. Add an external semantic-guardrail interface with explicit fail-open/fail-closed, latency budget, streaming chunk/turn behavior, and privacy mode. Do not bundle a vendor catalog or policy database. Bifrost, Kong, Portkey, and APISIX show demand for this boundary, but also the scope risk of implementing it as a product. [Bifrost overview](https://docs.getbifrost.ai/overview) · [Kong AI Gateway](https://developer.konghq.com/ai-gateway/) + +10. **Keep credentials and tenancy outside the runtime by default.** Preserve `CredentialProvider`/resolver interfaces and virtual-key identifiers for host attribution, but make the default safe composition resolve secrets through the existing credential store/keyring rather than the SQLite store that persists provider keys. If the SQLite budget store remains, document it as an optional single-host convenience and provide a `SecretProvider` port. This protects Flux's runtime identity and reduces credential blast radius. `storage/budgets.go:35-68` + +11. **Publish migration diagnostics, not a migration control plane.** Add a CLI/library capability report that compares a target gateway config or OpenAI/Anthropic request against Flux's effective model, operations, supported fields, expected transformations, and loss warnings. Preserve provider-qualified model aliases and offer an import adapter for a small declarative format if justified. Follow Agent Router's compatibility discipline and Kong's V1→V2 migration tooling without adopting their control planes. [Agent Router rename compatibility](https://github.com/theagentrouter/agent-router#formerly-envoy-ai-gateway) · [Kong migration guide](https://developer.konghq.com/ai-gateway/v2-migration-guide/) + +#### Priority 2 — optional ecosystem work only after core parity + +12. **Evaluation and experimentation through ports.** Export privacy-safe trace events and deterministic request/response fixtures that a host can replay into its evaluator. A minimal `Evaluator` callback can score a completed response or stream summary, but Flux should not ship datasets, a prompt lab, an experiment database, or an evaluation UI. LiteLLM's Evals API and Helicone/Maxim demonstrate demand while also showing why evaluation is a separate product layer. [LiteLLM Evals](https://docs.litellm.ai/docs/evals_api) · [Bifrost overview](https://docs.getbifrost.ai/overview) + +13. **Signed/versioned catalog and capability provenance.** Catalog rows should identify source (`embedded`, `live`, operator override), fetched time, confidence, and capability provenance. Hosts can choose update policy; Flux should not silently run mutable remote routing policy. This extends the existing live discovery/catalog model without a service. + +14. **Wasm only after a stable Go hook ABI exists.** If trusted Go hooks are insufficient, evaluate Wasm for streaming-safe, capability-limited transforms with deterministic fuel/memory limits, no ambient credentials, schema-version negotiation, signed-module policy, and no management operations. Higress shows the value of sandboxed hot updates; Bifrost's CVE shows the danger of broad execution authority. [Higress repository](https://github.com/higress-group/higress) · [Bifrost security analysis](https://research.jfrog.com/vulnerabilities/bifrost-is-vulnerable-to-unauthenticated-remote-code-execution-via-mcp-stdio-client-registration-cve-2026-90898/) + +15. **Optional edge/serverless adapter, not a second runtime.** Expose a standards-based HTTP handler and OpenTelemetry export that a thin adapter can deploy to an edge runtime. Avoid provider-specific Workers bindings in the core unless a host opts in; Portkey demonstrates deployment portability, while Go/WASM portability still needs a deliberate ABI. + +#### Patterns Flux should deliberately reject + +| Pattern | Why it conflicts with Flux | Peer evidence | +|---|---|---| +| Web console, org/team administration, SSO/SCIM product, billing/reseller UI | Turns a host-neutral runtime into a control plane and creates a large security/maintenance surface. | LiteLLM, Portkey, New API, LLM Gateway. | +| Mandatory PostgreSQL/Redis/etcd/Kubernetes | Breaks embeddability and single-binary/local use; stateful infrastructure should be optional adapters. | [LiteLLM deployment](https://docs.litellm.ai/docs/proxy/deploy) · [APISIX](https://github.com/apache/apisix) · [Agent Router](https://github.com/theagentrouter/agent-router) | +| MCP/A2A gateway, tool registry, skills marketplace, vector stores, memory, agent runtime | Tool execution, permissions, and product semantics belong to Rho/hosts under Flux's documented boundary. | Portkey, Bifrost, Agent Router, LiteLLM, New API. | +| General-purpose plugin marketplace or unrestricted process execution | Enormous attack surface; Bifrost's MCP stdio RCE is a current cautionary implementation. | [CVE analysis](https://research.jfrog.com/vulnerabilities/bifrost-is-vulnerable-to-unauthenticated-remote-code-execution-via-mcp-stdio-client-registration-cve-2026-90898/) | +| Built-in semantic guardrail vendor catalog | Couples runtime policy to vendors and creates latency/privacy ambiguity; an external hook is sufficient. | Bifrost, Portkey, Kong, APISIX. | +| Default LLM-classifier auto-routing | Adds cost, nondeterminism, latency, and hard-to-explain behavior. OpenZiti's three-stage cascade should remain explicitly opt-in and measurable. | [OpenZiti repository](https://github.com/openziti/llm-gateway) | +| Huge endpoint zoo before compatibility correctness | Partial support presented as “drop-in” creates silent data/tool/stream loss. Flux should finish chat/Anthropic fidelity first. | LiteLLM and New API demonstrate breadth; Bifrost demonstrates why matrices and raw fidelity are needed. | +| Mutable remote policy by default | A provider runtime should not let an unreviewed service silently change model, region, privacy, or price behavior. Keep discovery/catalog data separate from operator-approved route policy. | +| Plaintext provider secrets in general-purpose runtime storage | Conflicts with credential abstraction and increases blast radius. Keep secrets in keyrings/external secret stores; store references and status only. `storage/budgets.go` | +| Marketing counts or latency numbers as roadmap priorities | Counts are volatile and benchmarks are not comparable without a common harness. Optimize verified correctness and tail behavior. | + +#### Recommended sequencing + +1. **Compatibility and conformance:** full OpenAI chat behavior, native Anthropic façade, operation capability matrix, raw-fidelity tests. +2. **Policy correctness:** per-attempt authorization/budget/region checks, route-attempt trace, fallback metadata. +3. **Streaming correctness:** timeout state machine, safe pre-first-event failover, cancellation/partial-output semantics. +4. **Observability correctness:** 2026 OTel GenAI migration, route events, privacy defaults. +5. **Selective optimization:** sticky/P2C, general batch port, external guardrail hook, migration diagnostics. +6. **Only then:** optional Wasm, edge adapter, richer evaluation hooks, or more modalities/endpoints. + +This sequence keeps Flux's differentiator intact: a small, composable, Go provider runtime that hosts can embed, while making its wire contracts, policy decisions, and stream behavior more production-grade than broad gateway products. + +### Gaps + +- The proposed compatibility work requires a deliberate choice of canonical semantics where OpenAI, Anthropic, Gemini, and provider-native features conflict. That design decision is not resolved by the competitor research. +- Current Flux source already has more policy, budget, batch, and guardrail functionality than `README.md` highlights. Before implementation, each roadmap item should be checked against all composition paths and public/internal package boundaries to avoid duplicating an existing feature under another package. +- A formal provider conformance harness, representative sanitized request corpus, and reproducible cross-language gateway benchmark do not exist yet in the reviewed Flux checkout; these are prerequisites for measuring the proposed gains. +- Whether to add `/v1/responses` immediately or only after Anthropic compatibility is a product prioritization decision, not something the landscape evidence alone can determine. + +## 4. Addendum (2026-09-27): meta-gateway and catalog donors + +### Takeaway + +Section 1 correctly excludes OpenRouter and Vercel AI Gateway from the OSS peer set because their routing services are hosted. They still matter as **design donors**: Flux ships an `openrouter` adapter and plans a Models.dev-generated catalog, and both hosted gateways expose request-scoped routing controls that Flux's deployment router could normalize without adopting a control plane. + +### Cited Findings + +All sources below were accessed on **2026-09-27**. + +- **Flux baseline (repo-confirmed).** `provider/adapters/openrouter.go` is a 39-line wrapper over the OpenAI-compatible client: it forwards `ChatOptions` unchanged and sends neither OpenRouter's `provider` routing object nor its attribution headers. `docs/plans/audit-remediation.md` (WP14 catalog-data) already plans a generator from Models.dev `api.json` to catalog v1 with a reviewed overlay. +- **OpenRouter provider routing (official-docs).** A request's `provider` object accepts `order`, `allow_fallbacks`, `only`, `ignore`, `sort` (price, throughput or latency), `require_parameters`, `data_collection`, `zdr`, `quantizations`, `max_price`, `preferred_min_throughput` and `preferred_max_latency`; the optional attribution headers are `HTTP-Referer` and `X-OpenRouter-Title`. [Provider selection](https://openrouter.ai/docs/guides/routing/provider-selection) +- **OpenRouter 2026 announcements (official).** "In-Region Routing: Keep your data in the US or EU" (2026-09-09), "Give any model a terminal and files" (2026-09-08) and "Batch API: half-price inference by bundling requests" (2026-09-22). [Announcements](https://openrouter.ai/announcements/all) +- **Vercel AI Gateway routing (official-docs, page updated 2026-09-10).** `providerOptions.gateway` takes `order`, `only` and `sort` (`cost`, `ttft` or `tps`), `caching: 'auto'` for providers that need explicit cache markers, and a request-scoped `byok` credential map; model fallbacks and per-provider timeouts are documented alongside. [Provider options](https://vercel.com/docs/ai-gateway/models-and-providers/provider-options) +- **Models.dev (repo/API-confirmed).** MIT-licensed, 7,011 stars, source pushed 2026-09-26. Data lives as TOML under `providers/` (serving details such as pricing) and `models/` (model facts, inherited with `base_model`), and is published as `https://models.dev/api.json` plus `models.json` and `catalog.json`. [Repository](https://github.com/anomalyco/models.dev) + +### Inferences + +1. **OpenRouter passthrough, not emulation.** Give the `openrouter` adapter an explicit, typed provider-preference option (order, only/ignore, fallbacks, sort, ZDR/data-collection, price ceiling) and opt-in attribution headers whose values the host supplies. Flux should never inject its own attribution or silently widen provider choice. +2. **A normalized routing hint for the deployment router.** A request-scoped `order`/`only`/`sort` hint, the shape both hosted gateways converged on, can reorder or narrow deployments the operator already approved. It must never add a deployment outside an explicit routing policy (the explicit-policy exclusivity rule in `docs/plans/audit-remediation.md`). +3. **Data residency as policy.** OpenRouter's in-region routing shows that region is a first-class request constraint. In Flux it belongs in the per-attempt policy checks (Priority 0), not in endpoint strings. +4. **Catalog provenance from Models.dev.** Implement WP14 as planned, and record the source and fetch time per field so Priority 2 item 13 (signed/versioned catalog provenance) has real data to carry. + +### Gaps + +- OpenRouter's in-region hostnames and request parameters were not verified; only the announcement title and date were. An "Auto router" and "Fusion" release reported by the 2026-09-26 research refresh did not appear in the May–September 2026 announcement listing and are left out. +- Neither hosted gateway's routing algorithm is public, so their `sort` semantics are inputs for API shape, not evidence of routing quality. + diff --git a/research_notes/Flux OSS landscape roadmap/product_engineering.md b/research_notes/Flux OSS landscape roadmap/product_engineering.md new file mode 100644 index 00000000..b25c9b31 --- /dev/null +++ b/research_notes/Flux OSS landscape roadmap/product_engineering.md @@ -0,0 +1,338 @@ +# Flux OSS product surface and engineering maturity audit + +**Audit snapshot:** 2026-09-24, commit `f9037a82e403001d50868e4100bcfa19d8a387af` (`main` / `origin/main`). +**Method:** read-only inspection of the complete tracked documentation, API/specification, examples, scripts, workflows, module metadata, public package docs, representative implementation and tests, plus local race/coverage/vet/security checks and read-only GitHub metadata queries. No repository source was edited. + +## 1. Product surface, documentation, OSS operations, and trust + +### Takeaway + +Flux has a credible internal engineering foundation and unusually broad deterministic test coverage, but its public product story is not yet coherent enough for independent adoption. A fresh user cannot reliably execute the advertised `engine` quickstart, several headline surfaces are internal or unwired, documentation has drifted materially from code, and the only current-module release is `v0.0.1` with 29 later commits on `main`; Flux should be positioned as a pre-1.0 provider-runtime library, fix a small set of trust blockers, and release a deliberately narrow `v0.1.0` before pursuing `v1.0`. + +### Cited Findings + +#### Overall adoption assessment + +| Dimension | Assessment | Evidence-based rationale | +|---|---|---| +| Core Go engineering | **Good, with important host-facade gaps** | The current tree passes the full race suite, vet, module verification, and `govulncheck`; test code is larger than production code. However, the stable `engine` path rebuilds stateful transports on every request, so headline reliability middleware is not persistent in normal use. [`engine/engine.go:54-68,283-327`](../../engine/engine.go#L54-L68), [`setup/deployment.go:94-110`](../../setup/deployment.go#L94-L110) | +| Onboarding / quickstart | **Not adoption-ready** | The README constructs `engine.New(Options{})` and immediately streams, while runtime loading requires a valid catalog cache; the advertised default remote catalog currently contains a field rejected by the strict decoder. [`README.md:67-105`](../../README.md#L67-L105), [`engine/state.go:86-103`](../../engine/state.go#L86-L103), [`catalog/v1.go:444-460,615-644`](../../catalog/v1.go#L444-L460) | +| API focus | **Weak** | Fifty-two main-module packages are listed by `go list`; the host-facing `llm.Provider` is a composition of seven broad facets, and important public packages expose many declarations without a documented stability tier. [`llm/provider.go:9-24`](../../llm/provider.go#L9-L24), [`engine/types.go:1-164`](../../engine/types.go#L1-L164) | +| OSS documentation | **Mixed** | The README, architecture boundary, dynamic discovery, and decentralized-routing honesty are useful, but provider counts, env names, examples, OpenAPI, cache/audit claims, release tooling, and SDK claims have drifted. [`README.md:31-44,178-235`](../../README.md#L31-L44), [`docs/guides/CREDENTIAL-SETUP-FLOW.md:7-27`](../../docs/guides/CREDENTIAL-SETUP-FLOW.md#L7-L27), [`api/openapi.yaml:118-340`](../../api/openapi.yaml#L118-L340) | +| Release / supply-chain trust | **Early** | The sole current-module release is signed and CI is strong, but no release-PR automation, Go API diff, SBOM, provenance/attestation, dependency bot, or vulnerability-alert setting exists in the repository/GitHub configuration. [`release.yml:1-31`](../../.github/workflows/release.yml#L1-L31), [v0.0.1 release](https://github.com/GrayCodeAI/flux/releases/tag/v0.0.1), [repository security settings](https://api.github.com/repos/GrayCodeAI/flux) | +| Extensibility | **Mixed** | Custom OpenAI-compatible gateways, injected credential stores, the four-method `core.Provider`, and lower-level HTTP-client injection are good seams. Native provider onboarding, custom routing strategies, custom engine transports, model publication, cache backends, and telemetry sinks are fragmented or inaccessible. [`provider/dynamic.go:11-50`](../../provider/dynamic.go#L11-L50), [`credentials/store.go:27-32`](../../credentials/store.go#L27-L32), [`router/strategy.go:10-31,107-168`](../../router/strategy.go#L10-L31) | + +#### Product identity and public story + +- The README has a clear one-sentence promise—authentication, model resolution, streaming, retries, rate limiting, and caching—and accurately positions Flux as the provider engine beneath the Rho product face. This is the strongest part of the product narrative. [`README.md:5-9,31-44`](../../README.md#L5-L9) +- The repository simultaneously presents Flux as a host-neutral universal runtime, a library, an HTTP gateway, a gRPC server, a conversation DAG, an SDK suite, and an enterprise foundation. The current project has no standalone `main` package, while the HTTP and gRPC servers live below `internal/`, so general Go consumers cannot import those delivery surfaces. [`Makefile:1-3,50-53`](../../Makefile#L1-L3), [`internal/api/server.go:1-16`](../../internal/api/server.go#L1-L16), [`internal/grpc/README.md:1-20`](../../internal/grpc/README.md#L1-L20) +- The architecture document says operators can run `flux serve `, but the Makefile explicitly says Flux is a library with no standalone binary. This is a direct product-surface contradiction. [`docs/ARCHITECTURE.md:52-59`](../../docs/ARCHITECTURE.md#L52-L59), [`Makefile:1-3`](../../Makefile#L1-L3) +- The architecture's package map includes `errors/`, but the current root tree has no tracked `errors/` package. The host-facing package list is therefore not a generated or checked source of truth. [`README.md:288-325`](../../README.md#L288-L325) +- The README says hosts may import exactly `engine`, `llm`, `graph`, and `tools`, but the host `llm.Provider` combines generation, catalog, credential, selection, gateway, catalog-maintenance, and native-compaction responsibilities. That is a coherent Rho composition boundary, but it is not the smallest API for a general Go adopter. [`README.md:46-62`](../../README.md#L46-L62), [`llm/provider.go:16-24,137-152,201-212,244-281,317-328,388-402`](../../llm/provider.go#L16-L24) + +#### Quickstart and first-run experience + +- The README quickstart says Go 1.26+, a configured credential, `engine.New(Options{})`, then `Stream`; it does not initialize a catalog, persist provider state, save a key, or call a catalog refresh. [`README.md:67-105`](../../README.md#L67-L105) +- `Engine.loadRuntimeState` always loads the catalog with `RequireCache: true`; a missing cache returns `ErrCatalogCacheRequired`, and the error text tells the caller to run `rho models refresh`, a command outside Flux and unavailable to independent adopters. [`engine/state.go:86-103`](../../engine/state.go#L86-L103), [`catalog/v1.go:615-644`](../../catalog/v1.go#L615-L644), [`catalog/errors.go:3-7`](../../catalog/errors.go#L3-L7) +- `engine.New` silently chooses `~/.flux/model_catalog.json` when no state directory is supplied, while its provider path still comes from Flux's process-default configuration resolver. The quickstart therefore depends on undocumented local state. [`engine/engine.go:70-117`](../../engine/engine.go#L70-L117), [`catalog/refresh.go:26-33`](../../catalog/refresh.go#L26-L33) +- The default remote catalog is hosted at `https://langdag.com/model-catalog/v1/catalog.json`, outside the documented GrayCode ecosystem and without an in-repository publication workflow. [`catalog/v1.go:44-50`](../../catalog/v1.go#L44-L50), [`DYNAMIC-MODEL-DISCOVERY.md:90-108`](../../docs/guides/DYNAMIC-MODEL-DISCOVERY.md#L90-L108) +- On 2026-09-24, that default catalog was reachable and declared `generated_at: 2026-09-23`, but its `openai-direct` deployment included `api_protocol_ids`, a field absent from Flux's `Deployment` type. Because `ParseCatalog` enables `DisallowUnknownFields`, the current default document is expected to fail decoding. This is a verified compatibility break between the shipped default source and the checkout, and CI does not test the live default URL. [current default catalog](https://langdag.com/model-catalog/v1/catalog.json), [`catalog/v1.go:85-96,444-460`](../../catalog/v1.go#L85-L96) +- The catalog's `Provenance` type is descriptive metadata only (`source`, `source_url`, `observed_at`); there is no signature, digest pin, or provenance verification in the catalog package. [`catalog/v1.go:287-291`](../../catalog/v1.go#L287-L291) +- The catalog decoder deliberately rejects every unknown field, which makes additive producer evolution immediately breaking for consumers and directly conflicts with the broader engine policy that future event/field evolution should be additive. [`catalog/v1.go:444-460`](../../catalog/v1.go#L444-L460), [`HOST-ENGINE-BOUNDARY.md:188-193`](../../docs/architecture/HOST-ENGINE-BOUNDARY.md#L188-L193) + +#### Documentation coherence + +- The README's 28-provider table is mostly clear and useful, but the credential guide still says “Supported providers (15)” and lists a stale `z-ai` ID while omitting 13 registered gateways. [`README.md:200-235`](../../README.md#L200-L235), [`CREDENTIAL-SETUP-FLOW.md:7-27`](../../docs/guides/CREDENTIAL-SETUP-FLOW.md#L7-L27) +- StepFun is `STEPFUN_API_KEY` in the README and `.env.example`, but the canonical registry and constructor use `STEP_API_KEY`. `.env.example` also includes noncanonical `KIMI_API_KEY` while the registry uses `MOONSHOT_API_KEY`. [`README.md:225-233`](../../README.md#L225-L233), [`.env.example:1-34`](../../.env.example#L1-L34), [`catalog/registry/providers.go:72-81,230-240`](../../catalog/registry/providers.go#L72-L81) +- The credential guide's “adding a provider” checklist names only registry data, a live fetcher, and remote catalog metadata. Actual provider onboarding also touches compatibility configuration, runtime construction, provider config/credential inference, tests, env templates, and often protocol-specific logic. [`CREDENTIAL-SETUP-FLOW.md:80-85`](../../docs/guides/CREDENTIAL-SETUP-FLOW.md#L80-L85), [`catalog/live/fetchers.go:62-100`](../../catalog/live/fetchers.go#L62-L100), [`setup/deployment.go:183-383`](../../setup/deployment.go#L183-L383), [`provider/provider_registry.go:74-180`](../../provider/provider_registry.go#L74-L180) +- The docs index advertises planned retry/fallback and caching/audit guides, but those files do not exist. It also shows the OpenAPI file under `docs/api/` in its tree even though the tracked path is root-level `api/openapi.yaml`. [`docs/README.md:31-46`](../../docs/README.md#L31-L46) +- The host-boundary and decentralized-routing documents are unusually candid about secret isolation, last-good route behavior, missing persistence, absent consensus, and the lack of production load/fault-injection evidence. These are high-quality OSS trust documents and should be retained. [`HOST-ENGINE-BOUNDARY.md:108-127,156-193`](../../docs/architecture/HOST-ENGINE-BOUNDARY.md#L108-L127), [`DECENTRALIZED-FLUX.md:64-88`](../../docs/architecture/DECENTRALIZED-FLUX.md#L64-L88) +- The 469-line enterprise design is explicitly a proposed 45-engine-week product expansion into SSO/RBAC, dashboards, HQL, prompt management, canaries, A2A, fine-tuning, and SLA queues. It is not an executable current roadmap and would distract from adoption-critical work. [`FLUX-ENTERPRISE.md:1-22,303-379,442-465`](../../docs/design/FLUX-ENTERPRISE.md#L1-L22) + +#### Examples and API documentation + +- All three examples import lower-level `provider` and `provider/core`, directly contradicting the README's instruction that hosts use the stable `engine` facade. [`examples/basic/main.go:8-20`](../../examples/basic/main.go#L8-L20), [`examples/streaming/main.go:8-20`](../../examples/streaming/main.go#L8-L20), [`examples/multi-provider/main.go:10-25`](../../examples/multi-provider/main.go#L10-L25) +- The basic example always requests `claude-sonnet-4-6` even though it auto-detects the provider, and it dereferences `resp.Usage` without a nil check. [`examples/basic/main.go:17-35`](../../examples/basic/main.go#L17-L35) +- The streaming example is described as having auto-continuation but calls `StreamChat`, not `StreamChatContinue`; the explicit continuation API is a different method. [`examples/streaming/main.go:1-6,26-33`](../../examples/streaming/main.go#L1-L26), [`provider/chat.go:45-81`](../../provider/chat.go#L45-L81) +- The multi-provider example is manual error handling, not a Flux fallback chain, despite its description. [`examples/multi-provider/main.go:1-7,31-44`](../../examples/multi-provider/main.go#L1-L7) +- OpenAPI documents 11 paths, while the server registers 15. It omits `/ready`, `/rerank`, and `/v1/chat/completions`, even though README and architecture list all three. The prompt schema also omits implemented `system_prompt`, `max_tokens`, and `tools` fields. [`api/openapi.yaml:30-43,118-340`](../../api/openapi.yaml#L30-L43), [`internal/api/server.go:101-121,163-170`](../../internal/api/server.go#L101-L121) +- The OpenAPI version is a hand-maintained literal `0.0.1`; no workflow validates it against routes, implementation, or the module version. [`api/openapi.yaml:1-10`](../../api/openapi.yaml#L1-L10), [`ci.yml:1-300`](../../.github/workflows/ci.yml#L1-L300) +- The TypeScript and Python SDKs live under `internal/sdk`, making their current import paths unusable by external Go/Python/TypeScript consumers. The Go SDK is additionally a nested module excluded from root `go test ./...`. [`internal/sdk/go/go.mod:1-3`](../../internal/sdk/go/go.mod#L1-L3), [`internal/sdk/typescript/flux.ts:1-175`](../../internal/sdk/typescript/flux.ts#L1-L175), [`internal/sdk/python/flux.py:1-188`](../../internal/sdk/python/flux.py#L1-L188) +- The TypeScript package points directly at an unexported `FluxClient` source file, has no build/test scripts, and has no lockfile. The Python client has no tests and expects JSON from the node-delete endpoint even though the API returns 204. These are internal stubs, not distributable SDKs. [`internal/sdk/typescript/package.json:1-8`](../../internal/sdk/typescript/package.json#L1-L8), [`internal/sdk/python/flux.py:68-75`](../../internal/sdk/python/flux.py#L68-L75), [`api/openapi.yaml:223-237`](../../api/openapi.yaml#L223-L237) +- A fresh main-module test run reports only 52 packages and does not include the nested Go SDK. Manual `go test` inside that nested module passed, confirming the code exists but is outside the root gate. [`internal/sdk/go/go.mod:1-3`](../../internal/sdk/go/go.mod#L1-L3), [`go.mod:1-41`](../../go.mod#L1-L41) + +#### GoDoc and public API maturity + +- `engine/doc.go` clearly states the host boundary, and most engine DTOs have useful comments and an explicit event vocabulary. This is a good foundation for a stable package. [`engine/doc.go:1-14`](../../engine/doc.go#L1-L14), [`engine/types.go:62-84`](../../engine/types.go#L62-L84) +- The lower-level `core.Provider` is admirably small: `Chat`, `StreamChat`, `Ping`, and `Name`. [`provider/core/core.go:21-33`](../../provider/core/core.go#L21-L33) +- Public comments still contain stale ecosystem terms: `core` refers to “eagle/llm,” while `llm.Provider` describes itself as “rho's rho-owned view.” That is confusing in a standalone OSS package. [`provider/core/core.go:1-12`](../../provider/core/core.go#L1-L12), [`llm/provider.go:9-15`](../../llm/provider.go#L9-L15) +- A local `go list -f '{{if not .Doc}}...'` check found 17 packages without package documentation, including public `provider/adapters`, `provider/cache`, `provider/observability`, `provider/testkit`, `catalog/registry`, and `tools`. [`provider/adapters/provider_registry.go:1-12`](../../provider/adapters/provider_registry.go#L1-L12), [`catalog/registry/registry.go:1-18`](../../catalog/registry/registry.go#L1-L18) +- A local `go doc -all` count found approximately 146 exported declarations in `engine`, 134 in `provider/core`, 122 in `catalog`, 65 in `llm`, 58 in `router`, and 48 each in `provider` and `credentials`. There is no stability annotation for these packages beyond prose. [`engine/types.go:1-164`](../../engine/types.go#L1-L164), [`HOST-ENGINE-BOUNDARY.md:188-193`](../../docs/architecture/HOST-ENGINE-BOUNDARY.md#L188-L193) +- The `engine.ContractVersion == "2"` marker exists before a meaningful independent compatibility baseline: the current module has only `v0.0.1`, and no `apidiff` or previous-tag API comparison job exists. [`engine/engine.go:22-23`](../../engine/engine.go#L22-L23), [`ci.yml:1-300`](../../.github/workflows/ci.yml#L1-L300) + +#### Contribution, security, community, and release processes + +- Positive process assets exist: MIT license, conventional-commit guidance, contribution/security/code-of-conduct documents, issue forms, PR template, CODEOWNERS, lefthook, release workflow, Scorecard, and protected `main`. [`CONTRIBUTING.md:1-136`](../../CONTRIBUTING.md#L1-L136), [`SECURITY.md:1-71`](../../SECURITY.md#L1-L71), [`.github/PULL_REQUEST_TEMPLATE.md:1-59`](../../.github/PULL_REQUEST_TEMPLATE.md#L1-L59) +- The PR template contradicts the contribution guide about changelog ownership: the template says to update `CHANGELOG.md`, while the guide says release automation must generate it and contributors must not edit it. [`PULL_REQUEST_TEMPLATE.md:46-59`](../../.github/PULL_REQUEST_TEMPLATE.md#L46-L59), [`CONTRIBUTING.md:98-109`](../../CONTRIBUTING.md#L98-L109) +- The contribution guide says `make ci` runs “everything CI runs,” but the Makefile gate omits deadcode, duplication, secrets, markdown, gRPC-tag testing, and the cross-platform build matrix. Conversely, `make ci` mutates the tree through tidy and format. [`CONTRIBUTING.md:14-36`](../../CONTRIBUTING.md#L14-L36), [`Makefile:47-108`](../../Makefile#L47-L108), [`ci.yml:34-300`](../../.github/workflows/ci.yml#L34-L300) +- There are two CODEOWNERS files. The `.github/CODEOWNERS` file that GitHub normally discovers has stale `/client/` coverage and no ownership for current `provider/`, `engine/`, or `credentials/` paths; the root duplicate lists `/client/` as well. Team validity and actual maintainer availability cannot be verified from the public repository. [`.github/CODEOWNERS:1-22`](../../.github/CODEOWNERS#L1-L22), [`CODEOWNERS:1-20`](../../CODEOWNERS#L1-L20) +- `SECURITY.md` claims release checksums come from GoReleaser and that Ruff/Mypy/Python lockfiles run in CI. The actual release workflow only creates a GitHub release, there is no GoReleaser config, no Python/TypeScript CI, no Ruff/Mypy job, and no pnpm lockfile. [`SECURITY.md:45-57`](../../SECURITY.md#L45-L57), [`release.yml:1-31`](../../.github/workflows/release.yml#L1-L31) +- The README and contribution guide repeatedly claim release-please automation, but no release-please workflow/config is tracked. [`README.md:358-369`](../../README.md#L358-L369), [`CONTRIBUTING.md:38-42,98-107`](../../CONTRIBUTING.md#L38-L42) +- Branch protection is materially better than average: all 15 CI contexts are required, strict up-to-date checks are enabled, force pushes/deletions are blocked, and admin enforcement is on. However, no required approving review, conversation resolution, required signed commits, or required signed pushes appeared in the public branch-protection response. [branch protection API](https://api.github.com/repos/GrayCodeAI/flux/branches/main/protection), [`CONTRIBUTING.md:111-120`](../../CONTRIBUTING.md#L111-L120) +- GitHub reports vulnerability alerts, Dependabot security updates, secret scanning, and secret-scanning push protection disabled. TruffleHog in CI provides useful coverage, but it is not equivalent to repository-native alert and push-protection services. [repository security settings](https://api.github.com/repos/GrayCodeAI/flux), [`ci.yml:228-238`](../../.github/workflows/ci.yml#L228-L238) +- Git log verification found one human contributor across all commits; GitHub currently reports 3 stars, 0 forks, and 0 issues. This is a very young project with no public adoption or issue-triage history yet. [repository](https://github.com/GrayCodeAI/flux) +- `v0.0.1` is an annotated, signed tag; local `git tag -v` succeeded with an ED25519 signature. This is a real positive trust signal. The release workflow itself neither requires nor verifies that signature. [v0.0.1 release](https://github.com/GrayCodeAI/flux/releases/tag/v0.0.1), [`release.yml:15-31`](../../.github/workflows/release.yml#L15-L31) +- The sole release points to commit `34010bf`; local history shows 29 commits from `v0.0.1` to current `main`. The generated release notes list many unrelated PRs, while the checked-in changelog has conflicting ordering: `0.0.1` appears before a later-dated `0.1.0` section and includes pre-rename history. [v0.0.1 release](https://github.com/GrayCodeAI/flux/releases/tag/v0.0.1), [`CHANGELOG.md:35-50,175-235`](../../CHANGELOG.md#L35-L50) +- The immediately previous module was `github.com/GrayCodeAI/eyrie` at version `0.6.0`; the current module is `github.com/GrayCodeAI/flux` at `0.0.1`. The changelog mentions old module paths but supplies no migration guide with import mapping, state migration ordering, compatibility matrix, or rollback advice. [`go.mod:1-3`](../../go.mod#L1-L3), [`VERSION:1`](../../VERSION#L1), [`CHANGELOG.md:35-50`](../../CHANGELOG.md#L35-L50) +- `go list -m -versions github.com/GrayCodeAI/flux` returned only `v0.0.1`, confirming minimal current-module release history. [module versions](https://pkg.go.dev/github.com/GrayCodeAI/flux?tab=versions) +- `go list -m -u all` found many available updates, including OTel `1.44.0 -> 1.46.0`, gRPC `1.82.1 -> 1.84.0`, tokenizer `0.8.0 -> 0.8.1`, SQLite `1.51.0 -> 1.59.0`, and several `golang.org/x/*` updates. No Dependabot or Renovate configuration is tracked. [`go.mod:5-41`](../../go.mod#L5-L41) +- The versioning policy is external to Flux, in the Rho repository. An independently adoptable library should own its compatibility and support policy rather than requiring consumers to follow a sibling product's document. [`CONTRIBUTING.md:3-5`](../../CONTRIBUTING.md#L3-L5), [`SECURITY.md:9-11`](../../SECURITY.md#L9-L11) + +### Inferences + +- Flux should present one product: **a Go library for normalized, resilient calls to heterogeneous LLM providers**. HTTP/gRPC serving, conversation storage, distributed control plane, SDKs, budgets, and enterprise administration should either be separately versioned opt-in modules or removed from the core promise until they are real external surfaces. The current mixed story makes it difficult for an adopter to know what is supported, stable, and maintained. [`README.md:31-44,154-176`](../../README.md#L31-L44) +- The fresh-install catalog requirement is a release blocker, not a documentation nicety. A new user currently needs state prepared by an ecosystem host, while the default remote catalog is already schema-incompatible. A minimal embedded bootstrap or explicit `engine.Bootstrap(ctx)`/local-cache constructor is needed before broader release claims. [`engine/state.go:86-103`](../../engine/state.go#L86-L103), [`catalog/v1.go:444-460`](../../catalog/v1.go#L444-L460) +- The documentation should be made partly generated and continuously checked rather than manually expanded: provider/env tables from `ProviderSpec`, OpenAPI route parity from a public handler, examples as external-package tests, and release metadata from `VERSION`. This directly prevents the drift already present in 28-vs-15 provider counts, env-name drift, and OpenAPI omissions. [`catalog/registry/spec.go:31-74`](../../catalog/registry/spec.go#L31-L74), [`internal/api/server.go:101-121`](../../internal/api/server.go#L101-L121) +- The existing `v0.0.1` tag should not be followed by a `v1.0.0` jump. Publish a truthful `v0.1.0` “library foundation” release after the quickstart, engine middleware persistence, docs, and external-consumer smoke gate are fixed; use `v0.2.x` for extension seams and reserve `v1.0.0` for an audited compatibility baseline. [`VERSION:1`](../../VERSION#L1), [`release.yml:1-31`](../../.github/workflows/release.yml#L1-L31) +- The enterprise RFC should remain an RFC/design archive, not a public roadmap. The highest-return work is provider conformance, onboarding, deterministic releases, and operational evidence—not 45 engineer-weeks of adjacent product surface. [`FLUX-ENTERPRISE.md:303-379,442-465`](../../docs/design/FLUX-ENTERPRISE.md#L303-L379) + +### Gaps + +- No provider API keys or organization credentials were available, so this audit did not verify live model discovery, credential probes, wire compatibility, regional endpoints, or billing behavior against any hosted provider. Existing tests overwhelmingly use `httptest` and mocks. [`catalog/provider_live_parity_test.go:11-33`](../../catalog/provider_live_parity_test.go#L11-L33) +- Public GitHub metadata does not reveal whether the CODEOWNERS teams exist, whether Discussions are enabled, or how responsive the sole maintainer is. [repository metadata](https://github.com/GrayCodeAI/flux) +- The governance, ownership, signing-key custody, incident response, and release authority for the `langdag.com` catalog are not documented in this repository. [default catalog](https://langdag.com/model-catalog/v1/catalog.json) +- There is no public usage telemetry, support history, downstream adopter list, or production SLO evidence from which to assess real operability. Flux itself explicitly says distributed routing has no production load or fault-injection baseline. [`DECENTRALIZED-FLUX.md:64-88`](../../docs/architecture/DECENTRALIZED-FLUX.md#L64-L88) +- The TypeScript check was inconclusive because the installed global TypeScript/DOM libraries failed before yielding a source-specific result; this should not be attributed to Flux without a pinned toolchain run. [`internal/sdk/typescript/tsconfig.json:1-11`](../../internal/sdk/typescript/tsconfig.json#L1-L11) + +## 2. Test maturity, critical-path confidence, and extension ergonomics + +### Takeaway + +Flux's deterministic test suite is a major strength: 2,162 test functions across 271 files, extensive provider/catalog coverage, a real race pass, 65.6% aggregate coverage, and 75 benchmarks. Confidence drops sharply at the boundaries that matter to adopters—the clean-machine quickstart, current remote catalog, actual `engine` transport composition, cross-request middleware state, real provider behavior, nested SDKs, API compatibility, distributed processes, and performance/soak behavior. + +### Cited Findings + +#### Verified command results + +| Check | Result on 2026-09-24 | Interpretation | +|---|---|---| +| `go version && go mod verify` | Go `1.26.6`; all modules verified | Dependency integrity for the checked-in graph is good. [`go.mod:1-41`](../../go.mod#L1-L41) | +| `go test -race -count=1 -shuffle=on -coverprofile=... -covermode=atomic ./...` | All main-module packages passed; aggregate coverage **65.6%** | Strong deterministic/race health; coverage is only 5.6 points above the CI floor. [`ci.yml:138-172`](../../.github/workflows/ci.yml#L138-L172) | +| `go vet ./...` | Passed with no output | Compiler/static baseline is clean. [`ci.yml:101-114`](../../.github/workflows/ci.yml#L101-L114) | +| `go test -tags=grpc -count=1 ./...` | Passed, including `internal/grpc` | The optional gRPC build works, although untagged CI does not run it. [`internal/grpc/server_grpc.go:1-3`](../../internal/grpc/server_grpc.go#L1-L3) | +| `govulncheck ./...` | `No vulnerabilities found` | No currently reachable known Go vulnerability was found; repository-native alerts remain disabled. [`ci.yml:174-194`](../../.github/workflows/ci.yml#L174-L194), [repository security settings](https://api.github.com/repos/GrayCodeAI/flux) | +| `gofumpt -l . && goimports -l .` | No files listed | Formatting matches the pinned tools. [`ci.yml:35-68`](../../.github/workflows/ci.yml#L35-L68) | +| `golangci-lint v2.1.0 run --timeout=5m` | Reported `0 issues`, then the local run timed out | Treat lint as locally inconclusive; the latest GitHub CI run completed successfully. [latest CI run](https://github.com/GrayCodeAI/flux/actions/runs/35504614597) | +| Nested Go SDK `go test ./... && go vet ./...` | Passed | Useful tests exist but are outside the main CI/module gate. [`internal/sdk/go/go.mod:1-3`](../../internal/sdk/go/go.mod#L1-L3) | +| Python `ruff check . && mypy flux.py` | Ruff found two unused imports; mypy passed | The stub is not under an enforced Python gate. [`internal/sdk/python/flux.py:1-9`](../../internal/sdk/python/flux.py#L1-L9) | +| TypeScript `tsc --noEmit` | Failed in installed DOM library declarations | Inconclusive source-quality result; no pinned project toolchain runs in CI. [`internal/sdk/typescript/tsconfig.json:1-11`](../../internal/sdk/typescript/tsconfig.json#L1-L11) | + +#### Test volume and distribution + +- The repository contains **271 `*_test.go` files, 2,162 `Test*` functions, 4 fuzz targets, and 75 benchmark functions**. Test files total about 56,127 lines versus 46,869 production Go lines. This is excellent test investment. [`fuzz_test.go:1-134`](../../provider/fuzz_test.go#L1-L134), [`benchmarks_test.go:1-281`](../../provider/benchmarks_test.go#L1-L281) +- Test functions are concentrated in `provider` (1,009), `catalog` (309), `internal` (170), `config` (135), `runtime` (121), `engine` (91), `setup` (84), `credentials` (74), and `router` (68). This is appropriate for a provider runtime, but the stable engine is not the largest test domain. [`engine/contract_e2e_test.go:1-184`](../../engine/contract_e2e_test.go#L1-L184) +- There are 191 test files using `t.Parallel` and 69 using `httptest`, showing substantial concurrency and wire-protocol simulation. [`stream_test.go:17-113`](../../provider/core/stream_test.go#L17-L113) +- Only six true Go `Example*` functions exist, all under `types` and one adaptive-rate-limit example; there are no executable engine quickstart examples. [`types/example_test.go:1-56`](../../types/example_test.go#L1-L56), [`adaptive_ratelimit_test.go:501-534`](../../provider/resilience/adaptive_ratelimit_test.go#L501-L534) +- One circuit-breaker timing test is explicitly skipped as flaky, with a manual deterministic replacement. This is transparent, but it identifies remaining wall-clock sensitivity. [`circuitbreaker_test.go:100-110`](../../router/circuitbreaker_test.go#L100-L110) + +#### Coverage strengths + +- Core provider wire tests cover Anthropic, OpenAI-compatible behavior, cloud transports, streaming, tools, thinking, usage, retries, errors, and security behavior using realistic local servers. The shared SSE tests cover multiline data, cancellation, tool calls, thinking deltas, and TTFT. [`stream_test.go:17-250`](../../provider/core/stream_test.go#L17-L250) +- Catalog tests are extensive across provider registration, protocol matrices, deployment env derivation, cache loading, live parsing, deprecation, and topology. [`provider_live_parity_test.go:11-84`](../../catalog/provider_live_parity_test.go#L11-L84), [`protocol_matrix_test.go:27-87`](../../catalog/registry/protocol_matrix_test.go#L27-L87) +- Security tests are unusually broad for secret persistence, state sanitization, atomic migration, environment conflict detection, keyring cancellation behavior, and no-secret-leak assertions. [`state_security_test.go:1-220`](../../engine/state_security_test.go#L1-L220), [`keyring_platform_test.go:12-161`](../../credentials/keyring_platform_test.go#L12-L161) +- Distributed control-plane tests validate atomic last-good application, revision conflicts, signed manifest tampering, and periodic refresh. They use an in-memory `RoundTripper`, not TLS or separate processes. [`controlplane_test.go:64-196,198-313`](../../router/controlplane/controlplane_test.go#L64-L196) +- A public `verify` package describes a behavioral conformance harness with canonical chat/tool cases and baseline diffs. In practice, its only uses are its own fake-provider tests; no adapter, provider, workflow, or example invokes it. [`verify/verify.go:1-206`](../../verify/verify.go#L1-L206), [`verify/cases.go:1-46`](../../verify/cases.go#L1-L46), [`verify_test.go:13-149`](../../verify/verify_test.go#L13-L149) +- The engine contract E2E test is valuable for normalized DTO/event behavior, but it replaces the private `resolveTransport` with a mock. It does not exercise catalog-to-adapter construction, real routing, middleware, credentials, or the default transport. [`engine/contract_e2e_test.go:43-113`](../../engine/contract_e2e_test.go#L43-L113) +- HTTP “integration” tests run an `httptest.Server` with local SQLite and mock providers. They are solid component integration tests, not deployable end-to-end tests. [`integration_test.go:20-116`](../../internal/api/integration_test.go#L20-L116) + +#### Coverage gaps and weak critical paths + +- Aggregate coverage hides uneven critical areas: `provider/resilience` is 28.3%, `provider/media` 26.2%, `provider/observability` 39.1%, `provider/extraction` 44.1%, `operationsgraph` 44.1%, `graph` 12.7%, untagged `internal/grpc` 20.0%, and the root/`llm`/testkit packages report 0% in their own package runs. These numbers came from the verified race/coverage command and align with the package layout. [`operationsgraph/projection.go:1-200`](../../operationsgraph/projection.go#L1-L200), [`provider/resilience/roles.go:1-120`](../../provider/resilience/roles.go#L1-L120) +- The CI coverage command does not use `-coverpkg=./...`, despite an audit plan claiming that was done. Its 65.6% gate is therefore vulnerable to package-boundary accounting and does not enforce minimums for reliability, engine transport, or operations-critical packages. [`ci.yml:154-166`](../../.github/workflows/ci.yml#L154-L166), [`audit-remediation.md:272-279`](../../docs/plans/audit-remediation.md#L272-L279) +- There is no test for a fresh `engine.New(...).Stream(...)` path with an empty temporary `StateDir`; tests explicitly write `catalog.SeedCatalog()` first. This misses the README's primary adoption path. [`engine/contract_e2e_test.go:43-62`](../../engine/contract_e2e_test.go#L43-L62) +- There is no CI test against the current default remote catalog. The remote-source tests inject local HTTP clients/servers, so a producer schema addition can break production while CI remains green. [`v1_test.go:83-234`](../../catalog/v1_test.go#L83-L234), [default catalog](https://langdag.com/model-catalog/v1/catalog.json) +- No engine test invokes `defaultTransport`, `DeploymentProviderFromState`, or verifies that cache/rate-limit/circuit-breaker state survives multiple requests. Searches found only tests of DTO conversion and injected private transports. [`engine/engine.go:283-327`](../../engine/engine.go#L283-L327), [`contract_e2e_test.go:61-62`](../../engine/contract_e2e_test.go#L61-L62) +- Fuzzing covers only message sanitation, role merging, cache-key determinism, and guardrails. It does not target SSE framing, JSON decoding, retry-after parsing, auth redirect handling, catalog parsing, state migration, tool-argument accumulation, or stream cancellation. [`fuzz_test.go:10-134`](../../provider/fuzz_test.go#L10-L134), [`ci.yml:253-273`](../../.github/workflows/ci.yml#L253-L273) +- Benchmarks exist for cache keys, request building, sanitization, guardrails, metrics, circuit breakers, and adaptive limiting, but CI never runs them and no committed baseline, `benchstat` comparison, regression threshold, or published result exists. [`Makefile:73-74`](../../Makefile#L73-L74), [`ci.yml:34-300`](../../.github/workflows/ci.yml#L34-L300) +- No test starts separate Flux processes, serves TLS, restarts a replica from persisted state, injects network partitions/partitions clocks, or runs a soak/load profile. The distributed-routing document acknowledges these missing tests. [`DECENTRALIZED-FLUX.md:64-88`](../../docs/architecture/DECENTRALIZED-FLUX.md#L64-L88) +- The only Go compatibility axis is one exact toolchain, Go `1.26.6`; CI varies OS/architecture but not supported Go versions or the previous released API. [`ci.yml:21-23,275-300`](../../.github/workflows/ci.yml#L21-L23) +- The nested Go SDK tests are not run by root CI. No workflow tests Python or TypeScript. The main CI has no clean external-module `go get`/compile test against the published tag. [`internal/sdk/go/go.mod:1-3`](../../internal/sdk/go/go.mod#L1-L3), [`ci.yml:1-300`](../../.github/workflows/ci.yml#L1-L300) +- No test compares OpenAPI paths/schemas to registered routes, and no API compatibility tool checks exported Go symbols against the prior release. [`api/openapi.yaml:118-340`](../../api/openapi.yaml#L118-L340), [`internal/api/server.go:101-121`](../../internal/api/server.go#L101-L121) + +#### Critical functional gaps exposed by test architecture + +- `Engine` stores no memoized transport. Every generation calls `defaultTransport`, which reloads state, builds a new deployment router, and wraps it in newly allocated rate-limit/cache decorators. Circuit-breaker, rate-window, in-flight, and cache state therefore reset each call. [`engine/engine.go:54-68,283-327`](../../engine/engine.go#L54-L68), [`setup/deployment.go:94-110`](../../setup/deployment.go#L94-L110) +- Even if the cache wrapper were retained, `NewCachedProvider` claims zero config fields are replaced by defaults but copies `cfg.Enabled` directly; a zero `CacheConfig` is disabled. Thus `engine.Options{EnableCaching: true}` with the documented zero config does not cache. [`engine/engine.go:45-51,323-325`](../../engine/engine.go#L45-L51), [`provider/cache/semantic_cache.go:15-38,73-92`](../../provider/cache/semantic_cache.go#L15-L38) +- `Engine` describes its cache as semantic, but the wrapper it creates is deterministic hash/LRU caching. A real embedding-similarity cache exists in `provider/embeddings`, but is not wired into `engine.Options`. [`engine/engine.go:45-51`](../../engine/engine.go#L45-L51), [`provider/cache/semantic_cache.go:52-68,296-345`](../../provider/cache/semantic_cache.go#L52-L68), [`provider/embeddings/cache.go:17-90`](../../provider/embeddings/cache.go#L17-L90) +- The lower-level transport has a shared `http.Transport`, but no proxy support, forced HTTP/2 setting, response-header timeout, or custom redirect policy. Tests assert pool identity and a few timeout values, not proxy, redirects, auth-header stripping, or response-header stalls. [`provider/core/transport.go:23-64`](../../provider/core/transport.go#L23-L64), [`transport_test.go:9-84`](../../provider/core/transport_test.go#L9-L84) +- The SSE parser has useful unit tests, but the default 2 MiB scanner cap, idle/stall behavior, cross-host redirect behavior, and cross-request resource usage are not covered by contract or performance tests. [`provider/core/stream.go:22-89`](../../provider/core/stream.go#L22-L89), [`stream_test.go:17-113`](../../provider/core/stream_test.go#L17-L113) + +#### Extension-seam assessment + +| Extension target | Current seam | Assessment | +|---|---|---| +| OpenAI-compatible provider/gateway | `FluxClient.RegisterCustomProvider` and `engine.Options.CustomGateways` validate names/URLs and remain instance-local. [`provider/dynamic.go:11-50`](../../provider/dynamic.go#L11-L50), [`engine/host_runtime.go:19-106`](../../engine/host_runtime.go#L19-L106) | **Good.** This is the strongest extension path. | +| Credential backend | Three-method `credentials.Store` is injected through `engine.Options.SecretStore`. [`credentials/store.go:27-32`](../../credentials/store.go#L27-L32), [`engine/engine.go:25-38`](../../engine/engine.go#L25-L38) | **Good core interface**, but service naming remains a mutable process global and provider state is string-account based. [`credentials/store.go:11-25`](../../credentials/store.go#L11-L25) | +| Lower-level HTTP transport | Adapters accept `core.WithHTTPClient`; adapters implement a common configurable setter surface. [`provider/core/options.go:9-24,66-94`](../../provider/core/options.go#L9-L24) | **Good for advanced users**, but not exposed through `engine.Options`. | +| Custom engine transport | The hook is the unexported `Engine.resolveTransport` function field, mutated only by same-package tests. [`engine/engine.go:54-68,283-300`](../../engine/engine.go#L54-L68) | **Poor.** Independent adapters cannot participate in normal engine resolution. | +| Native provider | `ProviderSpec`, compatibility maps, live fetcher registry, `setup` switch, `provider` switch, config inference, env docs, and tests are separate edits. [`catalog/registry/providers.go:16-342`](../../catalog/registry/providers.go#L16-L342), [`catalog/live/fetchers.go:62-100`](../../catalog/live/fetchers.go#L62-L100), [`setup/deployment.go:183-383`](../../setup/deployment.go#L183-L383) | **Poor-to-moderate.** Metadata is declarative, construction is not. | +| Model | Live provider fetchers plus a remote `model-catalog/v1` document. [`catalog/live/fetchers.go:52-100`](../../catalog/live/fetchers.go#L52-L100), [`catalog/v1.go:44-67`](../../catalog/v1.go#L44-L67) | **Weak for ecosystem contributors.** No in-repo generator/publication/signature workflow; strict decoding makes producer changes brittle. | +| Routing strategy | String enum plus a central `switch`; no strategy interface. [`router/strategy.go:10-31,107-168`](../../router/strategy.go#L10-L31) | **Poor for third-party extension.** New strategies require modifying Flux. | +| Cache backend | `CacheBackend` and a Redis skeleton exist only under `internal/cache`; neither is wired into engine caching. [`internal/cache/backend.go:23-33,97-139`](../../internal/cache/backend.go#L23-L33) | **Not a public extension seam.** Redis explicitly lacks pooling, auth, DB selection, and reconnection guarantees. | +| Audit sink | `AuditSink` and JSONL sink exist only under `internal/observability` and have no production caller outside their own tests. [`internal/observability/audit.go:1-100`](../../internal/observability/audit.go#L1-L100) | **Not a public extension seam and currently inert.** | +| Telemetry | A public lower-level OTel `TracingProvider` exists, but engine does not wire it and it emits custom `provider.name`/`usage.*` attributes rather than the repository's `gen_ai.*` vocabulary. [`provider/observability/tracing.go:12-68`](../../provider/observability/tracing.go#L12-L68), [`internal/observability/genai_semconv.go:1-80`](../../internal/observability/genai_semconv.go#L1-L80) | **Partial.** Suitable for manual lower-level wrapping, not a coherent engine sink contract. | +| HTTP/gRPC server | Implementations are internal and therefore unavailable to ordinary external importers. [`internal/api/server.go:1-16`](../../internal/api/server.go#L1-L16), [`internal/grpc/grpc.go:1-120`](../../internal/grpc/grpc.go#L1-L120) | **Not adoptable as a standalone OSS service today.** | + +### Inferences + +- The suite currently provides strong confidence that individual packages behave as their same-package unit tests define them. It provides only moderate confidence that the stable `engine` composition retains state, a fresh user can start, or adapters agree with current hosted APIs. The missing confidence is concentrated at seams, not inside leaf functions. [`engine/contract_e2e_test.go:43-113`](../../engine/contract_e2e_test.go#L43-L113) +- A percentage-only 60% coverage gate is not the best next step. Add package-specific floors for `engine/defaultTransport`, `provider/core` transport/SSE, `router`, `credentials`, and catalog parsing, plus one clean-machine external consumer test; do not spend the next cycle chasing aggregate coverage in peripheral packages. [`ci.yml:154-166`](../../.github/workflows/ci.yml#L154-L166) +- The existing `verify` package should become the common adapter contract suite. Every built-in adapter should run the same sanitized fixture corpus for blocking chat, streaming, tools, usage, errors, cancellation, and prompt/provider-block replay. Live provider calls should be a separate opt-in/nightly job, not the deterministic PR gate. [`verify/verify.go:1-15`](../../verify/verify.go#L1-L15) +- Provider breadth should become support-tiered rather than expanded blindly: Tier A for Anthropic/OpenAI/Gemini/Azure/Bedrock/Vertex with nightly live conformance; Tier B for compatible gateways with deterministic wire fixtures; custom OpenAI-compatible endpoints remain user responsibility. This improves trust more than adding another gateway name. [`catalog/provider_live_parity_test.go:11-33`](../../catalog/provider_live_parity_test.go#L11-L33) +- Routing, cache, telemetry, and native provider extension should be unified behind a small set of public interfaces or explicitly declared non-extensible. Today several README features are implemented as isolated lower-level or internal components without a reachable composition path. [`router/strategy.go:107-168`](../../router/strategy.go#L107-L168), [`internal/cache/backend.go:23-33`](../../internal/cache/backend.go#L23-L33), [`internal/observability/audit.go:38-41`](../../internal/observability/audit.go#L38-L41) +- The first performance baseline should be a reproducible client benchmark against a local fake provider and local OpenAI-compatible server, measuring construction, catalog load, keychain/cache hit/miss, stream TTFT overhead, allocations, and concurrency. It should not attempt GPU/provider throughput ranking. Flux's own distributed-routing document correctly calls for measured load/soak evidence. [`DECENTRALIZED-FLUX.md:81-88`](../../docs/architecture/DECENTRALIZED-FLUX.md#L81-L88) + +### Gaps + +- No live-provider credentials were available, so current hosted API compatibility, model availability, regional routing, auth-header behavior, and provider-specific error shapes were not empirically validated. [`catalog/live/fetchers.go:20-36`](../../catalog/live/fetchers.go#L20-L36) +- No prior API baseline exists for `apidiff`, and the current release is too young to distinguish intentional pre-1.0 breakage from regressions. [module versions](https://pkg.go.dev/github.com/GrayCodeAI/flux?tab=versions) +- The test suite does not reveal production call volume, concurrency, memory profile, long-lived stream count, or secret-store latency. The 102-second race run and unit benchmarks are not operating evidence. [`Makefile:58-74`](../../Makefile#L58-L74) +- The TypeScript result is environment-inconclusive; a pinned Node/npm/TypeScript matrix is needed before drawing a source-quality conclusion. [`internal/sdk/typescript/package.json:1-8`](../../internal/sdk/typescript/package.json#L1-L8) +- Chaos, restart, persistence, and TLS behavior of distributed routing remain unverified beyond in-process mocks. [`router/controlplane/controlplane_test.go:131-196`](../../router/controlplane/controlplane_test.go#L131-L196) + +## 3. Smallest high-leverage public API and release strategy + +### Takeaway + +The smallest credible OSS contract is not the current 50-plus-package, seven-facet facade. It is a root-level Flux client with `New`, `Generate`, `Stream`, and `Close`, provider-neutral request/response/stream/error DTOs, injected secrets, catalog/state, OpenAI-compatible gateways, and a public transport resolver; `core.Provider` remains the extension port. Publish this as `v0.1.0` only after a clean-machine consumer test and engine-state fix, then automate compatibility, provenance, dependency updates, and conformance before `v1.0`. + +### Cited Findings + +- Flux already has the right generation primitives: `engine.New`, `Engine.Generate`, `Engine.Stream`, a pull-based `EventStreamer`, and typed errors. These are the natural core of a smaller public API. [`engine/engine.go:22-23,70-75,126-188`](../../engine/engine.go#L22-L188), [`llm/provider.go:26-43`](../../llm/provider.go#L26-L43), [`engine/errors.go:10-68`](../../engine/errors.go#L10-L68) +- Flux also already has the right extension port: the four-method concurrent-safe `core.Provider`. [`provider/core/core.go:21-33`](../../provider/core/core.go#L21-L33) +- What prevents a small API today is that the canonical host port bundles generation with model catalog, credential management, selection, gateway inspection, maintenance, and compaction; the stable package also re-exports setup/media/credential helpers beyond that port. [`llm/provider.go:16-24,137-152,201-212,244-281,317-328,388-402`](../../llm/provider.go#L16-L24), [`engine/types.go:94-164`](../../engine/types.go#L94-L164) +- The release system currently begins only after a human pushes a `v*` tag and then generates notes; it has no release PR, changelog/version update, API compatibility gate, or module-proxy smoke check. [`release.yml:1-31`](../../.github/workflows/release.yml#L1-L31) +- The repository's own audit plan identifies the highest-value missing work as conformance fixtures, README/GoDoc reconciliation, API diff, dependency/security/release operations, and a transport memoization fix, but that plan is not a completed release policy. [`audit-remediation.md:256-279,293-309`](../../docs/plans/audit-remediation.md#L256-L279) + +### Inferences + +#### 1. Recommended smallest stable API + +Move the primary public import to the root module for discoverability, while keeping `engine` as a temporary compatibility alias if needed. The stable v1 candidate should conceptually be: + +```go +package flux + +type Options struct { + Secrets SecretStore + StateDir string + Catalog CatalogSource + Gateways []OpenAICompatibleGateway + Transport TransportResolver +} + +type Client struct { /* implementation */ } + +func New(Options) (*Client, error) +func (*Client) Generate(context.Context, Request) (Response, error) +func (*Client) Stream(context.Context, Request) (*Stream, error) +func (*Client) Close() error +``` + +The stable DTO set should be only: + +- `Request`, `Response`, `Stream`, `Event`, `Route`, `Usage`, `Error`; +- capability/preference/limits needed to make routing decisions; +- `SecretStore` (`Get`, `Set`, `Delete`); +- `TransportResolver` as the advanced adapter seam; +- `OpenAICompatibleGateway` for the common custom endpoint. + +This is smaller than the current `llm.Provider`, while preserving the existing pull-based stream and typed error design. The current generation API and `core.Provider` already supply most of the required shapes. [`engine/types.go:7-92,142-164`](../../engine/types.go#L7-L92), [`llm/provider.go:26-43,45-135`](../../llm/provider.go#L26-L43) + +Keep setup/control APIs concrete and secondary, for example on a `Controller` returned by an optional `NewController` or in a separate `engine/setup` package. Do not make every host implement a mega-interface merely to configure credentials. [`llm/provider.go:137-152,201-212,244-328`](../../llm/provider.go#L137-L152) + +#### 2. Stability tiers + +| Tier | Packages | Policy | +|---|---|---| +| Stable | root `flux`, generation DTOs/errors/stream, `SecretStore`, `TransportResolver` | SemVer-protected; additions may be minor, removals/semantic changes major; checked with `apidiff`. | +| Advanced | `provider/core`, decorator interfaces, selected catalog/router types | Explicitly documented extension surface; slower-moving, tested against compatibility fixtures. | +| Experimental | current `catalog`, `router`, `provider/*`, conversation/storage/control-plane packages | No stability promise until a use case and external consumer justify promotion. | +| Internal / separate module | HTTP, gRPC, SDKs, distributed control plane | Either move to independently versioned modules or stop presenting them as core Flux deliverables. | + +This follows the repository's own observation that lower-level packages are public for staged migration but not the Rho product boundary. [`HOST-ENGINE-BOUNDARY.md:188-193`](../../docs/architecture/HOST-ENGINE-BOUNDARY.md#L188-L193) + +#### 3. Provider and extension redesign + +- Replace the current metadata/factory split with one provider plugin descriptor owned by a dependency-neutral registry: identity, credential metadata, protocol, factory, model lister, compatibility flags, and conformance tier. The catalog registry should consume this descriptor rather than requiring parallel edits in `setup` and `provider`. Current duplication is visible across `ProviderSpec`, live `Registry`, compatibility maps, and the two construction switches. [`catalog/registry/spec.go:31-74`](../../catalog/registry/spec.go#L31-L74), [`catalog/live/fetchers.go:62-100`](../../catalog/live/fetchers.go#L62-L100), [`setup/deployment.go:183-383`](../../setup/deployment.go#L183-L383), [`provider/provider_registry.go:74-180`](../../provider/provider_registry.go#L74-L180) +- Define a `Selector` interface for routing strategies rather than a string switch. Keep weighted/rendezvous/least-request built-ins; do not expose mutable global strategy registries. [`router/strategy.go:107-168`](../../router/strategy.go#L107-L168) +- Either promote and productionize cache/telemetry extension contracts or remove the corresponding README claims. The current internal Redis skeleton and audit sink should not be marketed as distributed/auditable infrastructure. [`internal/cache/backend.go:97-139`](../../internal/cache/backend.go#L97-L139), [`internal/observability/audit.go:38-100`](../../internal/observability/audit.go#L38-L100) +- Use the OpenTelemetry API and `log/slog` as the default extension standards. A custom telemetry DTO/sink adds surface without ecosystem value; a wrapped `core.Provider` remains available for advanced decorators. [`provider/observability/tracing.go:12-68`](../../provider/observability/tracing.go#L12-L68) +- Make model publication reproducible: an in-repo generator, source manifest, schema compatibility test, signed release artifact or digest, last-known-good cache, and tolerant decoding of additive fields. The current default-source schema mismatch demonstrates why this is adoption-critical. [`catalog/v1.go:44-50,444-460`](../../catalog/v1.go#L44-L50), [default catalog](https://langdag.com/model-catalog/v1/catalog.json) + +#### 4. Release strategy + +**Phase A — truth release (`v0.1.0`, 1–2 weeks)** + +1. Make a clean `HOME`/temporary state execute a complete mock-backed quickstart without `rho` commands. +2. Fix or replace the default catalog source; test its current schema and publish a signed/digested catalog artifact. +3. Memoize the engine transport and preserve router/breaker/cache/rate state; either fix or remove broken engine options. +4. Rewrite all examples through the stable API and add them as external-package compile tests. +5. Reconcile provider/env docs and either extract or clearly mark HTTP/gRPC/SDK features as internal/opt-in. +6. Remove false release/security claims and archive the enterprise RFC as non-roadmap design. +7. Add a clean external consumer job that creates a temporary module, `go get github.com/GrayCodeAI/flux@`, compiles `New/Generate/Stream`, and runs a mock provider. + +**Release acceptance gates:** + +- candidate tag equals `VERSION` and embedded `flux.Version`; +- full race suite, gRPC-tag build, `go vet`, lint, format, `govulncheck`, and module verification pass; +- `apidiff old-tag candidate` has no unapproved incompatible changes; +- README/examples are executable tests; +- OpenAPI is either public and route-validated or removed from the core product promise; +- default catalog parses in CI and has a signed/digested source; +- signed annotated tag is verified by the release job; +- SBOM and GitHub build provenance/attestation are attached to the release; +- module proxy resolution is verified before announcement. + +**Phase B — extension release (`v0.2.x`, 3–6 weeks)** + +- introduce root stable API and public `TransportResolver`/provider plugin registry; +- run shared conformance fixtures for the Tier-A providers; +- enable opt-in nightly live-provider tests with redacted cassettes and budget limits; +- add Dependabot/Renovate, vulnerability alerts, secret scanning/push protection, and a Go 1.26 patch/minor compatibility matrix; +- publish API diffs and package stability labels; +- establish a small local-server benchmark/soak baseline and publish machine-readable results. + +**Phase C — 1.0 gate (after at least 90 days or three real design partners)** + +- no unresolved compatibility exceptions in the supported API; +- documented support window and security policy owned inside Flux, not delegated to Rho; +- reproducible release from a clean clone; +- provider support matrix with conformance tier and last-verified date; +- migration guide from Eyrie and from the current pre-1.0 API; +- one or more downstream maintainers and a documented release/security succession path. + +The current `v0.0.1` module and recent rename make this the cheapest moment to narrow the API; after more consumers adopt the current facade, narrowing becomes materially harder. [`CHANGELOG.md:35-50`](../../CHANGELOG.md#L35-L50) + +#### 5. Prioritized practical roadmap + +| Priority | Change | Why it matters | Done when | +|---|---|---|---| +| P0 | Fresh-state quickstart + external consumer test | First five-minute adoption gate | A clean temporary environment runs a mock-backed engine example with no Rho CLI or pre-existing files. | +| P0 | Fix default catalog compatibility and trust | Prevents startup failure and untrusted metadata dependence | CI validates current artifact; schema additions are tolerated; artifact has publisher identity/digest. | +| P0 | Persist engine transport and stateful middleware | Makes retry/rate/cache/circuit claims real | Two-request tests prove cache hit, retained breaker state, and shared rate window through `Engine`. | +| P0 | Truthful docs/examples/API scope | Prevents users depending on inaccessible or nonfunctional surfaces | Every README claim maps to an external import/run/test; internal SDK/server claims are removed or extracted. | +| P0 | Automated release PR + compatibility gate | Turns one signed snapshot into a repeatable OSS process | A conventional merge can produce a version/changelog PR and candidate tag with no manual file edits. | +| P1 | Provider plugin/factory consolidation | Lowers contribution cost and drift | A new OpenAI-family provider requires one descriptor, one adapter/lister, and tests—not parallel switches. | +| P1 | Shared conformance suite | Builds provider trust | Tier-A adapters pass identical sanitized chat/stream/tool/usage/error/cancellation fixtures. | +| P1 | Real TLS/process/restart/soak test | Establishes distributed-routing confidence | A signed manifest is served over TLS, fetched by another process, survives peer outage, and restarts from a persisted last-good snapshot—or the feature remains experimental. | +| P1 | Performance baseline | Prevents “production-ready” by assertion | Reproducible local benchmark reports allocations, latency, concurrency, and stream behavior with committed result schema. | +| P2 | HTTP/gRPC/SDK extraction | Makes optional surfaces genuinely adoptable | Each is a public, tested, independently documented module with release ownership—or is deleted. | +| Defer | Enterprise org/RBAC/dashboard/HQL/A2A/SLA | High scope, low confidence return for a provider runtime | Keep as RFC only until external demand justifies it. [`FLUX-ENTERPRISE.md:442-465`](../../docs/design/FLUX-ENTERPRISE.md#L442-L465) | + +### Gaps + +- The recommended API is an engineering judgment, not a validated demand study. Before freezing it, interview a small set of intended Go adopters and test the `New/Generate/Stream` shape in two external applications. The current repository has only one human contributor and no public issue history. [repository](https://github.com/GrayCodeAI/flux) +- A `v0.1.0` date and effort estimate depend on maintainer capacity and whether the catalog publication infrastructure is available; neither can be established from the public repository. [`DECENTRALIZED-FLUX.md:64-88`](../../docs/architecture/DECENTRALIZED-FLUX.md#L64-L88) +- GitHub artifact attestations and SBOM tooling were not selected or tested in this audit; they should be chosen based on the repository's actual release mechanics rather than treated as already available. [`release.yml:1-31`](../../.github/workflows/release.yml#L1-L31) +- The exact stable/advanced package split needs an `apidiff` inventory and downstream-import census before deprecating any current package. The current tree already has ecosystem consumers beyond the documented Rho boundary. [`HOST-ENGINE-BOUNDARY.md:188-193`](../../docs/architecture/HOST-ENGINE-BOUNDARY.md#L188-L193) diff --git a/research_notes/Flux OSS landscape roadmap/provider_coverage.md b/research_notes/Flux OSS landscape roadmap/provider_coverage.md new file mode 100644 index 00000000..82ee6197 --- /dev/null +++ b/research_notes/Flux OSS landscape roadmap/provider_coverage.md @@ -0,0 +1,111 @@ +# Flux provider coverage audit + +## Provider surface and construction paths + +### Takeaway +Flux has **28 registry entries**, but it does not have one uniform execution path for those entries. The direct `provider.Client` path, the static adapter map, and the `setup`/deployment path overlap but do not always select the same protocol implementation; Vertex and Concentrate are the clearest divergences, while the native Gemini client is not selected for the `gemini` provider (it is selected only by the `gemini-vertex` setup deployment). + +### Cited Findings +- **Observed — canonical inventory.** The registry test asserts 28 providers and 28 live-fetcher keys. The IDs are `anthropic`, `openai`, `gemini`, `deepseek`, `grok`, `kimi`, `zai_coding`, `zai_payg`, `xiaomi_mimo_token_plan`, `xiaomi_mimo_payg`, `minimax_token_plan`, `minimax_payg`, `azure`, `bedrock`, `vertex`, `openrouter`, `concentrate`, `opengateway`, `stepfun`, `agnes`, `longcat`, `fireworks`, `canopywave`, `poolside`, `groq`, `clinepass`, `opencodego`, and `ollama` ([catalog/registry/providers.go:16-341](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L16-L341); [catalog/registry/provider_spec_test.go:10-45](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/provider_spec_test.go#L10-L45)). +- **Observed — four registry protocol families.** The same inventory declares Anthropic Messages for `anthropic` and `bedrock`; Gemini Generate Content for `gemini` and `vertex`; OpenAI Responses for `concentrate`; and OpenAI Chat Completions for the other 23 IDs ([catalog/registry/providers.go:16-341](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L16-L341)). The catalog's protocol matrix test separately classifies several providers as OpenAI-primary and treats dual-protocol providers as single-protocol from the catalog perspective ([catalog/registry/protocol_matrix_test.go:9-25](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/protocol_matrix_test.go#L9-L25); [catalog/registry/protocol_matrix_test.go:51-73](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/protocol_matrix_test.go#L51-L73)). +- **Observed — direct and setup construction are separate switches.** `provider.getOrCreateProvider` uses dedicated cases for Anthropic, Azure, Bedrock, Vertex, Z.AI, MiMo, OpenCode Go, Poolside, and LongCat, then falls through to `NewOpenAIClient` for all other compatible entries ([provider/provider_registry.go:74-179](https://github.com/GrayCodeAI/flux/blob/main/provider/provider_registry.go#L74-L179)). `setup.providerForDeployment` has its own explicit switch, including the native Gemini Vertex case, the OpenAI-compatible Gemini direct case, the Ollama wrapper, and the Concentrate Responses client ([setup/deployment.go:183-383](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L183-L383)). +- **Observed — registry metadata is not a constructor dispatch table.** The static adapter map derives a `ProviderType` from `TransportKind`; an empty `TransportKind` becomes `ProviderTypeOpenAICompatible`, and every non-local entry is marked as supporting streaming and tools ([provider/adapters/provider_registry.go:42-72](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/provider_registry.go#L42-L72)). The direct switch then ignores `AdapterID` and chooses constructors by provider name/type, so the registry's protocol metadata and actual client selection can diverge ([provider/provider_registry.go:107-175](https://github.com/GrayCodeAI/flux/blob/main/provider/provider_registry.go#L107-L175)). + +| Registry IDs | Registry declaration | Direct `provider.Client` selection | `setup`/deployment selection | Status | +|---|---|---|---|---| +| `anthropic` | Anthropic Messages / `anthropic` | `NewAnthropicClient` | `anthropic-direct` → `NewAnthropicClient` | **Observed: aligned** ([providers.go:20-30](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L20-L30); [provider_registry.go:107-111](https://github.com/GrayCodeAI/flux/blob/main/provider/provider_registry.go#L107-L111); [deployment.go:193-199](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L193-L199)) | +| `openai` | OpenAI Chat / `openai` | Generic `NewOpenAIClient` | `openai-direct` → `NewOpenAIClient` | **Observed: aligned** ([providers.go:32-40](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L32-L40); [deployment.go:238-243](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L238-L243)) | +| `gemini` | Gemini Generate Content / `gemini` | Generic OpenAI-compatible client because `TransportKind` is empty | `gemini-direct` → `NewGeminiOpenAIClient` | **Partial: native `NewGeminiClient` is not selected by these paths** ([providers.go:42-50](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L42-L50); [provider_registry.go:46-70](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/provider_registry.go#L46-L70); [deployment.go:258-263](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L258-L263)) | +| `azure` | OpenAI Chat / `openai-azure` | `NewAzureClient` | `openai-azure` → `NewAzureClient` | **Observed: aligned** ([providers.go:166-174](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L166-L174); [provider_registry.go:111-117](https://github.com/GrayCodeAI/flux/blob/main/provider/provider_registry.go#L111-L117)) | +| `bedrock` | Anthropic Messages / `anthropic-bedrock` | `NewBedrockClient` | `anthropic-bedrock` → `NewBedrockClient` | **Observed: aligned** ([providers.go:176-184](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L176-L184); [provider_registry.go:118-128](https://github.com/GrayCodeAI/flux/blob/main/provider/provider_registry.go#L118-L128)) | +| `vertex` | Gemini Generate Content / `gemini-vertex` | `NewVertexClient`, whose `Name()` is `anthropic-vertex` and whose URL is Anthropic Vertex | `gemini-vertex` → native `NewGeminiClient` | **Observed: divergent implementations** ([providers.go:187-195](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L187-L195); [provider_registry.go:129-138](https://github.com/GrayCodeAI/flux/blob/main/provider/provider_registry.go#L129-L138); [adapters/vertex.go:42-46](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/vertex.go#L42-L46); [deployment.go:264-271](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L264-L271)) | +| `openrouter` | OpenAI Chat / `openrouter` | Generic OpenAI client with compat flags | Thin `NewOpenRouterClient` wrapper | **Observed: same wire family** ([providers.go:200-207](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L200-L207); [adapters/openrouter.go:10-39](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openrouter.go#L10-L39)) | +| `deepseek`, `grok`, `kimi`, `minimax_token_plan`, `minimax_payg`, `opengateway`, `stepfun`, `agnes`, `fireworks`, `canopywave`, `groq`, `clinepass` | OpenAI Chat, usually with a named thin wrapper or `openai` adapter ID | Generic `NewOpenAIClient` in the direct switch, except the named compat flags are attached in the static map | Named thin wrappers for most, generic OpenAI for some | **Observed: mostly one wire family, but not one constructor identity** ([providers.go:52-161](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L52-L161); [providers.go:220-310](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L220-L310); [provider_registry.go:139-175](https://github.com/GrayCodeAI/flux/blob/main/provider/provider_registry.go#L139-L175); [deployment.go:272-357](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L272-L357)) | +| `zai_coding`, `zai_payg` | OpenAI Chat / Z.AI compat IDs | `NewZAIClient` with OpenAI primary and Anthropic fallback | `newZAIDeploymentClient` with the same dual-protocol shape | **Observed: aligned dual-protocol path** ([providers.go:84-111](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L84-L111); [adapters/zai.go:15-78](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/zai.go#L15-L78); [deployment.go:321-324](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L321-L324)) | +| `xiaomi_mimo_token_plan`, `xiaomi_mimo_payg` | OpenAI Chat / `xiaomi_mimo` | `NewMiMoClient` | `newMiMoDeploymentClient` → `NewMiMoClient` | **Partial: current MiMo client is OpenAI-only despite setup documentation describing a dual surface** ([providers.go:114-139](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L114-L139); [adapters/mimo.go:16-46](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/mimo.go#L16-L46); [deployment.go:358-401](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L358-L401)) | +| `poolside` | OpenAI Chat / `poolside` | `NewPoolsideClient` | `NewPoolsideClient` | **Observed: specialized non-streaming recovery path, not a native stream adapter** ([providers.go:284-290](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L284-L290); [adapters/poolside.go:10-53](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/poolside.go#L10-L53)) | +| `opencodego` | OpenAI Chat / `opencodego` | `NewOpenCodeGoClient` with per-model protocol routing | `NewOpenCodeGoClient` | **Observed: aligned dual-protocol router** ([providers.go:312-320](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L312-L320); [adapters/opencodego.go:12-50](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/opencodego.go#L12-L50)) | +| `longcat` | OpenAI Chat / `openai` in the registry | `NewLongCatClient` with OpenAI primary and Anthropic fallback | `NewLongCatClient` | **Observed: registry says one primary protocol, runtime preserves both** ([providers.go:253-261](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L253-L261); [adapters/longcat.go:15-75](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/longcat.go#L15-L75); [deployment.go:346-351](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L346-L351)) | +| `concentrate` | OpenAI Responses / `concentrate-responses` | Generic `NewOpenAIClient` in the direct switch | `NewConcentrateResponsesClient` | **Observed: protocol-divergent path** ([providers.go:210-217](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L210-L217); [provider_registry.go:139-175](https://github.com/GrayCodeAI/flux/blob/main/provider/provider_registry.go#L139-L175); [deployment.go:374-380](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L374-L380)) | +| `ollama` | OpenAI Chat / `openai` | Generic `NewOpenAIClient` | `NewOllamaClient`, which is still an OpenAI-compatible wrapper | **Partial: chat adapter has no native Ollama tags/show/keep-alive path** ([providers.go:323-340](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L323-L340); [adapters/ollama.go:10-39](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/ollama.go#L10-L39); [deployment.go:325-327](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L325-L327)) | + +- **Observed — aliases are split across layers.** Registry lookup accepts `google`→`gemini` and `xai`→`grok`, while the legacy catalog canonicalizer maps `gemini`→`google` and `grok`→`xai`; runtime normalization adds hyphen/underscore variants and legacy Z.AI/MiMo names ([catalog/registry/registry.go:39-52](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/registry.go#L39-L52); [catalog/registry/derive.go:20-33](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/derive.go#L20-L33); [catalog/v1_defaults.go:87-108](https://github.com/GrayCodeAI/flux/blob/main/catalog/v1_defaults.go#L87-L108); [runtime/selection.go:391-427](https://github.com/GrayCodeAI/flux/blob/main/runtime/selection.go#L391-L427)). +- **Partial — client configuration is not fully applied.** `FluxConfig` declares `Model` and `MaxRetries`, but `provider.Client` copies only provider, API key, and base URL; the model is instead resolved from the catalog only when the per-call options omit it, and `MaxRetries` is not applied ([llm/types.go:20-27](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go#L20-L27); [provider/client.go:43-67](https://github.com/GrayCodeAI/flux/blob/main/provider/client.go#L43-L67); [provider/chat.go:11-42](https://github.com/GrayCodeAI/flux/blob/main/provider/chat.go#L11-L42)). Adapter option setters do store `defaultModel`, `defaultMaxTokens`, and `defaultTemperature`, but the request builders use `opts` directly and do not read those stored fields ([provider/adapters/adapter_config.go:38-45](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/adapter_config.go#L38-L45); [provider/adapters/openai.go:299-327](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai.go#L299-L327); [provider/adapters/anthropic.go:487-535](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic.go#L487-L535)). +- **Observed — documentation counts are inconsistent.** The README correctly says 28 gateways ([README.md:200-233](https://github.com/GrayCodeAI/flux/blob/main/README.md#L200-L233)), while the runtime package comment says 16 registered providers ([runtime/runtime.go:10-18](https://github.com/GrayCodeAI/flux/blob/main/runtime/runtime.go#L10-L18)), the credential guide says 15 and still names `z-ai`, and the repository instructions say 75+ providers ([docs/guides/CREDENTIAL-SETUP-FLOW.md:7-27](https://github.com/GrayCodeAI/flux/blob/main/docs/guides/CREDENTIAL-SETUP-FLOW.md#L7-L27); [AGENTS.md:11-11](https://github.com/GrayCodeAI/flux/blob/main/AGENTS.md#L11)). + +### Inferences +- The highest-value architectural fix is a single constructor resolver driven by `ProviderSpec`/deployment metadata, with direct-client and deployment-client parity tests. Today a provider can be “registered” yet route through a different protocol depending on the entry point ([catalog/registry/providers.go:31-341](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/providers.go#L31-L341); [provider/provider_registry.go:107-175](https://github.com/GrayCodeAI/flux/blob/main/provider/provider_registry.go#L107-L175); [setup/deployment.go:183-383](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L183-L383)). +- Provider aliases should be normalized once at the public boundary and tested through registry, catalog, runtime, and adapter lookup. The current three canonical vocabularies make a seemingly harmless spelling change capable of selecting a different transport ([catalog/registry/registry.go:39-52](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/registry.go#L39-L52); [catalog/v1_defaults.go:87-108](https://github.com/GrayCodeAI/flux/blob/main/catalog/v1_defaults.go#L87-L108); [runtime/selection.go:391-427](https://github.com/GrayCodeAI/flux/blob/main/runtime/selection.go#L391-L427)). +- The current provider count is a registry count, not proof of 28 independently tested protocol implementations. Thin wrappers and generic fallbacks explain why the package can have broad nominal coverage while direct and deployment paths still disagree ([provider/adapters/provider_registry.go:46-72](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/provider_registry.go#L46-L72); [provider/provider_registry.go:139-175](https://github.com/GrayCodeAI/flux/blob/main/provider/provider_registry.go#L139-L175)). + +### Gaps +- The audit did not call live provider APIs, so registry endpoint strings, credential requirements, model availability, and provider-specific wire behavior remain source/test evidence rather than current production evidence ([catalog/live/types.go:5-43](https://github.com/GrayCodeAI/flux/blob/main/catalog/live/types.go#L5-L43); [catalog/live_enrich.go:11-101](https://github.com/GrayCodeAI/flux/blob/main/catalog/live_enrich.go#L11-L101)). +- The source does not provide a generated table proving that every `AdapterID` in the registry resolves to the same constructor in every path. That relationship needs an explicit conformance test rather than an inference from names ([catalog/registry/spec.go:31-57](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/spec.go#L31-L57); [provider/adapters/provider_registry.go:46-72](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/provider_registry.go#L46-L72)). +- The precise reason the direct path does not select native Gemini or Responses is not documented; the switch behavior is observable, but the intended public entry-point contract is ambiguous ([provider/provider_registry.go:107-175](https://github.com/GrayCodeAI/flux/blob/main/provider/provider_registry.go#L107-L175); [setup/deployment.go:258-271](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L258-L271)). + +## Protocol and capability coverage + +### Takeaway +The core chat/stream contract is implemented broadly, with native Anthropic, native Gemini, OpenAI-compatible, Anthropic-on-cloud, and Concentrate Responses paths. Coverage is nevertheless **partial** for lossless provider replay, system/developer semantics, structured-output shape, model-specific reasoning, output media, and feature discovery; embeddings, Anthropic batch, moderation, and media operations are separate capability surfaces rather than parts of the host `Provider` contract. + +### Cited Findings +- **Observed — shared request/response primitives.** `core.Provider` exposes only `Chat`, `StreamChat`, `Ping`, and `Name`; the canonical DTOs include content parts, tools, usage, finish/stop reason, warnings, and opaque provider-block fields ([provider/core/core.go:21-33](https://github.com/GrayCodeAI/flux/blob/main/provider/core/core.go#L21-L33); [llm/types.go:29-75](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go#L29-L75); [llm/types.go:248-287](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go#L248-L287)). The shared SSE parser supports `data:`/`event:` lines, multi-line data, context cancellation, and scanner-error events, with a 2 MiB maximum line/event buffer ([provider/core/stream.go:16-89](https://github.com/GrayCodeAI/flux/blob/main/provider/core/stream.go#L16-L89)). +- **Observed — OpenAI-compatible request surface.** The shared OpenAI builder supports text, image URLs, input audio, tool definitions/results, tool choice, usage-in-stream options, top-p/stop/service-tier/user, reasoning formats, penalties, log probabilities, seed, store, metadata, modalities, audio config, prediction, and web-search options ([provider/adapters/openai.go:191-452](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai.go#L191-L452)). It posts to `/chat/completions`, sets bearer authentication and a 32 MiB request-size guard, and normalizes non-streaming and SSE responses ([provider/adapters/openai.go:459-578](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai.go#L459-L578)). +- **Partial — OpenAI system/developer semantics.** `ChatOptions.System` is part of the canonical contract ([llm/types.go:135-191](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go#L135-L191)), but the OpenAI message builder constructs roles/content without a branch for `opts.System`; the `SupportsDeveloperRole` compatibility flag is declared but not consumed by that builder ([provider/adapters/compat.go:3-35](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/compat.go#L3-L35); [provider/adapters/openai.go:197-299](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai.go#L197-L299)). A caller that supplies a system message may still work, but the engine's `SystemPrompt` mapping is not visibly preserved on this path ([engine/convert.go:15-21](https://github.com/GrayCodeAI/flux/blob/main/engine/convert.go#L15-L21)). +- **Partial — OpenAI structured output shape.** The builder emits `{"type":"json_schema","json_schema": }` ([provider/adapters/openai.go:334-342](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai.go#L334-L342)). OpenAI's current Chat Completions contract describes `json_schema` as a descriptor object containing `name`, `schema`, and `strict` ([OpenAI Chat Completions reference](https://platform.openai.com/docs/api-reference/chat/create#chat-create-response_format)). The local test checks only that `ResponseFormat` is non-nil, so this is a likely conformance gap rather than a live failure proven by the test suite ([provider/adapters/openai_test.go:268-277](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai_test.go#L268-L277)). +- **Observed — Anthropic Messages coverage.** Anthropic sends `X-Api-Key`, `Anthropic-Version`, system, tools, tool choice, thinking, metadata, service tier, image/audio parts, prompt-cache breakpoints, and a `/v1/messages` endpoint; its parser extracts text, thinking, tool calls, stop reason, request/organization IDs, and token usage ([provider/adapters/anthropic.go:60-87](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic.go#L60-L87); [provider/adapters/anthropic.go:257-300](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic.go#L257-L300); [provider/adapters/anthropic.go:469-559](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic.go#L469-L559)). The SSE processor handles text, thinking, tool-input accumulation, usage, stop reason, TTFT, and incomplete tool-call errors ([provider/core/stream.go:92-290](https://github.com/GrayCodeAI/flux/blob/main/provider/core/stream.go#L92-L290)). +- **Partial — Anthropic replay state.** The canonical contract promises signed thinking/redacted blocks and other provider-owned replay state ([llm/types.go:58-75](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go#L58-L75)), but the Anthropic parser discards `redacted_thinking` and only returns thinking text; no adapter assigns `ProviderBlocks` or emits `provider_block` events ([provider/adapters/anthropic.go:233-300](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic.go#L233-L300)). The tool contract declares `RawArguments` and `ProviderMetadata`, but the response/stream paths do not populate either field ([tools/tool.go:5-18](https://github.com/GrayCodeAI/flux/blob/main/tools/tool.go#L5-L18); [provider/adapters/anthropic.go:281-287](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic.go#L281-L287); [provider/core/stream.go:214-227](https://github.com/GrayCodeAI/flux/blob/main/provider/core/stream.go#L214-L227)). +- **Observed — native Gemini request/response coverage.** The native client uses `generateContent` and `streamGenerateContent?alt=sse`, API-key or Vertex bearer auth, request-size limits, image/audio inline data, function calls/results, tool choice modes, response MIME/schema, safety settings, penalties, log probabilities, usage, and normalized finish reasons ([provider/adapters/gemini.go:39-75](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L39-L75); [provider/adapters/gemini.go:194-278](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L194-L278); [provider/adapters/gemini.go:280-461](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L280-L461); [provider/adapters/gemini.go:491-554](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L491-L554)). +- **Partial — native Gemini option loss.** Gemini reads a system role from `messages`, but does not read `opts.System`; the request builder also declares `ThinkingConfig` without populating it from `ThinkingEnabled`, budget, or mode, and a `tool` choice other than `any`/`none` falls back to `AUTO` ([provider/adapters/gemini.go:280-374](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L280-L374); [provider/adapters/gemini.go:376-418](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L376-L418)). Its image URL path treats a non-base64 URL as PNG inline data rather than passing a provider-supported URL/reference ([provider/adapters/gemini.go:311-343](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L311-L343)). +- **Observed — Gemini streaming has two implementations.** The default path uses the shared SSE parser, while `FLUX_GEMINI_SHARED_PARSER=0` selects the legacy bespoke loop; the switch is an environment-controlled compatibility path, not a capability declaration ([provider/adapters/gemini.go:18-37](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L18-L37); [provider/adapters/gemini.go:131-169](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L131-L169)). The normal processor emits content/tool calls and terminal usage/finish events but has no thinking-event mapping in its chunk loop ([provider/adapters/gemini.go:556-650](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L556-L650)). +- **Observed — Concentrate Responses coverage.** The dedicated client posts `/responses`, maps function tools/results, `text.format` JSON schema, reasoning effort, metadata, tool choice, parallel calls, and emits output text/tool-call/usage/done events ([provider/adapters/concentrate_responses.go:19-117](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/concentrate_responses.go#L19-L117); [provider/adapters/concentrate_responses.go:164-246](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/concentrate_responses.go#L164-L246); [provider/adapters/concentrate_responses.go:273-343](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/concentrate_responses.go#L273-L343)). Its tests explicitly assert the Responses tool shape, tool-call IDs, streaming tool arguments, schema validation, and terminal events ([provider/adapters/concentrate_responses_test.go:46-221](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/concentrate_responses_test.go#L46-L221); [provider/adapters/concentrate_responses_test.go:223-312](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/concentrate_responses_test.go#L223-L312)). +- **Partial — reasoning is provider-specific and lossy at the canonical boundary.** OpenAI-compatible paths can send `reasoning_content` back, use provider-specific `thinking` shapes, and split inline `` text; Anthropic maps adaptive/enabled/disabled thinking; Concentrate maps effort. However, the engine drops `GenerationOptions.ThinkingEnabled` while mapping only the deprecated `GLMThinkingEnabled` alias ([engine/convert.go:22-30](https://github.com/GrayCodeAI/flux/blob/main/engine/convert.go#L22-L30); [llm/provider.go:104-135](https://github.com/GrayCodeAI/flux/blob/main/llm/provider.go#L104-L135); [provider/core/stream.go:519-542](https://github.com/GrayCodeAI/flux/blob/main/provider/core/stream.go#L519-L542)). The engine stream normalizer also has no provider-block or raw-tool-metadata mapping ([engine/stream.go:124-165](https://github.com/GrayCodeAI/flux/blob/main/engine/stream.go#L124-L165)). +- **Partial — multimodal input is broader than multimodal output.** OpenAI, Anthropic, and Gemini builders accept text/image/audio input parts, and the media package provides OpenAI-compatible image generation and audio transcription clients ([provider/adapters/openai.go:221-267](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai.go#L221-L267); [provider/adapters/anthropic.go:349-429](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic.go#L349-L429); [provider/media/media.go:18-25](https://github.com/GrayCodeAI/flux/blob/main/provider/media/media.go#L18-L25)). The canonical `FluxResponse` and stream event expose text, thinking, tools, usage, and finish data, but no image/audio/file/citation response parts ([llm/types.go:248-287](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go#L248-L287)). +- **Observed — usage and error normalization is substantial but not uniform.** `FluxError` preserves provider, operation, status, request ID, message, and cause, and retry defaults cover 429/500/502/503/529 with Retry-After/backoff handling ([provider/core/errors.go:11-68](https://github.com/GrayCodeAI/flux/blob/main/provider/core/errors.go#L11-L68); [provider/core/retry.go:15-76](https://github.com/GrayCodeAI/flux/blob/main/provider/core/retry.go#L15-L76)). Anthropic usage includes cache and thinking fields, Gemini includes thoughts/cached tokens, and Concentrate/OpenAI expose their native usage subsets ([provider/adapters/anthropic.go:289-299](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic.go#L289-L299); [provider/adapters/gemini.go:530-537](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L530-L537); [provider/adapters/concentrate_responses.go:457-470](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/concentrate_responses.go#L457-L470)). However, the contract says Anthropic-style prompt totals should include separately reported cache tokens, while the parser assigns `InputTokens` directly ([llm/types.go:219-233](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go#L219-L233); [provider/adapters/anthropic.go:292-299](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic.go#L292-L299)). +- **Absent from the host provider interface — embeddings, batch, moderation, and media are separate.** `core.Embedder` is separate from `core.Provider`, implemented by `OpenAIClient` against `/embeddings`; Anthropic batch is a separate `BatchClient`; moderation and guardrails are wrappers; image/audio operations are separate clients ([provider/core/embedding.go:1-30](https://github.com/GrayCodeAI/flux/blob/main/provider/core/embedding.go#L1-L30); [provider/adapters/openai_embedding.go:15-113](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai_embedding.go#L15-L113); [provider/batch/batch.go:16-58](https://github.com/GrayCodeAI/flux/blob/main/provider/batch/batch.go#L16-L58); [provider/resilience/moderation.go:10-100](https://github.com/GrayCodeAI/flux/blob/main/provider/resilience/moderation.go#L10-L100); [provider/media/media.go:18-196](https://github.com/GrayCodeAI/flux/blob/main/provider/media/media.go#L18-L196)). +- **Partial — structured-output helper is not a provider capability negotiation layer.** `media.ChatWithStructuredOutput` validates a simplified subset of JSON Schema and retries by adding prompt feedback, while `media.WithStructuredOutput` currently returns an empty option ([provider/media/structured.go:26-79](https://github.com/GrayCodeAI/flux/blob/main/provider/media/structured.go#L26-L79); [provider/media/structured.go:148-203](https://github.com/GrayCodeAI/flux/blob/main/provider/media/structured.go#L148-L203); [provider/media/structured.go:195-203](https://github.com/GrayCodeAI/flux/blob/main/provider/media/structured.go#L195-L203)). This is separate from native Anthropic `output_config`, Gemini `responseSchema`, and Concentrate `text.format` behavior ([provider/adapters/anthropic.go:123-131](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic.go#L123-L131); [provider/adapters/gemini.go:385-391](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L385-L391); [provider/adapters/concentrate_responses.go:322-336](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/concentrate_responses.go#L322-L336)). +- **Partial — model-level capability discovery is stronger than provider-level booleans.** Live catalog entries preserve context, pricing, raw JSON, thinking, effort, structured output, code execution, citations, PDF, and image capability evidence, and merge logic prefers live capability values ([catalog/live/types.go:5-43](https://github.com/GrayCodeAI/flux/blob/main/catalog/live/types.go#L5-L43); [catalog/v1.go:143-158](https://github.com/GrayCodeAI/flux/blob/main/catalog/v1.go#L143-L158); [catalog/discover/merge.go:156-202](https://github.com/GrayCodeAI/flux/blob/main/catalog/discover/merge.go#L156-L202)). In contrast, the static provider config marks every non-local provider as tool-capable and reasoning-capable without a model or endpoint qualifier ([provider/adapters/provider_registry.go:30-70](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/provider_registry.go#L30-L70)). +- **Partial — cache and cost behavior is approximate.** The response cache key includes model, system, temperature, and selected message/tool fields, but not all content parts or behavior-affecting options such as tool definitions, response format, reasoning, and top-p/top-k; the coalescer key is even narrower ([provider/cache/semantic_cache.go:296-339](https://github.com/GrayCodeAI/flux/blob/main/provider/cache/semantic_cache.go#L296-L339); [provider/resilience/coalesce.go:16-48](https://github.com/GrayCodeAI/flux/blob/main/provider/resilience/coalesce.go#L16-L48); [provider/chat.go:28-39](https://github.com/GrayCodeAI/flux/blob/main/provider/chat.go#L28-L39)). Cache analytics uses hardcoded model-name price heuristics rather than catalog rates ([provider/observability/cache_analytics.go:38-53](https://github.com/GrayCodeAI/flux/blob/main/provider/observability/cache_analytics.go#L38-L53); [provider/observability/cache_analytics.go:85-139](https://github.com/GrayCodeAI/flux/blob/main/provider/observability/cache_analytics.go#L85-L139)). + +### Inferences +- Flux should treat “provider support” as an operation × endpoint profile × model × runtime-version tuple. The catalog already has the right ingredients for provenance and capability states, but the runtime registry still exposes coarse booleans and the stable DTO cannot carry every native response part ([catalog/live/types.go:26-43](https://github.com/GrayCodeAI/flux/blob/main/catalog/live/types.go#L26-L43); [provider/adapters/provider_registry.go:30-70](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/provider_registry.go#L30-L70); [llm/types.go:248-287](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go#L248-L287)). +- The highest-risk parity gaps are not ordinary text streaming; they are the semantic edges: system/developer roles, JSON-schema descriptors, reasoning/replay state, tool-call provenance, cache-token totals, and output modalities. Each can silently produce a valid-looking response with incorrect provider state ([provider/adapters/openai.go:197-342](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai.go#L197-L342); [provider/adapters/anthropic.go:233-300](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic.go#L233-L300); [llm/types.go:58-75](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go#L58-L75)). +- Separate optional interfaces are the correct boundary for embeddings, batch, moderation, media generation/transcription, and future provider-native responses. Expanding the four-method chat contract would couple every provider to operations it does not implement and would make the host facade less stable ([provider/core/core.go:21-33](https://github.com/GrayCodeAI/flux/blob/main/provider/core/core.go#L21-L33); [provider/core/embedding.go:5-10](https://github.com/GrayCodeAI/flux/blob/main/provider/core/embedding.go#L5-L10); [provider/batch/batch.go:38-58](https://github.com/GrayCodeAI/flux/blob/main/provider/batch/batch.go#L38-L58)). +- Stream reliability should distinguish initial response, first visible token, inter-chunk idle, and total deadline. The current shared client has a 10-minute HTTP timeout and a shared SSE parser, but no provider-specific first-token/idle timeout contract in the reviewed transport path ([provider/core/transport.go:10-21](https://github.com/GrayCodeAI/flux/blob/main/provider/core/transport.go#L10-L21); [provider/core/stream.go:31-89](https://github.com/GrayCodeAI/flux/blob/main/provider/core/stream.go#L31-L89)). + +### Gaps +- No live API fixtures establish whether the OpenAI-compatible JSON-schema shape, Gemini URL handling, MiMo dual-protocol claims, or provider-specific reasoning fields are accepted by current production endpoints. These should be marked **uncertain** until a pinned conformance fixture or opt-in live probe passes ([provider/adapters/openai.go:334-342](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai.go#L334-L342); [provider/adapters/gemini.go:311-343](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L311-L343); [docs/guides/CREDENTIAL-SETUP-FLOW.md:29-57](https://github.com/GrayCodeAI/flux/blob/main/docs/guides/CREDENTIAL-SETUP-FLOW.md#L29-L57)). +- The reviewed DTOs do not define a complete representation for provider-native output audio, images, files, citations, refusal blocks, or arbitrary provider extensions, so those capabilities cannot be called end-to-end “supported” from the current contract ([llm/types.go:248-287](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go#L248-L287); [catalog/v1.go:143-158](https://github.com/GrayCodeAI/flux/blob/main/catalog/v1.go#L143-L158)). +- The absence of provider-block mappings is source-confirmed, but the intended policy for redacted thinking and signed replay state needs a product/security decision before implementation; simply forwarding all opaque provider data could violate the current redaction intent ([provider/adapters/anthropic.go:261-286](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic.go#L261-L286); [llm/types.go:58-75](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go#L58-L75)). + +## Test evidence and roadmap priorities + +### Takeaway +The repository is healthy at the package level: the local full test run and `go vet ./...` pass, and adapter/core/registry coverage is substantial. The evidence is uneven, however: core OpenAI/Anthropic/Gemini/Responses paths have meaningful protocol tests, many compatibility providers have only constructor/chat-path/ping smoke tests, and the most important contract-loss cases (engine option mapping, provider replay, cross-path construction parity, and full race/CI gates) are not yet demonstrated by the visible tests. + +### Cited Findings +- **Observed — local verification.** On 2026-09-24, `go test -cover ./...` passed and `go vet ./...` produced no diagnostics under Go 1.26.6. The repository’s documented test and vet commands match those checks ([Makefile:56-71](https://github.com/GrayCodeAI/flux/blob/main/Makefile#L56-L71); [Makefile:79-90](https://github.com/GrayCodeAI/flux/blob/main/Makefile#L79-L90)). This audit did not run every CI job. +- **Observed — package coverage.** The full run reported `provider` 58.2%, `provider/adapters` 81.0%, `provider/core` 59.6%, `provider/media` 26.2%, `provider/resilience` 28.3%, `catalog/registry` 85.6%, `engine` 72.1%, `setup` 64.7%, and `runtime` 56.9%. These are local command results rather than a checked-in coverage artifact; the CI workflow separately enforces a minimum aggregate 60% while running race tests ([.github/workflows/ci.yml:138-172](https://github.com/GrayCodeAI/flux/blob/main/.github/workflows/ci.yml#L138-L172)). +- **Observed — strong core protocol tests.** OpenAI tests cover request field mapping, tool calls/results, image parts, Kimi cache roles, thinking formats, streaming, retries, and errors ([provider/adapters/openai_test.go:184-277](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai_test.go#L184-L277); [provider/adapters/openai_test.go:358-520](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai_test.go#L358-L520)). Anthropic tests cover request/response parsing, thinking, redacted thinking, tools, content parts, and streams ([provider/adapters/anthropic_test.go:243-458](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic_test.go#L243-L458); [provider/adapters/anthropic_test.go:460-520](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic_test.go#L460-L520)). Gemini tests cover native request bodies, tools, structured output, Vertex auth, stream parsers, and legacy-parser behavior ([provider/adapters/gemini_test.go:397-493](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini_test.go#L397-L493); [provider/adapters/gemini_test.go:565-782](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini_test.go#L565-L782)). Concentrate tests cover the Responses contract, tool IDs, strict schema shape, streaming, retry, and structured errors ([provider/adapters/concentrate_responses_test.go:46-221](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/concentrate_responses_test.go#L46-L221); [provider/adapters/concentrate_responses_test.go:332-447](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/concentrate_responses_test.go#L332-L447)). +- **Observed — thin compatibility tests are mostly smoke tests.** Representative tests for OpenGateway, CanopyWave, Kimi, and MiniMax assert construction, a `/chat/completions` path, a basic response, and ping; they do not assert provider-specific headers, streaming dialects, tools, structured output, errors, cancellation, or usage semantics ([provider/adapters/opengateway_test.go:16-65](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/opengateway_test.go#L16-L65); [provider/adapters/canopywave_test.go:16-65](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/canopywave_test.go#L16-L65); [provider/adapters/kimi_test.go:16-65](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/kimi_test.go#L16-L65); [provider/adapters/minimax_test.go:16-65](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/minimax_test.go#L16-L65)). +- **Observed — registry tests validate metadata, not cross-path runtime parity.** The registry tests verify count, live-fetcher presence, selected protocol IDs, credential requirements, and a policy matrix, but do not compare the client selected by `provider.Client` with the client selected by `setup.ProviderForDeployment` ([catalog/registry/provider_spec_test.go:10-190](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/provider_spec_test.go#L10-L190); [catalog/registry/protocol_matrix_test.go:27-87](https://github.com/GrayCodeAI/flux/blob/main/catalog/registry/protocol_matrix_test.go#L27-L87)). +- **Partial — engine conversion tests cover many options but not the known loss edge.** The conversion test asserts basic fields and a long list of advanced options, yet it does not set or assert `ThinkingEnabled`; no visible engine test asserts `ProviderBlock`, `RawArguments`, or `ProviderMetadata` propagation ([engine/convert_test.go:24-167](https://github.com/GrayCodeAI/flux/blob/main/engine/convert_test.go#L24-L167); [engine/stream.go:124-165](https://github.com/GrayCodeAI/flux/blob/main/engine/stream.go#L124-L165)). The core stream test fallback for malformed tool JSON uses an `_raw` map entry rather than exercising the declared `RawArguments` field ([provider/core/stream.go:440-451](https://github.com/GrayCodeAI/flux/blob/main/provider/core/stream.go#L440-L451); [tools/tool.go:5-18](https://github.com/GrayCodeAI/flux/blob/main/tools/tool.go#L5-L18)). +- **Observed — CI is broader than the local audit.** CI runs formatting, module hygiene, vet, lint, race-plus-coverage tests, security scanning, dead-code analysis, duplication detection, secret scanning, Markdown linting, and fuzz jobs ([.github/workflows/ci.yml:34-71](https://github.com/GrayCodeAI/flux/blob/main/.github/workflows/ci.yml#L34-L71); [.github/workflows/ci.yml:100-135](https://github.com/GrayCodeAI/flux/blob/main/.github/workflows/ci.yml#L100-L135); [.github/workflows/ci.yml:197-260](https://github.com/GrayCodeAI/flux/blob/main/.github/workflows/ci.yml#L197-L260)). The local normal test run does not establish those additional gates. +- **Observed — resilience has broad implementation but low measured coverage.** The resilience package includes continuation, coalescing, rate limiting, health, guardrails, moderation, and policy/role logic, but the full run reports 28.3% coverage for the package ([provider/resilience/continuation.go:15-107](https://github.com/GrayCodeAI/flux/blob/main/provider/resilience/continuation.go#L15-L107); [provider/resilience/coalesce.go:73-173](https://github.com/GrayCodeAI/flux/blob/main/provider/resilience/coalesce.go#L73-L173); [provider/resilience/ratelimit.go:10-183](https://github.com/GrayCodeAI/flux/blob/main/provider/resilience/ratelimit.go#L10-L183)). This is a coverage signal, not proof that the untested paths are defective. + +### Inferences +- **P0 — unify construction and close silent protocol divergence.** Make registry/deployment metadata select one constructor path, then add a table-driven parity test for all 28 IDs. The first regression cases should be `gemini`, `vertex`, and `concentrate`, because source already demonstrates different client/protocol choices for the same provider identity ([provider/provider_registry.go:107-175](https://github.com/GrayCodeAI/flux/blob/main/provider/provider_registry.go#L107-L175); [setup/deployment.go:258-271](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L258-L271); [setup/deployment.go:374-380](https://github.com/GrayCodeAI/flux/blob/main/setup/deployment.go#L374-L380)). +- **P0 — make the stable contract lossless or explicitly lossy.** Map `GenerationOptions.ThinkingEnabled`; preserve provider blocks, raw tool arguments, provider metadata, warnings, and relevant request IDs; and emit `CallWarning` when a provider ignores or rewrites a setting. Add round-trip tests through `Generate`, `Stream`, and the next host turn ([engine/convert.go:15-68](https://github.com/GrayCodeAI/flux/blob/main/engine/convert.go#L15-L68); [engine/stream.go:124-165](https://github.com/GrayCodeAI/flux/blob/main/engine/stream.go#L124-L165); [llm/types.go:58-75](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go#L58-L75)). +- **P0 — add protocol conformance fixtures for the semantic edges.** At minimum, assert OpenAI system/developer messages and the full JSON-schema descriptor, Gemini system/thinking/tool-choice behavior, Anthropic cache-token totals and signed blocks, Concentrate Responses parity, and Ollama native-vs-compatible profile selection ([provider/adapters/openai.go:197-342](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai.go#L197-L342); [provider/adapters/gemini.go:280-418](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L280-L418); [provider/adapters/anthropic.go:469-559](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/anthropic.go#L469-L559)). +- **P1 — replace provider booleans with an operation/profile capability matrix.** Use explicit `supported`, `unsupported`, and `unknown` states tied to provider, deployment, endpoint dialect, model, and observation timestamp; retain raw live metadata and provenance. This prevents an OpenAI-compatible endpoint from being treated as full parity merely because it accepts a text request ([catalog/live/types.go:26-43](https://github.com/GrayCodeAI/flux/blob/main/catalog/live/types.go#L26-L43); [catalog/discover/merge.go:156-202](https://github.com/GrayCodeAI/flux/blob/main/catalog/discover/merge.go#L156-L202); [provider/adapters/provider_registry.go:62-70](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/provider_registry.go#L62-L70)). +- **P1 — turn thin wrapper tests into shared conformance profiles.** A reusable provider profile should assert auth headers, base URL/path, request-size policy, non-streaming shape, SSE framing/terminal event, usage, structured errors, retry status, cancellation, and Ping behavior. This scales better than duplicating three-test smoke suites for every gateway ([provider/adapters/opengateway_test.go:16-65](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/opengateway_test.go#L16-L65); [provider/core/stream.go:31-89](https://github.com/GrayCodeAI/flux/blob/main/provider/core/stream.go#L31-L89)). +- **P1 — define stream deadlines and terminal invariants.** Add separate initial-response, first-token, inter-chunk idle, and total-deadline policies; require every stream path to close its body, honor context cancellation, and emit exactly one terminal state. This is especially important for Bedrock’s custom EventStream loop and Concentrate’s custom Responses parser ([provider/adapters/bedrock.go:113-259](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/bedrock.go#L113-L259); [provider/adapters/concentrate_responses.go:207-246](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/concentrate_responses.go#L207-L246); [provider/core/stream.go:73-89](https://github.com/GrayCodeAI/flux/blob/main/provider/core/stream.go#L73-L89)). +- **P2 — harden cache identity and cost attribution.** Include all output-affecting options and multimodal parts in response-cache/coalescer keys, and calculate analytics from catalog/live pricing with an explicit unknown-price state instead of model-substring heuristics ([provider/cache/semantic_cache.go:296-339](https://github.com/GrayCodeAI/flux/blob/main/provider/cache/semantic_cache.go#L296-L339); [provider/resilience/coalesce.go:16-48](https://github.com/GrayCodeAI/flux/blob/main/provider/resilience/coalesce.go#L16-L48); [provider/observability/cache_analytics.go:85-139](https://github.com/GrayCodeAI/flux/blob/main/provider/observability/cache_analytics.go#L85-L139)). +- **P2 — generate provider-facing documentation from the registry.** Remove the 15/16/75+ count drift, publish canonical IDs and aliases, and document which operations are chat, embeddings, batch, media, or moderation ([README.md:200-233](https://github.com/GrayCodeAI/flux/blob/main/README.md#L200-L233); [runtime/runtime.go:10-18](https://github.com/GrayCodeAI/flux/blob/main/runtime/runtime.go#L10-L18); [docs/guides/CREDENTIAL-SETUP-FLOW.md:7-27](https://github.com/GrayCodeAI/flux/blob/main/docs/guides/CREDENTIAL-SETUP-FLOW.md#L7-L27)). +- **P2 — add opt-in live smoke tests outside normal CI.** A credential-gated, read-only discovery/Ping test can refresh endpoint and model evidence without making ordinary unit CI depend on vendor availability; failures should be reported as stale/uncertain capability evidence rather than silently changing the contract ([catalog/live_enrich.go:11-101](https://github.com/GrayCodeAI/flux/blob/main/catalog/live_enrich.go#L11-L101); [.github/workflows/ci.yml:138-172](https://github.com/GrayCodeAI/flux/blob/main/.github/workflows/ci.yml#L138-L172)). + +### Gaps +- This audit did not run `go test -race`, `golangci-lint`, `govulncheck`, `gosec`, `deadcode`, `jscpd`, Markdown lint, or the fuzz job; those remain separate verification gates ([.github/workflows/ci.yml:138-260](https://github.com/GrayCodeAI/flux/blob/main/.github/workflows/ci.yml#L138-L260)). +- No independent conformance or performance benchmark covers all 28 providers. The current tests are deterministic HTTP fixtures and unit tests, so claims about real provider reliability, model support, or end-to-end latency remain unverified ([provider/adapters/openai_test.go:50-133](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/openai_test.go#L50-L133); [provider/adapters/gemini_test.go:620-782](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini_test.go#L620-L782)). +- The low package coverage in `provider/resilience` and `provider/media` should be interpreted as a prioritization signal, not as a complete defect count; the audit did not inspect every branch in those packages ([provider/resilience/health.go:10-193](https://github.com/GrayCodeAI/flux/blob/main/provider/resilience/health.go#L10-L193); [provider/media/media.go:65-196](https://github.com/GrayCodeAI/flux/blob/main/provider/media/media.go#L65-L196)). +- The current evidence does not establish whether dead compatibility setters, legacy Gemini parsing, or deprecated continuation paths are intentionally retained. Their call sites and removal policy should be decided before deletion ([provider/adapters/adapter_config.go:38-45](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/adapter_config.go#L38-L45); [provider/adapters/gemini.go:18-37](https://github.com/GrayCodeAI/flux/blob/main/provider/adapters/gemini.go#L18-L37); [provider/resilience/continuation.go:89-106](https://github.com/GrayCodeAI/flux/blob/main/provider/resilience/continuation.go#L89-L106)). diff --git a/research_notes/Flux OSS landscape roadmap/reliability_security.md b/research_notes/Flux OSS landscape roadmap/reliability_security.md new file mode 100644 index 00000000..53fa27b7 --- /dev/null +++ b/research_notes/Flux OSS landscape roadmap/reliability_security.md @@ -0,0 +1,472 @@ +# Flux Reliability and Security Audit + +**Audit date:** 2026-09-24 +**Scope:** reliability, routing and resilience, streaming and continuation, credentials, storage, API boundaries, observability, catalog, batch, and adjacent security controls. +**Mode:** read-only source review; no production code was changed. + +## Executive Summary + +Flux builds cleanly and the normal, race, vet, formatting, lint, vulnerability, and tagged gRPC checks pass. The repository nevertheless has several production-significant gaps at the boundaries where requests become provider calls and streams: + +- The host engine reloads catalog/config and reconstructs the provider stack for each request, so opt-in circuit/rate/cache state does not survive and credential lookups can occur repeatedly. +- The conversation path drops structured tool calls, thinking/provider replay state, and Anthropic's separate usage events; a `max_tokens` response without usage can drive continuation indefinitely. +- Stream wrappers do not consistently preserve cancellation, and clean EOF can be presented as successful completion without a terminal `done` event. +- Budget checks and usage recording are not an atomic reservation/idempotency operation, and streaming wrappers can over- or under-record usage. +- The embedded HTTP server can be mounted without the bind-address safety check, and its default listener is plain HTTP. +- Routing can append automatic fallback even when the caller did not permit fallback, can silently strip tools, and does not reliably report the deployment that actually served a request. + +No single issue proves a universal critical vulnerability, but the combination of stream state loss, budget races, credential migration behavior, and weak admission/deadline defaults makes the current system unsuitable for an untrusted multi-tenant deployment without additional host controls. + +## Verification Performed + +All commands were run from `/Users/lakshmanpatel/Desktop/OSS2026/graycode-eco/flux` without changing source files. + +| Check | Result | +|---|---| +| `go test ./...` | PASS | +| `go test ./... -race -count=1 -shuffle=on -timeout=300s` | PASS | +| `go vet ./...` | PASS; no diagnostics | +| `govulncheck ./...` | PASS; `No vulnerabilities found.` | +| `gofumpt -d .` | PASS; no diff | +| `goimports -d .` | PASS; no diff | +| `golangci-lint run ./...` | PASS; `0 issues.` | +| `git diff --check` | PASS | +| `go test -tags grpc ./internal/grpc` | PASS | + +These checks establish build and static-analysis health; they do not cover the lifecycle, security, and concurrency gaps below. + +## Findings + +### F-01 — High — Per-request state reload resets runtime protections + +**Classification:** Confirmed when the engine's opt-in wrappers are enabled; the state reload itself occurs on the normal path. + +**Evidence** + +- Selection loads runtime state in `engine/engine.go:329-372`, then transport resolution loads it again in `engine/engine.go:303-326`. +- `loadRuntimeState` reads and compiles the catalog, reloads provider configuration, and scans credential accounts in `engine/state.go:86-129`. +- `DeploymentProviderFromState` constructs a new deployment router in `setup/deployment.go:86-110`. +- The adaptive limiter and response cache are constructed inside every `defaultTransport` call in `engine/engine.go:315-325`. + +**Impact** + +Circuit-breaker history, adaptive rate-limit windows, and response-cache entries are reset per request. Every request can also perform catalog/config I/O and multiple keychain lookups. Under load, this adds avoidable latency and can turn a healthy provider into an apparently unhealthy one because failure history never accumulates. Selection and transport can observe different state if files or credentials change between the two loads. + +**Recommendation** + +Build one immutable runtime snapshot containing the compiled catalog, provider configuration, resolved adapters, and middleware. Select and execute from the same snapshot. Recompile only after an explicit config/catalog mutation or refresh, with generation-based invalidation. Add a single request admission/runtime owner that persists limiter, breaker, and cache state. + +**Tests to add** + +Assert one state load per generation, stable limiter/breaker/cache instances across requests, and deterministic invalidation after config or credential rotation. + +### F-02 — High — Conversation streaming loses tools, usage, reasoning, and continuation state + +**Classification:** Confirmed. + +**Evidence** + +- `conversation.Engine.streamAndSave` only handles `content`, `done`, and `error` events in `conversation/engine.go:266-291`; `tool_call`, `thinking`, `usage`, `warning`, and provider-block events are ignored. +- It persists only text and token usage attached to the `done` event in `conversation/engine.go:315-335`. +- Anthropic streams prompt usage in `message_start` and output usage in `message_delta`, both as separate `usage` events, in `provider/core/stream.go:236-281`. +- Continuation is based only on `stopReason == "max_tokens"` and cumulative completion tokens in `conversation/engine.go:305-313`; there is no hard continuation-call limit. +- The API accepts tools and passes them into `PromptOpts` in `internal/api/server.go:163-187`, but the conversation path does not convert emitted tool calls into persisted tool-call nodes. + +**Impact** + +- Tool-using responses can be returned as empty or text-only responses, with the model request for execution silently lost. +- Anthropic streaming requests can persist zero token usage and incorrect cost data. +- A provider that emits `max_tokens` without usage leaves `cumulativeOut` at zero, so `cumulativeOut < groupBudget` remains true and the conversation goroutine can continue indefinitely. +- Thinking text and provider replay state are not retained in the conversation DAG. + +**Recommendation** + +Use a single event accumulator that understands the complete normalized event vocabulary. Persist tool-call/tool-result nodes, thinking metadata, provider blocks, and exactly one finalized usage record. Make continuation limits mandatory and count both continuations and total provider calls. Treat missing usage as a provider/accounting error when a terminal response claims usage-dependent behavior rather than as zero forever. + +**Tests to add** + +Cover Anthropic streaming usage, tool calls, warning events followed by `done`, a `max_tokens` response with no usage, hard continuation limits, and replaying a stored assistant turn. + +### F-03 — High — Stream terminal and cancellation contracts are inconsistent + +**Classification:** Confirmed. + +**Evidence** + +- `engine.Stream.forward` treats a closed source channel as a normal return and does not require a terminal event in `engine/stream.go:82-107`. +- `engine/continuation.go:62-70` synthesizes a `done` event when the upstream stream ended without one. +- OpenAI parsing calls `finish("")` on a clean channel close or `[DONE]` in `provider/core/stream.go:455-494`; a transport EOF is therefore indistinguishable from a provider-completed response. +- The SSE parser does not flush a final event when EOF arrives before a blank separator in `provider/core/stream.go:43-89`. +- Guardrails and tracing return stream literals without a cancel function in `provider/resilience/guardrails.go:83-118` and `provider/observability/tracing.go:91-135`. +- The wrapper contract in `llm/types.go:289-307` relies on `Close` invoking the cancellation function. + +**Impact** + +A truncated or cancelled provider stream can be reported as a successful empty response. A consumer that abandons a wrapped stream can leave the forwarding goroutine and upstream HTTP body alive because `Close` has no cancellation action. Provider-specific health warnings can also be lost or converted into fatal errors by different layers. + +**Recommendation** + +Define one stream terminal state machine: `done`, `error`, or `cancelled`; require an explicit terminal event or return a typed truncation error. Make every wrapper use a request-scoped child context, preserve the source cancel function, and close/drain the source when the consumer stops. Flush a final SSE event only when the protocol permits it; otherwise emit a typed error. + +**Tests to add** + +Use a blocking fake provider to assert that `Close` cancels the source, a truncated body never yields a clean `done`, a warning/error event reaches the expected normalized type, and EOF-only streams are classified consistently. + +### F-04 — High — Missing admission and overly broad default deadlines permit resource exhaustion + +**Classification:** Confirmed defaults; impact is highest when the HTTP API or another host exposes the engine to concurrent callers. + +**Evidence** + +- The shared transport has no `MaxConnsPerHost` or `MaxConns` in `provider/core/transport.go:28-45`. +- Provider clients use a ten-minute end-to-end timeout in `provider/core/transport.go:10-11,56-63`. +- `requestContext` applies a timeout only when the caller supplies one in `engine/engine.go:508-513`. +- The embedded server uses a ten-minute write timeout and has no request or stream admission limit in `internal/api/server.go:73-85`. +- `DeploymentRouter.StreamChat` starts a goroutine and returns a stream without a global queue or concurrency bound in `router/deployment_router.go:182-247`. +- The API config documents virtual-key budget enforcement as optional and unmetered fallback behavior in `internal/api/server.go:35-48`. + +**Impact** + +Slow or stalled upstreams can retain connections, goroutines, memory, and budget reservations for long periods. A caller can open many streams without a bounded number of in-flight requests, causing upstream throttling, local memory pressure, and noisy retry amplification. + +**Recommendation** + +Add a shared admission controller with per-provider, per-tenant, and global limits. Use separate bounded deadlines for connect, response headers, first event, inter-event idle time, and total request duration. Enforce queue size and cancellation while waiting. Make limits explicit in the host-facing engine rather than relying on the ten-minute provider default. + +### F-05 — High — Budget enforcement is not an atomic reservation or idempotent charge + +**Classification:** Confirmed. + +**Evidence** + +- `CheckBudget` only reads current totals in `storage/budgets.go:157-178`. +- `RecordUsage` performs the increment and ledger insert later in a separate transaction in `storage/budgets.go:181-209`. +- `BudgetProvider` checks before the provider call and records afterward in `provider/observability/budget_provider.go:73-91,94-142`. +- `UsageLimitProvider` performs a pre-check and records later in `provider/observability/usage_limit.go:53-68,70-105`. +- Streaming wrappers record every `usage` event and ignore `RecordUsage` errors in `provider/observability/budget_provider.go:116-123` and `provider/observability/usage_limit.go:83-94`. +- The SQLite ledger stores an empty model and has no request/idempotency key in `storage/budgets.go:99-108,202-205`. + +**Impact** + +Concurrent requests can all pass the same pre-check and exceed a limit. Retries, reconnects, or multiple/cumulative usage events can create duplicate ledger entries. Cancellation can cause a successful provider call to go uncharged because the record error is discarded. Cost attribution cannot reliably reconstruct which request or model incurred a charge. + +**Recommendation** + +Introduce a reservation operation that atomically reserves estimated spend, followed by an idempotent finalization keyed by a provider request ID or Flux operation ID. Define whether usage events are incremental or cumulative, finalize once per request, and make ledger/model attribution mandatory. Treat a failed charge as an operational error or a durable retryable outbox event, not an ignored value. + +### F-06 — High — Embedded API exposure has unsafe default and mounting behavior + +**Classification:** Confirmed API design; exploitability depends on how the host mounts or exposes it. + +**Evidence** + +- Auth bind validation is called only by `ListenAndServe` in `internal/api/server.go:73-85` and `internal/httputil/httputil.go:88-103`. +- `Server.ServeHTTP` directly serves the handler in `internal/api/server.go:69-71`; a host mounting that handler on a non-loopback listener bypasses the safety check. +- `ValidateAuthConfig` allows any non-loopback bind when an API key is configured, but `ListenAndServe` uses plain HTTP in `internal/api/server.go:77-85`. +- `NewServer` does not validate a provider/store/dependency combination and readiness only checks wrapper pointers in `internal/api/server.go:50-66,150-161`. +- Node, alias, analytics, and provider-health routes have no tenant/owner scope in `internal/api/server.go:101-121`. + +**Impact** + +A mounted handler with an empty API key can be exposed without the loopback safeguard. A non-loopback plain-HTTP deployment sends bearer credentials and prompts in cleartext unless an external TLS terminator is correctly configured. In a shared store, any authenticated caller can potentially inspect or mutate conversations not owned by that caller. A server with a non-nil conversation wrapper but unusable provider can report ready. + +**Recommendation** + +Make server construction validate required dependencies and return an error. Put authorization and tenant scope on every sensitive route, not only at the listener boundary. Require TLS termination or an explicit `InsecureLocalOnly` mode for plain HTTP; refuse a non-loopback insecure bind even when an API key is set. Make readiness probe provider/configuration readiness, not just pointer presence. + +### F-07 — High/Medium — Deployment fallback can violate caller policy and capability requirements + +**Classification:** Confirmed in the deployment router; whether a caller is affected depends on route configuration. + +**Evidence** + +- Explicit routing is always followed by automatic fallback in `router/deployment_router.go:320-350,683-707`; there is no `AllowFallback` field in `core.ChatOptions` or the deployment router call. +- If no route supports all requested server tools, `eligibleChoices` returns the non-tool-capable choices in `router/deployment_router.go:365-389`. +- `optsForOffering` then silently removes unsupported tools in `router/deployment_router.go:554-576`. +- `streamWithDeployment` buffers every non-output event before the first output without a bound in `router/deployment_router.go:441-503`. +- The router records every failure before checking whether the error is transient in `router/deployment_router.go:160-173,212-236`. + +**Impact** + +An exact-model request can fail over even when the caller intended no fallback. A request that requires a tool can be sent to a provider without that tool, producing a semantically different request. A non-transient auth or client error can open a deployment circuit. A provider that emits an unbounded number of pre-output metadata events can consume memory before the first visible token. + +**Recommendation** + +Propagate an explicit fallback policy and original requirements into the router. Reject a stage if required capabilities cannot be preserved; never silently drop tools. Bound metadata buffering and expose a typed pre-output overflow error. Classify cancellation and non-transient errors before updating circuit state. Treat budget/auth failures separately from endpoint availability. + +### F-08 — Medium — Route attribution and resilience telemetry describe the wrong attempt + +**Classification:** Confirmed. + +**Evidence** + +- The host route contract includes actual deployment and attempt fields in `llm/provider.go:235-246`. +- Engine attaches the initially selected route in `engine/convert.go:88-97` and `engine/engine.go:283-300`. +- Deployment routing selects and retries deployments in `router/deployment_router.go:134-179,182-247`, but does not populate the response route with the successful deployment or attempt count. +- `EventRouteChanged` is declared in `engine/types.go:65-80`, but no production emission was found. +- Usage wrappers construct new stream results without preserving the request ID in `provider/observability/types.go:21-26` and `provider/observability/budget_provider.go:116-133`. + +**Impact** + +Billing, telemetry, debugging, and cache attribution can point at the first route rather than the deployment that actually served the request. Retry/fallback behavior is difficult to audit, and provider request IDs disappear through budget/usage wrappers. + +**Recommendation** + +Return a route result from each deployment attempt and attach the successful deployment/attempt to both blocking and streaming events. Emit `route_changed` when a retry or fallback changes the serving deployment. Preserve request IDs in every wrapper, preferably through a context or immutable stream metadata object rather than manually reconstructing wrappers. + +### F-09 — Medium — Response cache correctness is not safe across providers, tenants, or response state + +**Classification:** Confirmed. + +**Evidence** + +- The documented default says `CacheConfig.Enabled` is true, but the zero value is false and `NewCachedProvider` never enables it; see `provider/cache/semantic_cache.go:15-37,73-92`. +- The key includes model, system, temperature, and basic message content/tool data only in `provider/cache/semantic_cache.go:296-340`. It omits provider/deployment, tenant, tool choice, response format, seed, multimodal parts, stop sequences, top-p, and other behavior-affecting options. +- `core.CopyResponse` copies usage and some tool argument maps but not route, warnings, provider blocks, or arbitrary nested typed values in `provider/core/copy.go:3-56`. +- Cached entries are returned directly through a shallow response copy in `provider/cache/semantic_cache.go:176-199`. + +**Impact** + +An engine fallback can serve a response cached by a different deployment under the same model key. A tenant or credential context can observe another context's response if the cache is shared. Mutating opaque tool arguments, route metadata, or provider state can race with another consumer. Enabling caching with a zero config silently does nothing, making the documented behavior misleading. + +**Recommendation** + +Define a canonical request identity containing provider/deployment, tenant, model, all generation-affecting options, tool declarations/choice, and normalized message parts. Make cache scope explicit and never share across tenants unless deliberately namespaced. Use a complete immutable response snapshot or structured clone for cache entries. Add collision and concurrent-mutation tests. + +### F-10 — Medium — Coalescing couples unrelated caller lifetimes and leaks state + +**Classification:** Confirmed. + +**Evidence** + +- The first caller's context becomes the shared request context in `provider/resilience/coalesce.go:129-151`. +- A creator-context cancellation cleanup goroutine returns without deleting the map entry in `provider/resilience/coalesce.go:153-167`. +- Waiters are incremented but never decremented in `provider/resilience/coalesce.go:113-126,176-189`. +- The same response pointer is returned to every waiter in `provider/resilience/coalesce.go:125-127,176-187`. +- The key includes only provider, model, messages, temperature, and max tokens in `provider/resilience/coalesce.go:16-48`; system/tools and other request-affecting options are not part of the key. + +**Impact** + +Cancelling the creator cancels work for otherwise live waiters. Completed entries can remain forever when the creator context is already canceled, and the waiter counter eventually rejects new work even when no callers are waiting. Shared mutable response/tool maps can be changed by one consumer. Requests with different system or tool settings can be incorrectly coalesced. + +**Recommendation** + +Use a request-owned context independent of any waiter, cancel it when the final waiter leaves, decrement waiter state on every exit, and remove completed entries deterministically. Include a complete canonical request identity and return deep copies to waiters. + +### F-11 — Medium — Provider-specific replay state is declared but not implemented + +**Classification:** Confirmed contract gap. + +**Evidence** + +- The host contract explicitly defines `ProviderBlocks` for Anthropic thinking signatures, redacted thinking, and OpenAI reasoning items in `llm/types.go:49-75,248-287`. +- A repository search found definitions but no production adapter assignment or replay consumer. +- Anthropic response parsing extracts thinking text but skips redacted blocks and does not preserve signatures/data in `provider/adapters/anthropic.go:233-300`. +- Anthropic stream parsing records block type only to route thinking deltas and does not emit provider-block events in `provider/core/stream.go:165-210`. +- Engine normalization has no provider-block branch in `engine/stream.go:124-165`. + +**Impact** + +A second turn cannot replay required signed reasoning state. The first turn may appear successful, while the follow-up is rejected or produces a different answer. Redacted thinking data is silently discarded, contrary to the stated replay contract. + +**Recommendation** + +Make provider blocks a required part of the adapter response/stream contract. Preserve the exact wire block and provider, emit a normalized provider-block event, persist it with the assistant node, and replay only matching protocol blocks. Treat unsupported replay as an explicit warning or error rather than silently dropping it. + +### F-12 — Medium — Credential migration can delete secrets after partial failure + +**Classification:** Confirmed. + +**Evidence** + +- `migrateEnvFileAt` reads all key/value pairs but migrates only recognized discovery keys in `credentials/migrate.go:50-81`. +- It removes the file when `len(secrets) > 0`, even if `migrated == 0` or some keychain writes failed, in `credentials/migrate.go:82-85`. +- Migration marker creation ignores filesystem errors in `credentials/migrate.go:22-24,31-47`. + +**Impact** + +An env file containing an unrecognized secret, a partially failing keychain write, or a key not present in the current discovery registry can be deleted with the secret still only in the file, or partially migrated into the keychain. Users can lose credentials during an upgrade. + +**Recommendation** + +Treat migration as a transaction. First enumerate and validate every supported key, write all destinations, verify writes, and only then remove or atomically rename the source file. Preserve unrecognized entries in a permission-restricted backup and return an actionable error on any failure. Treat marker creation as part of migration success. + +### F-13 — Medium — Credential isolation and endpoint policy are too implicit for multi-tenant use + +**Classification:** Confirmed code paths; highest impact when configuration is tenant-controlled. + +**Evidence** + +- `ServiceName` is mutable process-global state without synchronization in `credentials/store.go:11-25`; default store replacement is protected separately. +- `CombinedStore` keeps decrypted secrets in a process cache for two seconds in `credentials/combined.go:18-34,57-115`. +- Compatibility configuration writes provider secrets into process environment variables in `config/provider_env.go:315-334,651-670`. +- Custom gateway URL validation checks syntax but permits HTTP, arbitrary hosts, and private addresses in `engine/host_runtime.go:46-90`; the dynamic-provider helper has the same limitation in `provider/dynamic.go:54-58`. +- Probe HTTP follows redirects by default in `internal/probehttp/probehttp.go:49-79`. +- OIDC uses `http.DefaultClient` when no client is injected in `credentials/oidc.go:49-55,83-97`. + +**Impact** + +Process-global environment and service settings can leak across tenants or hosts. A custom endpoint controlled by a less-trusted actor can receive API keys or reach internal services. Redirects and DNS changes can bypass an initially safe URL check. Default OIDC requests have no client-level timeout when the caller supplies no deadline. + +**Recommendation** + +Use explicit per-engine credential namespaces and injected stores in host paths. Prohibit process-environment fallback in strict host mode. Validate URL schemes, resolved IPs, ports, redirects, and DNS rebinding at connection time; require explicit opt-in for private/local gateways. Use a bounded client and trusted endpoint allowlist for OIDC/probe requests. Treat cached plaintext secrets as a deliberate, configurable threat-model tradeoff. + +### F-14 — Medium — Storage permissions and secret-at-rest handling are incomplete + +**Classification:** Confirmed; impact depends on host filesystem and key-storage policy. + +**Evidence** + +- The budget database stores provider API keys in `virtual_key_secrets` and documents plaintext storage in `storage/budgets.go:44-68,83-87`. +- Permission errors for the budget database and WAL/SHM sidecars are ignored in `storage/budgets.go:58-68`. +- The conversation SQLite store protects the main database but does not protect WAL/SHM sidecars in `storage/sqlite.go:48-66`. +- Recorder cassettes persist request messages and response/tool data, and request hashes intentionally omit images, temperature, and other varying options in `provider/observability/recorder.go:99-140,279-340` and `provider/observability/cassette.go:26-41,92-121`. + +**Impact** + +A permissive umask or ignored chmod failure can expose database sidecars, provider keys, prompts, or tool arguments. Shared replay cassettes can retain sensitive user data and can match a different request because the hash omits important fields. + +**Recommendation** + +Prefer credential references or an encrypted/keychain-backed store over plaintext provider keys. Treat chmod failures as fatal, secure all SQLite sidecars, and test them under a permissive umask. Define cassette redaction, retention, encryption, and complete request identity before enabling recording outside tests. + +### F-15 — Medium — Health checks can report false healthy states and race during lifecycle transitions + +**Classification:** Confirmed. + +**Evidence** + +- Anthropic `Ping` returns nil for every status except 401 in `provider/adapters/anthropic.go:640-656`; Bedrock does the same for 401/403 in `provider/adapters/bedrock.go:262-281`. Other adapters follow similar patterns. +- `NewHealthChecker` replaces the entire supplied config with defaults when only `Interval` is zero in `internal/health/healthcheck.go:100-110`; partially specified configs otherwise retain zero timeout/thresholds. +- `Check` returns a result without persisting it in `internal/health/healthcheck.go:143-164`; only the background loop calls `updateStatus`. +- Failure counts are read and then incremented separately under different locks in `internal/health/healthcheck.go:262-329`. +- `Stop` clears `cancel` before waiting for `done` at `internal/health/healthcheck.go:181-218`; a concurrent `Start` can therefore begin a new loop while the old loop is still shutting down, allowing overlapping health checks. + +**Impact** + +A provider returning 500, 404, or rate-limit responses can be marked healthy. Immediate checks and dashboard results can disagree. Concurrent checks can lose failure transitions. A start/stop race can run overlapping health loops and duplicate pings or distort failure state. + +**Recommendation** + +Normalize `Ping` to treat every non-success status as an error, validate each config field independently, persist immediate checks, and update state atomically. Capture the per-run done channel locally, serialize Start/Stop transitions, and add jitter to periodic checks. + +### F-16 — Medium — OpenAI-compatible proxy is a lossy subset and reports failures as normal completion + +**Classification:** Confirmed compatibility limitation. + +**Evidence** + +- `openAIChatMessage.Content` is a string, so standard array/object multimodal content cannot be represented in `internal/api/openai_proxy.go:22-25`. +- `n`, `top_p`, stop, presence/frequency penalties, and other fields are explicitly accepted but ignored in `internal/api/openai_proxy.go:28-44`. +- `splitOpenAIMessages` flattens roles into transcript text in `internal/api/openai_proxy.go:300-331`. +- On a conversation error, the proxy emits an error data event, then still emits a finish chunk and `[DONE]` in `internal/api/openai_proxy.go:238-258`. +- Tool calls are accepted but the underlying conversation path does not surface them; see F-02. + +**Impact** + +Clients can receive successful-looking responses with missing options, malformed multimodal input, or no tool calls. A failed stream is easy for OpenAI-compatible clients to interpret as a completed generation. + +**Recommendation** + +Either explicitly document the supported subset and reject unsupported combinations, or add canonical conversion for multimodal content, tool-call events/results, accepted generation options, and a terminal error protocol that does not emit a normal finish after failure. Add contract tests against representative OpenAI requests. + +### F-17 — Medium/Low — Batch and live-catalog auxiliary paths have cancellation and resource-bound gaps + +**Classification:** Confirmed in opt-in/auxiliary paths. + +**Evidence** + +- Batch submit reuses one request after `http.Client.Do` consumes its body and sleeps with `time.Sleep` in `provider/batch/batch.go:97-121`. +- `WaitUntilDone` can use a Retry-After delay beyond the overall timeout and uses `time.After`; transient responses are not drained/closed before the next poll in `provider/batch/batch_async.go:87-132`. +- Poll result JSONL is bounded to 4 MiB per line but the whole result slice is retained in memory in `provider/batch/batch_async.go:147-183`. +- Live fetchers use a 30-second HTTP client, but `FetchFunc` has no context parameter and `FetchLiveProviderCatalog` loops sequentially in `catalog/live/fetchers.go:20-36,52-100` and `catalog/live_enrich.go:14-101`; a refresh can therefore accumulate many provider timeouts. +- Gemini model discovery places the API key in a query string in `catalog/live/fetchers_providers.go:464-478`. +- Remote catalog URLs are accepted from explicit input/environment and fetched without signature/content trust validation in `catalog/v1.go:669-720`. +- `MemoryBackend` has no maximum entry count or byte bound in `internal/cache/backend.go:43-87`; `CacheWarmer.Stop` can race before its loop captures the stop channel in `internal/cache/cache_warmer.go:91-145`. + +**Impact** + +Batch callers can exceed their context deadline, leak connections, or retry an unusable request. Catalog refresh can block for many provider timeouts and cannot be canceled through its public fetch API. API keys can appear in proxy logs or browser history. Shared memory caches can grow until process pressure, and warmer shutdown can leave a goroutine running. + +**Recommendation** + +Clone request bodies for every retry, use context-aware timers, cap every delay by the remaining deadline, close/drain error bodies, and bound result memory. Make live/catalog fetch functions context-aware and concurrent with a global deadline. Use headers rather than query secrets, authenticate/pin remote catalogs, bound cache bytes/entries, and make stop/start lifecycle ownership explicit. + +### F-18 — Low/Medium — Telemetry and audit data are not consistently lifecycle-safe or convention-aligned + +**Classification:** Confirmed instrumentation gaps; exporter absence depends on host integration. + +**Evidence** + +- Provider tracing uses custom keys such as `provider.name` and `usage.prompt_tokens` in `provider/observability/tracing.go:33-67,70-130`, while the repository separately defines `gen_ai.*` keys in `internal/observability/genai_semconv.go:21-58`; no production use of those canonical keys was found. +- `conversation.Prompt` ends its span immediately after returning the event channel in `conversation/engine.go:166-175`, before asynchronous work completes. +- Recorder request data is stored in `provider/observability/recorder.go:113-140,320-340`; the default redactor is nil. +- Callback hooks launch one goroutine per callback/event in `provider/observability/callbacks.go:117-177,194-219`. +- Audit hashes are unsalted SHA-256 values in `internal/observability/audit.go:93-100`; the same prompt can therefore be correlated across tenants/logs. +- No OTel SDK/exporter setup was found in the audited repository paths; API tracing relies on the global provider, so a host must install one. + +**Impact** + +Spans can end early, wrapper request IDs can be lost, callback traffic can amplify memory/CPU, and telemetry attributes may not aggregate with the shared GrayCodeAI conventions. Raw cassette data and unsalted hashes can create privacy or correlation risk. + +**Recommendation** + +Use the canonical `gen_ai.*` and `cost.usd` attributes, bind span completion to stream closure, preserve request IDs, bound asynchronous callback delivery, and make redaction mandatory for recorders. Use tenant-scoped keyed hashes or omit content hashes where correlation is unnecessary. Document exporter wiring as a host responsibility and test it with an in-memory exporter. + +## Confirmed Good Controls + +- Native JSON request decoding uses a body limit and rejects unknown fields in `internal/httputil/httputil.go:33-54`. +- Provider JSON response parsing generally caps error bodies, and live catalog responses are bounded. +- Control-plane manifests are copied before publication, revision conflicts are rejected, and remote peers require HTTPS/signature keys; peer redirects are disabled in `router/controlplane/snapshot.go:68-140,143-204` and `router/controlplane/peers.go:74-110`. +- Control-plane route validation caps retries at 32 in `router/controlplane/snapshot.go:159-203`. Direct `NewDeploymentRouter` callers do not receive equivalent validation, which is why that path remains a risk. +- `doWithMimoAuthRetry` and shared probe code use context-aware requests and bounded clients in several paths. +- Provider-specific API keys are not serialized into the normal sanitized provider state; setup migrates them into the injected store in `engine/state.go:23-83`. +- Race-enabled tests and static checks pass, but the test suite does not cover the lifecycle and adversarial cases listed below. + +## Prioritized Remediation Plan + +### P0 — Protect correctness and spend + +1. Add a single stream terminal/cancellation coordinator and require explicit terminal state. +2. Fix conversation accumulation for tool calls, provider blocks, thinking, and final usage; enforce a hard continuation limit. +3. Make credential migration transactional and never delete an unverified source file. +4. Replace budget check/record with an atomic reservation/finalization protocol and idempotency key. +5. Add admission/deadline limits before exposing the API to untrusted callers. +6. Make API construction fail closed for missing dependencies, auth, TLS mode, and tenant scope. + +### P1 — Make routing and middleware policy-correct + +1. Cache and version runtime snapshots; preserve limiter, breaker, and cache state. +2. Propagate fallback and capability requirements through deployment routing; preserve request identity and actual route attribution. +3. Canonicalize cache and coalescer keys and deep-copy returned responses. +4. Normalize provider replay state and health/Ping semantics. +5. Secure custom endpoints, credential namespaces, and all secret-at-rest paths. + +### P2 — Improve operational edges + +1. Complete telemetry lifecycle and shared OpenTelemetry conventions. +2. Make batch/catalog/cache warmer operations context-aware and bounded. +3. Add compatibility tests for the OpenAI proxy and clarify unsupported options. + +## Test Gaps + +The following tests are especially important before production use: + +- Credential migration with unrecognized keys, partial keychain failure, existing values, and unwritable source/marker files. +- Conversation streaming with Anthropic input/output usage events, tool calls, warning events, provider blocks, no-usage `max_tokens`, and hard continuation limits. +- Stream wrapper `Close` cancellation and goroutine/body cleanup for guardrails, tracing, cache, budget, usage-limit, and recorder wrappers. +- EOF, truncated SSE, scanner errors, missing terminal events, and provider cancellation normalized consistently across adapters. +- Concurrent budget reservations, duplicate finalization, request-ID idempotency, failed ledger writes, and multi-event/cumulative usage. +- Router fallback with `AllowFallback=false`, missing tool capability, non-transient failure, unbounded pre-output events, and actual route/attempt attribution. +- Cache collisions across provider/deployment/tenant, all generation-affecting options, and concurrent mutation of cached tool arguments. +- Coalescer creator cancellation, waiter decrement, TTL cleanup, response isolation, and complete request-key equality. +- API non-loopback plain HTTP, `ServeHTTP` mounting without listener validation, unknown virtual keys, dependency readiness, tenant isolation, and OpenAI multimodal/tool/error streams. +- URL/redirect/private-address policy for custom gateways, probes, remote catalogs, and OIDC endpoints. +- Provider `Ping` returning 404/429/5xx, health config validation, concurrent `Check`, and Start/Stop/Start transitions. +- Batch cancellation, body reuse, Retry-After beyond deadline, response-body closure, and result-size bounds. +- OTel span completion and canonical attributes with an in-memory exporter, bounded callbacks, and recorder redaction. + +## Final Assessment + +Flux has a strong provider abstraction and useful defensive controls, but its current reliability guarantees are mostly local to individual provider calls. Production readiness depends on making stream state, routing identity, budget accounting, credential migration, and admission control shared primitives rather than independent wrapper behavior. The highest-value next change is a small, versioned request-runtime/stream-terminal core used by engine, conversation, router, budget, and observability paths; it would address several findings without requiring a platform-wide rewrite. diff --git a/research_notes/Flux OSS landscape roadmap/sdk_frameworks.md b/research_notes/Flux OSS landscape roadmap/sdk_frameworks.md new file mode 100644 index 00000000..a79ddfeb --- /dev/null +++ b/research_notes/Flux OSS landscape roadmap/sdk_frameworks.md @@ -0,0 +1,644 @@ +# Flux OSS landscape roadmap: SDK and framework capability donors + +**Research cutoff and access date:** 2026-09-24. All web links in this document were checked on 2026-09-24 unless a page itself reports a different verification date. Repository counts and release labels are snapshots, not quality scores. + +**Evidence labels:** + +- **Shipped:** visible in current source, API reference, package metadata, or a stable documentation page. +- **Beta/experimental:** the project explicitly labels the feature beta, experimental, preview, or otherwise unstable. +- **Claim:** a project or vendor description that was not independently benchmarked in this research. +- **Roadmap:** future or planned behavior; it is not treated as an existing Flux capability. + +**Flux-specific observations** in this document come from the local checkout on 2026-09-24. Public repository links are included for navigability, but the local checkout is authoritative for the current branch. + +## Research scope and selection + +### Takeaway + +Flux should remain a provider-neutral generation engine, not become an agent framework, application server, or UI. The most valuable donors fall into three groups: (1) SDK/runtime contracts (LiteLLM, Vercel AI SDK, any-llm, Pydantic AI, OpenAI Agents SDK, LangChain, Google Genkit, and Google ADK); (2) host/orchestration boundaries (LangGraph, Microsoft Agent Framework and its AutoGen/Semantic Kernel lineage, Haystack, LlamaIndex, BAML, Mastra, Strands, DSPy, Agno, and CrewAI); and (3) product-boundary comparators (Dify and Open WebUI). The final top-20 below is a research shortlist, not a recommendation to embed all twenty projects. + +### Cited Findings + +- Flux's accepted boundary says that `engine`, `llm`, `graph`, and `tools` are host contracts, while credentials, catalog, routing, transports, resilience, usage, and provider telemetry remain engine concerns; the host owns UX, agent loops, tools, permissions, sessions, checkpoints, and product semantics. [Flux host–engine boundary](https://github.com/GrayCodeAI/flux/blob/main/docs/architecture/HOST-ENGINE-BOUNDARY.md) (accessed 2026-09-24). +- Flux's public facade is contract v2, exposes a pull-based cancellable stream, emits a route before provider events, treats unknown event types as additive, and explicitly leaves tool execution and the next model turn to the host. [Flux engine contract](https://github.com/GrayCodeAI/flux/blob/main/engine/engine.go) and [Flux stream implementation](https://github.com/GrayCodeAI/flux/blob/main/engine/stream.go) (accessed 2026-09-24). +- The project already has useful donor-like primitives: opaque provider blocks, raw tool arguments, normalized usage semantics, typed engine errors, model provenance fields, a pull stream, retry configuration, and a data-only graph vocabulary. [Flux DTOs](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go), [tool contracts](https://github.com/GrayCodeAI/flux/blob/main/tools/tool.go), [error types](https://github.com/GrayCodeAI/flux/blob/main/engine/errors.go), and [graph contracts](https://github.com/GrayCodeAI/flux/blob/main/graph/graph.go) (accessed 2026-09-24). +- GitHub repository pages and APIs were used to verify public URLs, language, license signals, and recent activity. A high star count was not used as the primary selection criterion; directness to Flux's boundary, distinct capability patterns, and current maintenance were more important. [GitHub API repository metadata example: Vercel AI](https://api.github.com/repos/vercel/ai) and [GitHub API repository metadata example: Flux](https://api.github.com/repos/GrayCodeAI/flux) (accessed 2026-09-24). + +### Final top-20 selection + +The order is by practical relevance to Flux, not by popularity. “Direct donor” means Flux can learn a contract or implementation pattern; “host donor” means the project is useful mainly to clarify what must stay outside Flux; “product comparator” means it is valuable for UI, deployment, governance, or product-boundary lessons rather than for a library dependency. + +| # | Project or family | Why it belongs in the final set | Flux use | Boundary signal | +|---:|---|---|---|---| +| 1 | [LiteLLM](https://github.com/BerriAI/litellm) | Best public reference for gateway routing, deployment fallback, spend/usage, cache behavior, and end-to-end GenAI OTel. | Router semantics, cost accounting, OTel privacy defaults, passthrough rules. | Proxy/admin/multi-tenant control plane stays host-side. | +| 2 | [Vercel AI SDK](https://github.com/vercel/ai) | Explicit language-model specification, provider registry, middleware hooks, typed tool loop, stream parts, and telemetry. | Model port, middleware, provider metadata, event lifecycle, privacy filtering. | Core is a donor; UI, RSC, harnesses, and agent loop are not. | +| 3 | [any-llm Python](https://github.com/mozilla-ai/any-llm) and [any-llm-go](https://github.com/mozilla-ai/any-llm-go) | Closest newly discovered Go-oriented provider abstraction, with optional capability interfaces and normalized errors. | Go package shape, optional provider interfaces, capability declarations, error sentinels. | Provider library only; gateway/platform is separate. | +| 4 | [Pydantic AI](https://github.com/pydantic/pydantic-ai) | Strongest typed message-part, tool-schema, validation, usage, test-model, and evaluation patterns. | Canonical parts, schema validation, offline conformance fixtures, OTel instrumentation. | Agent loop, harness, memory, durable runtime, and CLI stay host-side. | +| 5 | [OpenAI Agents SDK](https://github.com/openai/openai-agents-python) | Unusually clear stream-completion semantics, usage aggregation, raw usage preservation, tool lifecycle, and trace hierarchy. | Stream terminal rules, usage provenance, cancellation/interrupt semantics, model-adapter boundary. | Runner, handoffs, sessions, guardrails, and agent state are explicitly host concerns. | +| 6 | [LangChain](https://github.com/langchain-ai/langchain) | Broadest practical comparison for model profiles, provider-native content blocks, structured output, retries, and middleware. | Capability profiles, content-block normalization, model-call middleware, retry policy. | Agent memory, retrieval ecosystem, LangSmith, and agent loop are host concerns. | +| 7 | [LangGraph](https://github.com/langchain-ai/langgraph) | Best reference for durable stateful graph execution, checkpoint modes, interrupts, and typed stream projections. | Graph contract vocabulary and lifecycle vocabulary only. | No graph runtime, checkpointer, or scheduler in Flux. | +| 8 | [Google Genkit](https://github.com/genkit-ai/genkit) | Action/flow/plugin architecture is available in Go as well as TypeScript and Python, with typed schemas, streaming, interrupts, and developer tooling. | Go action/flow ergonomics, plugin lifecycle, typed stream outputs, trace boundaries. | Flows, agent runtime, Dev UI, and deployment are host concerns; Agents API is beta. | +| 9 | [Google ADK](https://github.com/google/adk-python) | Rich event semantics (`partial`, complete, interrupted, final response), callbacks/plugins, evaluation, artifacts, and multi-language support. | Event flags, callback/telemetry hooks, capability and artifact references, conformance cases. | Runner, sessions, memory, workflow agents, and web UI are host concerns. | +| 10 | [Microsoft Agent Framework](https://github.com/microsoft/agent-framework), with [AutoGen](https://github.com/microsoft/autogen) and [Semantic Kernel](https://github.com/microsoft/semantic-kernel) lineage | Current Microsoft successor combines AutoGen's simple agent patterns with Semantic Kernel's state, type safety, filters, telemetry, and graph workflows. | Layered model/function middleware, typed routing, workflow event vocabulary, OTel sensitivity controls. | Agent/session/workflow runtime remains host-side; AutoGen is now maintenance mode. | +| 11 | [Haystack](https://github.com/deepset-ai/haystack) | Mature component/pipeline contracts, validation, serialization, async cancellation, and pluggable tracing. | Component-like extension points, cancellation invariants, custom tracer interface, stream chunks. | Pipelines, RAG, document stores, loops, and agent components are host concerns. | +| 12 | [LlamaIndex](https://github.com/run-llama/llama_index) | Event-driven workflows with typed events, graph validation, streaming, state/resources, durable execution, and a large RAG ecosystem. | Event envelope, state/resource separation, workflow validation, data-aware capability boundaries. | Workflow execution, indexing, retrieval, and document processing are host concerns. | +| 13 | [BAML](https://github.com/BoundaryML/baml) | Schema-first structured output, generated contracts, streaming typed data, editor/test tooling, and multi-language clients. | Schema descriptor, validator/repair hook, golden structured-output fixtures. | DSL, prompt editor, optimizer, and evaluation application stay outside the engine. | +| 14 | [Mastra](https://github.com/mastra-ai/mastra) | Strong TypeScript reference for model routing, typed workflows, durable suspension, storage, observability, and evals. | Router, typed workflow boundary, lifecycle hooks, eval-friendly trace fields. | Core framework, Studio, memory, and server are host concerns; enterprise code is separately licensed. | +| 15 | [Strands Agents](https://github.com/strands-agents/sdk-python) | Minimal model-driven agent protocol with explicit `invoke_async`/`stream_async`, lifecycle hooks, sessions, OTel, MCP, and multi-agent patterns. | Minimal adapter protocol, hook points, stream event design, session-independent usage. | Agent loop, tool execution, sandboxes, and multi-agent runtime stay host-side. | +| 16 | [DSPy](https://github.com/stanfordnlp/dspy) | Best donor for evaluation metrics, program signatures, prompt/program optimization, and versioned optimization artifacts. | Trace/metric export and deterministic evaluation hooks, not the optimizer itself. | Prompt/program optimization, fine-tuning, and agent modules stay host/tooling-side. | +| 17 | [Agno](https://github.com/agno-agi/agno) | Useful separation between SDK, AgentOS runtime, and control plane; explicit storage, approval, RBAC, and OTel surfaces. | Storage/usage interfaces, approval event vocabulary, runtime-vs-library boundary. | AgentOS, control plane, memory, knowledge, UI, and scheduling stay host-side. | +| 18 | [CrewAI](https://github.com/crewAIInc/crewAI) | Clear separation between autonomous Crews and event-driven Flows with state, branching, routing, and config-as-code. | Host workflow event/state vocabulary and declarative configuration lessons. | Crews, tasks, roles, memory, and execution engine are host concerns. | +| 19 | [Dify](https://github.com/langgenius/dify) | Product reference for visual workflows, prompt versioning/testing, model management, observability, APIs, and deployment modes. | Prompt/catalog/eval UX requirements and provider configuration semantics. | UI, app builder, RAG, tenant gateway, and commercial control plane are explicitly outside Flux. | +| 20 | [Open WebUI](https://github.com/open-webui/open-webui) | Product reference for self-hosted model UX, plugins, RBAC, artifacts, analytics, and explicit tool-loop ownership. | Extension ownership rules, model connection semantics, usage/UI requirements, safety warnings. | UI, auth, RAG, tool execution, and user-facing application semantics are outside Flux. | + +### Activity, language, and license snapshot + +The following is a compact verification record. “Recent push” means the GitHub API or repository page showed a push in the last few days before the cutoff; a page can change after this snapshot. + +| Family | Language(s) | License/governance signal | Activity signal at cutoff | +|---|---|---|---| +| LiteLLM | Python plus Rust components | MIT outside `enterprise/`; enterprise directory separately licensed. [License](https://github.com/BerriAI/litellm/blob/main/LICENSE) (accessed 2026-09-24) | GitHub API showed a push on 2026-09-23. [API](https://api.github.com/repos/BerriAI/litellm) (accessed 2026-09-24) | +| Vercel AI SDK | TypeScript | Apache-2.0. [License](https://github.com/vercel/ai/blob/main/LICENSE) (accessed 2026-09-24) | API showed a push on 2026-09-23. [API](https://api.github.com/repos/vercel/ai) (accessed 2026-09-24) | +| any-llm | Python; new Go port | Apache-2.0 for both repositories. [Python license](https://github.com/mozilla-ai/any-llm/blob/main/LICENSE) and [Go license](https://github.com/mozilla-ai/any-llm-go/blob/main/LICENSE) (accessed 2026-09-24) | Python API push 2026-09-23; Go port was created in 2026 and had roughly 80 commits at the cutoff. [Python API](https://api.github.com/repos/mozilla-ai/any-llm), [Go repository](https://github.com/mozilla-ai/any-llm-go) (accessed 2026-09-24) | +| Pydantic AI | Python | MIT. [License](https://github.com/pydantic/pydantic-ai/blob/main/LICENSE) (accessed 2026-09-24) | API showed a push on 2026-09-23. [API](https://api.github.com/repos/pydantic/pydantic-ai) (accessed 2026-09-24) | +| OpenAI Agents SDK | Python and TypeScript | MIT. [Repository](https://github.com/openai/openai-agents-python) (accessed 2026-09-24) | Python API showed a push on 2026-09-23. [API](https://api.github.com/repos/openai/openai-agents-python) (accessed 2026-09-24) | +| LangChain / LangGraph | Python and JavaScript/TypeScript | MIT. [LangChain API](https://api.github.com/repos/langchain-ai/langchain), [LangGraph API](https://api.github.com/repos/langchain-ai/langgraph) (accessed 2026-09-24) | Both APIs showed pushes on 2026-09-23. | +| Google Genkit / ADK | TypeScript, Go, Python, Dart; ADK also Java/Kotlin | Apache-2.0 repositories. [Genkit API](https://api.github.com/repos/genkit-ai/genkit), [ADK API](https://api.github.com/repos/google/adk-python) (accessed 2026-09-24) | Both APIs showed pushes on 2026-09-23. | +| Microsoft lineage | C#, Python; Agent Framework also has .NET/Go workflow documentation | Agent Framework and Semantic Kernel repositories show MIT; AutoGen documentation/code has split CC-BY/MIT terms. [Agent Framework license](https://github.com/microsoft/agent-framework/blob/main/LICENSE), [AutoGen repository](https://github.com/microsoft/autogen), [Semantic Kernel API](https://api.github.com/repos/microsoft/semantic-kernel) (accessed 2026-09-24) | Semantic Kernel API push 2026-09-19; AutoGen README says maintenance mode and the API last push was 2026-04-15; Agent Framework is the recommended successor. | +| Haystack / LlamaIndex | Python | Apache-2.0 and MIT respectively. [Haystack API](https://api.github.com/repos/deepset-ai/haystack), [LlamaIndex API](https://api.github.com/repos/run-llama/llama_index) (accessed 2026-09-24) | Both APIs showed pushes on 2026-09-23. | +| BAML | Rust plus generated clients | Apache-2.0 repository. [API](https://api.github.com/repos/BoundaryML/baml) (accessed 2026-09-24) | API showed a push on 2026-09-23. | +| Mastra | TypeScript | Apache-2.0 core; `ee/` directories use the Mastra Enterprise License. [Repository licensing](https://github.com/mastra-ai/mastra/blob/main/README.md) (accessed 2026-09-24) | Repository page showed a highly active 2026 codebase; current package/repo pages showed recent releases. [Repository](https://github.com/mastra-ai/mastra) (accessed 2026-09-24) | +| Strands | Python and TypeScript | Apache-2.0. [Python repository](https://github.com/strands-agents/sdk-python), [harness licensing](https://github.com/strands-agents/harness-sdk/blob/main/LICENSE.APACHE) (accessed 2026-09-24) | Python repository and harness pages showed active 2026 development. | +| DSPy | Python | MIT. [License](https://github.com/stanfordnlp/dspy/blob/main/LICENSE) (accessed 2026-09-24) | Release page showed 3.3.1 on 2026-08-21 and a newer beta line. [Releases](https://github.com/stanfordnlp/dspy/releases) (accessed 2026-09-24) | +| Agno / CrewAI | Python | Apache-2.0 and MIT respectively. [Agno API](https://api.github.com/repos/agno-agi/agno), [CrewAI API](https://api.github.com/repos/crewAIInc/crewAI) (accessed 2026-09-24) | Both APIs showed pushes on 2026-09-23. | +| Dify / Open WebUI | TypeScript plus Python; Python backend/frontend stacks | Dify uses a modified Apache-2.0-based license with multi-tenant and branding conditions. Open WebUI uses a custom license with branding conditions; older contributions retain other terms. [Dify license](https://github.com/langgenius/dify/blob/main/LICENSE), [Open WebUI license](https://github.com/open-webui/open-webui/blob/main/LICENSE) (accessed 2026-09-24) | Both APIs showed pushes on 2026-09-23. | + +**Selection decisions for projects not in the final 20:** + +- **Portkey AI Gateway** is a useful secondary gateway comparison, with routing, retries, guardrails, logs, and a unified endpoint, but LiteLLM covers the same donor space with a more directly inspectable OSS/proxy boundary. [Portkey gateway](https://github.com/Portkey-AI/gateway) (accessed 2026-09-24). +- **Instructor** and **Atomic Agents** are excellent focused donors for schema validation and small composable agents, but BAML, Pydantic AI, LangChain, and Pydantic Graph cover those patterns with broader production context. [Instructor](https://github.com/567-labs/instructor), [Atomic Agents](https://github.com/Eigenwise/atomic-agents) (accessed 2026-09-24). +- **Letta** is a valuable memory/context-management and stateful-agent comparator, but its center of gravity is a persistent agent server and ADE rather than a provider-neutral engine. [Letta](https://github.com/letta-ai/letta) (accessed 2026-09-24). +- **Continue** is a host/product reference, but its repository is described as no longer actively maintained and read-only; it should not drive Flux architecture. [Continue repository](https://github.com/continuedev/continue) (accessed 2026-09-24). +- **MCP, AG-UI, OpenResponses, and workflow products** are protocols or host integration surfaces, not competing provider runtimes. They are useful compatibility targets, not top-20 framework donors. [MCP](https://modelcontextprotocol.io/), [Google ADK/AG-UI integration announcement](https://developers.googleblog.com/delight-users-by-combining-adk-agents-with-fancy-frontends-using-ag-ui) (accessed 2026-09-24). + +### Inferences + +- The strongest Flux roadmap is not “copy an agent framework”; it is to make the existing model-call contract more explicit, testable, provider-neutral, and observable. +- Go-specific donors are unusually important for Flux: any-llm-go, Genkit Go, Google ADK Go, and the Microsoft workflow documentation show that a Go engine can expose typed provider/flow primitives without adopting Python agent semantics. +- Dify and Open WebUI belong in the final set precisely because they make the non-goals visible: UI, auth, RAG, prompt management, multi-tenant control planes, and application deployment are separate products. +- AutoGen and Semantic Kernel should be studied as design lineages and migration precedents, but new capability work should follow Microsoft Agent Framework when comparing current Microsoft agent features. + +### Gaps + +- GitHub API rate limits prevented a uniform API snapshot for a few repositories; those rows use their official repository pages and package/release pages instead. Exact stars, forks, and issue counts should be refreshed before publication. +- Several projects market “production-ready,” “100+ providers,” “millions of downloads,” or performance numbers. Those are retained as project claims, not independently verified conclusions. +- The research did not clone and execute every project. Source/docs establish shipped interfaces and stated behavior, not correctness under all providers. +- Version labels can move quickly after the cutoff, especially for AI SDK 7, Pydantic AI Harness, Genkit Agents, Google ADK 2.0, and Microsoft Agent Framework workflows. + +## Project findings and capability patterns + +### 1. LiteLLM — gateway and runtime donor + +**Evidence status:** The Python SDK, gateway/proxy, router, caching, observability documentation, and OTel v2 documentation are shipped. OTel v2 is explicitly opt-in and off by default. [LiteLLM documentation](https://docs.litellm.ai/docs), [OTel v2](https://docs.litellm.ai/docs/observability/opentelemetry_v2) (accessed 2026-09-24). + +**Capability patterns:** + +- The documented request path is virtual-key/budget authorization, rate limiting, router load balancing/fallback/retry, provider translation, response, then asynchronous spend/logging updates. This is a useful decomposition for Flux's route, resilience, usage, and telemetry phases. [Life of a Request](https://docs.litellm.ai/docs/proxy/architecture) (accessed 2026-09-24). +- Retry and fallback have distinct scopes: retries stay within a model group/deployment set, while fallbacks move to another model group. Flux currently exposes route and attempt metadata; this distinction should remain explicit rather than collapsing both into “retry.” [LiteLLM routing architecture](https://docs.litellm.ai/docs/proxy/architecture) (accessed 2026-09-24). +- Image URL handling demonstrates an explicit compatibility policy: pass a URL through when supported, otherwise download and convert to base64 with a size cap and cache. Flux should copy the policy and observability, not silently fetch arbitrary URLs. [LiteLLM image URL handling](https://docs.litellm.ai/docs/proxy/architecture) (accessed 2026-09-24). +- OTel v2 uses canonical `gen_ai.*` attributes, content capture off by default, trace-level filtering, and separate span kinds for model, tool, datastore, guardrail, and MCP work. This is a strong privacy and interoperability baseline. [OTel v2 capture and span attributes](https://docs.litellm.ai/docs/observability/opentelemetry_v2) (accessed 2026-09-24). +- Caching documentation explicitly warns that provider-specific parameters can change output and therefore must be represented in cache keys; Anthropic-format and passthrough routes may bypass ordinary response caching. [LiteLLM caching](https://docs.litellm.ai/docs/proxy/caching) (accessed 2026-09-24). + +**Boundary:** LiteLLM's virtual keys, teams, budgets, admin UI, database, and multi-tenant gateway are platform concerns. Flux already owns provider credentials and routing but should not absorb an organization control plane merely because LiteLLM offers one. [LiteLLM architecture](https://docs.litellm.ai/docs/proxy/architecture) (accessed 2026-09-24). + +**Assessment:** Highest-value donor for route/attempt semantics, spend normalization, cache-key discipline, and privacy-safe OTel. Treat its OpenAI-shaped response as a compatibility surface, not as Flux's canonical internal contract. + +### 2. Vercel AI SDK — provider specification and middleware donor + +**Evidence status:** AI SDK 7 documentation and the `LanguageModelV4` TypeScript specification are shipped source. Agent, harness, workflow, and some newer surfaces are separate product areas; the Python SDK is labeled beta on the documentation site. [AI SDK Core](https://ai-sdk.dev/docs/ai-sdk-core), [language-model-v4.ts](https://github.com/vercel/ai/blob/main/packages/provider/src/language-model/v4/language-model-v4.ts) (accessed 2026-09-24). + +**Capability patterns:** + +- The provider layer is a small explicit model interface with a specification version, provider identity, generate/stream operations, and provider metadata. That is a better extension seam than exposing each vendor SDK throughout Flux. [LanguageModelV4 source](https://github.com/vercel/ai/blob/main/packages/provider/src/language-model/v4/language-model-v4.ts) (accessed 2026-09-24). +- `customProvider` and `createProviderRegistry` centralize aliases, model restrictions, default settings, fallback providers, files/skills interfaces, and model-kind-specific registries. This maps well to Flux's existing catalog/gateway/route split. [Provider and model management](https://ai-sdk.dev/docs/ai-sdk-core/provider-management) (accessed 2026-09-24). +- Middleware has three useful hooks: `transformParams`, `wrapGenerate`, and `wrapStream`. Middleware order is defined, and streaming middleware must preserve stream-part identity and block boundaries. Flux should expose a similarly narrow model-call middleware chain. [Language Model Middleware](https://ai-sdk.dev/docs/ai-sdk-core/middleware) (accessed 2026-09-24). +- Tool definitions separate schema, execution, strictness, examples, approval, and tool context. Multi-step calls expose a stopping condition and accumulated `responseMessages`; Flux should expose the same distinction between “model requested a tool” and “host executed it.” [Tool Calling](https://ai-sdk.dev/docs/ai-sdk-core/tools-and-tool-calling) (accessed 2026-09-24). +- Telemetry records operation, provider call, and tool spans, filters runtime/tool context before export, follows OTel GenAI conventions, and supports explicit content recording controls. The separation between full runtime context and telemetry-filtered context is directly relevant to Flux's secret-handling boundary. [Telemetry](https://ai-sdk.dev/docs/ai-sdk-core/telemetry) (accessed 2026-09-24). + +**Boundary:** AI SDK Core is a donor. AI SDK UI, React/Svelte/Vue hooks, RSC, Generative UI, terminal UI, harness adapters, and the built-in multi-step agent surface are host/product layers. Flux should not become a TypeScript UI toolkit or a harness runtime. + +**Assessment:** Best donor for a versioned provider port, composable middleware, provider metadata, and privacy-aware telemetry. Copy semantics, not UI or agent abstractions. + +### 3. any-llm — closest Go-oriented provider abstraction + +**Evidence status:** The Python library and the official Go port are shipped repositories. The Go port is new relative to Flux's mature codebase, so its provider matrix should be treated as a current snapshot rather than a guarantee of parity. [Python repository](https://github.com/mozilla-ai/any-llm), [Go repository](https://github.com/mozilla-ai/any-llm-go), [Go provider package](https://pkg.go.dev/github.com/mozilla-ai/any-llm-go/providers) (accessed 2026-09-24). + +**Capability patterns:** + +- Python re-exports OpenAI completion types and extends them where providers add reasoning content; streaming exposes completion chunks and reasoning deltas, while response format accepts schemas and tool calls remain visible. This is a pragmatic compatibility strategy, but it also demonstrates why Flux needs its own canonical parts and provider blocks. [Completion types](https://docs.mozilla.ai/api-reference/completion-1) and [AnyLLM interface](https://docs.mozilla.ai/api-reference/any-llm) (accessed 2026-09-24). +- The Python exception hierarchy is opt-in and preserves provider, status, code, parameter, and original exception information. Flux already has a structured provider error, so the donor is the stable category set and `errors.Is`/`errors.As` ergonomics rather than the exact class names. [Unified exceptions](https://docs.mozilla.ai/api-reference/exceptions) (accessed 2026-09-24). +- The Go port defines a small `Provider` interface with `Name`, `Completion`, and `CompletionStream`, then adds optional interfaces for embeddings, model listing, capabilities, and error conversion. This is an excellent pattern for avoiding a monolithic provider interface. [Go provider interface](https://pkg.go.dev/github.com/mozilla-ai/any-llm-go/providers) and [contributor interface assertions](https://github.com/mozilla-ai/any-llm-go/blob/main/CONTRIBUTING.md) (accessed 2026-09-24). +- The Go port uses channels plus a separate error channel for streaming, and its provider matrix explicitly marks unsupported combinations such as reasoning or embeddings instead of pretending every provider is equivalent. [Go README](https://github.com/mozilla-ai/any-llm-go) and [provider matrix](https://github.com/mozilla-ai/any-llm-go/blob/main/docs/providers.md) (accessed 2026-09-24). +- The optional `ProviderData`/provider-specific metadata escape hatch and the OpenAI-compatible base provider are useful compatibility patterns, provided they remain adapter-internal or explicitly versioned. [Go provider types](https://pkg.go.dev/github.com/mozilla-ai/any-llm-go/providers) and [OpenAI-compatible provider](https://pkg.go.dev/github.com/mozilla-ai/any-llm-go/providers/openai) (accessed 2026-09-24). + +**Boundary:** any-llm is a provider library, not an agent runtime. Its optional gateway, virtual keys, and hosted platform are separate products; Flux should borrow the library shape and capability declarations, not its control plane. + +**Assessment:** Highest-priority Go donor. The main risk is copying an OpenAI-shaped response too literally and losing provider-native reasoning, signed thinking blocks, or media semantics. + +### 4. Pydantic AI — typed messages, schemas, and testability donor + +**Evidence status:** Core typed agent/model/tool functionality, streaming, MCP, OTel instrumentation, and offline testing are documented. Durable execution and the Harness are integration/package surfaces; their availability should be checked per deployment rather than assumed to be part of the core. [Pydantic AI overview](https://pydantic.dev/docs/ai/overview/) and [repository](https://github.com/pydantic/pydantic-ai) (accessed 2026-09-24). + +**Capability patterns:** + +- Messages are represented as typed parts such as user prompt, text response, tool call, and tool return, with request/response usage, model name, timestamps, run ID, and conversation ID. This is a stronger model than Flux's current `Content` plus parallel `ToolUse`/`ToolResults` fields for new contracts. [Pydantic AI tools and message output](https://pydantic.dev/docs/ai/tools-toolsets/tools/) (accessed 2026-09-24). +- Function signatures and docstrings generate JSON Schema; arguments are validated before execution, and tool validation errors can be returned to the model for another attempt. Output types are validated and can trigger reflection/self-correction. [Function tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/) and [Pydantic AI overview](https://pydantic.dev/docs/ai/overview/) (accessed 2026-09-24). +- The `TestModel` runs without an API key, which is a strong pattern for deterministic provider-adapter and host-contract tests. Flux should have an equivalent offline fake model/stream source for conformance tests. [Pydantic AI overview](https://pydantic.dev/docs/ai/overview/) (accessed 2026-09-24). +- Instrumentation emits standard OTel spans for model and tool calls, and Pydantic Evals tests agent behavior separately from the core runtime. This supports a boundary where Flux exports normalized events/usage while hosts own evaluations. [Instrumentation and evaluations](https://pydantic.dev/docs/ai/overview/) (accessed 2026-09-24). +- `pydantic-graph` is deliberately a separate typed state-machine library. That separation is directly relevant: graph vocabulary can be shared without making the model engine own graph execution. [Pydantic Graph](https://pydantic.dev/docs/ai/graph/) (accessed 2026-09-24). + +**Boundary:** Pydantic AI's `Agent`, Runner-equivalent loop, Harness, memory, planning, subagents, CLI, built-in web/voice interfaces, durable integrations, and Evals product are host concerns. Flux should adopt typed DTO and validation patterns only. + +**Assessment:** Best donor for schema-first contracts and deterministic tests. Do not import Pydantic's agent semantics into `engine`. + +### 5. OpenAI Agents SDK — stream completion, usage, and tool lifecycle donor + +**Evidence status:** The Python SDK and TypeScript SDK are official. The docs call the Python package production-ready, but many features are OpenAI-specific or hosted; those are not portable capabilities. [Official SDK selection page](https://developers.openai.com/api/docs/guides/agents/sdk) and [Python repository](https://github.com/openai/openai-agents-python) (accessed 2026-09-24). + +**Capability patterns:** + +- The documentation makes the boundary unusually explicit: `Agent` plus `Runner` manages turns, tools, guardrails, handoffs, and sessions; applications that want to own the loop should use the lower-level Responses API instead. This is a direct endorsement of Flux's “engine emits requests, host owns the loop” rule. [Agents](https://openai.github.io/openai-agents-python/agents) (accessed 2026-09-24). +- Streaming has a crucial completion rule: callers must consume the event iterator to the end because session persistence, approval bookkeeping, or history compaction can finish after the last visible token. `cancel()` can stop immediately or after the current turn, and interrupted runs expose resumable state. [Streaming](https://openai.github.io/openai-agents-python/streaming) (accessed 2026-09-24). +- The SDK has explicit raw response, run-item, and agent-update stream events, plus nested-agent streaming and a distinction between a handoff request and a tool call. Flux should use stable event phases and correlation IDs rather than infer turn boundaries from text deltas. [Streaming](https://openai.github.io/openai-agents-python/streaming) and [agent tool stream events](https://openai.github.io/openai-agents-python/ref/agent) (accessed 2026-09-24). +- Usage is aggregated across model calls, tool calls, handoffs, and compaction. `preserve_raw_usage` retains provider-specific usage fields, while third-party adapters may require `include_usage=True`; omission and zero must remain distinguishable. [Usage](https://openai.github.io/openai-agents-python/usage) (accessed 2026-09-24). +- Tracing has workflow, task, turn, agent, generation, function, guardrail, handoff, and audio spans. Its default sensitive-data behavior is permissive, and tracing is unavailable under zero-data-retention organizations; Flux should default to metadata-only and make content capture explicit. [Tracing](https://openai.github.io/openai-agents-python/tracing) (accessed 2026-09-24). +- Tools distinguish local function tools, provider-executed tools, hosted MCP, computer/code tools, deferred approval, and output trimming. Flux should expose tool requests and provider state but not execute or approve them. [Tools](https://openai.github.io/openai-agents-python/tools) (accessed 2026-09-24). + +**Boundary:** Runner, handoffs, guardrails, sessions, hosted tools, computer use, voice/realtime, and agent tracing are host or platform concerns. The portable donor is the model adapter and stream/usage contract. + +**Assessment:** High-value donor for terminal stream semantics, raw usage preservation, and cancellation state. Treat OpenAI-only feature claims as non-portable. + +### 6. LangChain — broad model and middleware comparison + +**Evidence status:** The current LangChain docs and integrations are shipped; the ecosystem is large and changes quickly. [LangChain models](https://docs.langchain.com/oss/python/langchain/models) and [repository](https://github.com/langchain-ai/langchain) (accessed 2026-09-24). + +**Capability patterns:** + +- The standard chat-model interface supports invoke, stream, batch, tool calling, structured output, multimodal blocks, reasoning, and provider-specific parameters. This is a useful capability matrix for Flux's model catalog and adapter contract. [Models](https://docs.langchain.com/oss/python/langchain/models) (accessed 2026-09-24). +- `model.profile` exposes context limits, image inputs, reasoning output, tool calling, and related capabilities, with much of the data sourced from models.dev. Flux already has context, pricing, capability, source, and live metadata fields; a typed, provenance-bearing profile is the next step. [Model profiles](https://docs.langchain.com/oss/python/langchain/models) and [models.dev](https://models.dev/) (accessed 2026-09-24). +- Message content may be a string or provider-native content blocks, while the framework exposes a normalized content-block view. This supports a dual representation: canonical parts for hosts plus opaque provider-native blocks for replay. [Messages](https://docs.langchain.com/oss/python/langchain/messages) (accessed 2026-09-24). +- Default model retries use exponential backoff for network, 429, and 5xx errors, but not ordinary client errors; the number of retries is configurable. This is a useful default, with Flux retaining provider-specific retryability and retry budgets. [Models: connection resilience](https://docs.langchain.com/oss/python/langchain/models) (accessed 2026-09-24). +- Middleware offers before/after-agent, before/after-model, and wrap-model-call hooks, including model/tool retries, tool-error conversion, content transformation, and stream transformers. Flux should adopt only the model-call layer, not agent state or tool execution middleware. [Custom middleware](https://docs.langchain.com/oss/python/langchain/middleware/custom) and [built-in middleware](https://docs.langchain.com/oss/python/langchain/middleware/built-in) (accessed 2026-09-24). + +**Boundary:** LangChain agents, LangSmith, retrieval/vector stores, memory, and deployment tooling are host/platform concerns. Flux can learn the standardized model interface and model-call middleware without depending on the LangChain package graph. + +**Assessment:** Broadest comparative reference for catalog capabilities, content blocks, structured output, and retry/middleware policy. Avoid importing its broad integration surface. + +### 7. LangGraph — durable graph and host-state donor + +**Evidence status:** LangGraph is a low-level orchestration runtime for stateful, long-running workflows; its persistence and streaming features are shipped. [LangGraph overview](https://docs.langchain.com/oss/python/langgraph/overview) and [repository](https://github.com/langchain-ai/langgraph) (accessed 2026-09-24). + +**Capability patterns:** + +- The framework explicitly separates deterministic steps from LLM-driven steps and focuses on durable execution, streaming, human-in-the-loop, and persistence rather than prompts or model abstraction. [LangGraph overview](https://docs.langchain.com/oss/python/langgraph/overview) (accessed 2026-09-24). +- Durable execution uses a checkpointer, thread identifier, and three durability modes: `exit`, `async`, and `sync`. The docs emphasize deterministic, idempotent tasks around non-deterministic side effects. [Durable execution](https://docs.langchain.com/oss/python/langgraph/durable-execution) (accessed 2026-09-24). +- Streaming exposes updates, messages, and custom projections, with a newer typed event-streaming API. This supports the idea that Flux's normalized model stream should be composable with host graph events without becoming a graph runtime. [LangChain streaming](https://docs.langchain.com/oss/python/langchain/streaming) and [LangGraph streaming](https://docs.langchain.com/oss/python/langgraph/streaming) (accessed 2026-09-24). + +**Boundary:** Checkpointers, thread stores, reducers, interrupt/resume commands, graph compilers, and durable execution belong in the host or workflow engine. Flux's `graph` package should remain contracts/projections only. + +**Assessment:** Important negative boundary donor: it demonstrates exactly how much stateful machinery a host may need around a stateless model call. + +### 8. Google Genkit — Go action/flow/plugin donor + +**Evidence status:** Genkit flows, actions, plugins, typed schemas, streaming, and developer tooling are shipped in multiple languages. The Agents API is explicitly beta and may make breaking changes in minor releases. [Genkit Go flows](https://genkit.dev/docs/go/flows) and [Genkit agents](https://genkit.dev/docs/go/agents/overview) (accessed 2026-09-24). + +**Capability patterns:** + +- A flow is a typed function with schema-validated input/output, partial streaming, trace visibility, and HTTP deployment adapters. This is a clean separation between model calls and host workflows. [Defining AI workflows in Go](https://genkit.dev/docs/go/flows) (accessed 2026-09-24). +- Actions, models, tools, prompts, retrievers, and evaluators are registry-like extension points. Go code can define tools, structured data, interrupts, and streaming flows without a UI dependency. [Genkit repository](https://github.com/genkit-ai/genkit) and [Genkit Go documentation](https://genkit.dev/docs/go/get-started) (accessed 2026-09-24). +- The developer CLI/UI can run flows, evaluate them, inspect traces, and expose model/tool/prompt components. This is a useful host-tooling boundary, not a reason for Flux to ship a Dev UI. [Genkit developer tools](https://genkit.dev/docs/dart/devtools/) (accessed 2026-09-24). +- Tool interrupts and restartable tools model approval/resume as explicit protocol states. Flux can normalize an `interrupt`/approval event for host consumption without implementing approval policy. [Genkit JS API reference](https://js.api.genkit.dev/) (accessed 2026-09-24). + +**Boundary:** Genkit flows, agents, Dev UI, evaluation UI, and deployment adapters are host concerns. The portable donor is the action/flow separation and typed streaming discipline; the beta Agents API should not drive Flux API commitments. + +**Assessment:** One of the best Go-language donors because it treats flows and model actions as separate contracts while still supporting rich streaming. + +### 9. Google ADK — event semantics and evaluation donor + +**Evidence status:** ADK is an open-source code-first agent toolkit in Python, TypeScript, Go, Java, and Kotlin. The docs identify graph-based workflows and other agent features as current, while some streaming/tutorial pages label future support. [ADK agents](https://google.github.io/adk-docs/agents) and [repository](https://github.com/google/adk-python) (accessed 2026-09-24). + +**Capability patterns:** + +- ADK's event stream distinguishes partial text, complete text, turn completion, interruption, tool calls, state deltas, and final-response filtering. These are better than treating every provider chunk as an untyped delta. [ADK events](https://google.github.io/adk-docs/events) (accessed 2026-09-24). +- Streaming is documented as unbuffered and event-driven, with explicit handling for audio, transcription, partial flags, and interrupted turns. Flux should preserve partial/complete semantics in its canonical stream even when it does not implement realtime audio. [ADK streaming guide](https://google.github.io/adk-docs/streaming/dev-guide/part3) (accessed 2026-09-24). +- Workflow agents include sequential, parallel, and loop primitives; plugins and callbacks add logging, monitoring, and side effects without changing the core agent. [ADK agents](https://google.github.io/adk-docs/agents) and [callbacks](https://google.github.io/adk-docs/callbacks) (accessed 2026-09-24). +- ADK exposes sessions, state, memory, artifacts, callbacks, evaluation, and deployment integrations. Those are excellent examples of host-owned surfaces around a model runtime. [ADK technical overview](https://google.github.io/adk-docs/get-started/about) (accessed 2026-09-24). + +**Boundary:** Runner, agent classes, workflow agents, sessions, memory, artifacts service, web interface, and deployment belong to the host/platform. Flux may borrow event flags and callback points, not the runtime. + +**Assessment:** Strong donor for event conformance and multimodal/artifact references. The framework itself is too host-oriented for Flux. + +### 10. Microsoft Agent Framework, AutoGen, and Semantic Kernel — convergence and middleware donor + +**Evidence status:** Microsoft Agent Framework is the direct successor to AutoGen and Semantic Kernel. AutoGen's current README says it is in maintenance mode and directs new users to Agent Framework; Semantic Kernel remains an active repository, while its agent/process capabilities have historically included preview surfaces. [Microsoft overview](https://learn.microsoft.com/en-us/agent-framework/overview), [AutoGen README](https://github.com/microsoft/autogen/blob/main/README.md), and [Semantic Kernel repository](https://github.com/microsoft/semantic-kernel) (accessed 2026-09-24). + +**Capability patterns:** + +- Agent Framework combines AutoGen-style single/multi-agent abstractions with Semantic Kernel-style session state, type safety, filters, telemetry, and graph-based workflows. [Microsoft overview](https://learn.microsoft.com/en-us/agent-framework/overview) (accessed 2026-09-24). +- Middleware is layered into agent-run, function-calling, and chat-client boundaries. Execution order is documented, and each layer can inspect/modify input, output, and control flow. Flux should expose only the chat/model-call layer, while hosts own agent and function middleware. [Agent middleware](https://learn.microsoft.com/en-us/agent-framework/agents/middleware) (accessed 2026-09-24). +- Workflows are directed graphs of executors and edges, support streaming and non-streaming execution, and document Go workflow packages in addition to Python/.NET. This reinforces the value of a portable graph vocabulary without a Flux-owned executor. [Workflow builder and execution](https://learn.microsoft.com/en-us/agent-framework/workflows/workflows) (accessed 2026-09-24). +- Workflow observability emits build/run/executor spans, can disable sensitive data, and exposes fine-grained telemetry options. [Workflow observability](https://learn.microsoft.com/en-us/agent-framework/workflows/observability) (accessed 2026-09-24). +- The migration guide explicitly maps AutoGen's event-driven team model to Agent Framework's typed graph workflow and calls out differences in distributed execution, sessions, and tool iteration. Those differences are useful boundary evidence, not a reason to merge the frameworks. [AutoGen migration guide](https://learn.microsoft.com/en-us/agent-framework/migration-guide/from-autogen) (accessed 2026-09-24). + +**Boundary:** Agent sessions, graph execution, checkpoints, context providers, orchestration patterns, and hosting are host concerns. Flux should not implement AutoGen/MAF agents or Semantic Kernel process runtime. + +**Assessment:** Use Agent Framework as the current Microsoft reference; use AutoGen and Semantic Kernel for historical patterns and migration lessons. AutoGen itself should not be a new Flux dependency. + +### 11. Haystack — component validation, cancellation, and tracing donor + +**Evidence status:** Haystack 3.1 documentation and source describe directed multigraph pipelines, async streaming, breakpoints, tools, and pluggable tracing. [Haystack pipelines](https://docs.haystack.deepset.ai/docs/pipelines) and [repository](https://github.com/deepset-ai/haystack) (accessed 2026-09-24). + +**Capability patterns:** + +- Pipelines validate component names, input/output sockets, and types at connection time; serialization allows a graph/configuration to be inspected and reproduced. This is a strong model for provider adapter manifests and conformance fixtures. [Pipelines](https://docs.haystack.deepset.ai/docs/pipelines) (accessed 2026-09-24). +- Async pipelines can stream `StreamingChunk` values, cap concurrency, and cancel/drain sibling tasks when a component fails or a consumer stops early. The docs also warn that synchronous components offloaded to threads cannot be interrupted and may finish side effects in the background. Flux should make cancellation guarantees explicit per adapter and avoid claiming universal cancellation. [Pipelines: async execution and cancellation](https://docs.haystack.deepset.ai/docs/pipelines) (accessed 2026-09-24). +- Tracing is opt-in, supports OpenTelemetry and multiple backends, and has a custom `Tracer` interface. Content tracing is disabled by default. These are directly reusable observability patterns. [Haystack tracing](https://docs.haystack.deepset.ai/docs/tracing) and [custom tracer](https://docs.haystack.deepset.ai/docs/tracing-custom-tracer) (accessed 2026-09-24). +- `PipelineTool` exposes a whole pipeline as a tool with input/output mapping and async invocation, illustrating why tool schemas and workflow execution should remain separate. [PipelineTool](https://docs.haystack.deepset.ai/docs/3.0/pipelinetool) (accessed 2026-09-24). + +**Boundary:** Pipelines, loops, RAG, document stores, and agent components are host products. Flux can borrow validation, cancellation, stream chunks, and tracer interfaces only. + +**Assessment:** Excellent donor for “explicit contracts, validate before run, content tracing off by default” rather than for a new Flux pipeline layer. + +### 12. LlamaIndex — typed event workflows and data boundary donor + +**Evidence status:** LlamaIndex Workflows are shipped as a standalone event-driven library and are automatically instrumented when supported integrations are used. [Workflows](https://docs.llamaindex.ai/en/stable/module_guides/workflow) and [repository](https://github.com/run-llama/llama_index) (accessed 2026-09-24). + +**Capability patterns:** + +- A workflow step receives a typed event and returns a typed event; branches are ordinary conditions, loops return events, and concurrency can be expressed with event lists. The framework validates the event graph before running. [Workflow introduction](https://docs.llamaindex.ai/en/stable/module_guides/workflow) (accessed 2026-09-24). +- `WorkflowHandler.stream_events()` exposes events while the final result remains awaitable, and `Context` separates per-run state from resources such as clients, indexes, and models. This is a useful distinction for Flux's stateless generation facade. [Workflow introduction](https://docs.llamaindex.ai/en/stable/module_guides/workflow) (accessed 2026-09-24). +- Workflow documentation includes branches/loops, error handling/retry steps, human-in-the-loop, durable workflows, testing, observability, and server deployment. These are host orchestration features, not provider-runtime features. [LlamaAgents workflow documentation](https://developers.llamaindex.ai/python/llamaagents/workflows/) (accessed 2026-09-24). +- The agent model supports FunctionAgent, ReAct, and CodeAct variants, while the broader framework is explicitly data/RAG-centric. This reinforces that agent loop and retrieval belong outside Flux. [Building an agent](https://developers.llamaindex.ai/python/framework/understanding/agent/) and [LlamaIndex overview](https://docs.llamaindex.ai/) (accessed 2026-09-24). + +**Boundary:** Workflow execution, checkpoints, RAG indexes, document loaders, agents, and deployment server are host concerns. Flux should expose normalized generation events and catalog facts, not the data plane. + +**Assessment:** Strong donor for typed event envelopes, validation, and state/resource separation; not a provider-runtime candidate. + +### 13. BAML — schema-first structured output donor + +**Evidence status:** BAML's core repository, generated clients, streaming structured data, and editor/test tooling are documented. Performance and compatibility claims are vendor claims unless independently reproduced. [BAML documentation](https://docs.boundaryml.com/home) and [repository](https://github.com/BoundaryML/baml) (accessed 2026-09-24). + +**Capability patterns:** + +- BAML makes prompts and output schemas first-class typed artifacts, with generated language bindings and streaming typed values. [BAML home](https://docs.boundaryml.com/home) (accessed 2026-09-24). +- The project exposes Go and other language clients, plus OpenAPI integration, which suggests a schema descriptor can be more portable than a language-specific response wrapper. [OpenAPI and multi-language announcement](https://boundaryml.com/blog/announcing-openapi-support) and [Go package](https://pkg.go.dev/github.com/boundaryml/baml-go) (accessed 2026-09-24). +- The editor/playground, linting, preview, and test concepts are valuable for Flux conformance fixtures and schema regression tests, even though the editor itself is not an engine responsibility. [BAML home](https://docs.boundaryml.com/home) (accessed 2026-09-24). + +**Boundary:** BAML's DSL, prompt files, editor, optimizer, and evaluation application are host/tooling concerns. Flux should expose a schema descriptor and validator/repair interface, not a second prompt language. + +**Assessment:** Best focused donor for structured output correctness and schema evolution. It complements rather than replaces the provider/runtime abstraction. + +### 14. Mastra — TypeScript workflow, observability, and evaluation donor + +**Evidence status:** Mastra's core framework and built-in agent/workflow/observability/eval surfaces are shipped; the repository uses a dual-license model with `ee/` code under the Mastra Enterprise License. [Mastra repository licensing](https://github.com/mastra-ai/mastra/blob/main/README.md) and [Mastra framework](https://mastra.ai/ai-agent-framework) (accessed 2026-09-24). + +**Capability patterns:** + +- Mastra separates open-ended agents from explicit workflows, with typed steps, branching, parallel execution, loops, suspension/resumption, and storage-backed state. [Mastra agents](https://mastra.ai/docs/agents/overview.md) and [Mastra workflows](https://mastra.ai/docs/workflows/overview) (accessed 2026-09-24). +- Workflow steps can call agents or deterministic tools, and structured output schemas validate handoffs. This is a useful host-side pattern for consuming Flux's normalized responses. [Agents and tools in workflows](https://mastra.ai/docs/workflows/agents-and-tools) (accessed 2026-09-24). +- Built-in scorers, datasets, experiments, traces, metrics, and logs make evaluation a first-class product concern rather than an implicit model-runtime feature. [Mastra framework](https://mastra.ai/ai-agent-framework) (accessed 2026-09-24). +- The repository advertises a unified model router; exact provider counts differ between README and product pages, so only the existence of a router—not a specific count—should be treated as a donor fact. [Mastra repository](https://github.com/mastra-ai/mastra) (accessed 2026-09-24). + +**Boundary:** Mastra agents, workflows, Studio, memory, server, and eval product are host concerns. Flux can borrow typed router and lifecycle ideas, but should not import a TypeScript framework or enterprise runtime. + +**Assessment:** Good donor for how an ecosystem can separate model calls from workflow/evaluation/UI layers. Use it as a TypeScript boundary reference, not a dependency. + +### 15. Strands Agents — minimal model-driven protocol donor + +**Evidence status:** Strands has Python and TypeScript SDKs, a model-driven agent loop, provider integrations, MCP, streaming, structured output, hooks, sessions, and OTel. [Strands documentation](https://strandsagents.com/docs/) and [Python repository](https://github.com/strands-agents/sdk-python) (accessed 2026-09-24). + +**Capability patterns:** + +- The minimal `AgentBase` protocol exposes asynchronous invoke and stream operations, making it easy to see the small surface a provider runtime needs. [Strands AgentBase API](https://strandsagents.com/docs/api/python/strands.agent.base) (accessed 2026-09-24). +- Provider support includes Bedrock, Anthropic, OpenAI, Gemini, Ollama, LiteLLM, and custom providers, with structured output and token/execution metrics. [Strands quickstart/features](https://strandsagents.com/docs/) (accessed 2026-09-24). +- Hooks/plugins cover lifecycle events around model calls and tool calls; sessions are pluggable, and OTel configuration supports tracing and metrics. [Strands repository guidance](https://github.com/strands-agents/sdk-python/blob/main/AGENTS.md) and [telemetry API](https://strandsagents.com/docs/api/python/strands.telemetry.config) (accessed 2026-09-24). +- Graph, Swarm, Workflow, and A2A patterns are explicit multi-agent/host concepts. [Strands multi-agent documentation](https://strandsagents.com/docs/) (accessed 2026-09-24). + +**Boundary:** Agent loop, tool execution, sandboxes, sessions, multi-agent patterns, and MCP client belong to the host. Flux can borrow the minimal model protocol and hook vocabulary, but should not add Strands' agent runtime. + +**Assessment:** Useful counterweight to large frameworks: a small model-first protocol plus explicit hooks can be enough for a provider engine. + +### 16. DSPy — evaluation, optimization, and program-artifact donor + +**Evidence status:** DSPy's signatures/modules, adapters, metrics, evaluation, and optimizers are shipped; the project has active 2026 releases and a beta line. [DSPy documentation](https://dspy.ai/), [releases](https://github.com/stanfordnlp/dspy/releases) (accessed 2026-09-24). + +**Capability patterns:** + +- DSPy programs are expressed as typed signatures and composable modules rather than hand-managed prompts. This is a useful conceptual model for stable test cases and prompt-independent provider conformance. [DSPy overview](https://dspy.ai/) (accessed 2026-09-24). +- Optimizers such as GEPA, MIPRO, SIMBA, and bootstrap approaches optimize instructions, demonstrations, or weights against explicit metrics and datasets. Flux should expose evaluation inputs/outputs and trace data, not own optimization. [DSPy optimizers](https://dspy.ai/3.0.2/learn/optimization/optimizers/) (accessed 2026-09-24). +- Release notes describe active work on typed provider-neutral model systems, tool-aware ReAct, MCP compatibility, and interpreter isolation. These are roadmap/current-release signals, not reasons to make Flux a DSPy host. [DSPy 3.3.1 release](https://github.com/stanfordnlp/dspy/releases) (accessed 2026-09-24). +- A March 2026 issue documents a LiteLLM supply-chain incident and the resulting dependency pin/remediation. This is a concrete reason to keep dependency and lockfile hygiene visible in Flux's conformance/release process, even though it is not a DSPy architecture feature. [DSPy issue 9500](https://github.com/stanfordnlp/dspy/issues/9500) (accessed 2026-09-24). + +**Boundary:** Prompt/program optimization, fine-tuning, datasets, and optimizer CLIs are host/tooling concerns. Flux can provide normalized events, usage, and evaluation hooks. + +**Assessment:** Essential donor for evaluation and versioned prompt/program artifacts, but explicitly not a runtime framework to embed. + +### 17. Agno — SDK/runtime/control-plane separation donor + +**Evidence status:** Agno's SDK, AgentOS runtime, storage, control plane, tracing, approvals, scheduling, and UI surfaces are documented. [Agno framework](https://www.agno.com/agent-framework), [repository](https://github.com/agno-agi/agno) (accessed 2026-09-24). + +**Capability patterns:** + +- The project explicitly separates building agents with the SDK from running them with AgentOS and managing them with a control plane. This is a useful warning against allowing a model SDK to absorb all three layers. [Agno repository](https://github.com/agno-agi/agno) (accessed 2026-09-24). +- Storage owns sessions, memory, knowledge, and traces; the runtime exposes SSE/WebSocket interfaces, approvals, scheduling, and human review. These are host/platform concerns, but their event and usage vocabulary can inform Flux contracts. [Agno framework](https://www.agno.com/agent-framework) (accessed 2026-09-24). +- OTel tracing, audit logs, JWT RBAC, and multi-user isolation are explicitly called out. Flux should borrow privacy-safe trace hooks and safe status projections, not implement an admin control plane. [Agno repository](https://github.com/agno-agi/agno) (accessed 2026-09-24). +- Performance and customer numbers on the product site are vendor claims and were not used as evidence for Flux priorities. [Agno about page](https://www.agno.com/about) (accessed 2026-09-24). + +**Boundary:** AgentOS, memory/knowledge, scheduling, approvals, RBAC, UI, and control plane are host concerns. Flux remains the transport/catalog/resilience layer. + +**Assessment:** Valuable organizational-boundary donor, less valuable as a direct provider-contract donor. + +### 18. CrewAI — event-driven host workflow donor + +**Evidence status:** CrewAI's Crews and Flows are shipped in the MIT-licensed repository. Flows provide event-driven state, branching, routing, and config-as-code; Crews provide higher-level autonomous multi-agent behavior. [CrewAI repository](https://github.com/crewAIInc/crewAI) and [Flows](https://docs.crewai.com/concepts/flows) (accessed 2026-09-24). + +**Capability patterns:** + +- Flows use start/listen/router decorators, persistent state, logical branching (`or_`/`and_`), and asynchronous composition. This is a clear example of host-owned orchestration around ordinary model/tool calls. [Flows](https://docs.crewai.com/concepts/flows) (accessed 2026-09-24). +- Crews and tasks are configurable in YAML/JSON and support streaming, tools, knowledge, and planning. The configuration model is useful for host prompt/version management, not Flux provider selection. [Crews](https://docs.crewai.com/concepts/crews) (accessed 2026-09-24). +- The framework's high-level agent abstraction makes it a poor dependency for a provider-neutral engine; its value is in showing what should remain above the engine boundary. [CrewAI repository](https://github.com/crewAIInc/crewAI) (accessed 2026-09-24). + +**Boundary:** Crews, tasks, roles, planning, memory, and flow execution belong to the host. Flux should not expose Crew/Task concepts. + +**Assessment:** Useful secondary donor for declarative host workflows, but not a direct Flux runtime candidate. + +### 19. Dify — product and prompt-management comparator + +**Evidence status:** Dify is an actively maintained application platform for visual agentic workflows, RAG, model/tool integrations, APIs, and self-hosted/cloud deployment. [Dify introduction](https://docs.dify.ai/introduction) and [repository](https://github.com/langgenius/dify) (accessed 2026-09-24). + +**Capability patterns:** + +- Dify's product surface combines visual workflow design, agent construction, prompt engineering, model management, RAG, monitoring, and published APIs. This is useful for understanding the application layer a host must own around a model engine. [Dify introduction](https://docs.dify.ai/introduction) (accessed 2026-09-24). +- Its documentation and product comparisons emphasize prompt crafting/versioning/testing, provider routing, cost tracking, and deployment modes. These are product/LLMOps requirements, not provider-neutral wire contracts. [Open WebUI comparison page](https://docs.openwebui.com/alternatives/dify) and [Dify repository](https://github.com/langgenius/dify) (accessed 2026-09-24). +- The Dify license is a modified Apache-2.0-based license with conditions around multi-tenant operation and frontend branding; it is not a clean permissive dependency for a reusable engine. [Dify license](https://github.com/langgenius/dify/blob/main/LICENSE) and [Dify license policy](https://docs.dify.ai/policies/open-source) (accessed 2026-09-24). + +**Boundary:** UI, visual workflow canvas, prompt IDE, RAG, application publishing, multi-tenant gateway, and commercial licensing stay outside Flux. Flux can expose catalog/preflight/usage primitives that a Dify-like host consumes. + +**Assessment:** Include to make the host/product boundary concrete, not to copy code or dependencies. + +### 20. Open WebUI — UI, extension, and tool-ownership comparator + +**Evidence status:** Open WebUI is an actively maintained, self-hosted AI interface with OpenAI-compatible/Ollama connections, RAG, voice/video, plugins, tools, RBAC, artifacts, and analytics. [Open WebUI documentation](https://docs.openwebui.com/) and [repository](https://github.com/open-webui/open-webui) (accessed 2026-09-24). + +**Capability patterns:** + +- Open WebUI supports multiple model/API connections, local inference, multi-model conversations, RAG, voice/video, artifact storage, usage analytics, and broad extension points. These are strong examples of host-owned UX and data-plane scope. [Open WebUI repository](https://github.com/open-webui/open-webui) (accessed 2026-09-24). +- Pipelines is a UI-agnostic OpenAI-compatible plugin framework for offloading custom logic. Its documentation warns that pipelines execute arbitrary code and must be trusted. [Pipelines](https://github.com/open-webui/pipelines) and [Pipelines documentation](https://docs.openwebui.com/features/extensibility/pipelines) (accessed 2026-09-24). +- The Pipe documentation gives a crucial tool-loop ownership rule: if the pipeline owns and executes a tool internally, it must not emit `delta.tool_calls`, or Open WebUI will execute the same tool again; if Open WebUI should execute it, the pipeline emits tool calls and terminates with `finish_reason: "tool_calls"`. [Pipe tool-call semantics](https://docs.openwebui.com/features/extensibility/pipelines/pipes) (accessed 2026-09-24). +- The current Open WebUI license adds branding restrictions for larger deployments; older code has different license history. This makes it a product comparator rather than a neutral library dependency. [Open WebUI license](https://docs.openwebui.com/license) and [license file](https://github.com/open-webui/open-webui/blob/main/LICENSE) (accessed 2026-09-24). + +**Boundary:** UI, auth, RBAC, RAG, tool execution, artifacts, model conversations, and user-facing analytics are host concerns. Flux should provide explicit tool ownership semantics and safe extension points, not a chat UI. + +**Assessment:** Important negative donor: it demonstrates how easily a provider runtime turns into an application/control plane, and how dangerous ambiguous tool-loop ownership is. + +### Cross-project synthesis + +#### Provider and normalized content patterns + +The most portable pattern is a two-layer contract: canonical parts for portable host logic and opaque provider blocks/metadata for lossless replay. LangChain and Vercel expose provider-native blocks; Pydantic uses typed request/response parts; OpenAI Agents and Google ADK expose explicit raw/run-item/partial event categories; Flux already has `ProviderBlock` but should ensure it survives every normalization path. [LangChain messages](https://docs.langchain.com/oss/python/langchain/messages), [Pydantic tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/), [OpenAI streaming](https://openai.github.io/openai-agents-python/streaming), [ADK events](https://google.github.io/adk-docs/events), [Flux DTOs](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go) (accessed 2026-09-24). + +#### Stream and tool-loop patterns + +All serious projects distinguish text deltas, tool-call deltas, complete tool calls, usage, errors, and terminal state. OpenAI adds the non-obvious rule that a stream is not complete until post-processing ends; Vercel exposes start/delta/end and input lifecycle hooks; ADK exposes partial/turn-complete/interrupted flags; Haystack defines cancellation and terminal stream behavior. Flux should make these distinctions normative in its event contract. [OpenAI streaming](https://openai.github.io/openai-agents-python/streaming), [Vercel tools](https://ai-sdk.dev/docs/ai-sdk-core/tools-and-tool-calling), [ADK events](https://google.github.io/adk-docs/events), [Haystack pipelines](https://docs.haystack.deepset.ai/docs/pipelines) (accessed 2026-09-24). + +#### Metadata and capability patterns + +Provider abstractions converge on capability declarations rather than a single “supports tools” boolean. LangChain's model profile, any-llm's capability matrix/optional interfaces, Vercel's provider registry, LiteLLM's model pricing/cost data, and Flux's `Model` fields each address a different part of routing. The durable pattern is a typed profile with provenance, freshness, and an explicit `unknown` state. [LangChain model profiles](https://docs.langchain.com/oss/python/langchain/models), [any-llm provider matrix](https://github.com/mozilla-ai/any-llm-go/blob/main/docs/providers.md), [Vercel provider management](https://ai-sdk.dev/docs/ai-sdk-core/provider-management), [LiteLLM models/pricing](https://docs.litellm.ai/docs/completion/input#model-list), [Flux model contract](https://github.com/GrayCodeAI/flux/blob/main/llm/provider.go) (accessed 2026-09-24). + +#### Reliability and middleware patterns + +Transport retries belong near the provider call, while tool retries and model fallbacks need host policy because tool side effects may not be idempotent. LangChain separates model and tool retry middleware; Microsoft separates chat, function, and agent middleware; Vercel wraps generate and stream independently; LiteLLM distinguishes within-group retries from cross-group fallback. Flux should keep the transport decision precise and expose metadata rather than silently retrying the whole turn. [LangChain middleware](https://docs.langchain.com/oss/python/langchain/middleware/built-in), [Microsoft middleware](https://learn.microsoft.com/en-us/agent-framework/agents/middleware), [Vercel middleware](https://ai-sdk.dev/docs/ai-sdk-core/middleware), [LiteLLM architecture](https://docs.litellm.ai/docs/proxy/architecture) (accessed 2026-09-24). + +#### Evaluation, prompt/version, and observability patterns + +DSPy, Pydantic Evals, Haystack tracing, OpenAI tracing, Pydantic instrumentation, LiteLLM OTel, and Vercel telemetry all separate runtime behavior from measurement. The shared lesson is that Flux should emit stable events, usage, route, and error metadata; hosts should own datasets, prompt registries, optimizer loops, and dashboards. [DSPy optimizers](https://dspy.ai/3.0.2/learn/optimization/optimizers/), [Pydantic overview](https://pydantic.dev/docs/ai/overview/), [Haystack tracing](https://docs.haystack.deepset.ai/docs/tracing), [OpenAI tracing](https://openai.github.io/openai-agents-python/tracing), [Vercel telemetry](https://ai-sdk.dev/docs/ai-sdk-core/telemetry) (accessed 2026-09-24). + +#### UI, deployment, and governance patterns + +Dify, Open WebUI, Agno, and Mastra show that UI, Studio, admin, multi-tenancy, storage, scheduling, and deployment are separable products. Their licenses also demonstrate why a provider engine should not copy a product's source-available terms or branding constraints. [Dify license](https://github.com/langgenius/dify/blob/main/LICENSE), [Open WebUI license](https://github.com/open-webui/open-webui/blob/main/LICENSE), [Mastra licensing](https://github.com/mastra-ai/mastra/blob/main/README.md), [Agno repository](https://github.com/agno-agi/agno) (accessed 2026-09-24). + +## Flux recommendations and boundary clarifications + +### Takeaway + +Flux's next capability work should be a **contract-hardening and conformance roadmap**, not an agent-framework roadmap. The highest-value additions are explicit stream phases and terminal invariants, lossless preservation of raw/provider state, typed capability and schema descriptors, provider-neutral structured-output validation, privacy-safe OTel, and a cross-provider conformance suite. Agent loops, tools, sessions, prompts, evaluations, RAG, UI, and durable workflows should remain host contracts. + +### Cited Findings + +#### P0: make the existing host contract lossless and testable + +Flux already defines `ContentPart`, `ProviderBlock`, `ToolCall.RawArguments`, `ToolCall.ProviderMetadata`, typed usage, typed errors, and a pull-based `EventStreamer`. [Flux DTOs](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go) and [Flux tools](https://github.com/GrayCodeAI/flux/blob/main/tools/tool.go) (accessed 2026-09-24). The local normalization path currently copies only tool ID, name, and decoded arguments when mapping a core tool event, while the DTO also carries raw arguments and provider metadata; the continuation path appends assistant text and a generic “Continue.” message. [Flux stream normalization](https://github.com/GrayCodeAI/flux/blob/main/engine/stream.go) and [Flux continuation](https://github.com/GrayCodeAI/flux/blob/main/engine/continuation.go) (accessed 2026-09-24). These are concrete reasons to add conformance tests before adding new framework-like features. + +**Recommended contract additions (additive where possible):** + +1. Add a stream event envelope with `Sequence`, `CallID`, `TurnID`/`SegmentID`, `Phase` (`start`, `delta`, `complete`, `error`, `cancelled`), `Index` for parallel/tool blocks, and an optional `ProviderBlock`/provider metadata field. Preserve `ToolCall.RawArguments` and `ProviderMetadata` through `engine.normalizeEvent`. +2. Define a terminal invariant: exactly one terminal outcome per call/continuation segment (`done`, `error`, or `cancelled`), with usage and route attached consistently. A stream must not appear complete merely because the provider stopped sending text; this follows the OpenAI completion lesson and is more robust than current implicit channel closure. [OpenAI streaming](https://openai.github.io/openai-agents-python/streaming) (accessed 2026-09-24). +3. Preserve provider-native reasoning/signed blocks and raw tool arguments on continuation requests. Make continuation segments individually addressable and aggregate usage across segments; do not let continuation silently drop provider state. [Flux provider blocks](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go) and [OpenAI usage](https://openai.github.io/openai-agents-python/usage) (accessed 2026-09-24). +4. Add explicit `ToolInputStart`, `ToolInputDelta`, `ToolInputComplete`, `ToolCallComplete`, `ToolError`, and optional `ToolApprovalRequest`/`ToolApprovalResponse` event types, while keeping tool execution in the host. [Vercel tool lifecycle](https://ai-sdk.dev/docs/ai-sdk-core/tools-and-tool-calling) and [OpenAI streaming](https://openai.github.io/openai-agents-python/streaming) (accessed 2026-09-24). +5. Make cancellation observable and testable: `Close` remains idempotent, context cancellation maps to a typed cancellation error, provider HTTP bodies are replayable for transport retries, and tests assert no goroutine/channel leak. Flux's current stream has an explicit close path and cancellation-aware forwarding. [Flux stream](https://github.com/GrayCodeAI/flux/blob/main/engine/stream.go) and [Flux retry](https://github.com/GrayCodeAI/flux/blob/main/provider/core/retry.go) (accessed 2026-09-24). + +#### P0: add a provider conformance suite, not an agent evaluation suite + +Build a black-box suite run against every adapter using local `httptest` servers and golden fixtures. It should test the **Flux contract**, not an agent loop. At minimum: + +| Conformance layer | Required cases | +|---|---| +| Request translation | text, system prompt, image URL/data URI, audio, tools, tool choice, limits, metadata, provider options, and secret redaction. | +| Blocking response | content, reasoning, finish reason, request ID, provider blocks, warnings, usage semantics, route/deployment, and structured output. | +| Streaming | text start/delta/end, reasoning, parallel tool calls, partial JSON arguments, malformed arguments, usage-before/after-finish, done/error/cancel, unknown provider events, and continuation segments. | +| Errors | auth, permission, rate limit, timeout, context exceeded, content filter, invalid request, model not found, 5xx, Retry-After, and cancellation; verify typed code plus retryability without parsing prose. | +| Replay | append normalized assistant message, replay signed/provider blocks, send tool result on the next host turn, and verify no state loss. | +| Media | URL passthrough, explicit conversion, MIME/size limits, unsupported modality, and no accidental secret/URL leakage in traces. | +| Structured output | native schema mode, function-call fallback only when explicitly enabled, JSON mode, validation failure, repair policy, raw output retention, and provider adjustment warning. | +| Reliability | retry budget, Retry-After, fallback scope, idempotency boundary, backoff jitter, and no retry for auth/client errors. | +| Compatibility | unknown event types are ignored safely; additive fields round-trip; provider-specific options require a capability declaration. | + +Use a capability matrix so unsupported cases are explicit `skip` results rather than false passes. Add property/fuzz tests for SSE framing, JSON argument fragments, unknown blocks, and oversized events. Haystack's connection validation, LlamaIndex's event-graph validation, Pydantic's offline `TestModel`, and Vercel's provider specification are useful precedents. [Haystack pipelines](https://docs.haystack.deepset.ai/docs/pipelines), [LlamaIndex workflows](https://docs.llamaindex.ai/en/stable/module_guides/workflow), [Pydantic overview](https://pydantic.dev/docs/ai/overview/), [Vercel model specification](https://github.com/vercel/ai/blob/main/packages/provider/src/language-model/v4/language-model-v4.ts) (accessed 2026-09-24). + +#### P0/P1: strengthen normalized content, tool, and schema types + +Flux's current `ContentPart` is a string-typed union with text, image URL, and base64 audio; `FluxTool.Parameters` is an untyped map; and `ResponseFormat` has only a type and schema string. [Flux content and tool types](https://github.com/GrayCodeAI/flux/blob/main/llm/types.go) (accessed 2026-09-24). Add an additive, versioned model such as: + +```text +ContentPart = text | image | audio | video | file | reasoning | tool_call | tool_result | provider_block +``` + +Each part should carry a stable type discriminator, MIME/format where relevant, source (`url`, `data`, `provider_reference`), and optional opaque provider data. Keep existing `Images` and `Content` fields for compatibility during the transition. This follows the normalized-plus-native pattern in LangChain, Vercel, Pydantic, and ADK. [LangChain messages](https://docs.langchain.com/oss/python/langchain/messages), [Pydantic tools](https://pydantic.dev/docs/ai/tools-toolsets/tools/), [ADK events](https://google.github.io/adk-docs/events) (accessed 2026-09-24). + +Add to tool definitions, without executing them: + +- `OutputSchema`/structured result descriptor. +- `Strict` and schema dialect/version. +- `ReadOnly`/`SideEffectClass` as advisory metadata for host policy, not an authorization decision. +- Stable tool namespace/version and compatibility metadata. + +Add to structured output: + +- schema name/version and JSON Schema dialect. +- `Strict` capability requirement. +- output mode (`native_schema`, `tool_call`, `json_mode`, `parse_only`) and the actual mode used. +- raw provider output plus validation error/repair count. +- explicit unsupported/capability-mismatch behavior. + +Flux currently maps any nonempty `OutputSchema` to a provider `json_schema` option in `engine/convert.go`; capability negotiation and a warning/error path are safer than assuming every provider accepts the same mode. [Flux conversion](https://github.com/GrayCodeAI/flux/blob/main/engine/convert.go) (accessed 2026-09-24). LangChain's distinction among `json_schema`, `function_calling`, and `json_mode` is a useful donor. [LangChain structured output](https://docs.langchain.com/oss/python/langchain/models) (accessed 2026-09-24). + +#### P1: add a narrow model-call middleware layer + +Expose a Flux-owned, provider-neutral middleware contract around one model call, not an agent loop. The minimum hooks are: + +- `BeforeRequest`/`TransformParams`. +- `WrapGenerate`. +- `WrapStream`. +- `AfterResponse`/`AfterStream`. +- typed context carrying cancellation, correlation IDs, and redacted telemetry fields. + +Define ordering, short-circuit behavior, and stream-backpressure guarantees. The middleware must be able to: + +- add/adjust normalized messages and provider options. +- apply caching or deterministic transforms. +- attach a guardrail or policy hook that can return a typed error. +- observe tool requests without executing them. +- redact secrets and sensitive content before telemetry. + +This follows Vercel's `transformParams`/`wrapGenerate`/`wrapStream` and LangChain's model-call wrapping, while intentionally excluding Microsoft-style agent/function middleware and host tool execution. [Vercel middleware](https://ai-sdk.dev/docs/ai-sdk-core/middleware), [LangChain custom middleware](https://docs.langchain.com/oss/python/langchain/middleware/custom), [Microsoft middleware](https://learn.microsoft.com/en-us/agent-framework/agents/middleware) (accessed 2026-09-24). + +#### P1: make model metadata capability-aware and provenance-bearing + +Flux already separates `Owner`, `ProviderID`, `GatewayID`, `CanonicalID`, `Source`, and `LiveMetadata`, and has a compiled catalog plus live/public discovery. [Flux model contract](https://github.com/GrayCodeAI/flux/blob/main/llm/provider.go) and [Flux catalog](https://github.com/GrayCodeAI/flux/blob/main/engine/catalog.go) (accessed 2026-09-24). Evolve this into a typed profile with: + +- capability values `supported`, `unsupported`, `unknown`, or `provider_defined`. +- input/output modalities, tool/structured-output/reasoning modes, context/output limits, and cache support. +- provider, model, deployment, and gateway provenance. +- pricing source/version/as-of timestamp and confidence. +- observed/live/compiled source and last-refresh time. +- provider option schema/version and passthrough support. +- actual served model/deployment versus requested alias. + +Borrow the profile idea from LangChain/models.dev, optional capability interfaces from any-llm, provider registries from Vercel, and pricing/spend separation from LiteLLM. Do not infer a capability merely because a provider accepts an OpenAI-shaped request. [LangChain model profiles](https://docs.langchain.com/oss/python/langchain/models), [any-llm provider matrix](https://github.com/mozilla-ai/any-llm-go/blob/main/docs/providers.md), [Vercel provider management](https://ai-sdk.dev/docs/ai-sdk-core/provider-management), [LiteLLM models/pricing](https://docs.litellm.ai/docs/completion/input#model-list) (accessed 2026-09-24). + +#### P1: separate usage, cost, and budget policy + +Keep `FluxUsage` as the normalized token fact, but add fields or a companion result for: + +- usage source: provider-reported, estimated, partial, or unavailable. +- input, output, reasoning, audio, cache-read, cache-creation, and uncached token detail. +- actual provider model/deployment and request ID. +- cost amount, currency, pricing version/as-of time, and whether cost is estimated. +- cache status and prompt-cache accounting. +- continuation segment and total-call usage. + +Aggregate usage across Flux-owned continuation calls, but do not aggregate host-owned agent turns. Never fabricate zero for an omitted provider field: distinguish omitted, unknown, and reported zero as the OpenAI Agents SDK does. [OpenAI usage](https://openai.github.io/openai-agents-python/usage) (accessed 2026-09-24). Emit usage as OTel GenAI attributes with content capture off by default, following LiteLLM and Vercel. [LiteLLM OTel v2](https://docs.litellm.ai/docs/observability/opentelemetry_v2), [Vercel telemetry](https://ai-sdk.dev/docs/ai-sdk-core/telemetry) (accessed 2026-09-24). + +#### P1: harden caching as an explicit engine policy + +Flux already offers response and semantic caching. The roadmap should make cache behavior inspectable: + +- canonical cache key includes normalized messages/parts, system prompt, tool definitions, schema, generation options, selected route/model/deployment, and provider-specific options that affect output. +- credentials, raw secret-bearing headers, and tenant secrets never enter keys or logs. +- cache result emits `hit`, `miss`, `bypass`, `stale`, or `error` metadata. +- streaming cache behavior is explicit; do not claim a cache hit before a complete response is available unless replay is supported. +- semantic similarity threshold and policy remain host-configurable; exact deterministic caching is safer as an engine default. +- prompt-cache tokens are accounted separately from application response-cache hits. + +LiteLLM's documentation specifically demonstrates provider-parameter cache-key and route-bypass concerns. [LiteLLM caching](https://docs.litellm.ai/docs/proxy/caching) (accessed 2026-09-24). + +#### P1: define safe provider-specific passthrough + +Provide two deliberately separate paths: + +1. **Normalized options:** allowlisted, versioned, provider-namespaced options that adapters validate and translate. +2. **Raw native passthrough:** an explicit opt-in escape hatch for callers who need a provider-native request/response, with a separate interface and clear warning that it is not portable. + +Never put untyped vendor SDK objects into the stable `llm`/`engine` contract. Preserve only opaque provider state needed for replay (`ProviderBlock`, thought signatures, provider metadata) unless a host explicitly requests raw passthrough. This is the safe middle ground between Vercel's provider options, any-llm's provider data, and Open WebUI's OpenAI-compatible plugin boundary. [Vercel provider options](https://ai-sdk.dev/docs/foundations/provider-options), [any-llm provider data](https://pkg.go.dev/github.com/mozilla-ai/any-llm-go/providers), [Open WebUI pipelines](https://docs.openwebui.com/features/extensibility/pipelines) (accessed 2026-09-24). + +#### P1: expand multimodal support without making Flux a media product + +Keep chat input parts provider-neutral and add file/document/video/audio references with MIME, URI/data source, and size limits. Keep separate, optional media operations for image generation and transcription if Flux already exposes them; do not add a gallery, asset manager, or browser UI. LiteLLM's explicit URL-to-base64 compatibility rule is a useful donor, but URL fetching must be size-limited, opt-in where appropriate, and never silently turn into an SSRF primitive. [LiteLLM image URL handling](https://docs.litellm.ai/docs/proxy/architecture), [Flux media facade](https://github.com/GrayCodeAI/flux/blob/main/engine/media.go) (accessed 2026-09-24). + +#### P2: expose optional engine capabilities, not a new framework + +If the roadmap later adds embeddings, reranking, batch, moderation, files, or provider-native responses, expose each as a small optional interface with capability negotiation and conformance tests. Any-llm-go's optional `EmbeddingProvider`, `ModelLister`, and `ErrorConverter` are a good Go pattern. [any-llm-go providers](https://pkg.go.dev/github.com/mozilla-ai/any-llm-go/providers) (accessed 2026-09-24). Do not add an agent, graph executor, session store, or UI merely because a framework bundles those features. + +### Minimal host-contract primitives Flux should expose + +These are the smallest additions that let a host build Rho-like semantics without importing Flux internals: + +1. **Generation request/response DTOs:** typed parts, tool schema, tool call/result, provider blocks, schema output, limits, route, warnings, usage, and error metadata. +2. **Event envelope:** call/turn/segment IDs, monotonic sequence, phase, block index, terminal status, raw/provider state, and cancellation state. +3. **Capability profile:** model/deployment/provider capabilities, limits, provenance, freshness, and output modes. +4. **Tool request contract:** schema, call ID, raw arguments, provider metadata, and result/error association; no executor. +5. **Structured-output contract:** schema descriptor, validation result, raw output, repair count, and provider mode. +6. **Extension contract:** versioned model-call middleware and provider adapter registration; no raw SDK types. +7. **Observability contract:** OTel-compatible spans, usage/cost fields, redaction policy, and injectable tracer/sink. +8. **Readiness contract:** preflight, catalog health, credential status, route preview, and live availability without leaking secrets. +9. **Graph vocabulary:** node/edge/event/provenance/idempotency fields only; no scheduler/checkpointer. +10. **Tool versioning vocabulary:** closed namespaces/behavior versions already present in `tools`; no tool policy engine. + +These primitives preserve the existing Flux ownership split: Flux owns credentials, catalog, route, transport, resilience, normalized usage, and provider telemetry; the host owns conversation history, permissions, tool execution, sessions, checkpoints, product memory, and user-visible behavior. [Flux boundary](https://github.com/GrayCodeAI/flux/blob/main/docs/architecture/HOST-ENGINE-BOUNDARY.md) and [Flux host port](https://github.com/GrayCodeAI/flux/blob/main/llm/provider.go) (accessed 2026-09-24). + +### Explicit capabilities Flux should reject + +| Reject from Flux | Reason / host owner | +|---|---| +| Agent classes, planners, handoffs, subagents, autonomous loops | OpenAI Agents, LangGraph, Pydantic AI, Google ADK, CrewAI, and Agno all place this above the model call. | +| Tool execution, permissions, approval policy, sandbox, MCP client | Open WebUI's tool-ownership rule and Vercel/OpenAI approval flows show that execution and policy are host/product semantics. | +| Conversation/session store, memory, checkpoints, replay runtime | LangGraph, Pydantic durable integrations, Letta, Agno, and ADK own durable state and recovery. | +| Graph compiler, scheduler, queue, distributed executor | LangGraph, Haystack, LlamaIndex, Microsoft Agent Framework, and ADK are workflow engines. Flux's `graph` package is explicitly data contracts only. | +| Prompt registry, prompt optimizer, fine-tuning, evaluation datasets | DSPy, Pydantic Evals, Dify, and BAML's editor/test tooling own this lifecycle. | +| RAG/vector/document ingestion | LlamaIndex, Haystack, Dify, and Open WebUI are data/application planes. | +| React/Svelte/Vue hooks, chat UI, generative UI, Studio, admin console | Vercel UI, Dify, Open WebUI, Genkit Dev UI, and Agno AgentOS are host products. | +| Virtual-key team billing/admin/multi-tenant control plane | LiteLLM Proxy, Dify, Agno, and Open WebUI provide platform services beyond a provider engine. | +| Untyped provider SDK leakage | Keep vendor types inside adapters; expose only normalized DTOs, opaque replay blocks, or explicit raw passthrough. | +| Automatic semantic caching or automatic tool-loop execution | Both change product semantics and can duplicate side effects; require explicit host policy. | + +### Recommended sequencing + +**P0: correctness before breadth** + +- Freeze a contract-v2 compatibility test matrix. +- Add stream phases, sequence/call/segment IDs, terminal invariants, typed stream errors, and raw/provider-state preservation. +- Add structured-output capability negotiation and validator/repair result types. +- Add cross-provider golden fixtures and fuzz tests for SSE/JSON/tool arguments. +- Make cancellation and retry scopes explicit in adapter documentation and tests. + +**P1: capability depth without framework creep** + +- Add typed multimodal parts and optional media capabilities. +- Add capability/provenance/freshness metadata to the catalog. +- Add narrow model-call middleware with streaming-safe ordering and redaction. +- Add usage provenance, cost fields, cache status, and OTel GenAI span hierarchy. +- Add versioned provider-options/passthrough interfaces. + +**P2: ecosystem integrations** + +- Add optional embeddings, reranking, batch, moderation, and file interfaces only when a host need is demonstrated. +- Add protocol bridges outside the core contract where useful, such as an OpenAI-compatible proxy or host-owned MCP adapter. +- Publish conformance reports and compatibility badges per provider/model route. + +**No-go:** no Flux `Agent`, `Runner`, `GraphRuntime`, `SessionStore`, `Memory`, `ToolExecutor`, `PromptStore`, `EvalRunner`, `AdminUI`, or chat frontend in the stable host packages. The existing boundary and contract version should be the guardrail. [Flux boundary](https://github.com/GrayCodeAI/flux/blob/main/docs/architecture/HOST-ENGINE-BOUNDARY.md) (accessed 2026-09-24). + +### Inferences + +- The most important “capability donor” is not a particular framework; it is the union of small, explicit contracts: model profile, typed parts, event envelope, schema descriptor, usage provenance, and adapter manifest. +- Flux's existing `ProviderBlock`, `RawArguments`, route/attempt metadata, pull stream, and graph-as-data decisions are unusually aligned with the strongest OSS patterns and should be preserved rather than replaced. +- A conformance suite will expose more value to hosts than adding another provider-specific convenience API: hosts can trust the same semantics across 28+ gateways and future adapters. +- Middleware, tracing, caching, and structured output can be engine capabilities only if their policy knobs remain explicit and host-owned where they affect side effects, privacy, or user experience. +- The best long-term competitive position for Flux is “the narrow, auditable provider engine with excellent conformance and lossless normalization,” not “the biggest agent framework.” + +### Gaps + +- This research did not verify every provider's current wire-level behavior against live credentials. The conformance suite should turn the recommendations into executable facts. +- Exact model capability and pricing data are time-sensitive; `Source`, `LiveMetadata`, freshness, and provenance should be treated as first-class fields rather than static assumptions. +- The local Flux checkout may contain unreleased contract-v2 changes. Before implementation, reconcile these recommendations with the current branch's tests and any sibling Rho pin. +- A final legal review is required before copying code, schemas, generated bindings, or documentation from Apache/MIT projects; Dify, Open WebUI, LiteLLM enterprise, and Mastra enterprise have additional terms. +- Benchmark claims, download counts, star counts, and vendor “production-ready” labels should not be used as roadmap priorities without independent measurements. diff --git a/research_notes/Flux OSS landscape roadmap/selection_methodology.md b/research_notes/Flux OSS landscape roadmap/selection_methodology.md new file mode 100644 index 00000000..08ca0a25 --- /dev/null +++ b/research_notes/Flux OSS landscape roadmap/selection_methodology.md @@ -0,0 +1,328 @@ +# Flux OSS Landscape: Top-20 Selection Methodology as of 2026-09-24 + +## Selection universe and inclusion/exclusion criteria + +### Takeaway + +Use a bounded, architecture-first universe rather than a global stars ranking. The defensible comparison set mixes direct provider-runtime/gateway peers with carefully capped serving runtimes, protocol SDKs, and observability standards so that one category cannot dominate merely because it has more popular repositories. + +### Cited Findings + +- Flux's comparison target is a host-neutral provider runtime: the public README says Flux owns credentials, model resolution, provider transports, normalized streams, retry/fallback, usage, and provider telemetry, while a host owns product UX, agent orchestration, tools, permissions, and sessions. [Flux README](https://github.com/GrayCodeAI/flux#what-is-flux) +- The discovery universe was the union of four GitHub repository searches, current canonical redirects, and a seeded set of technically relevant serving/provider/runtime projects: + - [`"llm gateway" in:name,description stars:>100 archived:false`](https://github.com/search?q=%22llm+gateway%22+in%3Aname%2Cdescription+stars%3A%3E100+archived%3Afalse&type=repositories) + - [`"ai gateway" in:name,description stars:>100 archived:false`](https://github.com/search?q=%22ai+gateway%22+in%3Aname%2Cdescription+stars%3A%3E100+archived%3Afalse&type=repositories) + - [`"llm router" in:name,description stars:>100 archived:false`](https://github.com/search?q=%22llm+router%22+in%3Aname%2Cdescription+stars%3A%3E100+archived%3Afalse&type=repositories) + - [`llm observability in:name,description stars:>100 archived:false`](https://github.com/search?q=llm+observability+in%3Aname%2Cdescription+stars%3A%3E100+archived%3Afalse&type=repositories) + - Seed strata: high-throughput LLM serving, local model runtimes, official provider SDKs, Go-native provider clients/runtimes, and OpenTelemetry/LLMOps standards. This catches technically important low-star projects that keyword ranking can miss, such as Go and protocol-specific projects. +- Repository identity was canonicalized before scoring. Most importantly, the former `envoyproxy/ai-gateway` project is now [`theagentrouter/agent-router`](https://github.com/theagentrouter/agent-router); the official documentation says it is the same code and maintainers renamed from Envoy AI Gateway and retained the existing CRD/API/CLI names. [Agent Router documentation](https://theagentrouter.ai/docs/) +- Hard inclusion gates were: + 1. Public, substantive source code under a canonical GitHub repository. + 2. A non-archived repository with a default-branch commit no older than 180 days on the retrieval date. + 3. An OSI-approved license for the core runtime, with mixed enterprise or third-party subtrees explicitly disclosed. + 4. Direct evidence for at least one Flux-critical concern: provider portability, routing, protocol translation, streaming, structured output, tools, reliability, serving, observability, or developer experience. + 5. First-party architecture documentation, source, releases, or substantive engineering discussions sufficient to validate the classification. +- Hard exclusions were: + - Agent applications, chat UI products, RAG builders, and workflow products that make Flux look like an app rather than a provider engine. + - Model-weight repositories, training frameworks, and generic compute projects without meaningful provider-runtime behavior. + - Archived projects and projects whose default branch has been inactive for more than 180 days. + - Source-only or source-available products without an OSI-approved core repository. + - General API gateways whose LLM-specific behavior is too small a part of the project to serve as a close Flux comparison. + - Additional official SDKs that merely duplicate a protocol already represented in a more influential SDK, unless they add a distinct language or wire-format lesson. +- To prevent the portfolio from degenerating into a stars list, the final set is capped at **8 direct peers, 6 model-serving runtimes, 4 provider SDKs, and 2 adjacent donors**. These are portfolio constraints, not claims that lower-scoring projects are unimportant. +- License status is evaluated from repository text, not only GitHub's SPDX field. MIT, Apache-2.0, and BSD-3-Clause are OSI-approved; a repository that is only source-available is not treated as OSS. [OSI MIT](https://opensource.org/license/mit); [OSI Apache-2.0](https://opensource.org/license/apache-2-0); [OSI BSD-3-Clause](https://opensource.org/license/bsd-3-clause) +- Three selected repositories use mixed trees but retain an OSI-approved core: + - LiteLLM is MIT outside `enterprise/`; that directory has a separate license. [LiteLLM LICENSE](https://github.com/BerriAI/litellm/blob/main/LICENSE) + - Langfuse is MIT outside its enterprise directories; those directories have a separate license. [Langfuse LICENSE](https://github.com/langfuse/langfuse/blob/main/LICENSE) + - TensorRT-LLM's root project is Apache-2.0, but the license file also identifies bundled third-party components and a separately licensed LTX-2 subtree. [TensorRT-LLM LICENSE](https://github.com/NVIDIA/TensorRT-LLM/blob/main/LICENSE) +- No final selection is source-available-only. Every selected repository has an OSI-approved core; mixed subtrees are called out rather than hidden. + +### Inferences + +- The right comparison object is a **portfolio**, not a single leaderboard: gateway peers establish the product boundary, serving engines establish downstream runtime behavior, SDKs establish wire semantics and DX, and telemetry projects establish normalized evidence. +- Go and embeddability deserve explicit weight because Flux is a Go library intended for host ownership, but they should not erase broader ecosystem influence. Bifrost, AxonHub, GoModel, LocalAI, Ollama, and `go-openai` therefore gain materially from the Go/embeddability criterion. +- Current canonical names matter for a roadmap. A comparison document that still points only to `envoyproxy/ai-gateway` would miss the project's 2026 identity and governance move. + +### Gaps + +- This is a reproducible, bounded research universe, not an exhaustive census of every GitHub repository. GitHub search is rank-based, localized, and can omit low-star or unusually described projects. +- The 180-day activity gate is a maintenance screen, not a quality guarantee. It does not measure contributor concentration, review latency, dependency health, or bus factor. +- License conclusions are engineering-due-diligence notes, not legal advice. + +## Exact ranked top 20 and verified GitHub metadata + +### Takeaway + +The final set contains **8 direct peers, 6 model-serving runtimes, 4 provider SDKs, and 2 adjacent donors**. The ranking is Flux-specific: architectural fit and Go/embeddability matter more than raw stars, while popularity remains a capped 15-point influence signal. + +### Cited Findings + +- **Snapshot rule:** GitHub REST/GraphQL metadata was retrieved on **2026-09-24**. Stars and forks are a live snapshot and will drift. “Latest release” is the latest published non-draft release; where a project intentionally publishes rolling prerelease/build tags, that status is marked. “Default-branch tip” is the exact commit observed during retrieval. [GitHub REST repositories API](https://docs.github.com/en/rest/repos/repos#get-a-repository) +- **Raw commit-count rule:** the 90-day column counts commits on the default branch since `2026-06-25T00:00:00Z`. It is a maintenance signal only; bots, generated commits, and merge policy can make raw counts non-comparable. [GitHub commits API](https://docs.github.com/en/rest/commits/commits#list-commits) + +### Flux-specific ranked set + +Score vector: **F/C/I/M/G/O/E** = architectural fit / capability overlap / influence / maintenance / Go-and-embeddability / OSS-license quality / evidence quality. Each is scored from 0–5 and weighted in the next section. + +| Rank | Project | Class | Score | F/C/I/M/G/O/E | Why it belongs | +|---:|---|---|---:|---|---| +| 1 | [maximhq/bifrost](https://github.com/maximhq/bifrost) | Direct peer | **97.0** | 5/5/4/5/5/5/5 | Closest current Go comparison: a unified provider API plus Go SDK, automatic fallback, load balancing, semantic caching, MCP, observability, and provider-native integrations. [Official overview](https://docs.getbifrost.ai/overview) | +| 2 | [BerriAI/litellm](https://github.com/BerriAI/litellm) | Direct peer | **95.0** | 5/5/5/5/3/4/5 | The broadest canonical gateway benchmark: unified provider calls, auth/budgets, routing, retries/fallbacks, caching, logging, and an OpenAI-compatible gateway. Its Rust core plus Python SDK is not a Go/host-neutral match, which keeps it below Bifrost for Flux. [Request architecture](https://docs.litellm.ai/docs/proxy/architecture) | +| 3 | [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute) | Direct peer | **94.0** | 5/5/5/5/2/5/5 | Exceptionally active and influential in 2026; implements a local-first multi-provider API, model routing, retries/fallbacks, circuit breakers, caching, tools, streaming, and native protocol passthrough. Its broad CLI/dashboard/protocol surface is less library-like than Flux. [Architecture index](https://github.com/diegosouzapw/OmniRoute/blob/main/docs/README.md); [OpenAPI](https://github.com/diegosouzapw/OmniRoute/blob/main/docs/openapi.yaml) | +| 4 | [looplj/axonhub](https://github.com/looplj/axonhub) | Direct peer | **91.5** | 5/4.5/3/5/5/5/4.5 | A Go gateway with inbound/outbound transformer layers, unified internal requests, multi-channel load balancing, failover, tracing, cost, and native OpenAI/Anthropic/Gemini surfaces. [Repository README](https://github.com/looplj/axonhub) | +| 5 | [ENTERPILOT/GoModel](https://github.com/ENTERPILOT/GoModel) | Direct peer | **90.0** | 5/4.5/2.5/5/5/5/4.5 | A compact Go runtime/gateway with OpenAI- and Anthropic-compatible APIs, model aliases, provider passthrough, retries/circuit breakers, scoped policies, caching, budgets, audit, and usage. [Official documentation](https://gomodel.enterpilot.io/) | +| 6 | [mudler/LocalAI](https://github.com/mudler/LocalAI) | Model-serving runtime | **85.5** | 3.5/4/4.5/5/5/5/5 | The strongest Go local-runtime donor: a swappable-backend inference engine exposing OpenAI, Anthropic, Ollama, and other compatible APIs, with streaming, tools, structured output, and backend lifecycle management. [Architecture](https://localai.io/docs/reference/architecture/index.html) | +| 7 | [theagentrouter/agent-router](https://github.com/theagentrouter/agent-router) | Direct peer | **84.5** | 4.5/4/3/4.5/4.5/5/5 | The canonical successor to Envoy AI Gateway: a Go control plane and Envoy data plane for unified schemas, provider fallback, token-aware limits, model virtualization, and MCP routing. [Capabilities](https://theagentrouter.ai/docs/capabilities/); [API reference](https://theagentrouter.ai/docs/api/) | +| 8 | [mnfst/llm-gateway](https://github.com/mnfst/llm-gateway) | Direct peer | **83.5** | 4.5/4/4/5/2/5/4.5 | Manifest is an active, influential routing gateway with OpenAI-compatible translation, local/cloud provider mixing, fallback, cost tracking, and a documented multidimensional routing algorithm. [Routing documentation](https://mnfst-manifest.mintlify.app/concepts/routing) | +| 9 | [ollama/ollama](https://github.com/ollama/ollama) | Model-serving runtime | **82.0** | 3/3.5/5/5/5/5/5 | A Go local-model runtime with model lifecycle management and OpenAI-compatible streaming, tools, JSON mode, vision, and Responses API support. [OpenAI compatibility](https://docs.ollama.com/api/openai-compatibility) | +| 10 | [Portkey-AI/gateway](https://github.com/Portkey-AI/gateway) | Direct peer | **81.5** | 4.5/5/4/3/2/5/4.5 | Still a highly influential direct benchmark for universal APIs, routing, retries, fallbacks, caching, circuit breakers, budgets, and observability. It remains in the set on influence and architectural fit, but its GitHub activity is a maintenance warning. [AI Gateway documentation](https://docs.portkey.ai/docs/product/ai-gateway) | +| 11 | [sgl-project/sglang](https://github.com/sgl-project/sglang) | Model-serving runtime | **74.5** | 2.5/4/4.5/5/2.5/5/5 | A high-performance serving framework whose model gateway adds worker discovery, cache-aware routing, circuit breakers, retries, rate limiting, health checks, OpenAI proxying, and tool execution. [Model Gateway architecture](https://github.com/sgl-project/sglang/blob/main/docs/advanced_features/sgl_model_gateway.md) | +| 12 | [vllm-project/vllm](https://github.com/vllm-project/vllm) | Model-serving runtime | **74.0** | 2.5/3.5/5/5/2.5/5/5 | The highest-volume serving donor for API-server separation, asynchronous streaming, scheduling, KV-cache behavior, OpenAI compatibility, and production process topology. [Architecture overview](https://github.com/vllm-project/vllm/blob/main/docs/design/arch_overview.md) | +| 13 | [sashabaranov/go-openai](https://github.com/sashabaranov/go-openai) | Provider SDK | **73.0** | 3/3/4/4/5/5/4 | The influential community Go reference for provider transport, typed requests, streaming, tools, Responses API migration, and compatibility-oriented DX. [README](https://github.com/sashabaranov/go-openai) | +| 14 | [langfuse/langfuse](https://github.com/langfuse/langfuse) | Adjacent capability donor | **71.5** | 2.5/3.5/4.5/5/2.5/4/5 | The strongest product-level observability/evaluation donor for causal traces, token/cost/latency data, sessions, scores, and OpenTelemetry-compatible ingestion. [Observability overview](https://langfuse.com/docs/observability/overview) | +| 15 | [ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp) | Model-serving runtime | **69.0** | 2/3/5/5/2.5/5/5 | The low-level local inference donor for constrained JSON, tool use, parallel decoding, continuous batching, multimodal serving, and OpenAI/Anthropic-compatible server routes. [Server documentation](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) | +| 16 | [openai/openai-python](https://github.com/openai/openai-python) | Provider SDK | **66.0** | 2/3/5/5/1/5/5 | The canonical OpenAI wire/DX donor for typed SSE streams, sync/async parity, response iteration, tool events, structured request types, and raw streaming-response control. [Repository README](https://github.com/openai/openai-python); [streaming guide](https://developers.openai.com/api/docs/guides/streaming-responses) | +| 17 | [NVIDIA/TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM) | Model-serving runtime | **65.0** | 2/3/4/5/2.5/4/5 | A high-performance NVIDIA serving donor with an OpenAI-compatible online server, health/metrics endpoints, batching, distributed execution, and multimodal serving. [Quick start](https://nvidia.github.io/TensorRT-LLM/quick-start-guide.html) | +| 18 | [anthropics/anthropic-sdk-python](https://github.com/anthropics/anthropic-sdk-python) | Provider SDK | **62.0** | 2/3.5/3/5/1/5/5 | The canonical Anthropic Messages protocol donor for typed SSE events, streaming accumulation helpers, partial JSON, final-message reconstruction, sync/async parity, and cloud-provider variants. [Python SDK documentation](https://platform.claude.com/docs/en/cli-sdks-libraries/sdks/python); [streaming helpers](https://github.com/anthropics/anthropic-sdk-python/blob/main/helpers.md) | +| 19 | [googleapis/python-genai](https://github.com/googleapis/python-genai) | Provider SDK | **61.5** | 2/3.5/3/5/1/5/4.5 | A distinct Google protocol donor for Gemini/Vertex request modeling, streaming, tools, multimodal parts, and generated SDK consistency. [Repository README](https://github.com/googleapis/python-genai) | +| 20 | [open-telemetry/semantic-conventions-genai](https://github.com/open-telemetry/semantic-conventions-genai) | Adjacent capability donor | **58.0** | 1.5/3/2.5/4.5/3/5/5 | The standards donor for portable `gen_ai.*` spans, token usage, streaming flags, finish reasons, model/provider identity, privacy-sensitive content capture, agent/tool conventions, and provider-specific extensions. [GenAI conventions](https://github.com/open-telemetry/semantic-conventions-genai); [GenAI spans](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md) | + +### Verified metadata snapshot + +| Project | Stars / forks | Default-branch commits, 90d | Latest release observed | Exact default-branch tip | License status | Primary language | +|---|---:|---:|---|---|---|---| +| [maximhq/bifrost](https://github.com/maximhq/bifrost) | [8,286 / 1,272](https://api.github.com/repos/maximhq/bifrost) | [2,031](https://api.github.com/repos/maximhq/bifrost/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [helm-chart-v2.1.43, 2026-09-23](https://github.com/maximhq/bifrost/releases/tag/helm-chart-v2.1.43) | [`ebd194d`, 2026-09-23](https://github.com/maximhq/bifrost/commit/ebd194db94b1d6c2476a058826d7670ecfa3102b) | [Apache-2.0, OSI](https://api.github.com/repos/maximhq/bifrost/license) | Go | +| [BerriAI/litellm](https://github.com/BerriAI/litellm) | [59,489 / 11,715](https://api.github.com/repos/BerriAI/litellm) | [12,971](https://api.github.com/repos/BerriAI/litellm/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v1.99.3, 2026-09-23](https://github.com/BerriAI/litellm/releases/tag/v1.99.3) | [`5beac4f`, 2026-09-23](https://github.com/BerriAI/litellm/commit/5beac4f18d83ee763ccb7de1216f50cff3a4b1de) | [MIT core, OSI; enterprise separate](https://github.com/BerriAI/litellm/blob/main/LICENSE) | Python; repository advertises a Rust core and Python SDK | +| [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute) | [69,585 / 9,876](https://api.github.com/repos/diegosouzapw/OmniRoute) | [4,590](https://api.github.com/repos/diegosouzapw/OmniRoute/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v3.8.50, 2026-08-26](https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.50) | [`18bbb10`, 2026-09-23](https://github.com/diegosouzapw/OmniRoute/commit/18bbb101980c639cfabdc5663213836d917c3269) | [MIT, OSI](https://api.github.com/repos/diegosouzapw/OmniRoute/license) | TypeScript | +| [looplj/axonhub](https://github.com/looplj/axonhub) | [5,289 / 714](https://api.github.com/repos/looplj/axonhub) | [346](https://api.github.com/repos/looplj/axonhub/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v1.0.0-beta10, 2026-09-06](https://github.com/looplj/axonhub/releases/tag/v1.0.0-beta10) | [`8094707`, 2026-09-23](https://github.com/looplj/axonhub/commit/809470775720976864a299f6d7d44cf464ccaa18) | [Apache-2.0, OSI](https://api.github.com/repos/looplj/axonhub/license) | Go | +| [ENTERPILOT/GoModel](https://github.com/ENTERPILOT/GoModel) | [1,181 / 104](https://api.github.com/repos/ENTERPILOT/GoModel) | [527](https://api.github.com/repos/ENTERPILOT/GoModel/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v0.1.96, 2026-09-22](https://github.com/ENTERPILOT/GoModel/releases/tag/v0.1.96) | [`d6f8924`, 2026-09-23](https://github.com/ENTERPILOT/GoModel/commit/d6f8924a59ac8f85759ae9db31c18e41d795c0d3) | [MIT, OSI](https://api.github.com/repos/ENTERPILOT/GoModel/license) | Go | +| [mudler/LocalAI](https://github.com/mudler/LocalAI) | [49,242 / 4,464](https://api.github.com/repos/mudler/LocalAI) | [1,279](https://api.github.com/repos/mudler/LocalAI/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v4.10.0, 2026-09-17](https://github.com/mudler/LocalAI/releases/tag/v4.10.0) | [`9ad18c8`, 2026-09-23](https://github.com/mudler/LocalAI/commit/9ad18c8c674671678f315ec901abab55886f97de) | [MIT, OSI](https://api.github.com/repos/mudler/LocalAI/license) | Go | +| [theagentrouter/agent-router](https://github.com/theagentrouter/agent-router) | [2,131 / 386](https://api.github.com/repos/theagentrouter/agent-router) | [166](https://api.github.com/repos/theagentrouter/agent-router/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v1.1.0, 2026-08-21](https://github.com/theagentrouter/agent-router/releases/tag/v1.1.0) | [`7d7c07f`, 2026-09-21](https://github.com/theagentrouter/agent-router/commit/7d7c07ffdb14241362def2f7d3f76cbd06d518cc) | [Apache-2.0, OSI](https://api.github.com/repos/theagentrouter/agent-router/license) | Go | +| [mnfst/llm-gateway](https://github.com/mnfst/llm-gateway) | [7,540 / 508](https://api.github.com/repos/mnfst/llm-gateway) | [823](https://api.github.com/repos/mnfst/llm-gateway/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [manifest@6.26.0, 2026-09-23](https://github.com/mnfst/llm-gateway/releases/tag/manifest%406.26.0) | [`3304b52`, 2026-09-23](https://github.com/mnfst/llm-gateway/commit/3304b52c3e2dac94b83c29b229730310bfa81dea) | [MIT, OSI](https://api.github.com/repos/mnfst/llm-gateway/license) | TypeScript | +| [ollama/ollama](https://github.com/ollama/ollama) | [181,528 / 17,982](https://api.github.com/repos/ollama/ollama) | [301](https://api.github.com/repos/ollama/ollama/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v0.34.3, 2026-09-19](https://github.com/ollama/ollama/releases/tag/v0.34.3) | [`b9cd4b1`, 2026-09-23](https://github.com/ollama/ollama/commit/b9cd4b1efbbf7bd1b9d202baafc64f0ebe63970e) | [MIT, OSI](https://api.github.com/repos/ollama/ollama/license) | Go | +| [Portkey-AI/gateway](https://github.com/Portkey-AI/gateway) | [13,069 / 1,310](https://api.github.com/repos/Portkey-AI/gateway) | [0](https://api.github.com/repos/Portkey-AI/gateway/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v1.15.2, 2026-01-12](https://github.com/Portkey-AI/gateway/releases/tag/v1.15.2) | [`669825c`, 2026-05-25](https://github.com/Portkey-AI/gateway/commit/669825cbe89ee51569918b8f78a9db486fd69dd4) | [MIT, OSI](https://api.github.com/repos/Portkey-AI/gateway/license) | TypeScript | +| [sgl-project/sglang](https://github.com/sgl-project/sglang) | [36,379 / 9,099](https://api.github.com/repos/sgl-project/sglang) | [4,306](https://api.github.com/repos/sgl-project/sglang/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v0.5.20, 2026-09-18](https://github.com/sgl-project/sglang/releases/tag/v0.5.20) | [`9544585`, 2026-09-23](https://github.com/sgl-project/sglang/commit/954458567e10257f1d8d5ff808d2dbc74fc26967) | [Apache-2.0, OSI](https://api.github.com/repos/sgl-project/sglang/license) | Python | +| [vllm-project/vllm](https://github.com/vllm-project/vllm) | [92,533 / 22,575](https://api.github.com/repos/vllm-project/vllm) | [3,810](https://api.github.com/repos/vllm-project/vllm/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v0.30.0, 2026-09-22](https://github.com/vllm-project/vllm/releases/tag/v0.30.0) | [`97b1b12`, 2026-09-23](https://github.com/vllm-project/vllm/commit/97b1b121176d405bcec861960dc605e9e41227e2) | [Apache-2.0, OSI](https://api.github.com/repos/vllm-project/vllm/license) | Python | +| [sashabaranov/go-openai](https://github.com/sashabaranov/go-openai) | [10,775 / 1,716](https://api.github.com/repos/sashabaranov/go-openai) | [12](https://api.github.com/repos/sashabaranov/go-openai/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v1.42.1, 2026-09-11](https://github.com/sashabaranov/go-openai/releases/tag/v1.42.1) | [`e0289b3`, 2026-09-22](https://github.com/sashabaranov/go-openai/commit/e0289b36c9611ec380b900490c3007078eb936b8) | [Apache-2.0, OSI](https://api.github.com/repos/sashabaranov/go-openai/license) | Go | +| [langfuse/langfuse](https://github.com/langfuse/langfuse) | [34,977 / 3,837](https://api.github.com/repos/langfuse/langfuse) | [2,030](https://api.github.com/repos/langfuse/langfuse/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v3.225.10, 2026-09-23](https://github.com/langfuse/langfuse/releases/tag/v3.225.10) | [`7ad9df`, 2026-09-23](https://github.com/langfuse/langfuse/commit/7ad9dfed444102f4ae678b7661d5c4ba9f934428) | [MIT core, OSI; enterprise separate](https://github.com/langfuse/langfuse/blob/main/LICENSE) | TypeScript | +| [ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp) | [129,334 / 23,648](https://api.github.com/repos/ggml-org/llama.cpp) | [1,365](https://api.github.com/repos/ggml-org/llama.cpp/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [b11147, 2026-09-23, rolling prerelease build](https://github.com/ggml-org/llama.cpp/releases/tag/b11147) | [`d2e5458`, 2026-09-23](https://github.com/ggml-org/llama.cpp/commit/d2e54583c7452353eb35d40431281f6ee984332f) | [MIT, OSI](https://api.github.com/repos/ggml-org/llama.cpp/license) | C++ | +| [openai/openai-python](https://github.com/openai/openai-python) | [31,679 / 5,957](https://api.github.com/repos/openai/openai-python) | [232](https://api.github.com/repos/openai/openai-python/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v3.19.1, 2026-09-23](https://github.com/openai/openai-python/releases/tag/v3.19.1) | [`4d12746`, 2026-09-23](https://github.com/openai/openai-python/commit/4d1274681db7f34f671ff9e71adc1461d22a927b) | [Apache-2.0, OSI](https://api.github.com/repos/openai/openai-python/license) | Python | +| [NVIDIA/TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM) | [14,703 / 2,771](https://api.github.com/repos/NVIDIA/TensorRT-LLM) | [2,497](https://api.github.com/repos/NVIDIA/TensorRT-LLM/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v1.3.0rc28, 2026-09-23, prerelease](https://github.com/NVIDIA/TensorRT-LLM/releases/tag/v1.3.0rc28) | [`f5ca105`, 2026-09-23](https://github.com/NVIDIA/TensorRT-LLM/commit/f5ca10543fc45cf3864703b31439f75070ccb2cc) | [Apache-2.0 core, OSI; mixed third-party notices](https://github.com/NVIDIA/TensorRT-LLM/blob/main/LICENSE) | Python | +| [anthropics/anthropic-sdk-python](https://github.com/anthropics/anthropic-sdk-python) | [3,915 / 862](https://api.github.com/repos/anthropics/anthropic-sdk-python) | [259](https://api.github.com/repos/anthropics/anthropic-sdk-python/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v1.8.0, 2026-09-22](https://github.com/anthropics/anthropic-sdk-python/releases/tag/v1.8.0) | [`4421d56`, 2026-09-22](https://github.com/anthropics/anthropic-sdk-python/commit/4421d56a4dd23550c7097c9b7ab5668bd11e09c4) | [MIT, OSI](https://api.github.com/repos/anthropics/anthropic-sdk-python/license) | Python | +| [googleapis/python-genai](https://github.com/googleapis/python-genai) | [3,989 / 1,018](https://api.github.com/repos/googleapis/python-genai) | [181](https://api.github.com/repos/googleapis/python-genai/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | [v2.25.0, 2026-09-22](https://github.com/googleapis/python-genai/releases/tag/v2.25.0) | [`4742c9a`, 2026-09-23](https://github.com/googleapis/python-genai/commit/4742c9a5c213a587add126a500a271824e2f0add) | [Apache-2.0, OSI](https://api.github.com/repos/googleapis/python-genai/license) | Python | +| [open-telemetry/semantic-conventions-genai](https://github.com/open-telemetry/semantic-conventions-genai) | [388 / 105](https://api.github.com/repos/open-telemetry/semantic-conventions-genai) | [82](https://api.github.com/repos/open-telemetry/semantic-conventions-genai/commits?since=2026-06-25T00%3A00%3A00Z&per_page=1) | No GitHub release observed | [`8ffdf56`, 2026-09-22](https://github.com/open-telemetry/semantic-conventions-genai/commit/8ffdf568e1b4391a99adb081db16e8102e36918e) | [Apache-2.0, OSI](https://api.github.com/repos/open-telemetry/semantic-conventions-genai/license) | Python; standards span generated/reference languages | + +### Metadata interpretation + +- The ranking is demonstrably not stars-only: OmniRoute has more stars than LiteLLM, while Bifrost ranks first because its Go implementation, SDK, provider abstraction, and current activity align more closely with Flux. [Bifrost repository](https://github.com/maximhq/bifrost); [OmniRoute repository](https://github.com/diegosouzapw/OmniRoute) +- Portkey is a **maintenance-watch selection**. Its direct architectural relevance and 13,069 stars justify inclusion, but the exact default-branch tip is 2026-05-25 and the 90-day count since 2026-06-25 is zero. [GitHub API](https://api.github.com/repos/Portkey-AI/gateway) +- Bifrost's release feed is component-scoped: the newest release is a Helm chart while gateway/transport components also released that day. Its exact current default-branch commit is the stronger maintenance signal. [Bifrost releases](https://github.com/maximhq/bifrost/releases) +- llama.cpp and TensorRT-LLM intentionally expose rolling prerelease tags, so their latest build/release candidate and same-day default-branch commits are more meaningful than a conventional stable-version cadence. [llama.cpp releases](https://github.com/ggml-org/llama.cpp/releases); [TensorRT-LLM releases](https://github.com/NVIDIA/TensorRT-LLM/releases) +- OpenTelemetry's GenAI conventions repository is new enough to have no observed GitHub release, but it is active and is the domain-specific successor target for Flux's `gen_ai.*` telemetry contract. [Repository README](https://github.com/open-telemetry/semantic-conventions-genai) + +### Inferences + +- The set captures both current 2026 entrants and established standards without letting raw popularity decide the first rank. +- Current maintenance is strong for 19 projects. Portkey is retained as a historically and architecturally important benchmark but should not be treated as evidence of current implementation velocity. +- Release tags alone are insufficient: projects with continuous build tags, monorepo component tags, or no first release require the exact commit signal. + +### Gaps + +- GitHub star/fork values are live counters, not archival facts. The values above are a 2026-09-24 snapshot. +- A repository-level release does not prove that every listed feature is production-ready. Beta and prerelease status is called out where observed. +- This metadata pass did not audit generated SDK provenance, dependency vulnerabilities, build reproducibility, or the proportion of commits produced by bots. + +## Direct peers versus strategically useful donors + +### Takeaway + +Only the eight gateway/runtime projects are true product-boundary peers. The other twelve are strategically useful donors: six teach serving behavior, four teach protocol and SDK behavior, and two teach evidence and telemetry behavior. + +### Cited Findings + +#### Classification definitions + +- **Direct peer:** exposes or implements a multi-provider runtime/gateway boundary comparable to Flux: provider translation, model resolution, routing, streaming, retries/fallbacks, and/or operational controls. It need not share Flux's library-first deployment model. +- **Model-serving runtime:** primarily executes models or manages inference workers. It is a downstream provider target and a donor for stream lifecycle, scheduling, batching, cancellation, health, and model lifecycle—not a direct substitute for Flux. +- **Provider SDK:** owns one provider's protocol and developer experience. It is a donor for wire fidelity, typed streams, errors, tools, structured output, and compatibility behavior, not for universal routing. +- **Adjacent capability donor:** contributes a cross-cutting capability Flux consumes or should interoperate with, especially traces, evaluations, or stable telemetry vocabulary. + +#### Eight direct peers + +| Project | Strategic role versus Flux | +|---|---| +| [Bifrost](https://github.com/maximhq/bifrost) | The closest current implementation peer: Go, provider abstraction, direct SDK plus gateway modes, routing, failover, semantic cache, MCP, and telemetry. [Official overview](https://docs.getbifrost.ai/overview) | +| [LiteLLM](https://github.com/BerriAI/litellm) | The broadest feature and ecosystem benchmark; compare router semantics, provider transforms, cost/spend, guardrails, and the SDK/proxy split. Its hybrid Python/Rust architecture is a deliberate contrast with Flux's Go host-facing facade. [Architecture](https://docs.litellm.ai/docs/proxy/architecture) | +| [OmniRoute](https://github.com/diegosouzapw/OmniRoute) | A high-velocity, high-influence gateway benchmark for provider fallback, model aliases, local/cloud mixing, native passthrough, and resilience. It is broader and more product/CLI-heavy than Flux. [Architecture index](https://github.com/diegosouzapw/OmniRoute/blob/main/docs/README.md) | +| [AxonHub](https://github.com/looplj/axonhub) | The clearest transformer-pipeline peer for inbound dialect → unified request → outbound provider → reverse stream/error transform. [Repository README](https://github.com/looplj/axonhub) | +| [GoModel](https://github.com/ENTERPILOT/GoModel) | A compact Go peer for scoped workflow policies, model aliases, provider passthrough, cache/budget/rate-limit composition, and embedded operational DX. [Official documentation](https://gomodel.enterpilot.io/) | +| [Agent Router](https://github.com/theagentrouter/agent-router) | A Go/Kubernetes/Envoy peer for declarative provider failover, request/response mutation, token-aware policy, model virtualization, and MCP routing. It is infrastructure-heavy rather than embeddable. [Capabilities](https://theagentrouter.ai/docs/capabilities/) | +| [Manifest (`mnfst/llm-gateway`)](https://github.com/mnfst/llm-gateway) | A routing-specific peer for prompt-complexity scoring, model-tier resolution, local/cloud provider mixing, sticky/session behavior, and cost-aware fallback. [Routing documentation](https://mnfst-manifest.mintlify.app/concepts/routing) | +| [Portkey Gateway](https://github.com/Portkey-AI/gateway) | The canonical feature checklist for universal API, cache, routing, retries, circuit breaker, load balancing, budgets, and canaries; current maintenance requires caution. [AI Gateway documentation](https://docs.portkey.ai/docs/product/ai-gateway) | + +#### Six model-serving runtime donors + +- [LocalAI](https://github.com/mudler/LocalAI) is the most relevant Go serving donor because it combines a stable API shim with swappable inference backends. [Architecture](https://localai.io/docs/reference/architecture/index.html) +- [Ollama](https://github.com/ollama/ollama) is the local model lifecycle and OpenAI-compatibility donor. Its official matrix shows streaming, tools, JSON mode, vision, and non-stateful Responses API support. [Compatibility matrix](https://docs.ollama.com/api/openai-compatibility) +- [SGLang](https://github.com/sgl-project/sglang) contributes a sophisticated model gateway and cache-aware worker-routing design in addition to its serving engine. [Model Gateway](https://github.com/sgl-project/sglang/blob/main/docs/advanced_features/sgl_model_gateway.md) +- [vLLM](https://github.com/vllm-project/vllm) contributes the clearest separation of API server, scheduler/KV cache, and GPU workers, making it valuable for lifecycle, throughput, cancellation, and streaming analysis. [Architecture](https://github.com/vllm-project/vllm/blob/main/docs/design/arch_overview.md) +- [llama.cpp](https://github.com/ggml-org/llama.cpp) contributes low-level structured output, tool use, parallel decoding, continuous batching, and protocol-compatible server behavior. [Server README](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) +- [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM) contributes production NVIDIA serving behavior, in-flight batching, distributed execution, OpenAI-compatible endpoints, health, and metrics. [Quick start](https://nvidia.github.io/TensorRT-LLM/quick-start-guide.html) + +#### Four provider SDK donors + +- [`sashabaranov/go-openai`](https://github.com/sashabaranov/go-openai) is the primary Go DX and community-compatibility donor. +- [`openai/openai-python`](https://github.com/openai/openai-python) is the canonical typed-SSE and Responses/Chat Completions wire donor. +- [`anthropics/anthropic-sdk-python`](https://github.com/anthropics/anthropic-sdk-python) is the canonical Anthropic event/accumulation protocol donor. +- [`googleapis/python-genai`](https://github.com/googleapis/python-genai) is the distinct Gemini/Vertex typed-content and multimodal protocol donor. + +#### Two adjacent capability donors + +- [Langfuse](https://langfuse/langfuse) is the product-level observability and evaluation donor. Its trace model records prompts, responses, tool/retrieval steps, token usage, latency, sessions, costs, and scores, and it accepts OpenTelemetry data. [Observability documentation](https://langfuse.com/docs/observability/overview) +- [OpenTelemetry GenAI semantic conventions](https://github.com/open-telemetry/semantic-conventions-genai) is the interoperability donor. It defines model/provider identity, operation names, token usage, streaming flags, finish reasons, content, agent/tool conventions, and provider extensions. [GenAI span specification](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md) + +### Comparison axes for using the set + +The comparison should be organized around Flux's actual boundaries rather than README feature counts: + +1. **Provider and wire portability:** inbound dialects, outbound transforms, native passthrough, provider-specific pass-through, custom endpoints, and lossless unsupported-feature behavior. +2. **Model resolution:** aliases, logical tiers, live catalogs, capability metadata, context limits, pricing, deprecation, and model-role slots. +3. **Streaming lifecycle:** typed semantic events, reasoning deltas, tool argument deltas, partial structured output, cancellation, closure, continuation, and backpressure. +4. **Reliability:** retry classification, backoff/`Retry-After`, circuit breaking, provider/model fallback, health, rate limits, queueing, and sticky sessions. +5. **Operational controls:** usage/cost attribution, budgets, audit privacy, caching policy, observability, readiness, and OpenTelemetry compatibility. +6. **Host neutrality:** stable facade, internal provider composition, Go package boundaries, optional proxy/gRPC surfaces, and avoidance of agent/product semantics. + +These axes follow Flux's public boundaries and architecture. [Flux README](https://github.com/GrayCodeAI/flux#ecosystem-boundaries); [Flux architecture](https://github.com/GrayCodeAI/flux/blob/main/docs/ARCHITECTURE.md) + +### Inferences + +- There is no exact open-source twin of Flux: no selected project combines Flux's Go 1.26 library-first host facade, internal provider registry, host-neutral engine boundary, and full routing/reliability stack in the same shape. +- Bifrost, AxonHub, and GoModel are the highest-priority implementation peers. LiteLLM is the highest-priority feature and ecosystem benchmark. OmniRoute, Manifest, Portkey, and Agent Router should be compared selectively by subsystem rather than treated as identical products. +- Serving engines and provider SDKs should not inflate a competitive feature total for Flux; they are reference implementations for specific contracts. + +### Gaps + +- Official documentation establishes intended behavior, not full wire compatibility. A provider-by-provider, field-by-field compatibility audit is still required. +- Several gateways publish benchmark claims, but no independent, reproducible benchmark was run for this selection. +- Classification is based on current repository architecture and public docs; projects can change category after future releases. + +## Scoring, ranking, and plausible alternatives + +### Takeaway + +The ranking is a transparent decision model, not an objective universal truth. Architecture fit (30), capability overlap (20), influence (15), maintenance (15), Go/embeddability (10), OSS quality (5), and evidence quality (5) produce a Flux-specific order in which stars can inform but never decide the result. + +### Cited Findings + +#### Formula and weights + +For each project: + +`Score = 6F + 4C + 3I + 3M + 2G + O + E` + +Each component is scored from 0 to 5 in half-point increments. + +| Component | Weight | 5-point meaning | +|---|---:|---| +| **F — Flux architecture fit** | 30 | Universal provider abstraction and a host-neutral runtime boundary closely matching Flux's purpose | +| **C — Capability overlap** | 20 | Strong coverage of routing, protocol translation, streams, retries/fallbacks, caching, tools/structured output, and operations | +| **I — Influence** | 15 | Strong adoption and ecosystem reach, anchored by stars/forks plus integrations, governance, release cadence, and production relevance; never stars alone | +| **M — Maintenance** | 15 | Current commit, recent release, visible engineering cadence, active issues/PRs, and credible ownership | +| **G — Go and embeddability** | 10 | Go implementation and/or a reusable library/facade that hosts can compose without adopting a large control plane | +| **O — OSS/legal quality** | 5 | OSI-approved core with clear, stable licensing and minimal mixed-license ambiguity | +| **E — evidence quality** | 5 | First-party architecture docs, source, release notes, tests, and substantive engineering discussions | + +- Influence receives only 15% of the score. The influence rubric uses broad bands—50k+ stars plus major ecosystem use; 10k–50k; 2k–10k; 500–2k; 100–500—then adjusts for governance, integration, and production evidence. It is not a direct sort by stars. Current metadata is available in the project API snapshots above. +- Category caps are applied after eligibility and scoring: direct peers ≤8, model-serving runtimes ≤6, provider SDKs ≤4, adjacent donors ≤2. This preserves comparison coverage without allowing highly starred serving repositories to crowd out the actual product boundary. +- A project can be technically strong and still be excluded because it is archived, inactive, source-available-only, not publicly available at the queried canonical repository, or redundant under a category cap. + +#### Strong near misses and specific exclusion reasons + +| Alternative | Verified 2026-09-24 signal | Why it is not in the final 20 | +|---|---|---| +| [theopenco/llmgateway](https://github.com/theopenco/llmgateway) | 1,658 stars / 190 forks; active tip 2026-09-23; [v1.18.0 on 2026-09-21](https://github.com/theopenco/llmgateway/releases/tag/v1.18.0); TypeScript; AGPL-3.0 core. [API](https://api.github.com/repos/theopenco/llmgateway) | Very plausible direct peer, but the eight-peer cap is already occupied by stronger influence/fit combinations. AGPL is OSI-approved and was not the exclusion reason. [LICENSE](https://github.com/theopenco/llmgateway/blob/main/LICENSE); [OSI AGPL-3.0](https://opensource.org/license/agpl-v3) | +| [smg-project/smg](https://github.com/smg-project/smg) | 543 / 172; active tip 2026-09-23; [v1.10.1 on 2026-08-27](https://github.com/smg-project/smg/releases/tag/v1.10.1); Rust; Apache-2.0. [API](https://api.github.com/repos/smg-project/smg) | Technically credible direct peer, but smaller adoption and no Go/embeddability advantage; displaced by the direct-peer cap. | +| [mozilla-ai/otari](https://github.com/mozilla-ai/otari) | 490 / 57; active tip and [v0.9.0 release](https://github.com/mozilla-ai/otari/releases/tag/v0.9.0) on 2026-09-23; Python; Apache-2.0. [API](https://api.github.com/repos/mozilla-ai/otari) | Active OSS direct peer, but lower adoption and weaker Go/host-library fit. | +| [llm-d/llm-d](https://github.com/llm-d/llm-d) | 4,640 / 788; active tip 2026-09-23; [v0.9.0 on 2026-08-17](https://github.com/llm-d/llm-d/releases/tag/v0.9.0); Shell-led Kubernetes orchestration; Apache-2.0. [API](https://api.github.com/repos/llm-d/llm-d) | Valuable distributed-inference/scheduling donor, but it orchestrates serving systems rather than implementing Flux's portable provider boundary; serving cap applies. | +| [kserve/kserve](https://github.com/kserve/kserve) | 5,991 / 1,700; active tip 2026-09-23; [v0.20.0 on 2026-08-06](https://github.com/kserve/kserve/releases/tag/v0.20.0); Go; Apache-2.0. [API](https://api.github.com/repos/kserve/kserve) | Influential and credible, but primarily a Kubernetes inference platform; too broad for a top-20 Flux runtime comparison after six stronger serving/runtime donors. | +| [triton-inference-server/server](https://github.com/triton-inference-server/server) | 11,003 / 1,839; active tip 2026-09-22; [v2.72.0 on 2026-08-31](https://github.com/triton-inference-server/server/releases/tag/v2.72.0); Python; BSD-3-Clause. [API](https://api.github.com/repos/triton-inference-server/server) | Excellent general inference-server donor, but less LLM/provider-portability-specific than vLLM/SGLang/LocalAI/Ollama/TensorRT and no Go advantage. | +| [open-telemetry/opentelemetry-go](https://github.com/open-telemetry/opentelemetry-go) | 6,558 / 1,482; active tip 2026-09-23; [v1.46.0 on 2026-08-25](https://github.com/open-telemetry/opentelemetry-go/releases/tag/v1.46.0); Go; Apache-2.0. [API](https://api.github.com/repos/open-telemetry/opentelemetry-go) | A foundational dependency rather than an LLM comparison target. The new GenAI conventions repository supplies the more strategic donor for Flux's provider telemetry contract. | +| [Helicone/ai-gateway](https://github.com/Helicone/ai-gateway) | 631 / 57; last tip 2025-11-21; Rust; GPL-3.0. [API](https://api.github.com/repos/Helicone/ai-gateway) | Fails the 180-day activity gate. The more active [Helicone platform](https://github.com/Helicone/helicone) is observability/product-heavy rather than a provider-runtime peer. | +| [lm-sys/RouteLLM](https://github.com/lm-sys/RouteLLM) | 5,537 / 435; last tip 2024-08-10; no observed release; Apache-2.0. [API](https://api.github.com/repos/lm-sys/RouteLLM) | Influential routing research/code, but inactive for more than two years and therefore fails the maintenance gate. | +| [huggingface/text-generation-inference](https://github.com/huggingface/text-generation-inference) | 10,884 / 1,291; repository archived; last tip 2026-03-21; [v3.3.7 on 2025-12-19](https://github.com/huggingface/text-generation-inference/releases/tag/v3.3.7). [API](https://api.github.com/repos/huggingface/text-generation-inference) | Archived, so it is excluded regardless of historical influence. | +| [coaidev/coai](https://github.com/coaidev/coai) | 9,318 / 1,224; last tip 2026-03-12; [v4.0.0 on 2025-10-23](https://github.com/coaidev/coai/releases/tag/v4.0.0); TypeScript; Apache-2.0. [API](https://api.github.com/repos/coaidev/coai) | A multi-tenant admin/billing/chat product as well as a gateway; broader product semantics and weaker current maintenance than selected peers. | +| [bentoml/BentoML](https://github.com/bentoml/BentoML) | 8,856 / 1,035; active tip 2026-09-07; [v1.4.39 on 2026-05-07](https://github.com/bentoml/BentoML/releases/tag/v1.4.39); Python; Apache-2.0. [API](https://api.github.com/repos/bentoml/BentoML) | Important serving/packaging donor, but more model-serving app framework than LLM provider runtime; serving cap applies. | + +#### Broad alternatives excluded by scope + +- General API/AI gateways such as [Kong](https://github.com/Kong/kong), [Apache APISIX](https://github.com/apache/apisix), [Higress](https://github.com/higress-group/higress), [kgateway](https://github.com/kgateway-dev/kgateway), and [Traefik](https://github.com/traefik/traefik) are credible routing/proxy donors, but their LLM-specific provider normalization is a smaller part of a much broader API gateway. [Higress API](https://api.github.com/repos/higress-group/higress); [kgateway API](https://api.github.com/repos/kgateway-dev/kgateway) +- Official Go SDKs [openai/openai-go](https://github.com/openai/openai-go) and [anthropics/anthropic-sdk-go](https://github.com/anthropics/anthropic-sdk-go) are credible, but their wire behavior duplicates the selected OpenAI/Anthropic SDKs while the selected `go-openai` adds stronger community adoption and Go DX evidence. [openai-go API](https://api.github.com/repos/openai/openai-go); [anthropic-sdk-go API](https://api.github.com/repos/anthropics/anthropic-sdk-go) +- Broad cloud SDKs such as [aws-sdk-go-v2](https://github.com/aws/aws-sdk-go-v2) and [azure-sdk-for-go](https://github.com/Azure/azure-sdk-for-go) are operationally important but too large to serve as focused LLM protocol/developer-experience comparisons. [AWS SDK API](https://api.github.com/repos/aws/aws-sdk-go-v2); [Azure SDK API](https://api.github.com/repos/Azure/azure-sdk-for-go) +- Agent/RAG/UI application frameworks were excluded because they compete with a host such as Rho, not with Flux's provider engine. This includes the architectural pattern represented by LangChain, LlamaIndex, Pydantic AI, Semantic Kernel, Dify, Open WebUI, and similar products. +- Source-available-only routers were excluded even when functionally relevant. For example, [NadirClaw](https://github.com/NadirRouter/NadirClaw) describes itself as source-available, so it fails the OSI-core gate. [Repository](https://github.com/NadirRouter/NadirClaw) +- The exact queried repositories `truefoundry/llm-gateway`, `lunary-ai/lunary`, `cloudflare/ai-gateway`, `kubernetes-sigs/llm-gateway`, and `humanloop/llm-router` did not resolve to public canonical repositories on the retrieval date. They therefore fail the public-repository evidence gate and are not presented as OSS comparison projects. [TrueFoundry lookup](https://api.github.com/repos/truefoundry/llm-gateway); [Lunary lookup](https://api.github.com/repos/lunary-ai/lunary); [Cloudflare lookup](https://api.github.com/repos/cloudflare/ai-gateway); [Kubernetes lookup](https://api.github.com/repos/kubernetes-sigs/llm-gateway); [Humanloop lookup](https://api.github.com/repos/humanloop/llm-router) + +### Inferences + +- The selected list should be periodically refreshed, but not automatically replaced by the newest stars. In particular, a future dormant or archived direct peer should fall behind active alternatives even if its historical star count is larger. +- Portkey is the clearest example of why maintenance must be visible: strong feature fit and influence earned inclusion, but its stale public repository prevents a higher maintenance score. +- Category caps are justified by comparison utility. A set containing 10 model servers and two gateways would answer “which inference engine is popular?” rather than “what should a universal Go provider runtime learn from?” + +### Gaps + +- Half-point scores involve judgment. The weights, anchors, caps, and live metadata make disagreements reviewable, but another reviewer could reasonably alter individual scores by 1–3 points. +- The study did not normalize misleading or rapidly changing GitHub metrics such as star velocity, contributor count, fork activity, or release-download counts. +- Excluded direct peers should be reconsidered if they add a capability absent from the selected set—for example, a formally specified lossless provider-passthrough contract or a proven host-neutral SDK boundary. + +## Strongest evidence and evidence limitations + +### Takeaway + +The most defensible evidence is triangulated: GitHub API and exact commits for current facts; repository license text for legal classification; first-party architecture docs for product boundaries; release feeds for maintenance; and engineering issues/RFCs for the hard problems that marketing pages omit. + +### Cited Findings + +#### Evidence hierarchy + +| Evidence tier | Strongest sources | What it establishes | +|---|---|---| +| Current repository facts | [GitHub repository API](https://docs.github.com/en/rest/repos/repos#get-a-repository), exact commit links, release links, and language/license fields in the metadata table | Stars, forks, archive status, primary language, repository activity, release cadence, and canonical identity as of 2026-09-24 | +| License facts | Repository license texts linked in the metadata table; [OSI license pages](https://opensource.org/licenses) | Which selected cores are genuinely OSI-approved and which subtrees are mixed or separately licensed | +| Direct-peer architecture | [LiteLLM request flow](https://docs.litellm.ai/docs/proxy/architecture), [Bifrost overview](https://docs.getbifrost.ai/overview), [Agent Router architecture/capabilities](https://theagentrouter.ai/docs/capabilities/), [OmniRoute architecture](https://github.com/diegosouzapw/OmniRoute/blob/main/docs/ARCHITECTURE.md), [AxonHub repository](https://github.com/looplj/axonhub), [GoModel documentation](https://gomodel.enterpilot.io/), [Portkey gateway](https://docs.portkey.ai/docs/product/ai-gateway), [Manifest routing](https://mnfst-manifest.mintlify.app/concepts/routing) | Provider normalization, routing, resilience, caching, operational control, and deployment-boundary differences | +| Serving behavior | [vLLM architecture](https://github.com/vllm-project/vllm/blob/main/docs/design/arch_overview.md), [SGLang Model Gateway](https://github.com/sgl-project/sglang/blob/main/docs/advanced_features/sgl_model_gateway.md), [Ollama compatibility matrix](https://docs.ollama.com/api/openai-compatibility), [llama.cpp server](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md), [LocalAI architecture](https://localai.io/docs/reference/architecture/index.html), [TensorRT-LLM quick start](https://nvidia.github.io/TensorRT-LLM/quick-start-guide.html) | Stream lifecycle, worker/model lifecycle, scheduling, batching, cancellation, health, tools, structured output, and API compatibility | +| Protocol and SDK behavior | [OpenAI streaming](https://developers.openai.com/api/docs/guides/streaming-responses), [Anthropic streaming helpers](https://github.com/anthropics/anthropic-sdk-python/blob/main/helpers.md), [Google GenAI repository](https://github.com/googleapis/python-genai), [`go-openai`](https://github.com/sashabaranov/go-openai) | Typed event semantics, partial accumulation, sync/async behavior, tools, raw stream control, and migration pressure | +| Observability and standards | [Langfuse observability](https://langfuse.com/docs/observability/overview), [OpenTelemetry GenAI spans](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-spans.md) | Trace data model, cost/latency/quality evidence, provider-neutral telemetry fields, streaming/finish semantics, and content privacy | +| Release evidence | Each exact release link in the metadata table | A real published artifact/tag, release date, prerelease status, and current release cadence | + +#### High-signal engineering discussions + +These are not treated as proof that a design is correct; they are evidence of where maintainers are wrestling with real interoperability and reliability problems. + +- **Provider protocol loss and passthrough:** Agent Router's [pass-through routing proposal](https://github.com/theagentrouter/agent-router/issues/948) and [native Anthropic API issue](https://github.com/theagentrouter/agent-router/issues/847) show the distinction between provider-native fidelity and lossy normalization. +- **Normalization edge cases:** LiteLLM's [reasoning-stream concatenation issue](https://github.com/BerriAI/litellm/issues/11000), [tool-call finish-reason issue](https://github.com/BerriAI/litellm/issues/12481), and [Ollama structured-output issue](https://github.com/BerriAI/litellm/issues/7131) demonstrate why semantic stream events cannot be treated as plain text concatenation. +- **Adaptive routing pressure:** Agent Router's [latency-aware routing issue](https://github.com/theagentrouter/agent-router/issues/812) captures the need to combine policy, health, latency, and model selection rather than use one static strategy. +- **Serving and scheduling pressure:** SGLang's [distributed KV-cache roadmap](https://github.com/sgl-project/sglang/issues/21846) and [tokenizer-to-scheduler RFC](https://github.com/sgl-project/sglang/issues/16787) show how queueing, cache locality, and handoff affect streaming latency and reliability. +- **API-surface pressure:** vLLM's [Responses API issue](https://github.com/vllm-project/vllm/issues/14721) and [multimodality RFC](https://github.com/vllm-project/vllm/issues/4194) document the difficulty of matching a fast-evolving client-facing API in a serving engine. +- **Telemetry evolution and privacy:** OpenTelemetry GenAI's [vendor-neutral reasoning discussion](https://github.com/open-telemetry/semantic-conventions-genai/issues/192) and [content-capture environment-variable proposal](https://github.com/open-telemetry/semantic-conventions-genai/issues/497) directly inform Flux's normalized telemetry and privacy projection. + +#### Strongest repository/source combinations by research question + +- **Closest implementation comparison:** Bifrost source plus official Go SDK docs; AxonHub's transformer package; GoModel's provider/workflow/retry source. These should be compared directly to Flux's `provider/core`, `provider/adapters`, `router`, `runtime`, and `engine` boundaries. [Flux architecture](https://github.com/GrayCodeAI/flux/tree/main/provider) +- **Feature breadth comparison:** LiteLLM's architecture and request flow, Portkey's gateway strategy docs, and OmniRoute's architecture/OpenAPI. [Flux features](https://github.com/GrayCodeAI/flux#features) +- **Protocol correctness:** official SDK streaming helpers and the selected serving servers' compatibility matrices, followed by source-level fixture comparisons for reasoning, tools, structured output, and finish reasons. +- **Reliability:** Agent Router/SGLang fallback and health designs, LiteLLM's retry/fallback flow, and the selected serving projects' scheduler/cancellation behavior. +- **Observability:** Flux's current operations graph plus OpenTelemetry GenAI conventions; Langfuse is the product/UX donor for causal trace inspection and evaluation. [Flux operations graph](https://github.com/GrayCodeAI/flux/tree/main/operationsgraph) + +### Inferences + +- Claims corroborated by source plus architecture docs are stronger than README feature bullets. Claims visible in issues are useful for identifying unresolved pressure points, not for declaring a project superior. +- Release and exact-commit evidence should be reviewed together. Rolling tags, monorepo component tags, and prerelease-only projects make any single “latest version” rule misleading. +- The selected set provides a balanced roadmap lens: direct peers reveal product gaps, serving runtimes reveal lifecycle and performance contracts, SDKs reveal wire/DX contracts, and observability donors reveal evidence contracts. + +### Gaps + +- No source-level feature-by-feature audit was performed for all 20 projects; the selection identifies what to compare, not a completed competitive matrix. +- No independent benchmark was run. Vendor statements such as Bifrost's overhead claims were not used as ranking evidence and should be independently reproduced before roadmap decisions. [Bifrost overview](https://docs.getbifrost.ai/overview) +- Maintenance scoring does not yet include issue-close latency, PR acceptance rate, release artifact downloads, maintainer concentration, security response, or dependency-update cadence. +- Some engineering proposals remain open or experimental. They identify direction and unresolved semantics, not shipped guarantees. +- License observations should be rechecked at implementation time, especially mixed enterprise/third-party trees and any future license changes. diff --git a/router/deployment_router.go b/router/deployment_router.go index bce1a0a3..e9f4accb 100644 --- a/router/deployment_router.go +++ b/router/deployment_router.go @@ -2,6 +2,7 @@ package router import ( "context" + "errors" "fmt" "math/rand/v2" "sort" @@ -137,6 +138,7 @@ func (r *DeploymentRouter) Chat(ctx context.Context, messages []core.FluxMessage return nil, err } var lastErr error + attemptsMade := 0 for stageIndex, stage := range r.routeFor(target.canonicalModelID) { choices := r.eligibleChoices(target, stage, opts) if len(choices) == 0 { @@ -157,13 +159,20 @@ func (r *DeploymentRouter) Chat(ctx context.Context, messages []core.FluxMessage lastErr = fmt.Errorf("stage %d has no available deployments", stageIndex) break } + attemptsMade++ resp, err := r.chatWithDeployment(ctx, messages, opts, target, choice.DeploymentID) if err == nil { r.recordSuccess(choice.DeploymentID) + attachResponseRoute(resp, deploymentRoute(opts, target, choice.DeploymentID, attemptsMade)) return resp, nil } lastErr = err - r.recordFailure(choice.DeploymentID) + if ctx.Err() != nil { + // The caller cancelled or its deadline passed: stop, and do not + // count the failure against the deployment. + return nil, err + } + r.recordFailure(choice.DeploymentID, err) if !IsTransient(err) { if ShouldTryNextDeployment(err) { break @@ -189,6 +198,8 @@ func (r *DeploymentRouter) StreamChat(ctx context.Context, messages []core.FluxM go func() { defer close(out) var lastErr error + var lastRoute *llm.ResolvedRoute + attemptsMade := 0 for stageIndex, stage := range r.routeFor(target.canonicalModelID) { choices := r.eligibleChoices(target, stage, opts) if len(choices) == 0 { @@ -209,28 +220,38 @@ func (r *DeploymentRouter) StreamChat(ctx context.Context, messages []core.FluxM lastErr = fmt.Errorf("stage %d has no available deployments", stageIndex) break } - fallback, err := r.streamWithDeployment(streamCtx, out, messages, opts, target, choice.DeploymentID) + attemptsMade++ + route := deploymentRoute(opts, target, choice.DeploymentID, attemptsMade) + lastRoute = route + if !sendRouterEvent(streamCtx, out, core.FluxStreamEvent{Type: "route_changed", Route: route}) { + return + } + fallback, err := r.streamWithDeployment(streamCtx, out, messages, opts, target, choice.DeploymentID, *route) if err == nil { r.recordSuccess(choice.DeploymentID) return } + if streamCtx.Err() != nil { + // The caller cancelled or its deadline passed. Only the + // caller's context decides this: an upstream timeout also + // matches context.DeadlineExceeded but must fail over. No + // failover and no breaker failure; the coordinated wrapper + // around this stream emits the terminal cancelled event. + return + } lastErr = err - r.recordFailure(choice.DeploymentID) + r.recordFailure(choice.DeploymentID, err) if !fallback { - select { - case out <- core.FluxStreamEvent{Type: "error", Error: err.Error()}: - case <-streamCtx.Done(): + if !streamFailureForwarded(err) { + sendRouterEvent(streamCtx, out, routerErrorEvent(err, route)) } return } - if !IsTransient(err) { + if !isRetryableStreamFailure(err) { if ShouldTryNextDeployment(err) { break } - select { - case out <- core.FluxStreamEvent{Type: "error", Error: err.Error()}: - case <-streamCtx.Done(): - } + sendRouterEvent(streamCtx, out, routerErrorEvent(err, route)) return } recentlyFailed = choice.DeploymentID @@ -239,12 +260,11 @@ func (r *DeploymentRouter) StreamChat(ctx context.Context, messages []core.FluxM if lastErr == nil { lastErr = fmt.Errorf("no route configured") } - select { - case out <- core.FluxStreamEvent{Type: "error", Error: fmt.Sprintf("deployment router: all deployments failed for %q: %v", target.canonicalModelID, lastErr)}: - case <-streamCtx.Done(): - } + sendRouterEvent(streamCtx, out, routerErrorEvent( + fmt.Errorf("deployment router: all deployments failed for %q: %w", target.canonicalModelID, lastErr), lastRoute, + )) }() - return llm.NewStreamResult(out, "", cancel), nil + return core.CoordinateStreamResult(ctx, llm.NewStreamResult(out, "", cancel)), nil } func (r *DeploymentRouter) Stats() map[string]int64 { @@ -438,7 +458,152 @@ func (r *DeploymentRouter) chatWithDeployment(ctx context.Context, messages []co return adapter.Provider.Chat(ctx, messages, nativeOpts) } -func (r *DeploymentRouter) streamWithDeployment(ctx context.Context, out chan<- core.FluxStreamEvent, messages []core.FluxMessage, opts core.ChatOptions, target deploymentTarget, deploymentID string) (fallback bool, err error) { +// deploymentRoute describes the deployment that served (or is serving) a +// request, so hosts can attribute usage and price to the actual backend. +func deploymentRoute(opts core.ChatOptions, target deploymentTarget, deploymentID string, attempts int) *llm.ResolvedRoute { + provider := strings.TrimSpace(opts.Provider) + if provider == "" { + provider = ownerProviderID(target.canonicalModelID) + } + model := strings.TrimSpace(opts.Model) + if model == "" { + model = target.canonicalModelID + } + return &llm.ResolvedRoute{ + Provider: provider, Model: model, DeploymentRouting: true, + DeploymentID: deploymentID, Attempts: attempts, + } +} + +// attachResponseRoute fills the response route from the router's view without +// overwriting anything the deployment already reported. +func attachResponseRoute(resp *core.FluxResponse, route *llm.ResolvedRoute) { + if resp == nil || route == nil { + return + } + if resp.Route == nil { + resp.Route = route + return + } + merged := *resp.Route + if merged.Provider == "" { + merged.Provider = route.Provider + } + if merged.Model == "" { + merged.Model = route.Model + } + if !merged.DeploymentRouting { + merged.DeploymentRouting = route.DeploymentRouting + } + if merged.DeploymentID == "" { + merged.DeploymentID = route.DeploymentID + } + if merged.Attempts < route.Attempts { + merged.Attempts = route.Attempts + } + resp.Route = &merged +} + +func sendRouterEvent(ctx context.Context, out chan<- core.FluxStreamEvent, event core.FluxStreamEvent) bool { + select { + case out <- event: + return true + case <-ctx.Done(): + return false + } +} + +// deploymentStreamError is a stream failure a deployment reported. It keeps +// the event's ErrorInfo so breaker accounting and the router's own terminal +// event do not have to re-parse the message. +type deploymentStreamError struct { + message string + info *llm.StreamErrorInfo + // forwarded is set when the failing event was already sent downstream. + forwarded bool +} + +func (e *deploymentStreamError) Error() string { return e.message } + +func streamFailureForwarded(err error) bool { + var streamErr *deploymentStreamError + return errors.As(err, &streamErr) && streamErr.forwarded +} + +// isRetryableStreamFailure extends IsTransient with the kinds a deployment +// reported in-band: an upstream timeout or an unavailable deployment is worth +// trying elsewhere even when its message matches no transient pattern. +func isRetryableStreamFailure(err error) bool { + var streamErr *deploymentStreamError + if errors.As(err, &streamErr) && streamErr.info != nil { + switch streamErr.info.Kind { + case llm.ErrKindTimeout, llm.ErrKindUnavailable: + return true + } + } + return IsTransient(err) +} + +// routerErrorEvent builds the router's terminal error event. It is only used +// while the caller's context is live, so it never reports a cancellation. +func routerErrorEvent(err error, route *llm.ResolvedRoute) core.FluxStreamEvent { + event := core.FluxStreamEvent{Type: "error", Route: cloneResolvedRoute(route)} + if err != nil { + event.Error = err.Error() + } + var streamErr *deploymentStreamError + if errors.As(err, &streamErr) && streamErr.info != nil { + info := *streamErr.info + event.ErrorInfo = &info + } + return event +} + +// isStreamFailureEvent reports whether a deployment event ends its stream +// with a failure. Warning-marked "error" events are non-fatal diagnostics +// (for example a reasoning-only response) that precede the real terminal, +// matching provider/core and the engine. +func isStreamFailureEvent(event core.FluxStreamEvent) bool { + return event.Type == "error" && event.Warning == "" || event.Type == "cancelled" || event.Type == "canceled" +} + +// providerFailureEvent normalizes a failure a deployment reported while the +// caller's context is still live. A "cancelled" event or a canceled kind here +// comes from a context the deployment owns (for example an adapter-side +// timeout), so it is an upstream failure the router may fail over from, not +// the caller's cancellation. +func providerFailureEvent(event core.FluxStreamEvent, deploymentID string) core.FluxStreamEvent { + cancelled := event.Type == "cancelled" || event.Type == "canceled" + event.Type = "error" + if event.Error == "" { + event.Error = fmt.Sprintf("deployment %q stream failed", deploymentID) + } + if event.ErrorInfo == nil && !cancelled { + return event + } + info := llm.StreamErrorInfo{Kind: llm.ErrKindCanceled} + if event.ErrorInfo != nil { + info = *event.ErrorInfo + } + switch info.Kind { + case llm.ErrKindCanceled: + info.Kind, info.Retryable = llm.ErrKindUnavailable, true + case llm.ErrKindTimeout: + info.Retryable = true + } + event.ErrorInfo = &info + return event +} + +func cloneResolvedRoute(route *llm.ResolvedRoute) *llm.ResolvedRoute { + if route == nil { + return nil + } + cloned := *route + return &cloned +} + +func (r *DeploymentRouter) streamWithDeployment(ctx context.Context, out chan<- core.FluxStreamEvent, messages []core.FluxMessage, opts core.ChatOptions, target deploymentTarget, deploymentID string, route llm.ResolvedRoute) (fallback bool, err error) { offering, adapter, err := r.resolveOffering(target, deploymentID) if err != nil { return true, err @@ -448,58 +613,81 @@ func (r *DeploymentRouter) streamWithDeployment(ctx context.Context, out chan<- if err != nil { return true, err } + if stream == nil { + return true, fmt.Errorf("deployment %q returned a nil stream", deploymentID) + } defer stream.Close() emitted := false var buffered []core.FluxStreamEvent - flush := func() { + annotate := func(event core.FluxStreamEvent) core.FluxStreamEvent { + if event.Route == nil { + event.Route = cloneResolvedRoute(&route) + } + return event + } + flush := func() bool { for _, event := range buffered { - select { - case out <- event: - case <-ctx.Done(): - return + if !sendRouterEvent(ctx, out, event) { + return false } } buffered = nil + return true } for event := range stream.Events { - if event.Type == "error" { - if emitted { - select { - case out <- event: - case <-ctx.Done(): - } - return false, fmt.Errorf("%s", event.Error) + event = annotate(event) + if isStreamFailureEvent(event) { + if ctx.Err() != nil { + // The caller cancelled or its deadline passed; the deployment + // is only reporting the consequence. + return false, ctx.Err() } - if event.Error == "" { - return true, fmt.Errorf("deployment %q stream failed before output", deploymentID) + event = providerFailureEvent(event, deploymentID) + failure := &deploymentStreamError{message: event.Error, info: event.ErrorInfo} + if emitted { + failure.forwarded = sendRouterEvent(ctx, out, event) + return false, failure } - return true, fmt.Errorf("%s", event.Error) + return true, failure } if isOutputEvent(event) { emitted = true - flush() - select { - case out <- event: - case <-ctx.Done(): + if !flush() || !sendRouterEvent(ctx, out, event) { return false, ctx.Err() } continue } - if emitted || event.Type == "done" { - flush() - select { - case out <- event: - case <-ctx.Done(): + if event.Type == "done" { + if !flush() || !sendRouterEvent(ctx, out, event) { return false, ctx.Err() } return false, nil } + if emitted { + // Usage, TTFT and provider-block events arrive between output + // events; only done ends a successful stream. + if !sendRouterEvent(ctx, out, event) { + return false, ctx.Err() + } + continue + } + // Before output, hold non-output events so a failover leaves no + // trace of the failed deployment. buffered = append(buffered, event) } + if ctx.Err() != nil { + return false, ctx.Err() + } if emitted { - return false, fmt.Errorf("deployment %q stream ended after output without done", deploymentID) + return false, &deploymentStreamError{ + message: fmt.Sprintf("deployment %q stream ended after output without done", deploymentID), + info: &llm.StreamErrorInfo{Kind: llm.ErrKindTruncated, Retryable: true}, + } + } + return true, &deploymentStreamError{ + message: fmt.Sprintf("deployment %q stream ended before output", deploymentID), + info: &llm.StreamErrorInfo{Kind: llm.ErrKindUnavailable, Retryable: true}, } - return true, fmt.Errorf("deployment %q stream ended before output", deploymentID) } func (r *DeploymentRouter) resolveOffering(target deploymentTarget, deploymentID string) (catalog.ModelOffering, DeploymentAdapter, error) { @@ -760,9 +948,60 @@ func (r *DeploymentRouter) recordSuccess(deploymentID string) { } } -// recordFailure records a deployment failure on its circuit breaker. -func (r *DeploymentRouter) recordFailure(deploymentID string) { +// recordFailure records a deployment failure on its circuit breaker when the +// error says something about the deployment's health. Callers must not pass a +// failure caused by the caller's own cancellation or deadline. +func (r *DeploymentRouter) recordFailure(deploymentID string, err error) { + if !shouldRecordBreakerFailure(err) { + return + } if cb := r.getCircuitBreaker(deploymentID); cb != nil { cb.Failure() } } + +// shouldRecordBreakerFailure reports whether err describes the deployment's +// health. A deadline that reaches it is the deployment timing out (net/http's +// Client.Timeout matches context.DeadlineExceeded), because callers filter out +// their own cancellation first. +func shouldRecordBreakerFailure(err error) bool { + if err == nil || errors.Is(err, context.Canceled) { + return false + } + if errors.Is(err, context.DeadlineExceeded) { + return true + } + var streamErr *deploymentStreamError + if errors.As(err, &streamErr) && streamErr.info != nil { + switch streamErr.info.Kind { + case llm.ErrKindTimeout, llm.ErrKindUnavailable, llm.ErrKindTruncated: + return true + case llm.ErrKindRateLimited, llm.ErrKindAuth, llm.ErrKindContextExceeded, + llm.ErrKindContentFiltered, llm.ErrKindInvalidRequest, llm.ErrKindCanceled: + return false + } + // Internal or unknown kinds fall through to the message checks. + } + var providerErr *core.FluxError + if errors.As(err, &providerErr) { + if providerErr.StatusCode == 0 { + return true + } + switch providerErr.StatusCode { + case 500, 502, 503, 504, 529: + return true + default: + return false + } + } + message := strings.ToLower(err.Error()) + for _, code := range []string{"500", "502", "503", "504", "529"} { + if strings.Contains(message, "http "+code) || strings.Contains(message, "http/"+code) || strings.Contains(message, "status "+code) || strings.Contains(message, "code "+code) { + return true + } + } + if strings.Contains(message, "429") || strings.Contains(message, "rate limit") { + return false + } + return strings.Contains(message, "connection") || strings.Contains(message, "transport") || strings.Contains(message, "unavailable") || strings.Contains(message, "bad gateway") || strings.Contains(message, "service unavailable") +} diff --git a/router/deployment_router_lifecycle_test.go b/router/deployment_router_lifecycle_test.go new file mode 100644 index 00000000..ce6f9458 --- /dev/null +++ b/router/deployment_router_lifecycle_test.go @@ -0,0 +1,362 @@ +package router + +import ( + "context" + "errors" + "fmt" + "sync/atomic" + "testing" + "time" + + "github.com/GrayCodeAI/flux/llm" + "github.com/GrayCodeAI/flux/provider/core" +) + +// scriptedStreamProvider streams a fixed script through +// core.CoordinateStreamResult, the way real adapters wrap their streams. With +// hold set, the stream stays open after the script until its context ends. +type scriptedStreamProvider struct { + name string + events []core.FluxStreamEvent + openErr error + hold bool + calls atomic.Int32 +} + +func (p *scriptedStreamProvider) Chat(ctx context.Context, _ []core.FluxMessage, _ core.ChatOptions) (*core.FluxResponse, error) { + p.calls.Add(1) + if err := ctx.Err(); err != nil { + return nil, err + } + if p.openErr != nil { + return nil, p.openErr + } + return &core.FluxResponse{Content: "from " + p.name}, nil +} + +func (p *scriptedStreamProvider) StreamChat(ctx context.Context, _ []core.FluxMessage, _ core.ChatOptions) (*core.StreamResult, error) { + p.calls.Add(1) + if p.openErr != nil { + return nil, p.openErr + } + ch := make(chan core.FluxStreamEvent) + go func() { + defer close(ch) + for _, event := range p.events { + select { + case ch <- event: + case <-ctx.Done(): + return + } + } + if p.hold { + <-ctx.Done() + } + }() + return core.CoordinateStreamResult(ctx, llm.NewStreamResult(ch, p.name+"-request", nil)), nil +} + +func (p *scriptedStreamProvider) Ping(context.Context) error { return nil } +func (p *scriptedStreamProvider) Name() string { return p.name } + +func healthyScript(name string) []core.FluxStreamEvent { + return []core.FluxStreamEvent{ + {Type: "content", Content: "from " + name}, + {Type: "done", StopReason: "end_turn"}, + } +} + +func newTwoStageRouter(t *testing.T, primary, fallback core.Provider) *DeploymentRouter { + t.Helper() + r, err := NewDeploymentRouter(DeploymentRouterOptions{ + Catalog: testCompiledCatalog(t), + Deployments: map[string]DeploymentAdapter{ + "anthropic-direct": {Provider: primary}, + "anthropic-vertex": {Provider: fallback}, + }, + Routing: RoutingPolicy{Providers: map[string][]RoutingStage{"anthropic": { + {Deployments: []DeploymentChoice{{DeploymentID: "anthropic-direct", Weight: 100}}}, + {Deployments: []DeploymentChoice{{DeploymentID: "anthropic-vertex", Weight: 100}}}, + }}}, + }) + if err != nil { + t.Fatal(err) + } + return r +} + +func startRouterStream(t *testing.T, ctx context.Context, r *DeploymentRouter) *core.StreamResult { + t.Helper() + stream, err := r.StreamChat(ctx, []core.FluxMessage{{Role: "user", Content: "hi"}}, core.ChatOptions{Model: "anthropic/claude-sonnet-4-6"}) + if err != nil { + t.Fatal(err) + } + t.Cleanup(stream.Close) + return stream +} + +// drainStream reads until the channel closes, failing the test if it stalls. +func drainStream(t *testing.T, stream *core.StreamResult) []core.FluxStreamEvent { + t.Helper() + var events []core.FluxStreamEvent + deadline := time.After(5 * time.Second) + for { + select { + case event, ok := <-stream.Events: + if !ok { + return events + } + events = append(events, event) + case <-deadline: + t.Fatalf("stream did not close; events so far: %+v", events) + } + } +} + +func eventTypes(events []core.FluxStreamEvent) []string { + types := make([]string, len(events)) + for i, event := range events { + types[i] = event.Type + } + return types +} + +func breakerFailures(r *DeploymentRouter, deploymentID string) int { + cb := r.getCircuitBreaker(deploymentID) + cb.mu.Lock() + defer cb.mu.Unlock() + return cb.failureCount +} + +func assertFailedOver(t *testing.T, events []core.FluxStreamEvent, fallback *scriptedStreamProvider) { + t.Helper() + for _, event := range events { + if event.Type == "error" || event.Type == "cancelled" { + t.Fatalf("unexpected %s event %+v; events = %v", event.Type, event, eventTypes(events)) + } + } + if len(events) == 0 || events[len(events)-1].Type != "done" { + t.Fatalf("events = %v, want the fallback to finish with done", eventTypes(events)) + } + if got := fallback.calls.Load(); got != 1 { + t.Fatalf("fallback calls = %d, want 1", got) + } +} + +func TestDeploymentRouterUpstreamTimeoutFailsOver(t *testing.T) { + t.Parallel() + primary := &scriptedStreamProvider{name: "direct", events: []core.FluxStreamEvent{ + {Type: "error", Error: "context deadline exceeded (Client.Timeout exceeded while reading body)"}, + }} + fallback := &scriptedStreamProvider{name: "vertex", events: healthyScript("vertex")} + r := newTwoStageRouter(t, primary, fallback) + + events := drainStream(t, startRouterStream(t, context.Background(), r)) + assertFailedOver(t, events, fallback) + if got := breakerFailures(r, "anthropic-direct"); got != 1 { + t.Fatalf("primary breaker failures = %d, want 1 (an upstream timeout is a health signal)", got) + } +} + +func TestDeploymentRouterUpstreamCancelledEventFailsOver(t *testing.T) { + t.Parallel() + // A "cancelled" event while the caller's context is live comes from a + // context the deployment owns, such as an adapter-side timeout. + primary := &scriptedStreamProvider{name: "direct", events: []core.FluxStreamEvent{ + {Type: "cancelled", Error: "context canceled"}, + }} + fallback := &scriptedStreamProvider{name: "vertex", events: healthyScript("vertex")} + r := newTwoStageRouter(t, primary, fallback) + + assertFailedOver(t, drainStream(t, startRouterStream(t, context.Background(), r)), fallback) +} + +func TestDeploymentRouterStreamOpenTimeoutFailsOver(t *testing.T) { + t.Parallel() + // net/http's Client.Timeout errors match context.DeadlineExceeded. + primary := &scriptedStreamProvider{name: "direct", openErr: fmt.Errorf("post: %w", context.DeadlineExceeded)} + fallback := &scriptedStreamProvider{name: "vertex", events: healthyScript("vertex")} + r := newTwoStageRouter(t, primary, fallback) + + assertFailedOver(t, drainStream(t, startRouterStream(t, context.Background(), r)), fallback) + if got := breakerFailures(r, "anthropic-direct"); got != 1 { + t.Fatalf("primary breaker failures = %d, want 1", got) + } +} + +func TestDeploymentRouterCallerCancelEmitsSingleTerminal(t *testing.T) { + t.Parallel() + primary := &scriptedStreamProvider{name: "direct", hold: true, events: []core.FluxStreamEvent{ + {Type: "content", Content: "partial"}, + }} + fallback := &scriptedStreamProvider{name: "vertex", events: healthyScript("vertex")} + r := newTwoStageRouter(t, primary, fallback) + + ctx, cancel := context.WithCancel(context.Background()) + defer cancel() + stream := startRouterStream(t, ctx, r) + for event := range stream.Events { + if event.Type == "content" { + break + } + } + cancel() + + events := drainStream(t, stream) + if len(events) != 1 || events[0].Type != "cancelled" { + t.Fatalf("events after cancel = %v, want exactly one cancelled terminal", eventTypes(events)) + } + if info := events[0].ErrorInfo; info == nil || info.Kind != llm.ErrKindCanceled || info.Retryable { + t.Fatalf("terminal error info = %+v, want non-retryable canceled", info) + } + if got := fallback.calls.Load(); got != 0 { + t.Fatalf("fallback calls = %d, want no failover on caller cancellation", got) + } + if got := breakerFailures(r, "anthropic-direct"); got != 0 { + t.Fatalf("primary breaker failures = %d, want 0 for caller cancellation", got) + } +} + +func TestDeploymentRouterChatCallerDeadlineStopsFailover(t *testing.T) { + t.Parallel() + primary := &scriptedStreamProvider{name: "direct"} + fallback := &scriptedStreamProvider{name: "vertex"} + r := newTwoStageRouter(t, primary, fallback) + + ctx, cancel := context.WithDeadline(context.Background(), time.Now().Add(-time.Second)) + defer cancel() + _, err := r.Chat(ctx, []core.FluxMessage{{Role: "user", Content: "hi"}}, core.ChatOptions{Model: "anthropic/claude-sonnet-4-6"}) + if !errors.Is(err, context.DeadlineExceeded) { + t.Fatalf("error = %v, want the caller's deadline", err) + } + if got := fallback.calls.Load(); got != 0 { + t.Fatalf("fallback calls = %d, want no failover after the caller's deadline", got) + } + if got := breakerFailures(r, "anthropic-direct"); got != 0 { + t.Fatalf("primary breaker failures = %d, want 0 for the caller's deadline", got) + } +} + +func TestDeploymentRouterForwardsEventsAfterOutputUntilDone(t *testing.T) { + t.Parallel() + // Anthropic reports output usage in message_delta, after the text and + // before message_stop; signed thinking blocks also follow output. + primary := &scriptedStreamProvider{name: "direct", events: []core.FluxStreamEvent{ + {Type: "usage", Usage: &core.FluxUsage{PromptTokens: 10}}, + {Type: "content", Content: "a"}, + {Type: "usage", Usage: &core.FluxUsage{CompletionTokens: 3}}, + {Type: "provider_block", ProviderBlock: &llm.ProviderBlock{Provider: "anthropic", Type: "thinking"}}, + {Type: "content", Content: "b"}, + {Type: "done", StopReason: "end_turn"}, + }} + fallback := &scriptedStreamProvider{name: "vertex", events: healthyScript("vertex")} + r := newTwoStageRouter(t, primary, fallback) + + events := drainStream(t, startRouterStream(t, context.Background(), r)) + want := []string{"route_changed", "usage", "content", "usage", "provider_block", "content", "done"} + if got := eventTypes(events); fmt.Sprint(got) != fmt.Sprint(want) { + t.Fatalf("events = %v, want %v", got, want) + } + if got := fallback.calls.Load(); got != 0 { + t.Fatalf("fallback calls = %d, want 0", got) + } +} + +func TestDeploymentRouterWarningDiagnosticsAreNotFatal(t *testing.T) { + t.Parallel() + diagnostic := core.FluxStreamEvent{Type: "error", Error: "reasoning-only response", Warning: "reasoning-only response"} + tests := []struct { + name string + script []core.FluxStreamEvent + want []string + }{ + { + name: "after output", + script: []core.FluxStreamEvent{ + {Type: "content", Content: "a"}, + diagnostic, + {Type: "done", Usage: &core.FluxUsage{CompletionTokens: 1}}, + }, + want: []string{"route_changed", "content", "error", "done"}, + }, + { + name: "before output", + script: []core.FluxStreamEvent{diagnostic, {Type: "content", Content: "a"}, {Type: "done"}}, + want: []string{"route_changed", "error", "content", "done"}, + }, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + primary := &scriptedStreamProvider{name: "direct", events: tt.script} + fallback := &scriptedStreamProvider{name: "vertex", events: healthyScript("vertex")} + r := newTwoStageRouter(t, primary, fallback) + + events := drainStream(t, startRouterStream(t, context.Background(), r)) + if got := eventTypes(events); fmt.Sprint(got) != fmt.Sprint(tt.want) { + t.Fatalf("events = %v, want %v", got, tt.want) + } + for _, event := range events { + if event.Type == "error" && event.Warning == "" { + t.Fatalf("diagnostic lost its warning marker: %+v", event) + } + } + if got := fallback.calls.Load(); got != 0 { + t.Fatalf("fallback calls = %d, want 0 for a non-fatal diagnostic", got) + } + }) + } +} + +func TestDeploymentRouterErrorEventsCarryErrorInfo(t *testing.T) { + t.Parallel() + tests := []struct { + name string + primary []core.FluxStreamEvent + fallback []core.FluxStreamEvent + wantKind string + }{ + { + name: "error after output keeps the adapter's kind", + primary: []core.FluxStreamEvent{{Type: "content", Content: "a"}, {Type: "error", Error: "429 Too Many Requests: rate limit exceeded"}}, + wantKind: llm.ErrKindRateLimited, + }, + { + name: "non-transient error before output", + primary: []core.FluxStreamEvent{{Type: "error", Error: "invalid_api_key: unauthorized"}}, + wantKind: llm.ErrKindAuth, + }, + { + name: "every deployment failed", + primary: []core.FluxStreamEvent{{Type: "error", Error: "503 service unavailable"}}, + fallback: []core.FluxStreamEvent{{Type: "error", Error: "503 service unavailable"}}, + wantKind: llm.ErrKindUnavailable, + }, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + primary := &scriptedStreamProvider{name: "direct", events: tt.primary} + fallback := &scriptedStreamProvider{name: "vertex", events: tt.fallback} + r := newTwoStageRouter(t, primary, fallback) + + events := drainStream(t, startRouterStream(t, context.Background(), r)) + var terminals []core.FluxStreamEvent + for _, event := range events { + if event.Type == "error" || event.Type == "cancelled" || event.Type == "done" { + terminals = append(terminals, event) + } + } + if len(terminals) != 1 || terminals[0].Type != "error" { + t.Fatalf("events = %v, want exactly one error terminal", eventTypes(events)) + } + terminal := terminals[0] + if terminal.ErrorInfo == nil || terminal.ErrorInfo.Kind != tt.wantKind { + t.Fatalf("terminal = %+v (info %+v), want kind %s", terminal, terminal.ErrorInfo, tt.wantKind) + } + if terminal.Route == nil || terminal.Route.DeploymentID == "" { + t.Fatalf("terminal route = %+v, want the deployment that failed", terminal.Route) + } + }) + } +} diff --git a/router/deployment_router_test.go b/router/deployment_router_test.go index e2420aff..abc97413 100644 --- a/router/deployment_router_test.go +++ b/router/deployment_router_test.go @@ -6,6 +6,7 @@ import ( "testing" "github.com/GrayCodeAI/flux/catalog" + "github.com/GrayCodeAI/flux/llm" "github.com/GrayCodeAI/flux/provider/core" ) @@ -118,6 +119,77 @@ func TestDeploymentRouterFallsBackAcrossStages(t *testing.T) { } } +func TestDeploymentRouterReportsActualRouteAndAttempts(t *testing.T) { + t.Parallel() + primary := &deploymentMockProvider{name: "direct", err: fmt.Errorf("HTTP 503 unavailable")} + fallback := &deploymentMockProvider{name: "vertex"} + r, err := NewDeploymentRouter(DeploymentRouterOptions{ + Catalog: testCompiledCatalog(t), + Deployments: map[string]DeploymentAdapter{ + "anthropic-direct": {Provider: primary}, + "anthropic-vertex": {Provider: fallback}, + }, + Routing: RoutingPolicy{Providers: map[string][]RoutingStage{"anthropic": { + {Deployments: []DeploymentChoice{{DeploymentID: "anthropic-direct", Weight: 100}}}, + {Deployments: []DeploymentChoice{{DeploymentID: "anthropic-vertex", Weight: 100}}}, + }}}, + }) + if err != nil { + t.Fatal(err) + } + resp, err := r.Chat(context.Background(), []core.FluxMessage{{Role: "user", Content: "hi"}}, core.ChatOptions{Model: "anthropic/claude-sonnet-4-6"}) + if err != nil { + t.Fatal(err) + } + if resp.Route == nil || resp.Route.DeploymentID != "anthropic-vertex" || resp.Route.Attempts != 2 { + t.Fatalf("route = %+v, want fallback deployment and two attempts", resp.Route) + } +} + +func TestDeploymentRouterStreamReportsRouteEvents(t *testing.T) { + t.Parallel() + primary := &deploymentMockProvider{name: "direct", streamErr: fmt.Errorf("HTTP 503")} + fallback := &deploymentMockProvider{name: "vertex"} + r, err := NewDeploymentRouter(DeploymentRouterOptions{ + Catalog: testCompiledCatalog(t), + Deployments: map[string]DeploymentAdapter{ + "anthropic-direct": {Provider: primary}, + "anthropic-vertex": {Provider: fallback}, + }, + Routing: RoutingPolicy{Providers: map[string][]RoutingStage{"anthropic": { + {Deployments: []DeploymentChoice{{DeploymentID: "anthropic-direct", Weight: 100}}}, + {Deployments: []DeploymentChoice{{DeploymentID: "anthropic-vertex", Weight: 100}}}, + }}}, + }) + if err != nil { + t.Fatal(err) + } + stream, err := r.StreamChat(context.Background(), []core.FluxMessage{{Role: "user", Content: "hi"}}, core.ChatOptions{Model: "anthropic/claude-sonnet-4-6"}) + if err != nil { + t.Fatal(err) + } + defer stream.Close() + var routes []string + var finalRoute string + for event := range stream.Events { + if event.Route != nil { + routes = append(routes, event.Route.DeploymentID) + } + if event.Type == "done" { + finalRoute = event.Route.DeploymentID + } + if event.Type == "error" { + t.Fatalf("unexpected stream error: %s", event.Error) + } + } + if len(routes) < 2 || routes[0] != "anthropic-direct" || routes[1] != "anthropic-vertex" { + t.Fatalf("route events = %v, want direct then vertex", routes) + } + if finalRoute != "anthropic-vertex" { + t.Fatalf("terminal route = %q, want vertex", finalRoute) + } +} + func TestShouldTryNextDeploymentCredits(t *testing.T) { t.Parallel() err := fmt.Errorf("requires more credits, or fewer max_tokens; can only afford 5705") @@ -409,3 +481,36 @@ func TestDeploymentRouterRetriesPreferDifferentEndpoint(t *testing.T) { t.Fatalf("healthy deployment called %d times; want 1", healthy.callCount) } } + +func TestShouldRecordBreakerFailure(t *testing.T) { + t.Parallel() + tests := []struct { + name string + err error + want bool + }{ + {"nil", nil, false}, + {"caller canceled", context.Canceled, false}, + {"upstream deadline", fmt.Errorf("post: %w", context.DeadlineExceeded), true}, + {"stream timeout kind", &deploymentStreamError{message: "read timed out", info: &llm.StreamErrorInfo{Kind: llm.ErrKindTimeout}}, true}, + {"stream auth kind", &deploymentStreamError{message: "HTTP 503 but auth", info: &llm.StreamErrorInfo{Kind: llm.ErrKindAuth}}, false}, + {"stream internal kind uses message", &deploymentStreamError{message: "connection reset", info: &llm.StreamErrorInfo{Kind: llm.ErrKindInternal}}, true}, + {"server error status", &core.FluxError{Provider: "p", Op: "chat", StatusCode: 503}, true}, + {"overloaded status", &core.FluxError{Provider: "p", Op: "chat", StatusCode: 529}, true}, + {"rate limited status", &core.FluxError{Provider: "p", Op: "chat", StatusCode: 429}, false}, + {"bad request status", &core.FluxError{Provider: "p", Op: "chat", StatusCode: 400}, false}, + {"transport error without status", &core.FluxError{Provider: "p", Op: "chat", Message: "dial tcp: connection refused"}, true}, + {"http 502 in message", fmt.Errorf("upstream returned HTTP 502"), true}, + {"rate limit message", fmt.Errorf("429 rate limit exceeded"), false}, + {"connection reset", fmt.Errorf("read: connection reset by peer"), true}, + {"invalid request message", fmt.Errorf("invalid_request_error: bad param"), false}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + if got := shouldRecordBreakerFailure(tt.err); got != tt.want { + t.Fatalf("shouldRecordBreakerFailure(%v) = %v, want %v", tt.err, got, tt.want) + } + }) + } +} diff --git a/router/router.go b/router/router.go index 0c9994db..60f8914a 100644 --- a/router/router.go +++ b/router/router.go @@ -144,7 +144,7 @@ func (r *Router) StreamChat(ctx context.Context, messages []core.FluxMessage, op r.stratState.endInFlight(provider.Name()) if err == nil { r.recordSuccess(provider.Name()) - return sr, nil + return core.CoordinateStreamResult(ctx, sr), nil } if !IsTransient(err) { return nil, err @@ -156,7 +156,7 @@ func (r *Router) StreamChat(ctx context.Context, messages []core.FluxMessage, op sr, err = fp.StreamChat(ctx, messages, opts) if err == nil { r.recordSuccess(fp.Name()) - return sr, nil + return core.CoordinateStreamResult(ctx, sr), nil } if !IsTransient(err) { return nil, err