Stop your docs from lying. A GitHub Action that detects when code changes contradict your documentation — and offers to fix them automatically.
See the changelog for release history.
Every engineering team has stale docs.
A developer refactors state management from Redux to Zustand. The code ships. The ARCHITECTURE.md still says "We use Redux Toolkit for all global state." Nobody notices — until a new hire wastes two days debugging the wrong mental model.
Knowledge Diff sits in your CI and plays the role of a vigilant tech writer — one that actually reads the diff.
On every pull request, Knowledge Diff:
- Reads the code diff — what functions, string literals, and lines were added/removed
- Finds relevant doc sections — matches changed files/symbols against project docs and AI-agent instructions such as
AGENTS.md,CLAUDE.md, and Copilot instructions - Asks an LLM once per changed file — batches the relevant sections into one structured drift check
- Comments on the PR — with specific, quote-level detail about what drifted
- Opens a patch PR (optional) — with suggested doc updates ready for your review
Definite contradiction: The code replaced Redux
createSlicewith Zustandcreate(), but the doc still describes Redux as the state management solution.Doc still says: "We use Redux Toolkit with createSlice for all global state."
Suggested update:
- We use Redux Toolkit with createSlice for all global state. + We use Zustand for client-side global state management.
Add this to .github/workflows/knowledge-diff.yml:
name: Knowledge Diff
on:
pull_request:
types: [opened, synchronize, reopened]
permissions:
contents: read # read the PR and documentation
pull-requests: write # required to post comments
jobs:
check-rationale-drift:
# Repository secrets are not available to pull requests from forks.
if: github.event.pull_request.head.repo.full_name == github.repository
runs-on: ubuntu-latest
steps:
- uses: oarisur/knowledge-diff@v1
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
openai-api-key: ${{ secrets.OPENAI_API_KEY }}That's it. Every PR now gets a documentation health check.
| Input | Required | Default | Description |
|---|---|---|---|
github-token |
✅ | — | GitHub token for posting comments. Use secrets.GITHUB_TOKEN. |
openai-api-key |
✅* | — | OpenAI API key. Required when llm-provider is openai. |
anthropic-api-key |
✅* | — | Anthropic API key. Required when llm-provider is anthropic. |
gemini-api-key |
✅* | — | Google Gemini API key. Required when llm-provider is gemini. |
llm-provider |
❌ | openai |
LLM backend: openai, anthropic, or gemini. |
llm-model |
❌ | gpt-4o-mini / claude-haiku-4-5-20251001 / gemini-2.5-flash |
Override the model. |
doc-files |
❌ | Common project docs and agent instructions | Comma-separated globs of docs to check. See default document patterns. |
code-extensions |
❌ | ts,tsx,js,jsx,py,go,rs,java,cpp,c,rb,php,swift,kt |
File extensions treated as code. |
sensitivity |
❌ | medium |
Drift threshold: low (definite only) / medium / high (includes ambiguities). |
auto-patch |
❌ | false |
Open a follow-up PR with suggested doc fixes when drift is detected. |
comment-mode |
❌ | update |
update = edit existing comment on re-push. new = always post fresh. |
max-files-per-run |
❌ | 20 |
Max code files to analyse per run (controls LLM cost). |
| Output | Description |
|---|---|
drift-detected |
"true" if any drift was found above the sensitivity threshold. |
drift-count |
Number of drift issues found. |
patch-pr-url |
URL of the auto-generated doc patch PR (empty if none created). |
analysis-complete |
"true" only when every selected candidate was analysed successfully. |
analysis-error-count |
Number of failures that made the analysis incomplete. |
Knowledge Diff checks normal Markdown documentation plus common instruction formats used by coding agents:
README.md,ARCHITECTURE.md,docs/**/*.md,
**/AGENTS.md,**/AGENTS.override.md,**/CLAUDE.md,**/GEMINI.md,
.github/copilot-instructions.md,.github/instructions/**/*.instructions.md,
.cursor/rules/**/*.mdc,.windsurfrules,.clinerules,.clinerules/**/*.md,
.roo/rules/**/*.md
Set doc-files explicitly to replace this list.
- uses: oarisur/knowledge-diff@v1
id: drift
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
openai-api-key: ${{ secrets.OPENAI_API_KEY }}
sensitivity: low # only definite contradictions
- name: Fail on drift
if: steps.drift.outputs.drift-detected == 'true'
run: |
echo "Definite documentation drift detected. Please update your docs."
exit 1- uses: oarisur/knowledge-diff@v1
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
llm-provider: anthropic- uses: oarisur/knowledge-diff@v1
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
gemini-api-key: ${{ secrets.GEMINI_API_KEY }}
llm-provider: gemini- uses: oarisur/knowledge-diff@v1
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
openai-api-key: ${{ secrets.OPENAI_API_KEY }}
auto-patch: "true"Auto-patching also requires contents: write. Some organizations separately disable pull-request creation by GitHub Actions; enable that repository or organization setting before using this option.
When drift is detected, a second PR like docs/knowledge-diff-42-a1b2c3d is opened targeting the same base branch — with the suggested text replacement applied. You review and merge (or discard) at your discretion.
- uses: oarisur/knowledge-diff@v1
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
openai-api-key: ${{ secrets.OPENAI_API_KEY }}
doc-files: "docs/architecture/*.md,CLAUDE.md"
sensitivity: highPR opened / push to PR
│
▼
[Fetch PR diff] ──► changed code files only (by extension)
│
▼
[Fetch PR-head docs] ─► project docs + AI-agent instruction files
│
▼
[Keyword index] ──► map: symbol/path → doc sections that mention it
│
▼
[LLM comparison] ──► one batched request per changed code file
│ evaluates up to 6 relevant doc sections
▼
[Drift found?]
├── YES ──► Post PR comment with quote-level explanation
│ └── auto-patch: true → open a doc-fix PR
└── NO ──► Post "✅ all clear" comment (updates existing one)
| Level | What gets flagged |
|---|---|
low |
Only definite contradictions — the doc says X, the code now does Y. |
medium (default) |
Definite contradictions + likely outdated statements. |
high |
All of the above + possible ambiguities. Err on the side of caution. |
Rather than sending entire files to the LLM (expensive, slow), Knowledge Diff:
- Splits each doc into sections by heading
- Builds a keyword index over all sections
- For each changed code file, looks up the top 6 most relevant sections by keyword overlap with the changed file path and symbol names
- Sends the code patch and those sections in one batched request per changed file
This keeps costs low and avoids irrelevant context diluting the analysis.
For comment-only mode, add:
permissions:
contents: read
pull-requests: writeFor auto-patch: "true", change contents to write. GitHub Actions must also be allowed to create pull requests in the repository or organization settings.
GitHub does not pass repository secrets (including LLM API keys) to normal pull_request workflows from forks, and the fork's GITHUB_TOKEN is normally read-only. Skip the job for fork PRs:
if: github.event.pull_request.head.repo.full_name == github.repositoryDo not switch to pull_request_target while checking out or executing untrusted pull-request code; that can expose secrets to malicious changes. A hosted GitHub App is the safer way to support untrusted fork PRs.
The repository now includes a deployable server-side GitHub App for teams that do not want LLM keys in repository Actions.
cp .env.hosted.example .env
docker compose up --buildThe hosted service provides:
- HMAC-SHA256 webhook verification against the raw request body
- short-lived GitHub installation-token authentication
- GitHub check runs and update-in-place hosted PR comments
- safe analysis of fork PRs without checking out or executing their code
- trusted-base
.github/knowledge-diff.ymlconfiguration - bounded concurrency, latest-event queueing, delivery deduplication, health checks, JSON logs, and graceful shutdown
- a non-root, read-only production container
Register the App with contents: read, pull_requests: write, and checks: write, subscribe to the Pull request event, and point its webhook to /api/github/webhooks. See the complete hosted GitHub App deployment guide and registration manifest template.
The included queue is intended for one service instance. Add a shared durable queue/idempotency backend before horizontal scaling.
Knowledge Diff includes a versioned benchmark with 30 hand-labeled cases: 15 real documentation contradictions and 15 relevant-but-harmless changes. It measures candidate retrieval separately from model classification so a mocked LLM cannot create a misleading quality score.
Run the deterministic retrieval gate used by CI:
npm run evaluate:gateRun the complete benchmark against a live provider after setting the corresponding OPENAI_API_KEY, ANTHROPIC_API_KEY, or GEMINI_API_KEY environment variable:
npm run evaluate -- --provider openai --gate --output evaluation/results/openai.json
npm run evaluate -- --provider anthropic --gate --output evaluation/results/anthropic.json
npm run evaluate -- --provider gemini --gate --output evaluation/results/gemini.jsonLive reports include candidate-level precision, recall, F1, false-positive rate, correct-target recall, latency, estimated token usage, and estimated cost. The default release thresholds are:
| Metric | Required |
|---|---|
| Retrieval recall@6 | ≥ 95% |
| Candidate precision | ≥ 85% |
| Candidate recall | ≥ 80% |
| Correct positive target recall | ≥ 80% |
| Provider failure rate | ≤ 5% |
Use --tag, --max-cases, custom threshold flags, or --model for focused comparisons. Custom models need --input-price and --output-price for cost estimates. See the evaluation guide for metric definitions and instructions for adding anonymized real-world cases.
The bundled cases are a reproducible starting benchmark, not a substitute for beta-repository evidence. Before a commercial launch, add anonymized mistakes and non-issues from real pull requests and keep the resulting provider reports as release artifacts.
Each PR run makes at most N LLM calls, where N is the number of changed code files (up to max-files-per-run). Up to six relevant documentation sections are batched into each call.
For a typical PR changing 5 files:
- Up to 5 model requests instead of 30 individual candidate requests
- Approximately 30,000 input tokens plus 6,000 output tokens in a section-heavy run
- Roughly $0.01 at current
gpt-4o-minipricing; actual usage depends on patch and section sizes
See official OpenAI model pricing. Set max-files-per-run: 10 to cap cost on large PRs.
git clone https://github.com/oarisur/knowledge-diff
cd knowledge-diff
npm install
npm run build # bundles to dist/index.js via ncc
npm test # run unit tests
npm run typecheck # verify TypeScript
npm run lint # ESLint checksMIT — see LICENSE.
PRs welcome. The action dogfoods itself — any change to src/ that contradicts this README.md will be caught by its own CI. 🧠