Skip to content

Latest commit

 

History

231 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

logo

Petsitter is an OpenAI-compatible proxy that layers smart harnesses on top of language models to give them capabilities they don't natively have. It also makes finicky behaviors reliable and dependable.

You install it, point it at a model, load a few example tricks, and suddenly things that model couldn't do before such as tool calling, structured JSON, multi-step reasoning start working. You can also protect secrets, have memory, share server instances across harnesses, and extend the tool trivially.

The built-in tricks are starting points. Tweak them, combine them, or use them as a reference to build something entirely different. Petsitter isn't a turnkey product; it's a kit.

How It Works

Petsitter_Intelligent_Proxy_-_Slide_2a

Petsitter intercepts every request/response pair and runs it through a pipeline of hooks. Each trick picks which hooks it needs:

  1. system_prompt - Inject instructions before the model sees the conversation
  2. pre_hook - Modify messages or inject tool definitions before the API call
  3. post_hook - Validate, retry, or transform the model's response
  4. info - Declare capabilities back to your application

Tricks also have lifecycle hooks (install, startup, shutdown, uninstall) for managing resources across their lifetime.

A trick can be as simple as appending a sentence to the system prompt, or as involved as routing subtasks to three different models in parallel. There's a GUI at / with tabs for managing tricksets and their tricks (Tricks / Models / Agents), a live activity log (Logs), and per-trickset logging configuration (Settings).

The Tricks tab lists your local tricks/*.py alongside community tricks published by other people, and the speech-bubble button in the header opens a Try It panel that sends a message through the pipeline so you can watch which tricks fire.

You can also edit tricks, reorder them, disable, add new ones, and filter them: 2026-07-04_15-13

Petsitter is part of the DAY50 suite of open-source tools for local AI workflows and constructing better agents.

The core goals of Petsitter are:

  • No model changes required - Works with any OpenAI-compatible endpoint
  • Pluggable architecture - Write your own tricks in Python. (Skills are included in .agents)
  • Transparent to your app - Point your existing code at petsitter instead of the model
  • Mix and match - Combine multiple tricks for compound effects

Quick Start

Quickest way:

$ uvx petsitter

Or you can do one off invocation:

# Run petsitter, reading settings from the default config file
# (~/.config/petsitter/config.json, or $PET_CONFIG_DIR)
petsitter -l localhost:8080

# Or point at a specific config file (model, tricksets, etc. all live there)
petsitter -c another_petsitter_config.conf.json -l localhost:8080

Configure the upstream model, tricksets, and modelset via the dashboard at http://localhost:8080 or the pet CLI — everything is persisted to the config file, so a plain petsitter starts the same way next time. pet accepts the same -c flag (before the subcommand, e.g. pet -c another_petsitter_config.conf.json ls) so both tools can target the same config area.

Either way, now you can point your AI applications to http://localhost:8080/v1 and you're going through the petsitter middleware.

Zero-Config Host Override (/p/)

The /p/ route is the easy way to proxy an existing endpoint: prefix whatever host you already use with http://localhost:8080/p/ and petsitter handles the rest. No trickset to create, no model config to swap.

# Instead of https://build.nvidia.com/...
http://localhost:8080/p/build.nvidia.com/...

Point your client's base_url at http://localhost:8080/p/<host> and petsitter forwards everything after the host to https://<host>/<rest> - whatever path the client appends. It's a dumb-client-friendly trick: the client just appends /chat/completions, /v1/models, or anything else to the base you give it, and petsitter proxies it through the normal trick pipeline.

Key behaviors:

  • HTTPS only - the upstream host is always assumed https://. This is a convenience feature for public endpoints.
  • Auth passthrough - your client's Authorization header is forwarded to the upstream, so each host's own API key works.
  • Model passthrough - the request's model field goes upstream as-is.
  • Trickset selection relies on the existing X-Title/Model filters - no special handling, the /p/ path only overrides the upstream host.
  • Chat completions (streaming included), model listings, and any other path under /p/ are proxied.
curl http://localhost:8080/p/build.nvidia.com/v1/chat/completions \
  -H "Authorization: Bearer <your-nvidia-key>" \
  -d '{"model":"meta/llama-3.3-70b-instruct","messages":[{"role":"user","content":"hi"}]}'

Config Diagnostic (__petsitter_config__)

Send a single user message containing exactly __petsitter_config__ and petsitter answers with a snapshot of its configuration instead of calling the upstream model:

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"my-model","messages":[{"role":"user","content":"__petsitter_config__"}]}'

The request is an exact copy of the real traversal: it runs the same keyword filtering, trickset selection, system-prompt injection, and pre-hooks, and builds the actual upstream URL, payload, and headers that a real request would use — then returns a snapshot in place of the upstream HTTP call. No upstream request is made, and your API key is never included (only its presence as "set", "bearer", or "none").

The snapshot includes:

  • model - configured upstream URL, model name, key presence, and the resolved target URL
  • request - X-Title, model, stream, plus original_messages vs transformed_messages (exactly what upstream would receive) and the would-be upstream payload/url/auth
  • trickset - the matched trickset and each active trick (class, display name, keywords, per-trick config)
  • tricksets - every loaded trickset and capabilities

It works on /v1/chat/completions and /p/<host>/.../chat/completions, and honors stream: true (returned as a normal chunked stream). It only triggers on an exact full-message match, so ordinary conversation is unaffected.

CLI Options

Option Short Description
--config -c Path to a config file (e.g., another_petsitter_config.conf.json) or a config directory. Defaults to $PET_CONFIG_DIR if set, else ~/.config/petsitter. Tricksets live in <base>/tricksets.
--listen -l Host:port to listen on (default: localhost:8080)
--version -v Show version and exit

pet subcommands

pet edits the same JSON files the dashboard writes, so the two always agree and neither needs the server running. pet --help lists everything; the ones worth knowing:

Command What it does
pet ts List tricksets; pet ts <name> for detail, pet ts <name> <trick> <param> <value> to set one
pet tricks List available local trick modules
pet add / pet rm Add or remove a trick from a trickset
pet search [query] Search the community index
pet cat <owner>/<name> Print a community trick's source without installing it
pet install <owner>/<name> Install from the index (a bare name instead runs a local trick's install() hook)
pet installed List tricks installed from the index
pet publish <trick> Publish a trick to the index
pet model Show or set model config
pet agents List, register, unregister harness agents

Community Tricks

Tricks are shareable. Anyone can publish one, and they show up in everyone's dashboard within the hour, in the Available Tricks list on the Tricks tab, next to your local ones.

There is no registry server. The index is a static index.json in day50-dev/tricks, rebuilt hourly by a GitHub Action that crawls public repos carrying the topic petsitter-trick. No accounts, no approval queue, nothing to keep running.

Installing

From the dashboard, hit Install on any community entry. It downloads, verifies the checksum, and adds it to the selected trickset. Or from the CLI:

pet search tool                  # search the index
pet cat dana/ollama-ctx          # read the source first
pet install dana/ollama-ctx --trickset opencode

Installed tricks land at <config>/tricks/<owner>/<slug>/<version>.py, and tricksets refer to them with a pkg: spec rather than a path:

{
  "name": "my-trickset",
  "tricks": [
    "tricks/json_mode.py",
    "pkg:dana/ollama-ctx@0.1.0"
  ]
}

The pkg: form is what makes a trickset portable. The same JSON works on another machine, where a /home/you/... path would not. Omit @version and the newest installed version is used.

Point at a different index (a private one for your org, say) with PET_REGISTRY_INDEX, either an https:// or a file:// URL. The index is cached for an hour; a stale cache is preferred to an error, so the list still works offline.

Publishing

Three steps, no ceremony:

  1. Put a __version__ on your Trick subclass. Semver, bumped whenever the file changes.
  2. Push it to a public GitHub repo. Root or a tricks/ directory, as many tricks per repo as you like.
  3. Add the topic: gh repo edit --add-topic petsitter-trick

pet publish tricks/my_trick.py runs steps 2 and 3 for you and checks step 1 first.

class OllamaCtxTrick(Trick):
    __version__ = "0.1.0"
    __brief__ = "Clamps num_ctx for ollama backends"
    __display_name__ = "Ollama Context Clamp"

Everything in the index is derived from that file and the GitHub API. You never type a checksum, a date, or an author:

Field Where it comes from
name your GitHub login + the filename, e.g. dana/ollama-ctx
version __version__
brief, display_name __brief__, __display_name__ (else the class name)
keywords, prompt_keyword, required_models the class attributes
url pinned to a commit SHA, so the bytes can never change under someone
sha256 computed from those bytes; pet install refuses on a mismatch
repo, stars, license, updated the GitHub API

Names can't collide between authors, because your GitHub login is the namespace, so there is nothing for anyone to adjudicate and publishing needs no permission. The crawler parses candidate files with ast; it never imports or executes them.

To update, bump __version__ and push. To unpublish, delete the repo or drop the topic. Anyone who already installed it keeps their copy, since the file is on their disk.

featured.json in the index repo controls which tricks appear before you click Show N more community tricks. It's promotion, not permission: nothing is ever kept out of the index for being unfeatured.

A trick is Python that runs inside petsitter with your API keys, the same trust model as any pip package. pet cat and the dashboard's Read button exist because a trick is one short file, a good deal more reviewable than the average dependency.

Try It

The speech-bubble button in the header opens a conversation panel docked over the dashboard. Type a message and it goes through chat_completions() exactly as a real client's would: same trickset matching, same keyword gating, same hooks, same upstream. It is not a simulation.

What comes back with each reply:

  • A pill per trick. Bright means it changed something, and the tooltip lists the stages it ran (Ran: system_prompt, post_hook). Dim means it was loaded but did nothing.
  • Why a trick stayed quiet. A keyword-gated trick that didn't fire reads Did not fire, needs keyword: banana.
  • Timing and tokens, next to the trickset that handled it.
  • The rows light up. Tricks that actually did something pulse in the Loaded Tricks list, so you can watch a reorder or a config change take effect.

Drag the panel by its header to move it, drag its corner to resize, and snaps it back to the bottom right. Whether it's open, where it sits, and the conversation itself are all remembered across refreshes.

It targets whichever trickset is selected in the pill bar, so switching tricksets switches what you're testing.

Creating Custom Tricks

flowchart TD
  A[Client POST] --> B
  A -.-> K[Prompt keyword scan]
  subgraph config[Reorderable via config]
    B[Trickset match] --> C[Keyword activate]
    C --> D[System prompt]
    D --> E[Pre-hook]
  end
  E --> L[LLM call]
  L --> F[Post-hook]
  F --> G[Capabilities]
  G --> Z[Client response]
  K -.-> Z
Loading

Tricks also have lifecycle hooks that run outside the request pipeline: install() on add, startup() on first concurrent use, shutdown() on last concurrent finish, and uninstall() on removal.

Here is a minimal trick that stops the model from using em-dashes (the long dash character that LLMs love to overuse) and replaces them with regular hyphens:

"""No Em-Dash trick - replaces em-dashes with hyphens."""

from petsitter.trick import Trick

EMDASH = "\u2014"

class NoEmDashTrick(Trick):
    __brief__ = "Replaces em-dashes with hyphens in model responses"
    __display_name__ = "No Em-Dash"

    def system_prompt(self, to_add: str) -> str:
        return "Do NOT use em-dashes. Use a regular hyphen (-) instead."

    def post_hook(self, context: list) -> list:
        if not context:
            return context
        last = context[-1]
        content = last.get("content", "")
        if EMDASH in content:
            content = content.replace(EMDASH, "-")
            last["content"] = content
        return context

The Trick class has four optional request hooks and optional keyword activation:

system_prompt(to_add: str) -> str

When: Called once per request, before any messages are sent to the model.

Purpose: Append instructions to the system prompt. This is how you "prime" the model to behave a certain way.

Example:

def system_prompt(self, to_add: str) -> str:
    return "IMPORTANT: Respond only in valid JSON. No markdown, no explanations."

By default the returned text is appended to any existing system prompt, deduplicated so repeated injection doesn't stack. If a trick genuinely needs to replace the whole system prompt (e.g. swapping in a complete harness), set replace_system_prompt = True on the class:

class SwapHarnessTrick(Trick):
    replace_system_prompt = True

    def system_prompt(self, to_add: str) -> str:
        return "FULL REPLACEMENT PROMPT"

pre_hook(context: list, params: dict) -> list

When: Called after the system prompt is set, before the model receives the messages.

Purpose: Modify the conversation context. You can inject tool definitions, add few-shot examples, or restructure messages.

Parameters:

  • context: List of message dicts ([{"role": "user", "content": "..."}])
  • params: Request parameters including tools, temperature, etc.

Example:

def pre_hook(self, context: list, params: dict) -> list:
    if "tools" in params:
        tools_json = json.dumps(params["tools"])
        context[0]["content"] += f"\n\nAvailable tools: {tools_json}"
    return context

post_hook(context: list) -> list

When: Called after the model responds, before the response goes back to your application.

Purpose: Validate, transform, or retry. This is where you can:

  • Parse the response and convert it to a different format
  • Detect when the model failed and call it again with feedback
  • Extract tool calls from natural language

Example (JSON validation with retry):

def post_hook(self, context: list) -> list:
    attempts = 3
    while attempts > 0:
        try:
            json.loads(context[-1]["content"])
            break
        except json.JSONDecodeError:
            attempts -= 1
            if attempts == 0:
                break
            context = callmodel(context, "That wasn't valid JSON. Try again.")
    return context

Example (Tool call detection):

def post_hook(self, context: list) -> list:
    content = context[-1]["content"]
    if self._looks_like_tool_call(content):
        context[-1]["tool_calls"] = [self._parse_tool_call(content)]
        context[-1]["content"] = None
    return context

info(capabilities: dict) -> dict

When: Called when building the response to your application.

Purpose: Declare what capabilities this trick provides. Some frameworks check for capabilities before using certain features.

Example:

def info(self, capabilities: dict) -> dict:
    capabilities["json_mode"] = True
    capabilities["tools_support"] = True
    return capabilities

Lifecycle Hooks

Every trick can implement up to 4 lifecycle hooks that the framework calls automatically:

install()

Called once when the trick is first added to a trickset. Use for one-time setup - clone repos, download files, create resources:

def install(self):
    self.cache_dir = Path("/tmp/my-trick-cache")
    self.cache_dir.mkdir(parents=True, exist_ok=True)
    download_model(self.cache_dir)

startup()

Called when the first concurrent request starts using this trick (the internal run counter goes 0→1). Use for per-session initialization. It open connections and preloads models:

def startup(self):
    self.session = httpx.Client()

shutdown()

Called when the last concurrent request finishes using this trick (run counter goes 1→0), or during server shutdown for all active tricks. Use for per-session cleanup. It closes connections and release resources:

def shutdown(self):
    self.session.close()

uninstall()

Called when the trick is removed from a trickset. Undo anything done during install():

def uninstall(self):
    import shutil
    shutil.rmtree(self.cache_dir, ignore_errors=True)

The startup/shutdown hooks use a reference counter so multiple concurrent requests to the same trick won't trigger repeated startup/shutdown calls - startup() fires once for the first request, and shutdown() fires when the last one finishes.

Keywords

Activation

Set keywords on your trick class to activate only when the user includes that word in their message - the keyword is stripped before the model sees it. See tricks/multiround.py for a working example.

# Trick fires when "multiround" is present
curl http://localhost:8080/v1/chat/completions \
  -d '{"messages":[{"role":"user","content":"multiround explain the CAP theorem"}]}'

# Trick does nothing without the keyword
curl http://localhost:8080/v1/chat/completions \
  -d '{"messages":[{"role":"user","content":"explain the CAP theorem"}]}'

Prompts

Prompt keywords let you inject commands to petsitter itself inline in your message using the format (<keyword>: <request>). The framework scans for registered keywords, strips the matching pattern before the model sees it, and routes the request to the appropriate handler.

The syntax is forgiving - a registered keyword can be triggered any of these ways:

  • (swapharness: opencode/claude.md) - parenthesized with a request
  • (swapharness:opencode/claude.md) - the space after the colon is optional
  • (swapharness:) or (swapharness) - empty request (e.g. list the harness tree)
  • just swapharness - a bare keyword alone in a message means an empty request

This is separate from trick keyword activation - keywords activate or deactivate tricks for the current request, while prompt keywords are commands to petsitter that bypass the model entirely.

How to register a prompt keyword

Set prompt_keyword on your Trick subclass:

class MyCommandTrick(Trick):
    prompt_keyword = "mycommand"
    __brief__ = "Handles (mycommand: ...) inline requests"

    def handle_prompt_keyword(self, request: str) -> dict | None:
        return {"role": "assistant", "content": f"You asked: {request}"}

The method receives the text after mycommand: and can return:

  • A message dict - injected as the model response (bypasses the upstream call)
  • None - the pattern is stripped but the normal pipeline continues

Notes

  • Execution goes in order of the prompt reference. Unrecognized prompt keywords are passed through and surface as a non-critical error in the response along with the rest of the response
  • The pattern (<keyword>: <request>) properly handles nested parentheses by tracking a depth counter.
  • Keyword matching is case-insensitive.
  • If the handler raises, an error message is returned as the assistant response.

Reference Templates

Tricks are managed with pet and grouped into tricksets; there are no per-trick command-line flags. Point petsitter at a model once, make a trickset, and add tricks to it:

pet model default url http://localhost:11434
pet model default model qwen3:8b

pet new mine                    # create a trickset (X-Title '*', Model '*')
pet add mine json_mode          # add a trick to it
petsitter                       # start; settings come from the config file

Every example below assumes that, so it only shows the pet add line. Swap mine for whichever trickset you're building, or use the dashboard's Available Tricks list instead.

Output Control

Capability Injection

  • Tool Calling - Add tool calling to models without native support
  • Conversational Tool - ANDYBOT persona tool calling for small/older models
  • MCP Tools - Inject tools from an mcp.json file into any harness

Pipeline

  • Kennel - Route cognitive subtasks to specialized models
  • Multi-Model Consultant - Two models cross-validate and improve each other's responses

Security

  • Secrets Protector - Detect and pseudonymize secrets/PII before they reach the model

Agent

  • Swap Harness - Browse and swap system prompts from AI tool repositories
  • Self-Improver - Runtime agent that can add, modify, and list tricks

Utility

  • Rules File - Inject a shared AGENTS.md-style rules file into the system prompt
  • Export It - Export conversation as llcat-compatible JSON

JSON Mode

tricks/json_mode.py

Enforces valid JSON output by adding formatting instructions to the system prompt, stripping markdown code blocks, and retrying with feedback if the response isn't valid JSON.

pet add mine json_mode

Code Validator

tricks/code_validator.py

After the model proposes a code change, asks it to describe what the change does, compares the description against the original user request, and retries with feedback if they don't match.

pet add mine code_validator

Tool Calling

tricks/tool_call.py

Enables tool calling for models without native support by injecting tool definitions into the prompt, parsing JSONRPC-style tool call responses, and converting them to OpenAI tool_calls format.

pet add mine tool_call

Conversational Tool

tricks/conversational_tool.py

A conversational approach to tool calling that uses the ANDYBOT persona instead of structured JSON output. The model says DEAR ANDYBOT, <FUNCTION> and ANDYBOT collects each parameter through dialogue:

  1. Model recognises it needs to call a tool and says DEAR ANDYBOT, GET_WEATHER
  2. ANDYBOT asks: "Can you provide location?"
  3. Model responds: Paris
  4. ANDYBOT builds the tool call and returns it to the application

This works well with small models (3B and under) and older models that struggle with reliable JSON output or native tool_calls. The conversational flow lets them express intent naturally instead of wrestling with syntax. It also supports inline arguments (DEAR ANDYBOT, GET_WEATHER location=Paris), optional parameters, and "I am confused"/"skip" recovery. The persona is only injected when the request actually carries tools.

pet add mine conversational_tool
pet add mine json_mode

MCP Tools

tricks/mcp_tools.py

Injects tools defined in an mcp.json file into any harness. Converts MCP tool definitions to OpenAI function-calling format and merges them into params["tools"]. Tools with name collisions take precedence over existing tool definitions.

Default path: ~/.config/petsitter/mcp.json. Use the mcp prompt keyword to switch files at runtime.

pet add mine mcp_tools          # reads ~/.config/petsitter/mcp.json

To point it at a different file, use the mcp prompt keyword in a message: (mcp: /path/to/my-tools.json).

The mcp.json format follows the MCP spec:

{
  "mcpSpec": "1.0.0",
  "server": { "name": "my-tools", "version": "1.0.0" },
  "tools": [
    {
      "name": "search_docs",
      "description": "Search documentation by query",
      "inputSchema": {
        "type": "object",
        "properties": {
          "query": { "type": "string" },
          "limit": { "type": "number", "default": 10 }
        },
        "required": ["query"]
      }
    }
  ]
}

Multi-Model Orchestration

A trick has full control of the request lifecycle - it can call any number of models, not just the one the user pointed at. This lets you decompose a problem into subtasks and route each one to the model best suited for it.

Petsitter supports this through model configs - JSON files that map role names to {url, model, key} objects. Tricks declare what roles they need; if a key is missing, petsitter prints a helpful error.

The model and key fields can be a string or boolean false - false means passthrough (don't set the field in the upstream request). This is distinct from "" which clears the value.

Example modelset.json:

{
    "default": {
        "url": "http://localhost:11434",
        "model": "Qwen3.5:8b"
    },
    "thinker": {
        "url": "http://localhost:11434",
        "model": "VibeThinker-3B-GGUF:q4_K_M"
    },
    "toolcall": {
        "url": "http://localhost:11434",
        "model": "lfm2.5:latest",
        "key": "sk-custom-key"
    }
}

Kennel

tricks/kennel.py is a reference implementation of the pattern above. It routes cognitive subtasks to three specialized models running in parallel - a thinker for chain-of-thought, a tool-caller for deciding which tools to invoke, and an emitter for generating the final response.

# Pull three small models that together fit on modest hardware (< 6B total)
ollama pull VibeThinker-3B    # reasoning / chain-of-thought
ollama pull LFM2.5-230M       # tool-calling (tiny, fast)
ollama pull Qwen3.5-2B        # response generation

# Each model sees a context optimized for its role
pet new kennel-demo
pet add kennel-demo kennel

# KennelTrick needs three model roles; scope them to this trickset
pet model thinker  url http://localhost:11434 --trickset kennel-demo
pet model thinker  model VibeThinker-3B-GGUF:q4_K_M --trickset kennel-demo
pet model toolcall url http://localhost:11434 --trickset kennel-demo
pet model toolcall model LFM2.5-230M --trickset kennel-demo
pet model default  url http://localhost:11434 --trickset kennel-demo
pet model default  model Qwen3.5-2B --trickset kennel-demo

The Models tab in the dashboard does the same thing with fewer keystrokes.

Pipeline:

  1. Thinker gets the conversation + "think step by step" → produces reasoning
  2. Tool-caller (if tools are present) gets context + reasoning + tool definitions → decides which tool to call
  3. Emitter receives the enriched context and generates the final response

Kennel is one architecture; you could write a trick that routes by language, by file type, by user role, or by anything else you can express in a post_hook.

Multi-Model Consultant

tricks/multiconsult.py

Cross-validates responses between two models through iterative refinement and voting. Requires a default model and a consultant model in the modelset.

Pipeline per round:

  1. model1's response (from the proxy call) is sent to model2 for improvement
  2. model2 generates a fresh response to the original prompt
  3. model1 improves model2's fresh response
  4. Both models vote on which improved output is better
  5. If they agree, return the winner; if not, repeat once more
  6. On second disagreement, randomly pick one as fallback
pet new consult
pet add consult multiconsult
pet model consultant url http://localhost:11434 --trickset consult
pet model consultant model qwen3:8b --trickset consult

Example modelset.json:

{
    "default": {
        "url": "http://localhost:11434",
        "model": "llama3:8b"
    },
    "consultant": {
        "url": "http://localhost:11434",
        "model": "qwen3:8b"
    }
}

Secrets Protector

tricks/secrets_protector.py

Detects and pseudonymizes sensitive information before it reaches the model, then restores original values in the response:

  • Detection - regex patterns for API keys (OpenAI, Anthropic, AWS, Google, Stripe), tokens (JWT, GitHub, Slack, Bearer), credentials (database URLs, private keys), and PII (emails, phones, SSNs, credit cards, IPs)
  • Format-preserving substitutes - realistic replacements (e.g., alice@example.comuser.0001@sanitized.local) that preserve token boundaries so the model's tokenizer doesn't conflate distinct entries
  • Bidirectional vault - consistent pseudonyms across the session (same secret → same substitute) with automatic restoration in both natural-language responses and tool call arguments
pet add mine secrets_protector

Swap Harness

tricks/swapharness.py

Browses and swaps system prompts from the system-prompts-and-models-of-ai-tools repository. On first use, it clones the repo into ~/.config/petsitter/harnesses/.

Use the swapharness prompt keyword to navigate the directory tree and select a system prompt file. The selected content is injected into the system prompt on every request until a different file is chosen or the trick is uninstalled.

pet add mine swapharness    # adding it runs install(), which clones the repo

Once installed, include (swapharness: path) in any user message to browse or select a harness:

User: (swapharness: Cursor Prompts)
Assistant: 📁 Cursor Prompts
           📄 Rules for All Models.md
           📄 Rules for Cursor.md
           📄 ...

User: (swapharness: Cursor Prompts/Rules for All Models.md)
Assistant: ✅ Harness set to Cursor Prompts/Rules for All Models.md (2847 chars)
           ────────────────────────────────────────────────
           You are Cursor, an advanced AI coding assistant...

The selected system prompt is prepended to every subsequent request. Run (swapharness: install) to clone the repo, or use the lifecycle CLI:

pet install swapharness             # clone the repo
pet uninstall swapharness           # remove the repo
pet lifecycle swapharness startup   # init per-session state
pet lifecycle swapharness shutdown  # cleanup session

Self-Improver

tricks/self_improver.py

Watches for the prompt keyword petsitter in your messages. When it sees (petsitter: <request>), it strips the tag and spawns an agent loop with the default model. The agent has tools to add, modify, and list trick files - it reads instructions from .agents/skills/self-improver/SKILL.md to understand the petsitter trick API and conventions.

This is a reference implementation for the prompt keywords pattern (see below).

pet add mine self_improver

Example usage:

User: (petsitter: add a trick that logs every request to a file)
Model: Creates tricks/request_logger.py and explains how to load it
User: explain the CAP theorem (petsitter: add a thinking mode)
Model: Explains CAP theorem (tag stripped, petsitter handled separately)

Export It

tricks/exportit.py

Exports the conversation history as an llcat-compatible JSON file. The output is the raw message array format used by OpenAI-compatible APIs, making it interoperable with llcat, prompt tools, and anything that speaks the Chat Completions message schema.

Use the exportit prompt keyword in any message to trigger the export:

pet add mine exportit
User: (exportit)
Assistant: Conversation exported to `/tmp/petsitter/convo-20260718-143022.json` (6 messages, llcat-compatible)

User: (exportit: backup before refactor)
Assistant: Conversation exported to `/tmp/petsitter/convo-20260718-143022.json` (6 messages, llcat-compatible)
Note: backup before refactor

The exported JSON is a plain array of messages in OpenAI Chat Completions format:

[
  { "role": "system", "content": "You are a helpful assistant." },
  { "role": "user", "content": "What is the CAP theorem?" },
  { "role": "assistant", "content": "The CAP theorem states...", "tool_calls": [] },
  { "role": "user", "content": "Can you give an example?" },
  { "role": "assistant", "content": "Sure! Consider a distributed..." }
]

Tool calls, reasoning (chain-of-thought), and tool results are all preserved in their standard formats. You can load the exported file directly with llcat -c convo.json or pipe it into any OpenAI-compatible tool.

Rules File

tricks/rules_file.py

Reads a plain-markdown rules file (AGENTS.md / CLAUDE.md style) and injects its content into the system prompt on every request. Because petsitter sits in front of any tool pointed at it, the same rules file applies across opencode, Claude Code, Codex, etc. - write the rules once and keep every harness consistent.

The rules path is configured per-trickset (the scope where petsitter config lives): set the rules_path config field on the trick via the dashboard, or switch files at runtime with the rules prompt keyword:

pet add mine rules_file
User: (rules: /path/to/rules.md)
Assistant: Loaded 123 chars of rules from /path/to/rules.md

User: (rules)
Assistant: Rules loaded from /path/to/rules.md (123 chars)

Content is cached and reloaded when the path changes, on startup, or on request. With no path configured the trick stays dormant, so requests pass through untouched.

Tricksets

A trickset bundles a group of tricks with routing filters. When a request comes in, petsitter matches the X-Title header and model field against each loaded trickset's filters, then runs only the tricks from matching sets.

Tricksets live as JSON files in the tricksets/ directory:

{
  "schema": "0.8.0",
  "name": "my-trickset",
  "filters": {
    "X-Title": "opencode*",
    "Model": "*"
  },
  "tricks": [
    "tricks/json_mode.py",
    "tricks/tool_call.py",
    "pkg:dana/ollama-ctx@0.1.0"
  ],
  "parameters": {},
  "models": {},
  "logfile": "~/.cache/petsitter/tricksets/my-trickset.log",
  "loglevel": "INFO"
}

Entries are either a path to a .py (absolute, or relative to the repo root) or a pkg:<owner>/<slug>@<version> spec pointing at a community trick you've installed.

The parameters field stores user-defined variables that tricks within the trickset can reference at runtime. The models field lets you override model routing for this trickset - each key maps to a {url, model, key} object (same format as the global model config), letting different tricksets use different models for the same role. Set model or key to false for passthrough. Manage both via the dashboard or the API.

Each loaded trickset is also exposed as a model named trickset/<name> (e.g., trickset/gemma4). Selecting this model in a client bypasses the filter matching and runs that trickset's tricks directly on every request.

The Models tab in the dashboard lets you configure model overrides per-trickset: select a trickset pill, then edit the model URL and name for each role. These overrides are stored in the trickset's models field and take precedence over the global model config when a trickset's tricks are running.

Using tricksets

pet new opencode --x-title 'opencode*' -t json_mode -t tool_call
petsitter

Every trickset in <config>/tricksets/ is loaded at startup, so there is nothing to pass on the command line. pet ts lists what you have.

Managing tricksets at runtime

The control panel at / has a full trickset manager. You can also use the API:

# List loaded tricksets
curl http://localhost:8080/api/tricksets

# List available trickset files
curl http://localhost:8080/api/tricksets/available

# Load a trickset
curl -X POST http://localhost:8080/api/tricksets/load \
  -d '{"path": "tricksets/gemma4.json"}'

# Update filters
curl -X PUT http://localhost:8080/api/tricksets/opencode \
  -d '{"filters": {"X-Title": "myagent*", "Model": "*"}}'

# Update model overrides for a trickset
curl -X PUT http://localhost:8080/api/tricksets/gemma4 \
  -d '{"models": {"default": "http://localhost:11434#m=llama3:8b", "toolcall": "http://localhost:11434#m=lfm2.5:latest"}}'

# Unload a trickset
curl -X POST http://localhost:8080/api/tricksets/unload \
  -d '{"name": "opencode"}'

How routing works

  1. Extract X-Title from the request header and model from the request body.
  2. For each loaded trickset, check if its filters match using fnmatch.
  3. Collect tricks from all matching sets, deduplicating by class name.
  4. Run the pipeline with only those tricks.

A trickset created without filters matches {"X-Title": "*", "Model": "*"}, so it acts as a catch-all and its tricks run on every request.

The schema field in a trickset JSON file records the petsitter version that wrote it. This tells tools how to interpret the file without needing an external lookup table.

Logging

Each trickset has its own log file so you can inspect what a specific set of tricks did. The logfile field sets the path (default ~/.cache/petsitter/tricksets/<name>.log) and loglevel sets the verbosity - DEBUG, INFO, WARNING, or ERROR (default INFO). Both are optional; if omitted, the defaults apply. Configure them from the Settings tab in the dashboard or via the API:

curl -X PUT http://localhost:8080/api/tricksets/my-trickset \
  -d '{"logfile": "~/.cache/petsitter/my-trickset.log", "loglevel": "DEBUG"}'

Every request through the pipeline is tagged with a short correlation id so you can follow it end-to-end. The tag appears in the matched trickset's log file and in the global activity log (Logs tab / GET /api/logs):

[ab12cd34] trickset 'gemma4' matched (X-Title='*' Model='gemma4*')
[ab12cd34] started multiround.py (run 0 -> 1)
[ab12cd34] calling upstream http://localhost:11434/v1/chat/completions model='gemma4'

Lifecycle events (install / uninstall / startup / shutdown) are written to the owning trickset's log file even when no request is running.

Agents

Petsitter has a one-click setup wizard for routing popular coding tools through the proxy. When you click Set up on an agent card in the Agents tab, petsitter:

  1. Detects your credentials (API keys, config files)
  2. Creates a trickset with the right tricks for that tool
  3. Patches the tool's config file to point at http://localhost:8080
  4. Saves the original config so it can be restored on shutdown

The exit button in the top-right restores every tool's original configuration and shuts petsitter down.

Available agents

Agent Config mechanism What gets patched
OpenCode ~/.config/opencode/opencode.json Provider baseURL
Claude Code ~/.claude/settings.json ANTHROPIC_BASE_URL in env block
Codex ~/.codex/config.toml openai_base_url

Each agent saves your original config to ~/.config/petsitter/registry.json and restores it on unregister or shutdown.

Adding agents

New agents live in agents/ and subclass Agent from agents/__init__. See .agents/skills/petsitter-create-agent/SKILL.md for the template and conventions.

API

# List available agents with detect status
curl http://localhost:8080/api/agents

# Register an agent (creates trickset, patches config)
curl -X POST http://localhost:8080/api/agents/claude-code/register

# Unregister an agent (restores original config)
curl -X POST http://localhost:8080/api/agents/claude-code/unregister

# Get registry state
curl http://localhost:8080/api/agents/registered

# Shutdown and restore all configurations
curl -X POST http://localhost:8080/api/shutdown

Model Configs

A model config JSON file lets you run multi-model tricks like Kennel that need different models for different subtasks. Each key maps to a {url, model, key} object:

{
    "default": {
        "url": "http://localhost:11434",
        "model": "Qwen3.5:8b"
    },
    "thinker": {
        "url": "http://localhost:11434",
        "model": "VibeThinker-3B-GGUF:q4_K_M"
    },
    "toolcall": {
        "url": "http://localhost:11434",
        "model": "lfm2.5:latest",
        "key": "sk-custom-key"
    }
}

The "default" key sets the primary model, the one used when a trick doesn't ask for a specific role. Tricks declare what keys they need - for example, KennelTrick requires ["default", "thinker", "toolcall"]. If a key is missing, petsitter prints a helpful error with the expected format.

The model and key fields accept:

  • A string - use as the model name / API key in upstream requests.
  • false (boolean) - passthrough, don't set the field at all.
  • "" (empty string) - explicitly clear the value.

Edit these from the Models tab, or with pet model:

pet model                                  # show every role as JSON
pet model thinker url http://localhost:11434
pet model thinker model VibeThinker-3B-GGUF:q4_K_M
pet model toolcall key false               # passthrough: use the client's key
pet model consultant --remove

Add --trickset <name> to scope a role to one trickset instead of the global config; those overrides live in the trickset's models field and win while that trickset's tricks are running.

Failure Modes

No global infinite-loop protection

post_hook receives the full context and returns a (potentially modified) context. The framework calls post_hooks once per request - it does not loop them. However, if a trick calls callmodel inside its own loop (as JSON Mode and Code Validator do), that loop is the trick's responsibility. None of the built-in tricks have unbounded loops, and custom tricks should follow the same pattern.

Examples solution: bounded retry loops

Two tricks loop internally: JSON Mode and Code Validator. Both default to 3 attempts, configurable via __init__. After exhausting attempts they give the model's best-effort output back to the user - they don't hang or cascade.

# Both accept max_attempts:
trick = JsonModeTrick(max_attempts=5)
trick = CodeValidatorTrick(max_attempts=5)

Network failures are not retried

callmodel and callmodel_sync make a single HTTP request to the upstream - no retry, no backoff. If the upstream is down, the error propagates as a 502 to the client. Add retry at the client level or wrap callmodel in your own try/except inside the trick. Errors are surfaced cleanly and thus easy to deal with.

Tool calls are client-driven

When a trick produces tool_calls in the response, petsitter returns them to your application. It does not execute the tool or re-invoke the model with the result - that's the client's job. If the client sends back a tool role message with the result, it enters the pipeline fresh on the next request.

Kennel sub-model failures

If a sub-model call in Kennel fails (e.g., the thinker model is unreachable), the exception propagates and the request fails. Kennel has no fallback - if you need resilience, wrap individual callmodel_sync calls in your own try/except.

API Endpoints

Petsitter exposes OpenAI-compatible endpoints plus management endpoints:

Proxy:

  • POST /v1/chat/completions - Chat completions (proxied + transformed)
  • GET /v1/models - List available models (proxied)
  • GET /health - Health check
  • * /p/{host}/{path} - Zero-config transparent proxy to https://{host}/{path} (any method)

Management:

  • GET /api/info - Server information
  • GET /api/tricks - List loaded tricks
  • GET /api/tricks/available - List available trick modules
  • POST /api/tricks/load - Load a trick
  • POST /api/tricks/unload - Unload a trick
  • POST /api/tricks/reorder - Reorder loaded tricks
  • GET /api/logs - Activity log
  • GET /api/tricksets - List loaded tricksets
  • GET /api/tricksets/available - List available trickset files
  • POST /api/tricksets/load - Load a trickset
  • POST /api/tricksets/unload - Unload a trickset
  • GET /api/tricksets/{name} - Get trickset details
  • PUT /api/tricksets/{name} - Update trickset filters, tricks, parameters, models, or logging config (logfile / loglevel)

Community index:

  • GET /api/registry - Search the index (?q=, ?all=1, ?refresh=1); entries are annotated with the installed version
  • POST /api/registry/install - Install {name, version, trickset} and optionally wire it into a trickset
  • GET /api/registry/source - Source of a trick (?name=&version=), from disk if installed

Playground:

  • POST /api/playground - Run {messages, trickset} through the real pipeline; returns the reply plus a trace of which tricks ran which hooks

A Swagger UI is available at /docs and the OpenAPI spec at /static/openapi.json.

Running Tests

# Activate virtual environment
source .venv/bin/activate

# Install test dependencies
pip install -e ".[test]"

# Run tests
pytest tests/

Two of the suites drive a real browser and need Playwright:

pip install playwright && playwright install chromium
python tests/test_registry_e2e.py     # index parsing, checksums, pkg: loading
python tests/test_playground_e2e.py   # boots a server against a stub upstream

Example: Using with an Agentic Framework

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8080/v1",
    api_key="not-needed"
)

response = client.chat.completions.create(
    model="any-model-name",
    messages=[{"role": "user", "content": "List files in /tmp"}],
    tools=[{"type": "function", "function": {"name": "get_weather", "parameters": ...}}]
)

License

MIT

About

Model and task specific harnesses sitting between inference and agent

Resources

Stars

39 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages