Version Packages - #818
Closed
blove wants to merge 1 commit into
Closed
Version Packages#818blove wants to merge 1 commit into
blove wants to merge 1 commit into
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 23, 2026 18:26
d888532 to
b9ec188
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 23, 2026 19:14
b9ec188 to
c3edc16
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 23, 2026 20:36
c3edc16 to
0a447ea
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 23, 2026 21:12
0a447ea to
e9232d3
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 23, 2026 21:20
e9232d3 to
a1c5b8a
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 24, 2026 01:20
a1c5b8a to
cf77c6e
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 24, 2026 04:17
cf77c6e to
863f2df
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 24, 2026 21:10
863f2df to
3b526e1
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 24, 2026 21:47
3b526e1 to
3897578
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 00:53
3897578 to
153f51b
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 02:43
153f51b to
c205ac6
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 03:19
c205ac6 to
d5df6e1
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 03:30
d5df6e1 to
260283e
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 08:11
8702994 to
339e526
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 09:34
339e526 to
252dd08
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 11:09
252dd08 to
a10dd69
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 15:28
a10dd69 to
8407861
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 17:06
8407861 to
7aaed99
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 17:07
7aaed99 to
a4106c8
Compare
blove
added a commit
to liveloveapp/hashbrown
that referenced
this pull request
Sep 25, 2026
…f rewriting them (#604) Pulls in the fix for cacheplane/b4run#778, which this example cannot get from a version bump. ## Why a bump is not enough #778 was fixed upstream by cacheplane/b4run#844: B4 now refuses at record time a recording that replay would reject, rather than writing a tape the next replay cannot open. It is merged but unreleased; 0.12.1 carries it and is waiting in cacheplane/b4run#818. Even once published, it would not reach this example. The invoicing eval harness keys its own fixtures in `server/evals/fixtures.ts` and overrides `getRecordedFixtures()`, because B4's built-in re-keying numbers turns by position and sets no `sequenceIndex`, which would let the LLM judge's prompt be answered by the app's first fixture. So #844's check never runs here. ## What changed - **Refuse, don't rewrite.** The record path used to store an assistant message with empty content as `{}`. #844 explains why that is the wrong call: aimock writes empty content for a genuinely empty reply, a refusal, and a truncated reply alike, so a tape cannot say which happened. `assertReplayable` now loads each tape through aimock's own validator before it is written. A refused case is reported, no file is written for it, and `--record` exits non-zero. Responses are kept exactly as recorded. - **Drop nine dead closing fixtures.** These tapes predate `render` becoming `returnDirect` and ended with the model's closing turn, stored as `{}` by the old rewrite. Since B4 0.12.0 a successful render ends the run, so that turn is never requested. Pure removals, 99 lines, no recorded response altered. - README and the B4 findings doc no longer describe the rewrite as current. ## Verification - Eval replay: 10 of 10 cases, mean 1.00, gate passed, with the trimmed tapes. - Every committed tape passes `assertReplayable`. - New tests written first and seen failing: responses are kept verbatim, a tape with an empty assistant message is refused naming the turn, a valid tape is accepted. - `nx run-many -t build,test,lint -p invoicing-server`: 248 tests pass, 0 lint errors, 23 pre-existing warnings. The version bump to 0.12.1 can follow once it is on npm; it changes nothing this PR depends on. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 17:37
a4106c8 to
9d418e0
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 17:47
9d418e0 to
dca4ab3
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 19:30
dca4ab3 to
bce4e7b
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 19:40
bce4e7b to
6d0976c
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 20:16
6d0976c to
d1a1ab3
Compare
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 21:50
d1a1ab3 to
222a0e9
Compare
Closed
github-actions
Bot
force-pushed
the
changeset-release/main
branch
from
September 25, 2026 22:28
222a0e9 to
2e72d5d
Compare
Contributor
Author
|
Stale: its changesets were consumed by the 0.13.0 release (#857). The bot recreates this PR on the next changeset. |
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR was opened by the Changesets release GitHub action. When you're ready to do a release, you can merge this and publish to npm yourself or setup this action to publish automatically. If you're not ready to do a release yet, that's fine, whenever you add more changesets to main, this PR will be updated.
Releases
@b4run/cli@0.12.1
Patch Changes
1f335b1:
createAgentHarnessanddefineEvalaccept aresponseSchema: the JSON Schema a Hashbrown client sends ashashbrown.responseSchema. The harness binds it on the root model exactly as the server does for an AG-UI run, with the same validation and provider-facing name. Scripted, live and recorded runs then send the model the sameresponse_formatproduction sends. Before, a recording made through the harness ran unconstrained, so it could capture replies production never produces, such as an empty final message. A route or provider that cannot constrain its output fails the run instead of running unconstrained.@b4run/cli/runtimenow exportsreadResponseFormat, the server's parser for that field.c301d77: Managed workspaces no longer re-read and re-verify the thread's whole captured source bundle before every filesystem or exec call. The bundle is read from storage and verified once when a session is admitted, and later calls on that session reuse the verified, frozen copy while the association still names the same operation and source digest. A new session, after idle reaping, an abort, or a failed call, reads and verifies the stored bundle again, so a tampered bundle is still refused at admission.
readInitialFilealso indexes the verified bundle once instead of re-hashing it for every file. On a 951-file, 10 MB workspace a warm backend call drops from about 340 ms to about 0.1 ms.5260ecb: Security fix: the
nodebuild target now carries the app's thread access policy in its build. Previouslyb4 buildleftsrc/thread-access.tsout of.b4/build/modules.mjsand the generatedserver.mjsprobed the disk for it at boot, where a missing file reads as "no policy" — so a built node app whose policy file was absent at boot (deleted, or left out of the image) served every thread endpoint ungated. The policy is now a static import of the manifest andserver.mjsrecords that the build saw one, so a missing policy file or a manifest older than the policy makes the server refuse to start instead of serving open, matching thehonoandverceltargets. Rebuild node apps that have a thread access policy. The built manifest now imports the policy when it loads, including inb4 checkafterb4 build, so a policy that requires an environment variable at import time needs it there too;b4 checknow says so when an app module throws while the manifest loads.3b489a5: A
sandbox.threadresolver may returnpermissions: that thread's permission gates then use its own allow-list in place of the app's, keep the app's mode and every denial (the app's and the thread's), and save an "Always" decision to the thread's record in the workspace installation, never to.b4/permissions.jsonor a configuredpermissions.store. A subagent runs under its parent thread's permissions.createThreadPermissionsStorebuilds such a store over anyPermissionsStoreand passes the store conformance suite. Empty and whitespace-only permission patterns are refused in a thread's lists. An "Always" whose pattern the thread's record cannot hold (longer thanMAX_THREAD_GRANT_LENGTH, empty or containing NUL) allows that call once with a warning instead of failing the run.0dd8fff:
sandbox.threaddecides each thread's whole sandbox (workspace, image and policy) once, at the thread's first admission. The image goes through the new optionalManagedWorkspaceProvider.resolveImageEnvironment, and its identity is recorded in the thread's creation intent; the image reference and the policy overrides are recorded beside the association in the same transaction, and every reconnect runs the thread's recorded policy over the app's.dockerSandbox({ images })bounds which images a thread may name and refuses anything else before any Docker call;imageis optional whenimagesis given. A thread may not open a network the app denies, andsecuritystays per app.b4 check,b4 buildand startup refuse unknownsandboxkeys andthreadbesideworkspace; a thread-sandbox app builds to a{ version: 2, kind: "thread" }artifact. The installation now stores a workspace's source only after its environment resolves, so a refused image leaves no source behind.Behaviour change:
b4 check,b4 buildand startup now refuse any key in thesandboxblock other thanworkspace,thread,provider,network,env,resources,securityandidleTimeoutMs. A misspelt key used to be ignored silently, which left every thread in a per-app sandbox; rename or remove any other key.90b68be: Recording now fails at record time when replay would reject the recording. Previously
getRecordedFixtures()returned whatever aimock captured, while replay loads fixtures through aimock's validator. A turn whose assistant message came back empty therefore produced a tape that the next replay refused withcontent is empty string.getRecordedFixtures()now applies the same validation and throws an error naming each rejected turn and its user message, with the likely causes.b4 eval --recordreports it asRefused to record <eval> › <case>with exit code 2 and does not write that case's fixture file. If the run's answer is a tool call,returnDirecton that tool ends the run on its result, so no empty closing turn is recorded.1da86ae: A sandbox with no
networksetting now defaults to{ mode: "allow" }. The previous default was{ mode: "allow", denylist: ["169.254.169.254"] }, but neither the Docker nor the Kubernetes provider enforces an allow-modedenylist. The entry claimed a block on the cloud metadata endpoint that never took effect, and runtime behavior is unchanged: allow-mode egress was open before and is open now. On a cloud VM, setnetwork: { mode: "deny" }when the sandbox does not need the network. Otherwise block the endpoint outside B4.run, with a host firewall or egress proxy for Docker, or with theb4-sandbox-infrachart's default-deny egress backstop or your own NetworkPolicy for Kubernetes. TheSandboxPolicy.networkJSDoc and the sandbox and configuration docs now say which lists each reference provider ignores.79c5f63: Hand a thread its workspace at creation.
sandbox.stagedWorkspacesservesPUT /workspace/sources/:digest(a content-addressedSourceBundleupload, verified against its digest, one at a time per process with429 upload_in_flightto a second, withinuploadTimeoutMs, default 120 s,408past it, and withinmaxStagedBytes, default 1 GiB,507past it) and acceptsworkspace: { sourceDigest, environmentLinks?, baseline? }onPOST /threads, which may name only an uploaded source (never one an admission stored for another thread) and is checked against the upload's recorded file paths without re-reading its bytes; the app's resolver (sandbox.thread, or a functionsandbox.workspace) receives it asthread.stagedat the thread's first admission, and sources nothing references are reclaimed afterretentionMs(default 24 hours), at boot and before each upload. The option needs a resolver and a thread-access policy;b4 check,b4 buildand boot refuse it otherwise.Behaviour changes:
ThreadAccessRequest.requestedWorkspaceis a new required field (ThreadAccessRequestedWorkspace | undefined, exported from@b4run/sdk):{ sourceDigest }on the newworkspace.source.putoperation (acreatewith no thread) and the whole reference plusuploadedByon athread.createthat names a workspace,undefinedeverywhere else. A stamp a policy returns onworkspace.source.putis kept as the upload's uploader and handed back inuploadedBy, so a policy can require a caller to choose only what it uploaded. Code that builds aThreadAccessRequestby hand needsrequestedWorkspace: undefined;@b4run/testing'screateThreadAccessHarnessaccepts it on a check.ThreadOperationgains"workspace.source.put", so an exhaustiveswitchover it needs a case. EnablingstagedWorkspacesmeans auditing the policy'screatehandler.POST /threadsrefuses aworkspacefield it will not serve (400 workspace_not_accepted, after the policy's decision) instead of ignoring it, in every app. In an app withstagedWorkspacesit also refuses a body over 1 MiB (413 payload_too_large, after a policy that refuses the caller has answered), and runs at most four creates naming a workspace at once (429 workspace_create_in_flight); every other app reads its create body as before.DELETE /threads/:thread_idforgets the thread's staged workspace before its row, and boot forgets the staged workspace of any thread whose row is gone.fcf6d83: Read a thread's workspace over HTTP.
sandbox.workspaceRead: "http"servesPOST /threads/:thread_id/workspace/inspect, a bounded read-only inventory of a thread's managed workspace, authorized by the app's thread-access policy as the newthread.workspaceoperation;b4 check,b4 buildand boot refuse it without a policy.sandbox.workspaceReadTimeoutMsbounds one read (default 120 s).readThreadWorkspacein@b4run/cli/workspaceis the client.inspectWorkspacegainsrootand throwsWorkspaceInspectionErrorwith a code (invalid_options,root_missing,refused,changed); B4.run's bounded reads throwWorkspaceReadLimitErrorwith their existing messages, and the Docker and Kubernetes batched walk's entry-limit refusal is aWorkspaceInspectionError.ThreadOperationgains a member, so an exhaustiveswitchover it needs a case. On Node, a request body a handler stops reading part-way is now discarded rather than resetting the connection, so a refusal such as a 413 reaches the client.Updated dependencies [f2ee6cf]
Updated dependencies [3b489a5]
Updated dependencies [0dd8fff]
Updated dependencies [1da86ae]
Updated dependencies [79c5f63]
Updated dependencies [0d06d72]
Updated dependencies [fcf6d83]
@b4run/core@0.12.1
Patch Changes
0d06d72: Workspace agents get an
editFiletool and rangedreadFile.editFile({ path, oldText, newText, replaceAll? })replaces an exact span of an existing file through the same permission-gated handle asreadFileandwriteFile, and refuses whenoldTextis missing or ambiguous (overlapping matches count) instead of guessing. It edits UTF-8 files only and refuses any other file without touching it.readFileaccepts optional 1-based, inclusivestartLine/endLineand prefixes a ranged read with a[<path> lines a-b of N]header; a read without a range is unchanged. The tool descriptions now steer models to read large files in ranges and change them witheditFilerather than rewriting them in full withwriteFile, which risked truncating large files.Tool scoping: a route that denies
writeFilealso loseseditFile, so existingdeny: ["writeFile"]routes stay read-only. NameeditFileinallowto keep it. An allow-list namingwriteFiledoes not granteditFile.b4 checkwarns whenapprovenames aneditFilethat awriteFiledeny withholds.Updated dependencies [f2ee6cf]
Updated dependencies [3b489a5]
Updated dependencies [0dd8fff]
Updated dependencies [1da86ae]
Updated dependencies [79c5f63]
Updated dependencies [fcf6d83]
create-b4-app@0.12.1
Patch Changes
@b4run/devkit@0.12.1
Patch Changes
EPIPE. The Docker client and the devkit test process helper now ignoreEPIPEon the child's stdin, where the exit status already reports the failure, and still surface any other stdin error. Batched workspace reads send their scripts over stdin, which made this reachable.@b4run/evals@0.12.1
Patch Changes
createAgentHarnessanddefineEvalaccept aresponseSchema: the JSON Schema a Hashbrown client sends ashashbrown.responseSchema. The harness binds it on the root model exactly as the server does for an AG-UI run, with the same validation and provider-facing name. Scripted, live and recorded runs then send the model the sameresponse_formatproduction sends. Before, a recording made through the harness ran unconstrained, so it could capture replies production never produces, such as an empty final message. A route or provider that cannot constrain its output fails the run instead of running unconstrained.@b4run/cli/runtimenow exportsreadResponseFormat, the server's parser for that field.@b4run/inspector@0.12.1
Patch Changes
@b4run/langchain@0.12.1
Patch Changes
@b4run/langgraph@0.12.1
Patch Changes
@b4run/memory@0.12.1
Patch Changes
@b4run/memory-pgvector@0.12.1
Patch Changes
@b4run/permissions@0.12.1
Patch Changes
sandbox.threadresolver may returnpermissions: that thread's permission gates then use its own allow-list in place of the app's, keep the app's mode and every denial (the app's and the thread's), and save an "Always" decision to the thread's record in the workspace installation, never to.b4/permissions.jsonor a configuredpermissions.store. A subagent runs under its parent thread's permissions.createThreadPermissionsStorebuilds such a store over anyPermissionsStoreand passes the store conformance suite. Empty and whitespace-only permission patterns are refused in a thread's lists. An "Always" whose pattern the thread's record cannot hold (longer thanMAX_THREAD_GRANT_LENGTH, empty or containing NUL) allows that call once with a warning instead of failing the run.@b4run/postgres-storage@0.12.1
Patch Changes
@b4run/sandbox@0.12.1
Patch Changes
f2ee6cf:
inspectWorkspaceno longer makes one backend call per entry when the backend can batch. Filesystem backends gain two optional methods.walkTreereturns every entry below a directory withlstatmetadata in one call, andreadBinaryFilesreads many files withreadBinaryFile's per-filemaxBytesbound. When a backend has both, inspection walks once, reads in a few calls, and applies exactly the same name, kind, size, and budget checks as before. A file that grew between the walk and the read is refused.The Docker backend and the Docker workspace reader implement both methods. Inspecting a container workspace now costs a few
docker execcalls instead of one perlistDir,lstat, and read. Over 1,214 entries that took inspection from 143 s to about 1 s. The walk is bounded inside the container, names travel as exact bytes, and a walk over the entry limit fails rather than being truncated.withFilesystemLoggingpasses both methods through. Other backends are unchanged and keep the per-entry path.0781125: The Kubernetes sandbox backend now implements
walkTreeandreadBinaryFiles, soinspectWorkspaceover a pod's workspace costs a few execs instead of one perlistDir,lstat, and read. Over 1,212 entries on a kind cluster, batched inspection took under a second, where each per-entry exec cost about 80 ms. The batch read's script goes tosh -son stdin rather than in argv, because a Kubernetes exec carries its command in the request URL, where an API server or proxy may cap the length. Batch read scripts also redirect each read's stdin from/dev/null, so no command in a script fed on stdin can consume the rest of it.0dd8fff:
sandbox.threaddecides each thread's whole sandbox (workspace, image and policy) once, at the thread's first admission. The image goes through the new optionalManagedWorkspaceProvider.resolveImageEnvironment, and its identity is recorded in the thread's creation intent; the image reference and the policy overrides are recorded beside the association in the same transaction, and every reconnect runs the thread's recorded policy over the app's.dockerSandbox({ images })bounds which images a thread may name and refuses anything else before any Docker call;imageis optional whenimagesis given. A thread may not open a network the app denies, andsecuritystays per app.b4 check,b4 buildand startup refuse unknownsandboxkeys andthreadbesideworkspace; a thread-sandbox app builds to a{ version: 2, kind: "thread" }artifact. The installation now stores a workspace's source only after its environment resolves, so a refused image leaves no source behind.Behaviour change:
b4 check,b4 buildand startup now refuse any key in thesandboxblock other thanworkspace,thread,provider,network,env,resources,securityandidleTimeoutMs. A misspelt key used to be ignored silently, which left every thread in a per-app sandbox; rename or remove any other key.1da86ae: A sandbox with no
networksetting now defaults to{ mode: "allow" }. The previous default was{ mode: "allow", denylist: ["169.254.169.254"] }, but neither the Docker nor the Kubernetes provider enforces an allow-modedenylist. The entry claimed a block on the cloud metadata endpoint that never took effect, and runtime behavior is unchanged: allow-mode egress was open before and is open now. On a cloud VM, setnetwork: { mode: "deny" }when the sandbox does not need the network. Otherwise block the endpoint outside B4.run, with a host firewall or egress proxy for Docker, or with theb4-sandbox-infrachart's default-deny egress backstop or your own NetworkPolicy for Kubernetes. TheSandboxPolicy.networkJSDoc and the sandbox and configuration docs now say which lists each reference provider ignores.03795da: A Docker command that exits before reading its stdin no longer crashes the process with an uncaught
EPIPE. The Docker client and the devkit test process helper now ignoreEPIPEon the child's stdin, where the exit status already reports the failure, and still surface any other stdin error. Batched workspace reads send their scripts over stdin, which made this reachable.fcf6d83: Read a thread's workspace over HTTP.
sandbox.workspaceRead: "http"servesPOST /threads/:thread_id/workspace/inspect, a bounded read-only inventory of a thread's managed workspace, authorized by the app's thread-access policy as the newthread.workspaceoperation;b4 check,b4 buildand boot refuse it without a policy.sandbox.workspaceReadTimeoutMsbounds one read (default 120 s).readThreadWorkspacein@b4run/cli/workspaceis the client.inspectWorkspacegainsrootand throwsWorkspaceInspectionErrorwith a code (invalid_options,root_missing,refused,changed); B4.run's bounded reads throwWorkspaceReadLimitErrorwith their existing messages, and the Docker and Kubernetes batched walk's entry-limit refusal is aWorkspaceInspectionError.ThreadOperationgains a member, so an exhaustiveswitchover it needs a case. On Node, a request body a handler stops reading part-way is now discarded rather than resetting the connection, so a refusal such as a 413 reaches the client.Updated dependencies [f2ee6cf]
Updated dependencies [3b489a5]
Updated dependencies [0dd8fff]
Updated dependencies [1da86ae]
Updated dependencies [79c5f63]
Updated dependencies [fcf6d83]
@b4run/sdk@0.12.1
Patch Changes
79c5f63: Hand a thread its workspace at creation.
sandbox.stagedWorkspacesservesPUT /workspace/sources/:digest(a content-addressedSourceBundleupload, verified against its digest, one at a time per process with429 upload_in_flightto a second, withinuploadTimeoutMs, default 120 s,408past it, and withinmaxStagedBytes, default 1 GiB,507past it) and acceptsworkspace: { sourceDigest, environmentLinks?, baseline? }onPOST /threads, which may name only an uploaded source (never one an admission stored for another thread) and is checked against the upload's recorded file paths without re-reading its bytes; the app's resolver (sandbox.thread, or a functionsandbox.workspace) receives it asthread.stagedat the thread's first admission, and sources nothing references are reclaimed afterretentionMs(default 24 hours), at boot and before each upload. The option needs a resolver and a thread-access policy;b4 check,b4 buildand boot refuse it otherwise.Behaviour changes:
ThreadAccessRequest.requestedWorkspaceis a new required field (ThreadAccessRequestedWorkspace | undefined, exported from@b4run/sdk):{ sourceDigest }on the newworkspace.source.putoperation (acreatewith no thread) and the whole reference plusuploadedByon athread.createthat names a workspace,undefinedeverywhere else. A stamp a policy returns onworkspace.source.putis kept as the upload's uploader and handed back inuploadedBy, so a policy can require a caller to choose only what it uploaded. Code that builds aThreadAccessRequestby hand needsrequestedWorkspace: undefined;@b4run/testing'screateThreadAccessHarnessaccepts it on a check.ThreadOperationgains"workspace.source.put", so an exhaustiveswitchover it needs a case. EnablingstagedWorkspacesmeans auditing the policy'screatehandler.POST /threadsrefuses aworkspacefield it will not serve (400 workspace_not_accepted, after the policy's decision) instead of ignoring it, in every app. In an app withstagedWorkspacesit also refuses a body over 1 MiB (413 payload_too_large, after a policy that refuses the caller has answered), and runs at most four creates naming a workspace at once (429 workspace_create_in_flight); every other app reads its create body as before.DELETE /threads/:thread_idforgets the thread's staged workspace before its row, and boot forgets the staged workspace of any thread whose row is gone.fcf6d83: Read a thread's workspace over HTTP.
sandbox.workspaceRead: "http"servesPOST /threads/:thread_id/workspace/inspect, a bounded read-only inventory of a thread's managed workspace, authorized by the app's thread-access policy as the newthread.workspaceoperation;b4 check,b4 buildand boot refuse it without a policy.sandbox.workspaceReadTimeoutMsbounds one read (default 120 s).readThreadWorkspacein@b4run/cli/workspaceis the client.inspectWorkspacegainsrootand throwsWorkspaceInspectionErrorwith a code (invalid_options,root_missing,refused,changed); B4.run's bounded reads throwWorkspaceReadLimitErrorwith their existing messages, and the Docker and Kubernetes batched walk's entry-limit refusal is aWorkspaceInspectionError.ThreadOperationgains a member, so an exhaustiveswitchover it needs a case. On Node, a request body a handler stops reading part-way is now discarded rather than resetting the connection, so a refusal such as a 413 reaches the client.@b4run/sqlite-storage@0.12.1
Patch Changes
3b489a5: A
sandbox.threadresolver may returnpermissions: that thread's permission gates then use its own allow-list in place of the app's, keep the app's mode and every denial (the app's and the thread's), and save an "Always" decision to the thread's record in the workspace installation, never to.b4/permissions.jsonor a configuredpermissions.store. A subagent runs under its parent thread's permissions.createThreadPermissionsStorebuilds such a store over anyPermissionsStoreand passes the store conformance suite. Empty and whitespace-only permission patterns are refused in a thread's lists. An "Always" whose pattern the thread's record cannot hold (longer thanMAX_THREAD_GRANT_LENGTH, empty or containing NUL) allows that call once with a warning instead of failing the run.0dd8fff:
sandbox.threaddecides each thread's whole sandbox (workspace, image and policy) once, at the thread's first admission. The image goes through the new optionalManagedWorkspaceProvider.resolveImageEnvironment, and its identity is recorded in the thread's creation intent; the image reference and the policy overrides are recorded beside the association in the same transaction, and every reconnect runs the thread's recorded policy over the app's.dockerSandbox({ images })bounds which images a thread may name and refuses anything else before any Docker call;imageis optional whenimagesis given. A thread may not open a network the app denies, andsecuritystays per app.b4 check,b4 buildand startup refuse unknownsandboxkeys andthreadbesideworkspace; a thread-sandbox app builds to a{ version: 2, kind: "thread" }artifact. The installation now stores a workspace's source only after its environment resolves, so a refused image leaves no source behind.Behaviour change:
b4 check,b4 buildand startup now refuse any key in thesandboxblock other thanworkspace,thread,provider,network,env,resources,securityandidleTimeoutMs. A misspelt key used to be ignored silently, which left every thread in a per-app sandbox; rename or remove any other key.79c5f63: Hand a thread its workspace at creation.
sandbox.stagedWorkspacesservesPUT /workspace/sources/:digest(a content-addressedSourceBundleupload, verified against its digest, one at a time per process with429 upload_in_flightto a second, withinuploadTimeoutMs, default 120 s,408past it, and withinmaxStagedBytes, default 1 GiB,507past it) and acceptsworkspace: { sourceDigest, environmentLinks?, baseline? }onPOST /threads, which may name only an uploaded source (never one an admission stored for another thread) and is checked against the upload's recorded file paths without re-reading its bytes; the app's resolver (sandbox.thread, or a functionsandbox.workspace) receives it asthread.stagedat the thread's first admission, and sources nothing references are reclaimed afterretentionMs(default 24 hours), at boot and before each upload. The option needs a resolver and a thread-access policy;b4 check,b4 buildand boot refuse it otherwise.Behaviour changes:
ThreadAccessRequest.requestedWorkspaceis a new required field (ThreadAccessRequestedWorkspace | undefined, exported from@b4run/sdk):{ sourceDigest }on the newworkspace.source.putoperation (acreatewith no thread) and the whole reference plusuploadedByon athread.createthat names a workspace,undefinedeverywhere else. A stamp a policy returns onworkspace.source.putis kept as the upload's uploader and handed back inuploadedBy, so a policy can require a caller to choose only what it uploaded. Code that builds aThreadAccessRequestby hand needsrequestedWorkspace: undefined;@b4run/testing'screateThreadAccessHarnessaccepts it on a check.ThreadOperationgains"workspace.source.put", so an exhaustiveswitchover it needs a case. EnablingstagedWorkspacesmeans auditing the policy'screatehandler.POST /threadsrefuses aworkspacefield it will not serve (400 workspace_not_accepted, after the policy's decision) instead of ignoring it, in every app. In an app withstagedWorkspacesit also refuses a body over 1 MiB (413 payload_too_large, after a policy that refuses the caller has answered), and runs at most four creates naming a workspace at once (429 workspace_create_in_flight); every other app reads its create body as before.DELETE /threads/:thread_idforgets the thread's staged workspace before its row, and boot forgets the staged workspace of any thread whose row is gone.Updated dependencies [f2ee6cf]
Updated dependencies [3b489a5]
Updated dependencies [0dd8fff]
Updated dependencies [1da86ae]
Updated dependencies [79c5f63]
Updated dependencies [fcf6d83]
@b4run/testing@0.12.1
Patch Changes
81021b2:
@b4run/testingnow depends on@copilotkit/aimock^1.43.0(was^1.37.4, resolving to 1.38.0). Recording over the OpenAI Responses API now keeps tool calls. Before, a turn that only called a tool was recorded as an empty assistant message with the call dropped. Recorded fixtures now also carry the upstream's tokenusage, so a replay reports the recorded counts instead of a length-based estimate. Re-recording an existing tape therefore adds ausagefield to each response.1f335b1:
createAgentHarnessanddefineEvalaccept aresponseSchema: the JSON Schema a Hashbrown client sends ashashbrown.responseSchema. The harness binds it on the root model exactly as the server does for an AG-UI run, with the same validation and provider-facing name. Scripted, live and recorded runs then send the model the sameresponse_formatproduction sends. Before, a recording made through the harness ran unconstrained, so it could capture replies production never produces, such as an empty final message. A route or provider that cannot constrain its output fails the run instead of running unconstrained.@b4run/cli/runtimenow exportsreadResponseFormat, the server's parser for that field.90b68be: Recording now fails at record time when replay would reject the recording. Previously
getRecordedFixtures()returned whatever aimock captured, while replay loads fixtures through aimock's validator. A turn whose assistant message came back empty therefore produced a tape that the next replay refused withcontent is empty string.getRecordedFixtures()now applies the same validation and throws an error naming each rejected turn and its user message, with the likely causes.b4 eval --recordreports it asRefused to record <eval> › <case>with exit code 2 and does not write that case's fixture file. If the run's answer is a tool call,returnDirecton that tool ends the run on its result, so no empty closing turn is recorded.79c5f63: Hand a thread its workspace at creation.
sandbox.stagedWorkspacesservesPUT /workspace/sources/:digest(a content-addressedSourceBundleupload, verified against its digest, one at a time per process with429 upload_in_flightto a second, withinuploadTimeoutMs, default 120 s,408past it, and withinmaxStagedBytes, default 1 GiB,507past it) and acceptsworkspace: { sourceDigest, environmentLinks?, baseline? }onPOST /threads, which may name only an uploaded source (never one an admission stored for another thread) and is checked against the upload's recorded file paths without re-reading its bytes; the app's resolver (sandbox.thread, or a functionsandbox.workspace) receives it asthread.stagedat the thread's first admission, and sources nothing references are reclaimed afterretentionMs(default 24 hours), at boot and before each upload. The option needs a resolver and a thread-access policy;b4 check,b4 buildand boot refuse it otherwise.Behaviour changes:
ThreadAccessRequest.requestedWorkspaceis a new required field (ThreadAccessRequestedWorkspace | undefined, exported from@b4run/sdk):{ sourceDigest }on the newworkspace.source.putoperation (acreatewith no thread) and the whole reference plusuploadedByon athread.createthat names a workspace,undefinedeverywhere else. A stamp a policy returns onworkspace.source.putis kept as the upload's uploader and handed back inuploadedBy, so a policy can require a caller to choose only what it uploaded. Code that builds aThreadAccessRequestby hand needsrequestedWorkspace: undefined;@b4run/testing'screateThreadAccessHarnessaccepts it on a check.ThreadOperationgains"workspace.source.put", so an exhaustiveswitchover it needs a case. EnablingstagedWorkspacesmeans auditing the policy'screatehandler.POST /threadsrefuses aworkspacefield it will not serve (400 workspace_not_accepted, after the policy's decision) instead of ignoring it, in every app. In an app withstagedWorkspacesit also refuses a body over 1 MiB (413 payload_too_large, after a policy that refuses the caller has answered), and runs at most four creates naming a workspace at once (429 workspace_create_in_flight); every other app reads its create body as before.DELETE /threads/:thread_idforgets the thread's staged workspace before its row, and boot forgets the staged workspace of any thread whose row is gone.Updated dependencies [f2ee6cf]
Updated dependencies [1f335b1]
Updated dependencies [c301d77]
Updated dependencies [5260ecb]
Updated dependencies [3b489a5]
Updated dependencies [0dd8fff]
Updated dependencies [90b68be]
Updated dependencies [1da86ae]
Updated dependencies [79c5f63]
Updated dependencies [0d06d72]
Updated dependencies [fcf6d83]
@b4run/vite-plugin@0.12.1
Patch Changes
@b4run/workspace@0.12.1
Patch Changes
f2ee6cf:
inspectWorkspaceno longer makes one backend call per entry when the backend can batch. Filesystem backends gain two optional methods.walkTreereturns every entry below a directory withlstatmetadata in one call, andreadBinaryFilesreads many files withreadBinaryFile's per-filemaxBytesbound. When a backend has both, inspection walks once, reads in a few calls, and applies exactly the same name, kind, size, and budget checks as before. A file that grew between the walk and the read is refused.The Docker backend and the Docker workspace reader implement both methods. Inspecting a container workspace now costs a few
docker execcalls instead of one perlistDir,lstat, and read. Over 1,214 entries that took inspection from 143 s to about 1 s. The walk is bounded inside the container, names travel as exact bytes, and a walk over the entry limit fails rather than being truncated.withFilesystemLoggingpasses both methods through. Other backends are unchanged and keep the per-entry path.3b489a5: A
sandbox.threadresolver may returnpermissions: that thread's permission gates then use its own allow-list in place of the app's, keep the app's mode and every denial (the app's and the thread's), and save an "Always" decision to the thread's record in the workspace installation, never to.b4/permissions.jsonor a configuredpermissions.store. A subagent runs under its parent thread's permissions.createThreadPermissionsStorebuilds such a store over anyPermissionsStoreand passes the store conformance suite. Empty and whitespace-only permission patterns are refused in a thread's lists. An "Always" whose pattern the thread's record cannot hold (longer thanMAX_THREAD_GRANT_LENGTH, empty or containing NUL) allows that call once with a warning instead of failing the run.0dd8fff:
sandbox.threaddecides each thread's whole sandbox (workspace, image and policy) once, at the thread's first admission. The image goes through the new optionalManagedWorkspaceProvider.resolveImageEnvironment, and its identity is recorded in the thread's creation intent; the image reference and the policy overrides are recorded beside the association in the same transaction, and every reconnect runs the thread's recorded policy over the app's.dockerSandbox({ images })bounds which images a thread may name and refuses anything else before any Docker call;imageis optional whenimagesis given. A thread may not open a network the app denies, andsecuritystays per app.b4 check,b4 buildand startup refuse unknownsandboxkeys andthreadbesideworkspace; a thread-sandbox app builds to a{ version: 2, kind: "thread" }artifact. The installation now stores a workspace's source only after its environment resolves, so a refused image leaves no source behind.Behaviour change:
b4 check,b4 buildand startup now refuse any key in thesandboxblock other thanworkspace,thread,provider,network,env,resources,securityandidleTimeoutMs. A misspelt key used to be ignored silently, which left every thread in a per-app sandbox; rename or remove any other key.1da86ae: A sandbox with no
networksetting now defaults to{ mode: "allow" }. The previous default was{ mode: "allow", denylist: ["169.254.169.254"] }, but neither the Docker nor the Kubernetes provider enforces an allow-modedenylist. The entry claimed a block on the cloud metadata endpoint that never took effect, and runtime behavior is unchanged: allow-mode egress was open before and is open now. On a cloud VM, setnetwork: { mode: "deny" }when the sandbox does not need the network. Otherwise block the endpoint outside B4.run, with a host firewall or egress proxy for Docker, or with theb4-sandbox-infrachart's default-deny egress backstop or your own NetworkPolicy for Kubernetes. TheSandboxPolicy.networkJSDoc and the sandbox and configuration docs now say which lists each reference provider ignores.79c5f63: Hand a thread its workspace at creation.
sandbox.stagedWorkspacesservesPUT /workspace/sources/:digest(a content-addressedSourceBundleupload, verified against its digest, one at a time per process with429 upload_in_flightto a second, withinuploadTimeoutMs, default 120 s,408past it, and withinmaxStagedBytes, default 1 GiB,507past it) and acceptsworkspace: { sourceDigest, environmentLinks?, baseline? }onPOST /threads, which may name only an uploaded source (never one an admission stored for another thread) and is checked against the upload's recorded file paths without re-reading its bytes; the app's resolver (sandbox.thread, or a functionsandbox.workspace) receives it asthread.stagedat the thread's first admission, and sources nothing references are reclaimed afterretentionMs(default 24 hours), at boot and before each upload. The option needs a resolver and a thread-access policy;b4 check,b4 buildand boot refuse it otherwise.Behaviour changes:
ThreadAccessRequest.requestedWorkspaceis a new required field (ThreadAccessRequestedWorkspace | undefined, exported from@b4run/sdk):{ sourceDigest }on the newworkspace.source.putoperation (acreatewith no thread) and the whole reference plusuploadedByon athread.createthat names a workspace,undefinedeverywhere else. A stamp a policy returns onworkspace.source.putis kept as the upload's uploader and handed back inuploadedBy, so a policy can require a caller to choose only what it uploaded. Code that builds aThreadAccessRequestby hand needsrequestedWorkspace: undefined;@b4run/testing'screateThreadAccessHarnessaccepts it on a check.ThreadOperationgains"workspace.source.put", so an exhaustiveswitchover it needs a case. EnablingstagedWorkspacesmeans auditing the policy'screatehandler.POST /threadsrefuses aworkspacefield it will not serve (400 workspace_not_accepted, after the policy's decision) instead of ignoring it, in every app. In an app withstagedWorkspacesit also refuses a body over 1 MiB (413 payload_too_large, after a policy that refuses the caller has answered), and runs at most four creates naming a workspace at once (429 workspace_create_in_flight); every other app reads its create body as before.DELETE /threads/:thread_idforgets the thread's staged workspace before its row, and boot forgets the staged workspace of any thread whose row is gone.fcf6d83: Read a thread's workspace over HTTP.
sandbox.workspaceRead: "http"servesPOST /threads/:thread_id/workspace/inspect, a bounded read-only inventory of a thread's managed workspace, authorized by the app's thread-access policy as the newthread.workspaceoperation;b4 check,b4 buildand boot refuse it without a policy.sandbox.workspaceReadTimeoutMsbounds one read (default 120 s).readThreadWorkspacein@b4run/cli/workspaceis the client.inspectWorkspacegainsrootand throwsWorkspaceInspectionErrorwith a code (invalid_options,root_missing,refused,changed); B4.run's bounded reads throwWorkspaceReadLimitErrorwith their existing messages, and the Docker and Kubernetes batched walk's entry-limit refusal is aWorkspaceInspectionError.ThreadOperationgains a member, so an exhaustiveswitchover it needs a case. On Node, a request body a handler stops reading part-way is now discarded rather than resetting the connection, so a refusal such as a 413 reaches the client.Updated dependencies [79c5f63]
Updated dependencies [fcf6d83]
@b4run/ag-ui@0.12.1
@b4run/config-biome@0.12.1
@b4run/config-typescript@0.12.1
@b4-example/chat-server@0.0.49
Patch Changes
@b4-example/chat-web@0.0.22
Patch Changes
@b4-example/code-fixer-server@0.0.12
Patch Changes
@b4-example/memory@0.0.35
Patch Changes
@b4-example/research-server@0.0.31
Patch Changes
@b4-example/research-web@0.0.22
Patch Changes
@b4-example/software-factory-controller@0.0.6
Patch Changes
@b4-example/software-factory-drafter@0.0.5
Patch Changes
@b4-example/software-factory-server@0.0.8
Patch Changes