Repository navigation
release: v6.9.7 - #67
Conversation
- Effort selector (low/medium/high/max) for GLM 5.3, GLM 5.3 Flash, Gemini 3.8 Flash and DeepSeek V4.1 Flash: /effort slider (Enter saves, s for this session only), /effort <level>, ←/→ in /model, shown in the status bar. Saved per model in config.json and read per request, so every chat on the machine uses it. Sent as X-MATTERAI-REASONING-EFFORT. - Confirm before resuming a session with a cold prompt cache (idle 5+ min): shows the share of the weekly limit from /axoncode/usage/estimate, for any 100k+ token session or when the cost reaches 1%. Estimates are prefetched while the /resume picker is open. - Model and effort persist across restarts: the model catalog is cached in ~/.orbcode/models-cache.json and loaded before settings. - Compaction recovers from context overflow, trims the summary request to fit the window, counts unreported tool output, and retries after a failure on the next turn. - read_file caps characters (100k per file, 200k per call). - A replaced conversation (/new, /resume, sign-in/out, MCP reload) no longer streams into the next one.
There was a problem hiding this comment.
🧪 PR Review is completed: Solid release — the new compaction, effort-selector and resume-cost logic is well covered by tests. Three issues found: describeResumeCost omits the $ its own test expects (test:resume fails), the resume-cost estimate prices the currently selected model instead of the resumed session's model, and measuredMessages double-counts the assistant message. Reviewed src/api/models.ts, src/auth/auth.ts, src/config/settings.ts, src/tools/executors/files.ts, src/index.tsx, src/api/client.ts, src/api/headers.ts, src/ui/components/EffortPicker.tsx, src/ui/components/ModelPicker.tsx, src/ui/components/ResumeConfirm.tsx, src/ui/components/StatusBar.tsx, package.json, test/compaction.test.ts, test/effort-picker.test.tsx, test/files.test.ts, test/model-effort.test.ts, test/model-persistence.test.ts, test/model-picker.test.tsx, test/resume-confirm.test.tsx: no issues found. (test/resume-cost.test.ts is correct as written — the src/core/resumeCost.ts comment is the fix for its failing assertion.)
Skipped files
CHANGELOG.md: Skipped file patternpackage-lock.json: Skipped file pattern
⬇️ Low Priority Suggestions (2)
src/ui/App.tsx (1 suggestion)
Location:
src/ui/App.tsx(Lines 1069-1081)🟡 Logic
Issue: The resume-cost estimate prices the currently selected model (
getModel(loadSettings().model)), but the resumed agent runs on the session's own model —resumeSessioncallscreateAgent(session)andcreateAgentpassesmodelId: current.model(line 860). When the two differ,inputPrice,billsPlan()and the backend share query (getGatewayModelId(model)) all use the wrong model, so the confirmation prompt can show a wrong dollar cost or plan-window share, and the prefetched share cache is keyed for the wrong model.Fix: Resolve the model from
session.model— per session inside theprefetchResumeSharesloop, and inhandleResume.Impact: The resume confirmation reflects what the resumed session will actually bill.
- (sessions: SessionData[]) => { - const model = getModel(loadSettings().model); - for (const session of sessions.slice(0, 10)) { - const cold = estimateResumeCost(session, model); - if (cold) void fetchResumeShare(model, cold.contextTokens); - } - }, - [fetchResumeShare], - ); - - const handleResume = useCallback( - async (session: SessionData) => { - const model = getModel(loadSettings().model); + (sessions: SessionData[]) => { + for (const session of sessions.slice(0, 10)) { + const model = getModel(session.model); + const cold = estimateResumeCost(session, model); + if (cold) void fetchResumeShare(model, cold.contextTokens); + } + }, + [fetchResumeShare], + ); + + const handleResume = useCallback( + async (session: SessionData) => { + const model = getModel(session.model);
src/core/agent.ts (1 suggestion)
Location:
src/core/agent.ts(Lines 1640-1640)🟡 Logic
Issue: The usage report sets
measuredMessages = this.messages.lengthwhile the stream loop is still running, butcontextTokens(input + output) already covers the assistant message thatrunStepappends after the loop. That message lands at indexmeasuredMessages, soestimatedContextTokens()addsmessageTokenEstimate()for it on top of theoutputTokensalready included incontextTokens— double-counting it every step. The field's own doc says only "tool results, new user input" should be estimated, and the inflation makes auto-compaction (and the "Context is X% full" notice) fire earlier than the intended 80% threshold.Fix: Mark the about-to-be-appended assistant message as covered by the report:
this.messages.length + 1.Impact: The context estimate matches the report's real coverage; compaction triggers at the intended threshold instead of prematurely dropping history.
- this.measuredMessages = this.messages.length + this.measuredMessages = this.messages.length + 1
Adds a warmup request that primes the gateway prompt cache with a task's system prompt + tools (same taskId, 1 output token, X-AxonCode-Warmup header) on launch and /new, so the first real turn starts warm. MCP startup is awaited so the tool list matches the first turn. Shipped behind CACHE_WARMUP_ENABLED = false until it has more testing; with the flag off agents are still created lazily on the first message.
|
✅ Reviewed the changes: Solid prompt-cache warmup implementation: the buildRequest refactor is behavior-faithful, warmup is best-effort with correct abort propagation and re-entrancy guards (warmupController abort/clear, messages re-check after whenStarted), the McpManager.whenStarted() addition is clean, and everything ships behind CACHE_WARMUP_ENABLED=false — no issues found in the new code. Note: the three previously flagged items (src/core/resumeCost.ts dollar sign, src/ui/App.tsx resume-cost model resolution, src/core/agent.ts measuredMessages) are not part of this diff's hunks and could not be re-verified or anchored against it, so they are neither re-emitted nor marked resolved. |
There was a problem hiding this comment.
🧪 PR Review is completed: Clean header-label refactor — withHeaderModel nicely replaces the 11-line findIndex block in switchModel; one session-override semantics question to verify on headerModelLabel. Note (pre-existing, not from this diff): @modelcontextprotocol/sdk ^1.29.0 carries a HIGH advisory (GHSA-6qxp-vccf-f47h, fixed in 1.31.0) — worth bumping before this release ships.
⬇️ Low Priority Suggestions (1)
src/ui/App.tsx (1 suggestion)
Location:
src/ui/App.tsx(Lines 406-406)🟡 Business Logic (needs verification)
Issue:
headerModelLabelresolves the displayed effort viagetModelEffort(loadModelEfforts(), model). IfloadModelEfforts()returns only the persisted per-model picks from~/.orbcode/config.json(with session-only overrides layered on separately viawithSessionModelEfforts), then the header refresh added aftersetSessionModelEffort(...)— thes/ "this session only" path — re-renders the header with the persisted level instead of the override just picked, making that refresh a no-op for the session-only flow (and likewise for/modelswitches to a model that has a live session override).Fix: Resolve session-aware efforts before reading the level:
getModelEffort(withSessionModelEfforts(loadModelEfforts()), model). IfloadModelEfforts()already includes session overrides, no change is needed — please verify which semantics apply.Impact: Keeps the header MODEL line consistent with the effort actually sent on the next request.
- const effort = getModelEffort(loadModelEfforts(), model); + const effort = getModelEffort(withSessionModelEfforts(loadModelEfforts()), model);
Stage new releases under ~/.orbcode/versions/<v>/, verify them (--version, --self-test) and flip a 'current' pointer; the launcher hands off to the staged version and falls back to itself on any failure or a failed start. Releases auto-install after 24h on npm; opt out with ORBCODE_DISABLE_AUTOUPDATE=1 or "autoUpdates": false.
There was a problem hiding this comment.
🧪 PR Review is completed: Well-structured background-update design (stage → verify → flip pointer, with health/attempts markers and a cross-process lock); two hardening fixes needed in src/utils/autoUpdate.ts — a TOCTOU in the stale-lock takeover and an unvalidated registry latest reaching path/Windows-shell contexts. Reviewed bin/orbcode.js: no issues found. Reviewed bin/select-version.js: no issues found. Reviewed src/index.tsx: no issues found. Reviewed src/ui/App.tsx: no issues found. Reviewed test/auto-update.test.ts: no issues found.
Skipped files
README.md: Skipped file patternRELEASE.md: Skipped file pattern
⬇️ Low Priority Suggestions (2)
src/utils/autoUpdate.ts (2 suggestions)
Location:
src/utils/autoUpdate.ts(Lines 208-214)🟠 Concurrency
Issue: The stale-lock takeover in
acquireLockhas a TOCTOU window. Two processes can both read the same stale lock content and then both callfs.unlinkSync(lockPath)— the second unlink removes the fresh lock the first process just created with itswxwrite (which lands between the second process's read and its unlink). Both then succeed on their retrywriteFileSync(..., { flag: "wx" })and proceed intostageVersionconcurrently, running twonpm installs into the same<version>.tmpdirectory and breaking the lock's mutual exclusion. The release path already guards with a content check (fs.readFileSync(lockPath, "utf8") === token), but the stale-removal path does not.Fix: Capture the lock content when reading it for the staleness check, and only unlink if the file still contains that same content; otherwise
continueand retry the exclusivewxwrite on the next loop iteration.Impact: Restores the lock's mutual-exclusion guarantee so concurrent sessions can't race an install into the same staging directory.
- try { - const [pidText, atText] = fs.readFileSync(lockPath, "utf8").split(":"); - const pid = Number(pidText); - const at = Number(atText); - const stale = !isPidAlive(pid) || !Number.isFinite(at) || Date.now() - at > LOCK_STALE_MS; - if (!stale) return null; - fs.unlinkSync(lockPath); + try { + const raw = fs.readFileSync(lockPath, "utf8"); + const [pidText, atText] = raw.split(":"); + const pid = Number(pidText); + const at = Number(atText); + const stale = !isPidAlive(pid) || !Number.isFinite(at) || Date.now() - at > LOCK_STALE_MS; + if (!stale) return null; + if (fs.readFileSync(lockPath, "utf8") !== raw) continue; + fs.unlinkSync(lockPath);Location:
src/utils/autoUpdate.ts(Lines 470-471)🟡 Security (defense-in-depth)
Issue:
info.latestcomes straight from the npm registry/latestresponse (fetchLatestNpmVersiononly checkstypeof === "string") and reachesstageVersionunvalidated. There it becomes a filesystem path viaversionDir(version)/${finalDir}.tmpand an npm argument spawned withshell: trueon Windows — so a malformed or hostile registry value (e.g.../../x, or one carrying shell metacharacters) could direct directory creation outside~/.orbcode/versionsor inject into the Windows npm command line. Every other version string in this flow (selfVersion, thecurrentpointer, staged versions, Node versions) is checked withisVersionString; the one value that actually drives installs is the only one that isn't.Fix: Reject a non-semver
latestbefore staging, mirroring the validation applied everywhere else.Impact: Guarantees all staged-version writes stay under
~/.orbcode/versions/<semver>and keeps registry-controlled data out of path/shell contexts.- const latest = info.latest; - if (!info.updateAvailable || !latest) return stagedNotice ?? info; + const latest = info.latest; + if (!info.updateAvailable || !latest) return stagedNotice ?? info; + if (!isVersionString(latest)) return stagedNotice ?? quiet;
- agent: mark the assistant message as measured only after it is pushed, so its output tokens aren't counted twice in the context estimate - autoUpdate: only remove a stale lock if it still holds the content we judged stale (TOCTOU), and reject a non-semver registry latest - bump @modelcontextprotocol/sdk to ^1.32.1 (GHSA-6qxp-vccf-f47h)
There was a problem hiding this comment.
🧪 PR Review is completed: Both prior autoUpdate.ts findings are correctly resolved (stale-lock content re-check, isVersionString gate on latest). Two new issues: usageReported is not cleared by rollbackForRetry (skews the compaction trigger's accounting after a mid-stream retry), and package-lock.json is missing the ^1.32.1 SDK bump. Reviewed src/utils/autoUpdate.ts: no new issues found.
Skipped files
package-lock.json: Skipped file pattern
⬇️ Low Priority Suggestions (2)
src/core/agent.ts (1 suggestion)
Location:
src/core/agent.ts(Lines 1616-1616)🟡 Correctness
Issue:
usageReportedis step-scoped stream state, butrollbackForRetry— which resets every other accumulator (assistantText,reasoningDetails,toolCallsByIndex,nextSyntheticIndex) — does not clear it. If a dropped attempt emitted ausagechunk and the retried attempt completes without one, the staletruemakes line 1717 setmeasuredMessages = this.messages.lengtheven though the last usage report never covered the retried assistant message. The 80% compaction trigger then under-counts context and fires late — the exact accounting this release is hardening.Fix: Clear the flag inside
rollbackForRetryalongside the other accumulators (the declaration comment below marks the invariant).Impact: Keeps
measuredMessagesaccurate across mid-stream retries so auto-compaction triggers at the right threshold.- let usageReported = false + // Cleared by rollbackForRetry with the other step accumulators — a usage + // chunk from a dropped attempt must not mark the retried assistant + // message as already measured. + let usageReported = false
package.json (1 suggestion)
Location:
package.json(Lines 56-56)🟠 Dependencies
Issue:
package.jsonnow requires@modelcontextprotocol/sdk^1.32.1, butpackage-lock.jsonis not part of this PR — the committed lockfile still resolves1.29.0(node_modules/@modelcontextprotocol/sdk→sdk-1.29.0.tgz).npm cifails when package.json and the lockfile are out of sync, and a plainnpm installwill silently rewrite the lockfile for every contributor after this release.Fix: Regenerate and commit the lockfile (
npm install) so it pins 1.32.1 together with this spec bump.Impact: Keeps
npm cireproducible and avoids lockfile churn for everyone pulling this release.- "@modelcontextprotocol/sdk": "^1.32.1", +
Summary
Effort selector for GLM 5.3, GLM 5.3 Flash, Gemini 3.8 Flash and DeepSeek V4.1 Flash:
/effortopens a Faster ↔ Smarter slider (low · medium · high · max). Enter saves it;sapplies it to this session only./effort <level>sets it directly, ←/→ in/modeladjusts the highlighted model, and the status bar shows the level in effect.~/.orbcode/config.jsonand read on every request, so every chat on the machine picks them up from its next turn.X-MATTERAI-REASONING-EFFORT; the backend maps it to each gateway's supported levels. Which models get the selector comes from the catalog'sreasoning_efforts, with a built-in fallback.Cold-cache resume prompt: resuming a session idle 5+ minutes (past the gateway's prompt-cache TTL) asks first, Claude Code style: idle time, context size, and the share of the weekly limit from
GET /axoncode/usage/estimate. It shows for any 100k+ token session, or smaller ones costing ≥1% of the window. Options: resume, or start a new conversation (the old one stays available via/task). Estimates are prefetched while the/resumepicker is open.Model and effort persist across restarts: catalog-only models (e.g. DeepSeek V4.1 Flash) used to read as unknown at startup, so the saved model fell back to GLM 5.3 Flash and its effort to medium. The catalog is now cached in
~/.orbcode/models-cache.jsonand loaded before settings, the saved model is restored once the live catalog arrives, and the plan-default logic reads the saved model from disk.Compaction hardening:
read_filecaps characters as well as lines (100k per file, 200k per call).Prompt-cache warmup (shipped disabled): on launch and
/new, the agent and its taskId can be created up front and a warmup request primes the gateway prompt cache with the task's system prompt + tools (same taskId for provider session affinity, 1 output token,X-AxonCode-Warmup: 1), so the first real turn starts warm. MCP startup is awaited so the tool list matches the first turn. Gated behindCACHE_WARMUP_ENABLED = falseinsrc/core/agent.tsuntil it has more testing; with the flag off, behavior is unchanged (agents are still created lazily on the first message).Fix: a replaced conversation (
/new,/resume, sign-in/out, MCP reload) no longer streams into the next one.Version bumped to 6.9.7; changelog updated.
Backend dependency
Needs gravity-console-backend
0d19216(already onmasterand deployed):/axoncode/usage/estimate, reasoning-effort mapping, and catalogreasoning_efforts. Against an older backend the resume prompt falls back to showing context size only, and the effort selector shows but has no effect.Warmup support (inactive while the flag is off) is in gravity-console-backend
9c469aaonmaster: warmup requests skip title generation, Auto routing and think-data storage. That commit also stops duplicate title generation across restarts and rollouts.Validation
npm run typecheckandnpm run buildpasstest:resume,test:effortandtest:compactiontest:warmuppasses 4/4: the warmup's system prompt and tools match the first turn exactly; it sends the header, taskId andmax_tokens: 1test:uipasses 27/27, including new effort-picker, model-picker and resume-confirm tests