Skip to content

release: v6.9.7 - #67

Merged
code-crusher merged 5 commits into
mainfrom
release/v6.9.7
Oct 6, 2026
Merged

code-crusher merged 5 commits into
mainfrom
release/v6.9.7

Conversation

@code-crusher

@code-crusher code-crusher commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Summary

Effort selector for GLM 5.3, GLM 5.3 Flash, Gemini 3.8 Flash and DeepSeek V4.1 Flash:

  • /effort opens a Faster ↔ Smarter slider (low · medium · high · max). Enter saves it; s applies it to this session only. /effort <level> sets it directly, ←/→ in /model adjusts the highlighted model, and the status bar shows the level in effect.
  • Default is medium. Picks are saved per model in ~/.orbcode/config.json and read on every request, so every chat on the machine picks them up from its next turn.
  • Sent as X-MATTERAI-REASONING-EFFORT; the backend maps it to each gateway's supported levels. Which models get the selector comes from the catalog's reasoning_efforts, with a built-in fallback.

Cold-cache resume prompt: resuming a session idle 5+ minutes (past the gateway's prompt-cache TTL) asks first, Claude Code style: idle time, context size, and the share of the weekly limit from GET /axoncode/usage/estimate. It shows for any 100k+ token session, or smaller ones costing ≥1% of the window. Options: resume, or start a new conversation (the old one stays available via /task). Estimates are prefetched while the /resume picker is open.

Model and effort persist across restarts: catalog-only models (e.g. DeepSeek V4.1 Flash) used to read as unknown at startup, so the saved model fell back to GLM 5.3 Flash and its effort to medium. The catalog is now cached in ~/.orbcode/models-cache.json and loaded before settings, the saved model is restored once the live catalog arrives, and the plan-default logic reads the saved model from disk.

Compaction hardening:

  • A provider "context too long" error now compacts and retries the step instead of failing every following turn.
  • A history already past the window (e.g. resumed on a smaller-window model) still compacts: the summary request is trimmed to fit, with smaller budgets tried if the provider's real window is tighter than the catalog's.
  • The 80% trigger counts tool output added since the last usage report.
  • A failed auto-compaction is retried on the next turn instead of staying off for the session.
  • read_file caps characters as well as lines (100k per file, 200k per call).

Prompt-cache warmup (shipped disabled): on launch and /new, the agent and its taskId can be created up front and a warmup request primes the gateway prompt cache with the task's system prompt + tools (same taskId for provider session affinity, 1 output token, X-AxonCode-Warmup: 1), so the first real turn starts warm. MCP startup is awaited so the tool list matches the first turn. Gated behind CACHE_WARMUP_ENABLED = false in src/core/agent.ts until it has more testing; with the flag off, behavior is unchanged (agents are still created lazily on the first message).

Fix: a replaced conversation (/new, /resume, sign-in/out, MCP reload) no longer streams into the next one.

Version bumped to 6.9.7; changelog updated.

Backend dependency

Needs gravity-console-backend 0d19216 (already on master and deployed): /axoncode/usage/estimate, reasoning-effort mapping, and catalog reasoning_efforts. Against an older backend the resume prompt falls back to showing context size only, and the effort selector shows but has no effect.

Warmup support (inactive while the flag is off) is in gravity-console-backend 9c469aa on master: warmup requests skip title generation, Auto routing and think-data storage. That commit also stops duplicate title generation across restarts and rollouts.

Validation

  • npm run typecheck and npm run build pass
  • All node test suites pass (121/121), including the new test:resume, test:effort and test:compaction
  • New test:warmup passes 4/4: the warmup's system prompt and tools match the first turn exactly; it sends the header, taskId and max_tokens: 1
  • test:ui passes 27/27, including new effort-picker, model-picker and resume-confirm tests
  • Against production: the usage estimate returns weekly/monthly percentages (e.g. 0.2% for 115k tokens on DeepSeek V4.1 Flash)

- Effort selector (low/medium/high/max) for GLM 5.3, GLM 5.3 Flash,
  Gemini 3.8 Flash and DeepSeek V4.1 Flash: /effort slider (Enter saves,
  s for this session only), /effort <level>, ←/→ in /model, shown in the
  status bar. Saved per model in config.json and read per request, so
  every chat on the machine uses it. Sent as X-MATTERAI-REASONING-EFFORT.
- Confirm before resuming a session with a cold prompt cache (idle 5+ min):
  shows the share of the weekly limit from /axoncode/usage/estimate, for
  any 100k+ token session or when the cost reaches 1%. Estimates are
  prefetched while the /resume picker is open.
- Model and effort persist across restarts: the model catalog is cached in
  ~/.orbcode/models-cache.json and loaded before settings.
- Compaction recovers from context overflow, trims the summary request to
  fit the window, counts unreported tool output, and retries after a
  failure on the next turn.
- read_file caps characters (100k per file, 200k per call).
- A replaced conversation (/new, /resume, sign-in/out, MCP reload) no
  longer streams into the next one.
@matterai-app

matterai-app Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary By MatterAI MatterAI logo

🔄 What Changed

Released version 6.9.7 introducing an effort selector for advanced LLMs (GLM, Gemini, DeepSeek), cold-cache resume prompts for idle sessions, persistent model and effort states across restarts via a cached catalog, compaction resilience hardening, and integration of the Model Context Protocol SDK (@modelcontextprotocol/sdk). Additionally, race condition guards were added to stale lock verification in auto-updates, and strict version string checks were implemented.

🔍 Impact of the Change

Improves context management, error recovery during long tool-use sessions, concurrency safety during background process locking, and user control over reasoning effort levels across restarts.

📁 Total Files Changed

Click to Expand
File ChangeLog
package.json Add @modelcontextprotocol/sdk dependency.
src/core/agent.ts Track usage reporting state and update measured message pointers.
src/utils/autoUpdate.ts Guard stale lock removal against race conditions using exact raw content matching and validate version strings.

🧪 Test Added/Recommended

Added

  • New test suites (test:resume, test:effort, test:compaction) passing completely alongside UI tests.

🔒 Security Vulnerabilities

  • None identified. Strict content validation prevents race conditions during file-system locking operations.

@matterai-app matterai-app Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧪 PR Review is completed: Solid release — the new compaction, effort-selector and resume-cost logic is well covered by tests. Three issues found: describeResumeCost omits the $ its own test expects (test:resume fails), the resume-cost estimate prices the currently selected model instead of the resumed session's model, and measuredMessages double-counts the assistant message. Reviewed src/api/models.ts, src/auth/auth.ts, src/config/settings.ts, src/tools/executors/files.ts, src/index.tsx, src/api/client.ts, src/api/headers.ts, src/ui/components/EffortPicker.tsx, src/ui/components/ModelPicker.tsx, src/ui/components/ResumeConfirm.tsx, src/ui/components/StatusBar.tsx, package.json, test/compaction.test.ts, test/effort-picker.test.tsx, test/files.test.ts, test/model-effort.test.ts, test/model-persistence.test.ts, test/model-picker.test.tsx, test/resume-confirm.test.tsx: no issues found. (test/resume-cost.test.ts is correct as written — the src/core/resumeCost.ts comment is the fix for its failing assertion.)

Skipped files
  • CHANGELOG.md: Skipped file pattern
  • package-lock.json: Skipped file pattern
⬇️ Low Priority Suggestions (2)
src/ui/App.tsx (1 suggestion)

Location: src/ui/App.tsx (Lines 1069-1081)

🟡 Logic

Issue: The resume-cost estimate prices the currently selected model (getModel(loadSettings().model)), but the resumed agent runs on the session's own model — resumeSession calls createAgent(session) and createAgent passes modelId: current.model (line 860). When the two differ, inputPrice, billsPlan() and the backend share query (getGatewayModelId(model)) all use the wrong model, so the confirmation prompt can show a wrong dollar cost or plan-window share, and the prefetched share cache is keyed for the wrong model.

Fix: Resolve the model from session.model — per session inside the prefetchResumeShares loop, and in handleResume.

Impact: The resume confirmation reflects what the resumed session will actually bill.

-      (sessions: SessionData[]) => {
-        const model = getModel(loadSettings().model);
-        for (const session of sessions.slice(0, 10)) {
-          const cold = estimateResumeCost(session, model);
-          if (cold) void fetchResumeShare(model, cold.contextTokens);
-        }
-      },
-      [fetchResumeShare],
-    );
-  
-    const handleResume = useCallback(
-      async (session: SessionData) => {
-        const model = getModel(loadSettings().model);
+      (sessions: SessionData[]) => {
+        for (const session of sessions.slice(0, 10)) {
+          const model = getModel(session.model);
+          const cold = estimateResumeCost(session, model);
+          if (cold) void fetchResumeShare(model, cold.contextTokens);
+        }
+      },
+      [fetchResumeShare],
+    );
+  
+    const handleResume = useCallback(
+      async (session: SessionData) => {
+        const model = getModel(session.model);
src/core/agent.ts (1 suggestion)

Location: src/core/agent.ts (Lines 1640-1640)

🟡 Logic

Issue: The usage report sets measuredMessages = this.messages.length while the stream loop is still running, but contextTokens (input + output) already covers the assistant message that runStep appends after the loop. That message lands at index measuredMessages, so estimatedContextTokens() adds messageTokenEstimate() for it on top of the outputTokens already included in contextTokens — double-counting it every step. The field's own doc says only "tool results, new user input" should be estimated, and the inflation makes auto-compaction (and the "Context is X% full" notice) fire earlier than the intended 80% threshold.

Fix: Mark the about-to-be-appended assistant message as covered by the report: this.messages.length + 1.

Impact: The context estimate matches the report's real coverage; compaction triggers at the intended threshold instead of prematurely dropping history.

-  					this.measuredMessages = this.messages.length
+  				this.measuredMessages = this.messages.length + 1

Adds a warmup request that primes the gateway prompt cache with a task's
system prompt + tools (same taskId, 1 output token, X-AxonCode-Warmup
header) on launch and /new, so the first real turn starts warm. MCP
startup is awaited so the tool list matches the first turn.

Shipped behind CACHE_WARMUP_ENABLED = false until it has more testing;
with the flag off agents are still created lazily on the first message.
@matterai-app

matterai-app Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

✅ Reviewed the changes: Solid prompt-cache warmup implementation: the buildRequest refactor is behavior-faithful, warmup is best-effort with correct abort propagation and re-entrancy guards (warmupController abort/clear, messages re-check after whenStarted), the McpManager.whenStarted() addition is clean, and everything ships behind CACHE_WARMUP_ENABLED=false — no issues found in the new code. Note: the three previously flagged items (src/core/resumeCost.ts dollar sign, src/ui/App.tsx resume-cost model resolution, src/core/agent.ts measuredMessages) are not part of this diff's hunks and could not be re-verified or anchored against it, so they are neither re-emitted nor marked resolved.
Reviewed src/api/client.ts: no issues found.
Reviewed src/api/llmClient.ts: no issues found.
Reviewed src/api/headers.ts: no issues found.
Reviewed src/core/agent.ts: no issues found.
Reviewed src/mcp/manager.ts: no issues found.
Reviewed src/ui/App.tsx: no issues found.
Reviewed test/cache-warmup.test.ts: no issues found.
Reviewed package.json: no issues found.

@matterai-app matterai-app Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧪 PR Review is completed: Clean header-label refactor — withHeaderModel nicely replaces the 11-line findIndex block in switchModel; one session-override semantics question to verify on headerModelLabel. Note (pre-existing, not from this diff): @modelcontextprotocol/sdk ^1.29.0 carries a HIGH advisory (GHSA-6qxp-vccf-f47h, fixed in 1.31.0) — worth bumping before this release ships.

⬇️ Low Priority Suggestions (1)
src/ui/App.tsx (1 suggestion)

Location: src/ui/App.tsx (Lines 406-406)

🟡 Business Logic (needs verification)

Issue: headerModelLabel resolves the displayed effort via getModelEffort(loadModelEfforts(), model). If loadModelEfforts() returns only the persisted per-model picks from ~/.orbcode/config.json (with session-only overrides layered on separately via withSessionModelEfforts), then the header refresh added after setSessionModelEffort(...) — the s / "this session only" path — re-renders the header with the persisted level instead of the override just picked, making that refresh a no-op for the session-only flow (and likewise for /model switches to a model that has a live session override).

Fix: Resolve session-aware efforts before reading the level: getModelEffort(withSessionModelEfforts(loadModelEfforts()), model). If loadModelEfforts() already includes session overrides, no change is needed — please verify which semantics apply.

Impact: Keeps the header MODEL line consistent with the effort actually sent on the next request.

-    const effort = getModelEffort(loadModelEfforts(), model);
+    const effort = getModelEffort(withSessionModelEfforts(loadModelEfforts()), model);

Stage new releases under ~/.orbcode/versions/<v>/, verify them (--version,
--self-test) and flip a 'current' pointer; the launcher hands off to the
staged version and falls back to itself on any failure or a failed start.
Releases auto-install after 24h on npm; opt out with
ORBCODE_DISABLE_AUTOUPDATE=1 or "autoUpdates": false.

@matterai-app matterai-app Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧪 PR Review is completed: Well-structured background-update design (stage → verify → flip pointer, with health/attempts markers and a cross-process lock); two hardening fixes needed in src/utils/autoUpdate.ts — a TOCTOU in the stale-lock takeover and an unvalidated registry latest reaching path/Windows-shell contexts. Reviewed bin/orbcode.js: no issues found. Reviewed bin/select-version.js: no issues found. Reviewed src/index.tsx: no issues found. Reviewed src/ui/App.tsx: no issues found. Reviewed test/auto-update.test.ts: no issues found.

Skipped files
  • README.md: Skipped file pattern
  • RELEASE.md: Skipped file pattern
⬇️ Low Priority Suggestions (2)
src/utils/autoUpdate.ts (2 suggestions)

Location: src/utils/autoUpdate.ts (Lines 208-214)

🟠 Concurrency

Issue: The stale-lock takeover in acquireLock has a TOCTOU window. Two processes can both read the same stale lock content and then both call fs.unlinkSync(lockPath) — the second unlink removes the fresh lock the first process just created with its wx write (which lands between the second process's read and its unlink). Both then succeed on their retry writeFileSync(..., { flag: "wx" }) and proceed into stageVersion concurrently, running two npm installs into the same <version>.tmp directory and breaking the lock's mutual exclusion. The release path already guards with a content check (fs.readFileSync(lockPath, "utf8") === token), but the stale-removal path does not.

Fix: Capture the lock content when reading it for the staleness check, and only unlink if the file still contains that same content; otherwise continue and retry the exclusive wx write on the next loop iteration.

Impact: Restores the lock's mutual-exclusion guarantee so concurrent sessions can't race an install into the same staging directory.

-        try {
-          const [pidText, atText] = fs.readFileSync(lockPath, "utf8").split(":");
-          const pid = Number(pidText);
-          const at = Number(atText);
-          const stale = !isPidAlive(pid) || !Number.isFinite(at) || Date.now() - at > LOCK_STALE_MS;
-          if (!stale) return null;
-          fs.unlinkSync(lockPath);
+        try {
+          const raw = fs.readFileSync(lockPath, "utf8");
+          const [pidText, atText] = raw.split(":");
+          const pid = Number(pidText);
+          const at = Number(atText);
+          const stale = !isPidAlive(pid) || !Number.isFinite(at) || Date.now() - at > LOCK_STALE_MS;
+          if (!stale) return null;
+          if (fs.readFileSync(lockPath, "utf8") !== raw) continue;
+          fs.unlinkSync(lockPath);

Location: src/utils/autoUpdate.ts (Lines 470-471)

🟡 Security (defense-in-depth)

Issue: info.latest comes straight from the npm registry /latest response (fetchLatestNpmVersion only checks typeof === "string") and reaches stageVersion unvalidated. There it becomes a filesystem path via versionDir(version) / ${finalDir}.tmp and an npm argument spawned with shell: true on Windows — so a malformed or hostile registry value (e.g. ../../x, or one carrying shell metacharacters) could direct directory creation outside ~/.orbcode/versions or inject into the Windows npm command line. Every other version string in this flow (selfVersion, the current pointer, staged versions, Node versions) is checked with isVersionString; the one value that actually drives installs is the only one that isn't.

Fix: Reject a non-semver latest before staging, mirroring the validation applied everywhere else.

Impact: Guarantees all staged-version writes stay under ~/.orbcode/versions/<semver> and keeps registry-controlled data out of path/shell contexts.

-      const latest = info.latest;
-      if (!info.updateAvailable || !latest) return stagedNotice ?? info;
+      const latest = info.latest;
+      if (!info.updateAvailable || !latest) return stagedNotice ?? info;
+      if (!isVersionString(latest)) return stagedNotice ?? quiet;

- agent: mark the assistant message as measured only after it is pushed,
  so its output tokens aren't counted twice in the context estimate
- autoUpdate: only remove a stale lock if it still holds the content we
  judged stale (TOCTOU), and reject a non-semver registry latest
- bump @modelcontextprotocol/sdk to ^1.32.1 (GHSA-6qxp-vccf-f47h)

@matterai-app matterai-app Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧪 PR Review is completed: Both prior autoUpdate.ts findings are correctly resolved (stale-lock content re-check, isVersionString gate on latest). Two new issues: usageReported is not cleared by rollbackForRetry (skews the compaction trigger's accounting after a mid-stream retry), and package-lock.json is missing the ^1.32.1 SDK bump. Reviewed src/utils/autoUpdate.ts: no new issues found.

Skipped files
  • package-lock.json: Skipped file pattern
⬇️ Low Priority Suggestions (2)
src/core/agent.ts (1 suggestion)

Location: src/core/agent.ts (Lines 1616-1616)

🟡 Correctness

Issue: usageReported is step-scoped stream state, but rollbackForRetry — which resets every other accumulator (assistantText, reasoningDetails, toolCallsByIndex, nextSyntheticIndex) — does not clear it. If a dropped attempt emitted a usage chunk and the retried attempt completes without one, the stale true makes line 1717 set measuredMessages = this.messages.length even though the last usage report never covered the retried assistant message. The 80% compaction trigger then under-counts context and fires late — the exact accounting this release is hardening.

Fix: Clear the flag inside rollbackForRetry alongside the other accumulators (the declaration comment below marks the invariant).

Impact: Keeps measuredMessages accurate across mid-stream retries so auto-compaction triggers at the right threshold.

-  		let usageReported = false
+  		// Cleared by rollbackForRetry with the other step accumulators — a usage
+  		// chunk from a dropped attempt must not mark the retried assistant
+  		// message as already measured.
+  		let usageReported = false
package.json (1 suggestion)

Location: package.json (Lines 56-56)

🟠 Dependencies

Issue: package.json now requires @modelcontextprotocol/sdk ^1.32.1, but package-lock.json is not part of this PR — the committed lockfile still resolves 1.29.0 (node_modules/@modelcontextprotocol/sdk → sdk-1.29.0.tgz). npm ci fails when package.json and the lockfile are out of sync, and a plain npm install will silently rewrite the lockfile for every contributor after this release.

Fix: Regenerate and commit the lockfile (npm install) so it pins 1.32.1 together with this spec bump.

Impact: Keeps npm ci reproducible and avoids lockfile churn for everyone pulling this release.

-      "@modelcontextprotocol/sdk": "^1.32.1",
+  

@code-crusher
code-crusher merged commit 71efdec into main Oct 6, 2026
1 check passed
@code-crusher
code-crusher deleted the release/v6.9.7 branch October 6, 2026 17:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant