Repository navigation
feat(sleep): add adversarial candidate probes - #263
Bogdan (Dan) Baciu (bogdanbaciu21) wants to merge 5 commits into
Conversation
|
Ready for review. This is one commit and nine files on the current base, at exact head The focused integration slice passed on Ubuntu, macOS, and Windows: 177 passed, 2 optional skips, and 6 subtests on each runner. The complete suites also passed with zero failures: Ubuntu and macOS 1,513 passed; Windows 1,468 passed. Strict docs passed on Python 3.12. The CLA check passed. The upstream CI run is marked |
|
Thanks for the unusually thorough test matrix and for making the feature opt-in. The direction is useful, but the blocking decision is not yet sound enough to gate adoption. The main issue is attribution. Please make blocking baseline-relative: evaluate baseline and candidate on identical source/probe pairs, then base rejection on a paired change such as the candidate gap worsening relative to the baseline gap. Evidence should retain all four scores so the decision is auditable. At minimum, add regressions where (1) baseline and candidate are equally frame-sensitive and adoption is not blocked, (2) the candidate improves both scores but retains a gap and is not blocked, and (3) only candidate-introduced degradation is blocked. There are two related validity problems:
Until baseline-relative comparison, stochastic robustness, and conservative transformations are covered, please remove or disable the blocking path and keep the probes advisory-only. The deterministic mock tests and cross-platform green suite verify plumbing, but they do not validate the blocking signal itself. |
|
Got it. Will do. These are incredibly helpful comments thank you so much for the time and care you took to provide them. It will take me a bit to make these changes, will reply when complete at a high quality level. |
|
Thanks again for the review. All three issues are addressed at head
Validation at the exact head across five native runners (Ubuntu 3.10/3.11/3.12, macOS arm64, Windows): full suite 1,538 passed on Ubuntu and macOS and 1,493 on Windows with zero failures; the focused slice is 202 passed everywhere; strict docs pass. On your closing point: we considered removing the blocking path entirely and keeping probes advisory-only. We kept blocking opt-in behind the three conditions you set (baseline-relative comparison, repeated-rollout consistency, conservative transformations), plus the rollout floor, and advisory remains the default. If you would still prefer advisory-only until the margin has been calibrated on a public scenario, we are happy to disable blocking in this PR and propose it separately with that calibration. |
|
Thanks for fixing baseline-relative comparison, repeated rollouts, and conservative request transformations. Re-reviewing
I reproduced this using the real cycle/group pipeline and the PR's deterministic robust backend, with: The ordinary consolidation runs, but the skill group is reported as Please forward the rollout setting through the grouped call and per-group diagnostic config, and add a real multi-skill cycle regression with blocking enabled and at least two rollouts. This can be fully offline; no new paid-model experiment is necessary for this wiring fix. |
…ation - blocking is baseline-relative: identical source/probe pairs scored under baseline and candidate docs; a row is brittle only when the candidate gap worsens beyond the margin in a strict majority of rollout indices - evidence retains all four aggregated scores plus per-rollout samples - dream_adversarial_rollouts (cap 8); blocking requires >= 2 - _strip_polite_frame restricted to politeness-marked requests with negative tests for ability/permission/desire forms - baseline documents are required arguments so no caller can silently compare against an unintended baseline
84bbde9 to
67b2f16
Compare
|
Thanks for the precise reproduction. I fixed the multi-skill omission and added the real offline regression you requested at head
Proof at No paid-provider run was added because this is the fully offline fanout wiring defect from your review. Ready for re-review. |
|
Re-reviewed This is the right kind of evidence for the wiring defect; I am not requesting a paid-provider run merely to verify argument propagation. Keeping the configured rollout count in diagnostics is also useful. Please refresh the validation summary in the PR description: its multi-platform table still says every job asserted The stated limitations on request-frame probes, proxy coverage, rollout cost, and opt-in blocking remain important. The focused propagation blocker is resolved; final review and the exact-head upstream CI gate remain. That CI run currently awaits maintainer approval, so this comment is not merge approval. |
|
Thanks. The validation summary is refreshed; the head is unchanged at
The stated limitations on request-frame probes, proxy coverage, rollout cost, and opt-in blocking stand as written. Nothing else changed. |
|
Thank you for addressing the multi-skill rollout handling and documenting the limits of the probe evidence. Re-reviewed One output mismatch remains at An independent real-cycle regression produces correct structured evidence: Its Please update the table columns and field names to expose the baseline-relative evidence, including |
|
Welcome back! Will do! |
|
Fixed the report mismatch in The table now shows baseline source/probe, candidate source/probe, and gap change. Four real-cycle regression cases cover brittle and robust candidates in advisory and blocking mode, asserting both the structured evidence and the exact numeric row in The reproduced brittle row now reads: Validation on that exact head: 51 adversarial tests pass locally; the five fork jobs (Ubuntu Python 3.10/3.11/3.12, macOS arm64 3.12, Windows 3.12) pass the focused integration slice and full suite. The full suite reports 1,547 passed on Ubuntu/macOS and 1,502 passed on Windows, with platform/optional skips reported separately. Strict docs pass on all three Python 3.12 platforms. The PR description now separates this receipt from the historical matrix. These are fork-runner results; upstream CI for this head still awaits repository approval. No paid-provider experiment was added for this rendering correction. This rendering fix is ready for re-review. Upstream workflow approval and the final merge decision remain with the project. |
|
Thank you for the report-rendering fix. The independent October 6 review of For a miner-produced The focused shipped suite passed 210 tests (2 skipped), but the independent contract suite had 2 failures and 3 passes, also reproduced on the simulated merge. The lost sample identity is inherited; this PR newly uses those purportedly independent samples as evidence for a blocking majority. Before merge, ensure supported tool/backend routes obtain distinct samples, or explicitly classify unsupported routes as unsupported/inconclusive rather than claiming repeated-sample confidence. Add a miner-derived cycle regression with actual provider-boundary call counts and a text positive control. This is a material evaluation issue, not a post-merge cosmetic follow-up. |
|
Okay. Lets try to get this settled since its an August PR and honestly I'm getting frustrated and discouraged. If there's needed changes or rework please be blunt and bury me in it comprehensively so I can address it all end to end. If my work is not up to par I will work to improve if brought to my attention. Yifan Yang (@Yif-Yang). Let's set the bar for a high quality contribution and repeatedly hit it. Thank you - dan |
The tool-aware replay path dropped sample_id: replay_one called attempt_with_tools without it, and the inherited marker fallback (used by Pi, Azure, and other backends without a real tool loop) called attempt at sample zero. Repeated rollouts of a tool_called task were therefore served from one cache entry. For adversarial probes, a single failed response counted as a majority and could block a candidate; for mainline dream_rollouts, K rollouts of a tool task made one provider call and reported K identical scores, so contrastive reflection saw no spread. - attempt_with_tools accepts sample_id on every shipped backend; the inherited fallback and the mock tool model forward it to attempt, the real tool loops already start a fresh uncached process per call, and DualBackend forwards it to the target - replay_one forwards sample_id on the tool route, keeping the historical call shape for sample zero - Backend.distinct_samples(tools=...) states whether a route yields distinct repeated samples; custom backends without a sample_id parameter report False - with rollouts > 1, probe rows on a route without distinct samples are not replayed and are reported as inconclusive; they never flag, and when no conclusive row remains blocking fails closed with block_reason=inconclusive_repeated_samples_unsupported - regressions: a miner-derived Pi cycle with provider-boundary call counts for the tool route and a text positive control, dream rollouts of a tool task through Pi, sample forwarding through the inherited fallback and DualBackend, the shipped-backend route contract, and legacy-route inconclusive evidence in the gate and report Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Thanks for the precise reproduction. Fixed at Root cause:
The regressions you asked for:
Validation at the exact head: five fork runners (Ubuntu 3.10/3.11/3.12, macOS arm64 3.12, Windows 3.12) each assert the SHA before testing. All pass the focused slice (218 passed). The complete suite reports 1,558 passed on Ubuntu/macOS and 1,513 on Windows, with platform skips reported separately. Strict docs pass on all three 3.12 runners. A simulated merge onto current |
What Problem This Solves
A candidate skill can improve its held-out validation score while depending on the exact wording or framing of harvested requests. It can therefore pass the ordinary gate and still break when the same request arrives with a harmless surface change.
The first revision of these probes scored only the candidate, so blocking mode could not distinguish brittleness the candidate introduced from prompt-frame sensitivity already present in the backend, and a single stochastic sample could mark a row brittle at the default margin.
Why This Change Was Made
This revision compares probe sensitivity with the baseline and adds a repeated-rollout consistency screen before any probe result can block adoption.
Identical source and probe pairs are scored under both the current documents and the candidate documents, each score is the mean of a configured number of repeated rollouts, and a row is brittle only when the candidate's probe minus source gap worsens beyond the margin relative to the baseline gap and the worsening holds in a strict majority of rollout indices. Blocking additionally requires at least two rollouts, so one stochastic sample can never reject a candidate; advisory runs may use one. The baseline documents are required arguments, so no caller can silently compare against an unintended baseline. Request-frame transformations are restricted to a defensible semantic-preservation contract: only explicitly politeness-marked requests are reframed, and ability, permission, and desire questions are never touched.
Repeated rollouts count as evidence only when they are distinct samples. Every replay route forwards the rollout index (
sample_id): the text route and the inheritedTOOL_CALL:fallback salt the attempt cache with it, and the real tool-loop backends start a fresh, uncached process per call. A custom backend that cannot acceptsample_idgetsinconclusiverows that never flag a candidate, so the evaluator never claims a repeated-sample confidence the route cannot deliver.Project Fit
User Impact
Operators can see which harmless request variation broke a staged candidate, and whether the breakage is candidate-introduced or pre-existing backend sensitivity. Advisory mode adds evidence without changing gate decisions; explicit blocking mode rejects only candidate-introduced degradation, after operators calibrate
dream_adversarial_marginon their task mix and setdream_adversarial_rolloutsto at least two.Proof
The deterministic proof covers the pre-registered contract scenarios, each comparing the candidate arm versus the baseline arm on identical source and probe pairs.
tool_calledtask replayed through the shipped Pi backend with three rollouts, where only the Pi child process is faked and one candidate-probe response fails, now shows that failure as one failed rollout out of three and accepts the candidate, matching the text-route positive control. Provider-boundary call counts are asserted per arm and role. The same forwarding gives mainlinedream_rolloutsof a tool task distinct samples: four provider calls instead of one.Regression tests also pin train, validation, test, and provenance isolation; target-only routing in dual-backend operation; the three-variant, 256-probe, and eight-rollout resource bounds; strict configuration validation including the blocking rollout floor; non-finite and zero-probe fail-closed behavior; redaction; and Markdown-safe reporting.
These are deterministic contract tests, not a live-provider performance claim, so no stochastic lift or confidence interval is claimed.
Academic Support
Testing
Current head
4f3a908adcf8a83f5dec3bc64461578258fd1967The current head addresses the repeated-sampling review.
replay_one()droppedsample_idon the tool route, and the inherited marker fallback calledattempt()at sample zero, so every nominal rollout of atool_calledtask reused one cache entry.attempt_with_tools()now acceptssample_idon every shipped backend. The inherited fallback and the mock tool model forward it toattempt(), the real tool loops already start a fresh uncached process per call, andDualBackendforwards it to the target.Backend.distinct_samples(tools=...)reports whether a route can provide distinct samples; with more than one rollout, rows on a route that cannot are reported asinconclusiveand are not replayed.New regressions: the miner-derived Pi cycle for the tool route and the text positive control, with provider-boundary call counts; dream rollouts of a tool task through Pi; forwarding through the inherited fallback and
DualBackend; the shipped-backend route contract; and legacy-route inconclusive evidence in the gate decision andreport.md. Against the prior head4d93cad055f9, the tool-route cycle case fails with the candidate blocked while the text control passes. Against currentmain, the tool-route dream-rollout case makes one provider call instead of four.Fork-runner validation below asserts
actual_sha == expected_sha == 4f3a908adcf8a83f5dec3bc64461578258fd1967before testing. These are contributor-run checks, not official upstream CI.Local macOS x86_64 / Python 3.12 validation also passed:
tests/test_adversarial_dream.pyreported 62 passed; the complete suite reported 1,559 passed, 11 skipped, 353 subtests; strict docs andgit diff --checkpassed. A simulated merge onto currentmain(343db22) reported 1,777 passed, 11 skipped, 359 subtests. No test was deselected and no failure was masked. The Windows count difference is platform/optional-dependency skips.No paid provider was called; only the provider process boundary is faked. Exact-head upstream CI awaits repository approval, and no official upstream CI success is claimed.
Historical evidence: superseded head 4d93cad, not current validation
The current head fixes the Markdown renderer to read all four baseline/candidate score fields and
gap_change. Four new real-cycle cases cover brittle and robust candidates in advisory and blocking mode. Each checks the structured scores and the exact numeric row in the generatedreport.md; all four failed on the prior head67b2f16441cfbefore the renderer fix.For the planted brittle candidate, the report now renders the baseline source/probe scores as
0.000 / 0.000, candidate source/probe as1.000 / 0.000, and gap change as-1.000, with statusbrittle.Fork-runner validation below asserts
actual_sha == expected_sha == 4d93cad055f9210438867d02c8e7c24236902040before testing. These are contributor-run checks, not official upstream CI.Local Linux/Python 3.12 validation also passed:
python3 -m pytest tests/test_adversarial_dream.py -qreported 51 passed;python3 -m pytest -qreported 1,547 passed, 12 skipped, 353 subtests; strict docs andgit diff --checkpassed. No test was deselected in the complete-suite runs and no failure was masked. The Windows count difference is reported as platform/optional-dependency skips.The exact-head upstream CI run is
action_requiredand awaits repository approval. No official upstream CI success is claimed.Historical evidence: superseded head 84bbde9, not current validation
The matrix below was collected on fork runners against
84bbde9ac191, the candidate reviewed before the multi-skill rollout-forwarding fix. It is retained for history only. It does not describe the current head and it is not official CI.Every job in that historical matrix asserted exact candidate
84bbde9ac191before testing. No test was deselected and no failure was masked. The Windows difference is explicit platform and optional dependency skips, not failures.Regression tests also pin train, validation, test, and provenance isolation; target-only routing in dual-backend operation; the three-variant, 256-probe, and eight-rollout resource bounds; strict configuration validation including the blocking rollout floor; non-finite and zero-probe fail-closed behavior; redaction; and Markdown-safe reporting.
Limitations & Negative Results
attemptorattempt_with_toolsdoes not acceptsample_idcannot provide distinct repeated samples. With more than one rollout its probe rows are reported asinconclusiveinstead of being scored, and blocking mode fails closed if no conclusive row remains.Reproduce It Yourself
Linux or macOS, shell in a fresh working directory: