Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,11 @@
`start --expect-source-revision` exits 1 when any expected revision is
mismatched or unavailable (the report is still printed), 0 when all match,
and 2 for invalid arguments.
- A separate synthetic revision benchmark scores retrieved records, cited
revisions, proposed values and outbound action proposals independently. Its
unkeyed record list includes superseded records, so copying a stale record
with the correct value fails the record and citation layers. First trials
(18, three models) all passed every layer; the evidence records this ceiling.

### Fixed
- The Hermes environment-filter control isolates inherited case variants of
Expand Down
16 changes: 16 additions & 0 deletions components/manifest.json
Original file line number Diff line number Diff line change
Expand Up @@ -1051,6 +1051,22 @@
{
"path": "docs/feedback-review-2026-10-04.md",
"policy": "development"
},
{
"path": "tests/fixtures/continuity/revision-invalidation.json",
"policy": "development"
},
{
"path": "docs/evidence/continuity-revision-2026-10-04/README.md",
"policy": "development"
},
{
"path": "docs/evidence/continuity-revision-2026-10-04/fence-normalization.jsonl",
"policy": "development"
},
{
"path": "docs/evidence/continuity-revision-2026-10-04/trials.jsonl",
"policy": "development"
}
]
},
Expand Down
58 changes: 58 additions & 0 deletions docs/continuity-benchmark.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,10 @@ run these commands there; the benchmark never needs your personal workspace.
- `contextos`: the same instructions plus canonical decisions, blockers, current
priorities, and an older session containing superseded ideas.

The separate revision scenario below has two supported profiles and an extended
response format. The original short and long prompts and historical scores
remain unchanged.

```bash
python scripts/continuity-benchmark.py prepare --profile handoff > handoff-prompt.txt
python scripts/continuity-benchmark.py prepare --profile contextos > contextos-prompt.txt
Expand Down Expand Up @@ -139,3 +143,57 @@ measure it. Installed-host conformance remains a separate test suite.

The [2026-09-23 long-sequence evidence](evidence/continuity-long-2026-09-23/README.md)
reports 27 fresh-session trials, category counts, rejected answers, and limits.

## Test revision attribution and proposed actions

The `revision` scenario tests two questions with synthetic retrieval records.
An old record and its reviewed replacement both select CSV, so returning CSV
alone cannot establish that invalidation worked. A second question keeps a
launch date unconfirmed and requires holding the launch announcement.

```bash
python scripts/continuity-benchmark.py prepare --scenario revision --profile contextos
python scripts/continuity-benchmark.py prepare --scenario revision --profile handoff
python scripts/continuity-benchmark.py score --scenario revision --profile contextos --response response.json
```

Each profile receives the same facts and current retrieval records; `handoff`
uses the concise handoff source. Use fresh sessions with tools disabled and
keep the answer key outside them, as with the other scenarios. The extended
answer contains `value`, `source`, `quote`, `record_id`, `source_revision`,
`supersedes`, `dependency_check`, and `action`. The prompt supplies an
unordered list of retrieval records with source revisions, supersession links
and dependency-check results; it is not keyed by question. In `contextos`, the
list also contains the superseded session records, whose own checks still
match, so copying a record passes only when the agent selects the current one.
Actions are proposed codes, never executed operations.

The scorer reports each layer independently:

| Layer | Passing evidence |
|---|---|
| Retrieved record | Correct current record ID, superseded ID and supplied dependency-check result |
| Cited revision | Current source path, full normalized-text SHA-256 and supporting sentence |
| Proposed value | Correct CSV choice or unresolved launch status |
| Outbound action | Outline CSV columns locally or hold the public announcement |

All four layers must pass for a question to count as grounded correct. A
correct value with a stale citation, missing supersession link, false check
result, or publishing proposal fails. Citation rejection is counted separately
from record and action failures.

`record --scenario revision` preserves raw responses and layer scores in JSONL.
`summarize` returns mean passing-question counts for each layer and the count
of format failures. Keep revision results in a separate JSONL from the legacy
scenarios. `instructions` and `handoff-sentences` are unsupported for this case.

The [2026-10-04 revision trials](evidence/continuity-revision-2026-10-04/README.md)
ran three models through Codex and Cursor. All 18 trials passed every layer,
so this version does not separate those models; a harder case is needed.

This is a response-format and attribution evaluation over supplied synthetic
evidence. It does not observe a host's retrieval, execute a dependency graph,
authenticate the reviewer, or send an outbound action. Unit controls validate
the scorer; they are not model trials or evidence of live stale-memory repair.
For actual selected-file checks, use [source revision expectations](continuity.md).
The remaining design recommendations are in the [feedback review](feedback-review-2026-10-04.md).
3 changes: 2 additions & 1 deletion docs/continuity.md
Original file line number Diff line number Diff line change
Expand Up @@ -104,7 +104,8 @@ receipt, under `files_changed`. Content writes use normalized-text
`sha256_before_raw`/`sha256_after_raw`. An absent file is represented by `null`.
The proposal digest identifies the reviewed proposal, not the post-apply
workspace. Apply rejects changed input snapshots before writing; a rejected
apply creates no successful receipt.
apply creates no successful receipt. The [revision benchmark](continuity-benchmark.md#test-revision-attribution-and-proposed-actions)
scores evidence and proposed actions separately.

For attached application repositories, supply the same explicit `--kernel-root`,
`--context-root`, and `--working-root` arguments used by lifecycle commands.
Expand Down
70 changes: 70 additions & 0 deletions docs/evidence/continuity-revision-2026-10-04/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
# Revision-scenario trials, 2026-10-04

These 18 trials use the synthetic Lantern revision fixture from source commit
`dae11f3cd55139f6456ae5ca582deaec0faded56` (`--scenario revision`). Each of three
models received each supported profile (`contextos` and `handoff`) in three
fresh sessions. The orchestrator (Claude Code) ran them on 2026-10-04.

- `gpt-6.1-sol` ran through the Codex CLI 0.160.0 (`codex-cli` in the records):
`codex exec --ephemeral -s read-only --json`, started in an empty temporary
directory with the prompt on standard input.
- `grok-4.7-low-fast` and `composer-2.5` ran through the Cursor agent CLI in
ask mode, each call in a new throwaway workspace (`cursor-agent` in the
records).

The model IDs are the ones requested. Neither CLI reported a provider-side
model ID in its output, so the records do not independently establish which
model served each call. The Codex event streams contained only the agent
message and turn events, with no command, file, MCP or web-search items. The
Cursor wrapper reports no tool events, so tool isolation there rests on the
empty workspace and ask mode, not on an observed event log.

The [trial records](trials.jsonl) retain every response, prompt hash, score,
latency and token count. Token counts include each CLI's own system prompt, so
they are much larger than the fixture (3,180 prompt characters for
`contextos`, 2,198 for `handoff`). The [normalization log](fence-normalization.jsonl)
marks five `composer-2.5` responses that arrived as valid JSON inside one
Markdown `json` fence. The orchestrator removed that fence before `record`, as
in the [long-sequence trials](../continuity-long-2026-09-23/README.md); the
stored `raw_response` is the unfenced body. No other response was changed, and
no stored response contains a carriage return.

Each `prompt_sha256` is the SHA-256 of `prepare(scenario, profile)` for the
revision scenario: `dfa69fbfbe9a…` for `contextos` and `c9ced04db08f…` for
`handoff`. The orchestrator recomputed both from this fixture and they match.

## Scores

The table was generated by
`python scripts/continuity-benchmark.py summarize --results docs/evidence/continuity-revision-2026-10-04/trials.jsonl`.
Each value is the mean number of the two questions passing that layer.

| Model | Profile | Trials | Record | Revision | Value | Action | All layers | Format failures |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| composer-2.5 | contextos | 3 | 2.00 | 2.00 | 2.00 | 2.00 | 2.00 | 0 |
| composer-2.5 | handoff | 3 | 2.00 | 2.00 | 2.00 | 2.00 | 2.00 | 0 |
| gpt-6.1-sol | contextos | 3 | 2.00 | 2.00 | 2.00 | 2.00 | 2.00 | 0 |
| gpt-6.1-sol | handoff | 3 | 2.00 | 2.00 | 2.00 | 2.00 | 2.00 | 0 |
| grok-4.7-low-fast | contextos | 3 | 2.00 | 2.00 | 2.00 | 2.00 | 2.00 | 0 |
| grok-4.7-low-fast | handoff | 3 | 2.00 | 2.00 | 2.00 | 2.00 | 2.00 | 0 |

Every trial passed every layer on both questions. In the `contextos` profile,
each model selected `export-002` and `launch-002` over the superseded session
records, cited the current source and its revision, reported the supplied
`mismatch` result and proposed the authorized action. No trial proposed the
launch announcement.

## What this does and does not show

At this difficulty the scenario does not separate these three models: there is
no variance to compare. That is a ceiling result, not evidence that the models
handle stale context in general. The fixture is small, the superseded records
are labelled as superseded in the session text, and the prompt tells the agent
to prefer the current source. A useful next version needs a harder case, for
example a superseded record whose source text does not say it is superseded,
a stale record in the `handoff` profile, or longer context with the replacement
far from the question.

These are supplied synthetic records scored offline. They do not observe host
retrieval, execute a dependency graph, authenticate a reviewer or send an
action. Three trials per cell and three models are a small sample.
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
{"model_id": "composer-2.5", "profile": "contextos", "trial_index": 1, "normalization": "removed one Markdown json fence"}
{"model_id": "composer-2.5", "profile": "contextos", "trial_index": 2, "normalization": "removed one Markdown json fence"}
{"model_id": "composer-2.5", "profile": "contextos", "trial_index": 3, "normalization": "removed one Markdown json fence"}
{"model_id": "composer-2.5", "profile": "handoff", "trial_index": 1, "normalization": "removed one Markdown json fence"}
{"model_id": "composer-2.5", "profile": "handoff", "trial_index": 3, "normalization": "removed one Markdown json fence"}
Loading
Loading