Skip to content

Commit f6c04ff

Browse files
conorbronsdonclaude
andcommitted
Record the first revision-scenario trials
Eighteen fresh-session trials: gpt-6.1-sol through the Codex CLI and grok-4.7-low-fast and composer-2.5 through the Cursor agent, three per profile. Every trial passed every layer, so this version of the scenario does not separate these models. The evidence README records the ceiling, the isolation actually observed, the five fence normalizations and the limits, and proposes harder cases. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ADCUQCfj16hLrYMgDc5MJo
1 parent dae11f3 commit f6c04ff

6 files changed

Lines changed: 111 additions & 1 deletion

File tree

‎CHANGELOG.md‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -11,7 +11,8 @@
1111
- A separate synthetic revision benchmark scores retrieved records, cited
1212
revisions, proposed values and outbound action proposals independently. Its
1313
unkeyed record list includes superseded records, so copying a stale record
14-
with the correct value fails the record and citation layers.
14+
with the correct value fails the record and citation layers. First trials
15+
(18, three models) all passed every layer; the evidence records this ceiling.
1516

1617
### Fixed
1718
- The Hermes environment-filter control isolates inherited case variants of

‎components/manifest.json‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1055,6 +1055,18 @@
10551055
{
10561056
"path": "tests/fixtures/continuity/revision-invalidation.json",
10571057
"policy": "development"
1058+
},
1059+
{
1060+
"path": "docs/evidence/continuity-revision-2026-10-04/README.md",
1061+
"policy": "development"
1062+
},
1063+
{
1064+
"path": "docs/evidence/continuity-revision-2026-10-04/fence-normalization.jsonl",
1065+
"policy": "development"
1066+
},
1067+
{
1068+
"path": "docs/evidence/continuity-revision-2026-10-04/trials.jsonl",
1069+
"policy": "development"
10581070
}
10591071
]
10601072
},

‎docs/continuity-benchmark.md‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -187,6 +187,10 @@ from record and action failures.
187187
of format failures. Keep revision results in a separate JSONL from the legacy
188188
scenarios. `instructions` and `handoff-sentences` are unsupported for this case.
189189

190+
The [2026-10-04 revision trials](evidence/continuity-revision-2026-10-04/README.md)
191+
ran three models through Codex and Cursor. All 18 trials passed every layer,
192+
so this version does not separate those models; a harder case is needed.
193+
190194
This is a response-format and attribution evaluation over supplied synthetic
191195
evidence. It does not observe a host's retrieval, execute a dependency graph,
192196
authenticate the reviewer, or send an outbound action. Unit controls validate
Lines changed: 70 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,70 @@
1+
# Revision-scenario trials, 2026-10-04
2+
3+
These 18 trials use the synthetic Lantern revision fixture from source commit
4+
`dae11f3cd55139f6456ae5ca582deaec0faded56` (`--scenario revision`). Each of three
5+
models received each supported profile (`contextos` and `handoff`) in three
6+
fresh sessions. The orchestrator (Claude Code) ran them on 2026-10-04.
7+
8+
- `gpt-6.1-sol` ran through the Codex CLI 0.160.0 (`codex-cli` in the records):
9+
`codex exec --ephemeral -s read-only --json`, started in an empty temporary
10+
directory with the prompt on standard input.
11+
- `grok-4.7-low-fast` and `composer-2.5` ran through the Cursor agent CLI in
12+
ask mode, each call in a new throwaway workspace (`cursor-agent` in the
13+
records).
14+
15+
The model IDs are the ones requested. Neither CLI reported a provider-side
16+
model ID in its output, so the records do not independently establish which
17+
model served each call. The Codex event streams contained only the agent
18+
message and turn events, with no command, file, MCP or web-search items. The
19+
Cursor wrapper reports no tool events, so tool isolation there rests on the
20+
empty workspace and ask mode, not on an observed event log.
21+
22+
The [trial records](trials.jsonl) retain every response, prompt hash, score,
23+
latency and token count. Token counts include each CLI's own system prompt, so
24+
they are much larger than the fixture (3,180 prompt characters for
25+
`contextos`, 2,198 for `handoff`). The [normalization log](fence-normalization.jsonl)
26+
marks five `composer-2.5` responses that arrived as valid JSON inside one
27+
Markdown `json` fence. The orchestrator removed that fence before `record`, as
28+
in the [long-sequence trials](../continuity-long-2026-09-23/README.md); the
29+
stored `raw_response` is the unfenced body. No other response was changed, and
30+
no stored response contains a carriage return.
31+
32+
Each `prompt_sha256` is the SHA-256 of `prepare(scenario, profile)` for the
33+
revision scenario: `dfa69fbfbe9a…` for `contextos` and `c9ced04db08f…` for
34+
`handoff`. The orchestrator recomputed both from this fixture and they match.
35+
36+
## Scores
37+
38+
The table was generated by
39+
`python scripts/continuity-benchmark.py summarize --results docs/evidence/continuity-revision-2026-10-04/trials.jsonl`.
40+
Each value is the mean number of the two questions passing that layer.
41+
42+
| Model | Profile | Trials | Record | Revision | Value | Action | All layers | Format failures |
43+
|---|---|---:|---:|---:|---:|---:|---:|---:|
44+
| composer-2.5 | contextos | 3 | 2.00 | 2.00 | 2.00 | 2.00 | 2.00 | 0 |
45+
| composer-2.5 | handoff | 3 | 2.00 | 2.00 | 2.00 | 2.00 | 2.00 | 0 |
46+
| gpt-6.1-sol | contextos | 3 | 2.00 | 2.00 | 2.00 | 2.00 | 2.00 | 0 |
47+
| gpt-6.1-sol | handoff | 3 | 2.00 | 2.00 | 2.00 | 2.00 | 2.00 | 0 |
48+
| grok-4.7-low-fast | contextos | 3 | 2.00 | 2.00 | 2.00 | 2.00 | 2.00 | 0 |
49+
| grok-4.7-low-fast | handoff | 3 | 2.00 | 2.00 | 2.00 | 2.00 | 2.00 | 0 |
50+
51+
Every trial passed every layer on both questions. In the `contextos` profile,
52+
each model selected `export-002` and `launch-002` over the superseded session
53+
records, cited the current source and its revision, reported the supplied
54+
`mismatch` result and proposed the authorized action. No trial proposed the
55+
launch announcement.
56+
57+
## What this does and does not show
58+
59+
At this difficulty the scenario does not separate these three models: there is
60+
no variance to compare. That is a ceiling result, not evidence that the models
61+
handle stale context in general. The fixture is small, the superseded records
62+
are labelled as superseded in the session text, and the prompt tells the agent
63+
to prefer the current source. A useful next version needs a harder case, for
64+
example a superseded record whose source text does not say it is superseded,
65+
a stale record in the `handoff` profile, or longer context with the replacement
66+
far from the question.
67+
68+
These are supplied synthetic records scored offline. They do not observe host
69+
retrieval, execute a dependency graph, authenticate a reviewer or send an
70+
action. Three trials per cell and three models are a small sample.
Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,5 @@
1+
{"model_id": "composer-2.5", "profile": "contextos", "trial_index": 1, "normalization": "removed one Markdown json fence"}
2+
{"model_id": "composer-2.5", "profile": "contextos", "trial_index": 2, "normalization": "removed one Markdown json fence"}
3+
{"model_id": "composer-2.5", "profile": "contextos", "trial_index": 3, "normalization": "removed one Markdown json fence"}
4+
{"model_id": "composer-2.5", "profile": "handoff", "trial_index": 1, "normalization": "removed one Markdown json fence"}
5+
{"model_id": "composer-2.5", "profile": "handoff", "trial_index": 3, "normalization": "removed one Markdown json fence"}

0 commit comments

Comments
 (0)