Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions doc/rfc/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@ Design documents and technical proposals, grouped by scope. Shared/cross-cutting
- [Extension Contract](submitqueue/extension-contract.md) - When extensions take orchestrator identity (request/batch) and resolve granular content themselves vs. take controller-resolved data; revises the BuildRunner base/head contract
- [Gateway Status and List APIs](submitqueue/status-list-api.md) - Gateway-owned request context, materialized current status, sqid or change-URI status lookup, and queue admission listing
- [Speculation](submitqueue/speculation.md) - Why SubmitQueue speculates, the path/tree model, and the two pluggable seams: speculation-tree enumeration and path selection
- [Outcome Scorer](submitqueue/outcome-scorer.md) - How likely a batch is to reach Succeeded: `Score(ctx, batch, paths)` as a logit-linear model — a base content price plus YAML weights on path and batch evidence (`pathPassed`, `pathFailed`, `landing`, `cancelling`)
- [Best-First Speculation Path Generation](submitqueue/speculation-generator-best-first.md) - The default Generator: per-head lazy streams of flip subsets merged best-first across heads, log-probability ranking, and the strict snapshot contract
- [Modular Queue Wiring](submitqueue/modular-queue-wiring.md) - Declare-don't-assemble engine (`pipeline.Construct`) that unifies topic registry, controller registration, DLQ pairing, and lifecycle ordering into one typed call; services self-declare via Deps struct + Stages slice, hosts own per-queue profiles and transport

Expand Down
126 changes: 126 additions & 0 deletions doc/rfc/submitqueue/outcome-scorer.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,126 @@
# Outcome Scorer
Comment thread
behinddwalls marked this conversation as resolved.

How likely a batch is to reach Succeeded, given the change's content plus what this speculate run has already observed.

See [speculation.md](speculation.md) for batches, paths, heads, and the Speculator. This document specifies the dependency probability the default Generator uses to rank paths.

## The idea

`bestfirst` ranks a path by the probability that every unresolved dependency assumption holds. For each dependency it calls **one** extension: `Scorer.Score(ctx, batch, paths)` — the probability of *succeeds*; it uses the complement for *fails*.

That number has two parts, composed as one scorer:

1. A **base** price for the change from content signals such as its size. Heuristic and composite supply this. They implement the same `Score` and ignore `paths`.
2. An **evidence** layer that revises the base with facts the speculate run already holds: a path *passed*, a path *failed*, the batch is *landing*, the batch is *cancelling*.

Evidence is the scorer the queue exposes. The base sits under it. There is no sibling Predictor factory.

This is a logit-linear model (a GLM with a logit link) with configured weights:

```
logit(p') = logit(p_base) + sum_i w_i x_i
```

`p_base` is the heuristic or composite price (the offset). `x_i` are binary features from the path set and batch state. `w_i = log(factor_i)` are YAML weights, not a fitted likelihood.

**Default is a no-op.** Every factor starts at `1` (`w = 0`), so `Score` returns the base price until someone sets a factor.

## What a factor is

A factor is an odds multiplier for one feature. It is not itself a probability: `10` does not mean `0.10`, and `0.3` does not mean the batch is 30% likely to succeed. Equivalently `w = log(factor)` on the logit.

| Value | Weight | Meaning |
| --- | --- | --- |
| `1` | `0` | Leave the base price alone (the default if the key is omitted) |
| greater than `1` | positive | More likely to reach Succeeded |
| between `0` and `1` (exclusive) | negative | Less likely to reach Succeeded |

Config accepts any finite value greater than `0`; there is no finite upper cap.

The unconfigured base prices every batch at `0.5`. From that price, one factor `f` produces `sigmoid(logit(0.5) + log(f)) = f / (1 + f)`:

| Factor | Price |
| --- | --- |
| `1` | 0.50 |
| `10` | ~0.91 |
| `12` | ~0.92 |
| `0.3` | ~0.23 |
| `0.25` | 0.20 |

A base price of `0.6` with `pathPassed: 10` becomes about `0.94`. `landing: 12` on top of that becomes about `0.995`.

`pathFailed: 0.3` from `0.5` becomes about `0.23`. It applies at most once because the path set has one current entry for the all-*succeeds* path; retry attempts replace that entry rather than adding evidence. `0` is rejected because it would pin matching batches at probability 0.

Odds revision keeps the result in `(0, 1)` and makes the same factor mean the same thing at any base price. Neutral factors return the base unchanged, including `0` or `1`. Adding to the probability provides neither property. A clamp at a small epsilon is a numerical guard around those endpoints, not part of the linear predictor.

This is **not** a fitted GLM: weights are configured, not trained; there is no extra intercept (`p_base` is the offset); features are hand-defined, not learned.

YAML. Evidence is the scorer; the content provider is `base`:

```yaml
scorer:
type: evidence
factors:
pathPassed: 10
pathFailed: 0.3
landing: 12
cancelling: 0.1
base:
type: heuristic
```

The example values above are guesses, for reading the tables. The shipped default is to omit `factors` (every factor `1`). Omitted `type` is `evidence`. Omitted `base` is the default heuristic. `type: heuristic` and `type: composite` belong on `base` (and on composite `components`), not at the top level.

Profiles may set `factors` under `defaults.scorer` and revise them per queue. An omitted key keeps the inherited value — from defaults, or `1` when neither side named it. A queue `scorer` block overlays named factor keys; a present `base` replaces the default base wholesale.

## Evidence

| YAML key | When it applies | Typical direction |
| --- | --- | --- |
| `pathPassed` | Once, if a path that assumes every dependency *succeeds* has *passed* | Up |
| `pathFailed` | Once, if the path that assumes every dependency *succeeds* has *failed* | Down |
| `landing` | While the batch is *landing* | Up |
| `cancelling` | While the batch is *cancelling* | Down |

`bestfirst` already treats a terminal batch as a fact (*Succeeded*, *Failed*, *Cancelled*). The scorer is not asked. *Landing* is not terminal: a land can still fail, so how much it is worth stays a price.

### Only the *succeeds* path counts

The batch being priced is itself a head, so the run may have built it more than once under different assumptions about *its* dependencies. Only one of those builds is evidence.

Take `C` depending on `B`, and `B` depending on `A`. Ranking `C`'s candidates needs the probability that `B` reaches Succeeded, so the Generator calls `Score` with `B` and `B`'s path set. That set can hold two finished builds:

| `B`'s path | What was compiled |
| --- | --- |
| `B` with `A` *succeeds* | `B` on top of `A`'s changes |
| `B` with `A` *fails* | `B` without them |

`B` lands after `A` does, so the first build is a build of the code that will actually land: if it *passed*, `B` is likely to land, and `pathPassed` applies.

The second is a different set of changes. `B` may call something `A` introduces and fail to compile on its own — a *failed* result that says nothing about `B` landing in the normal case. Counting it would push `B` down the ranking over a build it was never going to need, while a green build of the real combination sits in the same set.

So `pathPassed` and `pathFailed` both look only at paths that assume every dependency *succeeds*. Results on any other path are skipped. This is a filter on which results are evidence, not a check on whether an assumption came true — nothing here revisits that.

## Rejected alternatives

Design choices a reader might suggest after the sections above. Each names the alternative, why it fails here, and what this RFC does instead.

### A sibling Predictor factory

Keep `Score(ctx, batch)` for content and add `Predict(ctx, batch, paths)` as a second extension. Ranking only needs one number per unresolved batch; two factories duplicate the per-queue seam. **Instead:** one `Scorer.Score(ctx, batch, paths)`. Evidence is a scorer implementation that wraps a base.

### Put `paths` only on heuristic and composite

Every content backend reads the path set. We tried forwarding path sets through composite: components discarded them. **Instead:** heuristic and composite implement the same `Score` and ignore `paths`. Evidence is the layer that reads them.

### One fitted model for content and evidence

Train a single estimate over diff shape and build outcomes together. Content signals and situation signals change at different rates, need different amounts of data, and would force every queue onto the same content scorer. **Instead:** the base stays per-queue; evidence weights layer on in YAML. `p_base` remains the GLM offset if someone later fits `w`.

### Treat *landing* and *cancelling* as settled in the Generator

Rank a *landing* batch like Succeeded and a *cancelling* batch like Cancelled. We tried and reverted: a land can still fail, so the rank was wrong once outcomes diverged. **Instead:** only terminal states short-circuit in the Generator; *landing* and *cancelling* are scorer features (see [Evidence](#evidence)).

### Let the scorer read the path-set store

`Score` loads path sets from storage on each call — smaller API, fewer parameters. Each call can see a different snapshot mid-run (stale or split-brain relative to the rank the Generator is building). **Instead:** the speculate run reads all path sets as one snapshot and passes the matching set for each dependency.
8 changes: 4 additions & 4 deletions doc/rfc/submitqueue/speculation-generator-best-first.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,9 +47,9 @@ The batch being built is written before its assumptions. For example, `C [A succ

## The snapshot is a caller precondition

`Generate` receives the queue's live batches as a snapshot and takes it as given. A well-formed snapshot carries unique, non-empty batch IDs, includes every batch a head's direct dependencies reference, and gives no head an empty, duplicate, or self dependency. Those are preconditions the caller owns, established where the snapshot is assembled. The generator does not re-check them: it is on the hot path of every run, the checks it could make are the ones an assembled-correctly snapshot can never fail, and paying for them here only spreads the same contract across two places. A malformed snapshot yields undefined candidates rather than an error.
`Generate` receives the queue's live batches and path sets as one snapshot and takes it as given. Path sets are what each batch's builds have done so far — at most one per head, none for a batch nothing has speculated on. The speculate controller assembles both halves once per run and never re-reads mid-run; see [outcome-scorer.md](outcome-scorer.md) for why the scorer does not load them itself. A well-formed snapshot carries unique, non-empty batch IDs, includes every batch a head's direct dependencies reference, and gives no head an empty, duplicate, or self dependency. Those are preconditions the caller owns, established where the snapshot is assembled. The generator does not re-check them: it is on the hot path of every run, the checks it could make are the ones an assembled-correctly snapshot can never fail, and paying for them here only spreads the same contract across two places. A malformed snapshot yields undefined candidates rather than an error.

A dependency the generator cannot price is the one bad input it absorbs, because it arrives from the injected scorer rather than from the caller and there is no earlier point that could catch it. Three cases take the same 0.95 default: a score outside `[0, 1]` or `NaN`, a scorer call that returned an error, and a dependency the snapshot never carried. The default is optimistic on purpose, so a dependency nobody could estimate keeps its head's preferred path near the front instead of burying it or failing the whole run on one number — and failing the whole run is the real hazard, because `Generate` seeds the heap for every head at once, so one unpriceable dependency would otherwise cost the queue every candidate it had. A batch absent from the snapshot is never passed to the scorer at all: it would resolve to a zero-valued batch belonging to no queue, so scoring it would price some other batch entirely or fail on the empty queue name. Any deliberate defaulting still belongs to the scorer implementation, which knows what information it does and does not have; this is only the floor under it.
A dependency the generator cannot price is the one bad input it absorbs, because it arrives from the injected scorer rather than from the caller and there is no earlier point that could catch it. Three cases take the same 0.95 default: a probability outside `[0, 1]` or `NaN`, a scorer call that returned an error, and a dependency the snapshot never carried. The default is optimistic on purpose, so a dependency nobody could estimate keeps its head's preferred path near the front instead of burying it or failing the whole run on one number — and failing the whole run is the real hazard, because `Generate` seeds the heap for every head at once, so one unpriceable dependency would otherwise cost the queue every candidate it had. A batch absent from the snapshot is never passed to the scorer at all: it would resolve to a zero-valued batch belonging to no queue, so scoring it would price some other batch entirely or fail on the empty queue name. Any deliberate defaulting still belongs to the scorer implementation, which knows what information it does and does not have; this is only the floor under it.

## Step 1: `Generate` prepares each head

Expand Down Expand Up @@ -420,7 +420,7 @@ A and D tie at 1.0, so batch ID puts A first. Other exact ties prefer fewer flip

`Generate` must eagerly:

- score every unique unresolved direct dependency needed by an eligible head, substituting the default for any score that is not a probability;
- predict every unique unresolved direct dependency needed by an eligible head, substituting the default for an unusable probability;
- choose each unresolved dependency's preferred assumption and calculate its `flipCost`; and
- total the best score for every head.

Expand All @@ -447,7 +447,7 @@ The ordering stays the same. `CandidatePath.RankingScore` contains this logarith
- `Succeeded` fixes an assumption to succeeds.
- `Failed` or `Cancelled` fixes an assumption to fails.
- `Cancelling` remains undecided because cancellation may lose a race with completion.
- `Landing` also remains undecided, because a land can fail. It is tempting to treat it as committed to landing and skip the scorer call, but that puts a state-specific policy inside the search: whether a path betting against a landing batch is worth funding is a question of price, and price belongs to the scorer. The allocator draws the same line — "no batch state enters this decision" — and the generator holds it too. Nothing is lost by staying open: a single passed path still waits for the land result, while passed paths covering every outcome let the controller bypass the dependency (see [speculation.md](speculation.md)). Funding the unlikely side spends budget, which is the allocator's to ration.
- `Landing` also remains undecided, because a land can fail. It is tempting to treat it as committed to landing and skip scoring, but that would turn an uncertain state into a fact inside the search. How much *landing* changes the probability belongs to the scorer; the Generator only consumes that price, and the Allocator still does not interpret batch state. Nothing is lost by staying open: a single passed path still waits for the land result, while passed paths covering every outcome let the controller bypass the dependency (see [speculation.md](speculation.md)). Funding the unlikely side spends budget, which is the Allocator's to ration.
- A fixed assumption stays in the returned path but contributes probability 1 and has no flip.
- A shared dependency is scored once per run.

Expand Down
Loading
Loading