Decision Router is an experiment in authority routing for AI systems.
The architectural question it explores is not simply:
Which model should answer?
but:
Which component is allowed to make which kind of decision?
A generative model should not implicitly determine the semantic space, epistemic basis, decision authority, and final expression of the same operation. These responsibilities can instead be represented as separable stages that are independently constrained and validated. That separation — not any single model — is the main architectural idea of this repository, and it is presented here as a hypothesis under exploration, not as an experimentally proven universal law.
A TypeScript/React prototype of a customer-support refund system built as a composed pipeline with explicit stage boundaries, alongside the monolithic baselines it is compared against. The benchmark domain is a known, explicitly defined refund policy (30-day window, active-account requirement, repeat-refund escalation) over fixture facts — so the experiment tests whether the system can reliably apply a stated policy through separated stages, not whether any component invents correct business policy on its own.
The repository contains two pipelines — do not confuse them:
- Live benchmark pipeline (
src/lib/live/pipeline.ts): the architecture under evaluation. Semantic routing → contracts → verified facts → bounded judgment → response composition, with fail-closed branches. It forms decisions and generates responses. It executes no payments or refunds. - Demo pipeline (
src/lib/pipeline/composed.ts): an earlier, broader end-to-end sketch that does include simulated deterministic action execution (executeRefund, flag-for-review) against local mock data, using the heuristicmockJevAdapter. It demonstrates the wider architecture; it is not what the benchmark scores.
Each stage of the pipeline holds exactly one kind of authority, and each authority is constrained by the stages before it:
| Responsibility | Held by | Constrained by |
|---|---|---|
| Semantic interpretation / routing | Semantic Router (LLM) | Fixed route catalog, JSON contract |
Selection of a bounded decision_space |
Semantic Router (LLM) | Known policy domains (plan_change_policy, duplicate_charge, …) |
| Retrieval / establishment of verified facts | Deterministic / source layer | Fixture record (in production: DB/API) — never model memory |
| Bounded semantic judgment | Jev | Permitted decision space, typed questions, confidence threshold |
| Decision-contract validation | Deterministic code | Allowed decisions, reason codes, evidence requirements |
| Consequential action | Out of scope for the benchmark | (Demo pipeline only: deterministic executeRefund) |
| Natural-language generation | Response Composer (LLM) | Already-established decision + guardrail |
The point: no single model call spans interpretation, fact-finding, policy application, and prose at once. The LLM stages interpret and express; they do not authorize. Authorization lives in the judgment plus its deterministic validation envelope.
User Input
↓
Semantic Router (LLM: Big Pickle via the local OpenCode server)
↓
Route Contract (deterministic validation → typed contract or triage)
↓
Verified Facts (deterministic fixture lookup — never model memory)
↓
Jev Bounded Judgment (real Jev Decision API over typed questions)
↓
Decision Contract (deterministic validation: enums, reason codes,
confidence ≥ 0.55, evidence requirements)
↓
Response Composer (LLM: renders the already-made decision into prose)
↓
User Response
Each ↓ between stages is a validation boundary with its own fail-closed
branch (see §8). Stage implementations live in src/lib/live/pipeline.ts;
the dev-server credential boundary lives in server/live.ts.
The LLM classifies the request into a typed route plus a decision_space
(src/lib/live/pipeline.ts, routerPrompt). Routing is intent-only: a
refund request routes to refund_request even when eligibility is unknown,
because eligibility is decided later. The router does not make the
business decision — it never approves, denies, or applies policy.
validateRouter converts the router's JSON into an explicit structured
contract (route, decisionSpace, extractedEntities, ambiguity) or
fails. Unknown routes, malformed JSON, and human_triage routes halt the
run here (status triage/error) instead of forcing a decision downstream.
Facts (purchase_age_days, refund_window_days, account_status,
previous_refunds_count) come from the scenario fixture — the stand-in for
the authoritative system of record. If no facts exist, the pipeline halts
with facts_missing rather than letting a model guess undisputed facts.
Jev receives the user request, the route contract, and the verified facts,
plus two typed questions (refund_decision, reason_code) whose criteria
encode the stated refund policy (jevQuestions). It returns a choice +
confidence per question.
Jev provides a typed judgment layer in which the application's semantic decision space can be made explicit and bounded.
Jev does not "solve truth," and it is not merely a classifier bolted on for labels: it performs the bounded semantic judgment — applying the stated policy to the verified facts inside the permitted decision space — while deterministic validation constrains and checks that judgment. Precisely:
Deterministic systems establish verified facts and enforce contracts; Jev performs bounded semantic judgment within the permitted decision space.
Low-confidence judgments (< 0.55) are demoted to request_review /
irregular_review rather than put on the wire; unknown decisions or
decision/reason mismatches halt with jev_invalid_output; an unreachable
Jev halts with jev_error.
buildDecision plus the decision_contract stage validate the judgment
against the allowed decision set (approve_refund, deny_refund,
request_review), the allowed reason codes per decision, and the evidence.
Only a validated contract reaches the composer.
The LLM translates the already-established decision into a short customer
message (composerPrompt). It has no authority to change, reinterpret, or
override the judgment. A deterministic guardrail (composerGuard) rejects
any message that contradicts the decision (composer_contradiction) —
e.g. approval language under a denial, or any definitive outcome under a
request_review.
Monolithic setups give one model call implicit authority over every responsibility at once: it reads the request, recalls (or invents) facts, applies policy, and writes the email in a single opaque generation. When that call is wrong, there is no stage boundary at which to catch it, and no explicit state in which the system can refuse to decide.
The composed pipeline replaces that single implicit authority with separable, independently constrained stages. The observable payoff measured here is narrow and specific: the composed system has explicit states in which it refuses to produce a consequential decision instead of forcing an answer (§8). Nothing in this repository proves a general claim about AI safety — only that this architecture can represent and enforce explicit failure boundaries on the tested cases.
The eval harness (scripts/eval.ts, scripts/fault-eval.ts) runs a fixed
scenario set against three architectures:
- Router (live) — the pipeline in §3–4. LLM stages call Big Pickle
through the dev middleware (
/api/live/llm→ local OpenCode desktop server); judgment calls the real Jev Decision API (/api/live/jev). - Direct-LLM (CoT + JSON) — one model call over the user input plus the
support record, emitting
{ decision, reason, email }(scripts/lib/cot.ts). - ReAct agent — a tool-calling loop with
get_support_record/execute_refund(scripts/lib/react.ts). - Direct+validator (Tasks 2–3) — Direct-LLM followed by a deterministic,
zero-LLM-call structural validator (
scripts/lib/validator.ts).
Scenario set: 11 live-app fixtures (src/data/scenarios/*.json) plus 9
policy-boundary edge variants mined from the gold rules
(scripts/lib/scenarios.ts), for 20 scenarios. Gold expectations
(src/lib/live/gold.ts) are derived from the same stated policy the Jev
judge is told: approve iff within window on an active account, deny outside
it / when inactive, escalate repeat-refund / ambiguous cases to review, and
halt on corrupt or missing inputs. Scoring is symmetric: unnecessary review
on an unambiguous case fails just as a wrong decision does.
The headline v1 run: 20 scenarios × 3 architectures × 2 repeats (N = 120),
baselines on Gemini flash-lite. See eval_results_summary.md for the full
matrix, and docs/ for the detailed reports (EVAL_REPORT_STANDARD.md,
EVAL_REPORT_STE100.md), the fault-injection fairness record
(docs/fault-injection.md), the model-separation follow-up
(docs/model-separation.md), and the instrumentation notes (docs/rigor.md).
In the current benchmark, the composed architecture matched the direct-LLM baseline on ordinary decision accuracy while outperforming the tested baselines on explicitly defined fail-closed cases.
Concretely (eval-results-v1.json, preserved as eval-results.json):
| Router (composed) | Direct-LLM | ReAct | |
|---|---|---|---|
| Decision accuracy (gold = decide, n = 30) | 86.7% (26/30) | 86.7% (26/30) | 76.7% (23/30) |
| Fail-closed handling (gold = halt, n = 10) | 100% (10/10) | 0% (0/10) | 0% (0/10) |
| Overall (n = 40) | 90.0% (36/40) | 65.0% (26/40) | 57.5% (23/40) |
The overall-accuracy gap is arithmetic, not a second independent finding: it is fully explained by the fail-closed rows, where baselines force a decision on corrupt input. Do not present this as "Decision Router is 90% accurate while LLMs are only 65% accurate" — on well-formed decisions the two tied.
Foundation-model confound. The v1 run does not isolate architecture from
model choice: the router's LLM stages run Big Pickle and its judgment runs
Jev, while the v1 baselines run Gemini flash-lite. V1 results therefore
represent architecture + foundation-model configuration, not
architecture in isolation. A matched-model follow-up (Task 3,
docs/model-separation.md, repeat = 1) re-ran all baselines on Big Pickle:
fail-closed behavior was unchanged on both models (monolithic baselines
0/4–0/5 on faults regardless of model; the deterministic validator halted
4/4 on both), which is consistent with the halt difference being structural,
but a single repeat-1 draw does not establish this conclusively. Later
repeat-5 runs (eval-results-*-rep5.json) cover baselines only and are
reported in docs/rigor.md.
"Fail-closed handling" here means: on explicitly defined failure conditions,
the system halts (status error/triage, no consequential decision
emitted) instead of forcing an answer. The tested conditions are:
- ambiguous routing (
human_triage/ unknown route → halt to triage) - malformed or missing router output (
router_malformed_json) - missing verified facts (
facts_missing) - Jev unreachable (
jev_error) or invalid judgment (jev_invalid_output, including decision/reason mismatch and sub-threshold confidence) - composer output contradicting the decision (
composer_contradiction)
The meaningful architectural observation is that the composed system has explicit states in which it can refuse a consequential decision. Two qualifications apply:
- These cases are architecture-aware: several fault kinds (notably
jev-500, a judgment-stage outage) describe components the baselines do not have, and are excluded from baseline denominators (docs/fault-injection.md). The benchmark therefore demonstrates whether the architecture can represent and enforce explicit failure boundaries — it is not an unbiased universal safety comparison, and it proves nothing about general AI safety. - A prompt-only halt instruction did not change baseline behavior (0/4 told vs 0/4 not-told in the Task-2 fault test), while a deterministic structural validator halted 4/4 on both models. The evidence is consistent with fail-closed behavior coming from structure (validation boundaries), not from wording — at the tested scale.
Demo pipeline (src/lib/pipeline/composed.ts + mock adapters)
- demonstrates the broader end-to-end architecture, including simulated
deterministic action execution (executeRefund / flag-for-review)
- runs entirely locally on mockJevAdapter (keyword heuristic) + mockLlmAdapter
- is NOT scored by the benchmark and says nothing about real Jev
Benchmark pipeline (src/lib/live/pipeline.ts + server/live.ts)
- evaluates semantic routing, bounded judgment, decision contracts,
response composition, and fail-closed behavior
- calls real services (Big Pickle LLM stages, real Jev Decision API)
- executes no payments or refunds; "actions" never leave the trace
- does not constitute a production payment/refund execution system
npm install
npm run devUI stack: React 19 + TypeScript, styled with Tailwind CSS v4 and
DaisyUI (theme defined in src/index.css). Run a preset example, or type
a message and pick a customer — the trace shows which component handled each
step. Toggle Compare with LLM-only to see the same input shot through a
single opaque model call.
Live services need credentials (see .env.example): OPENCODE_SERVER_*
for the Big Pickle runtime, JEV_API_* for the real Jev Decision API,
GEMINI_* for the eval baselines. Without them, only the mock-adapter demo
path runs — and mock-adapter behavior must never be cited as evidence about
Jev.
Benchmark:
# baselines only (needs GEMINI_API_KEY)
npm run eval -- --pipelines cot,react --repeat 2
# including the live router (needs a running dev server + OPENCODE/JEV creds)
npm run eval -- --pipelines cot,react,router --repeat 2 --live-base http://localhost:5173
npm run fault-eval # fault-injection matrix for the baselines- Model configurations differ across architectures in v1 (Big Pickle + Jev vs Gemini flash-lite): results are architecture + foundation-model configuration, not architecture in isolation.
- Benchmark size is small: 20 scenarios (11 fixtures + 9 edge variants); v1 at 2 repeats (N = 120 total runs).
- Single-run results unless otherwise stated; no variance or confidence
interval is established from the v1 run (repeat-5/CI instrumentation is
in
docs/rigor.mdfor baselines). - Fail-closed cases are architecture-aware, not a universal safety test; baselines lack equivalent contract/abstention stages by construction.
- The policy domain is intentionally bounded: a stated 30-day refund policy over four fixture fields. The test is reliable routing + bounded judgment + policy-constrained decision formation — not discovery of business policy.
- The benchmark does not establish that Jev is more accurate than an LLM in general, nor that the composed architecture is causally superior in all settings.
- This is a prototype evaluation, not production reliability: fixture facts instead of live databases, single-turn inputs, no auth/payments infra, and a latency cost (~9.8 s mean for the composed run vs ~2.1 s for Direct-LLM in v1) that is out of scope for correctness claims.
Does establish (at the tested scale, with the stated confounds):
- The composed pipeline tied the direct-LLM baseline on well-formed refund decisions (86.7% / 26/30 each) and halted instead of deciding on all 10 tested failure runs, where the monolithic baselines decided in all 10.
- Explicit validation boundaries can convert corrupt/ambiguous inputs into observable halts rather than forced decisions — including against a prompt-only halt instruction, which changed nothing.
Does not establish:
- That the composed architecture is more accurate in general, or that Jev is intrinsically more intelligent than an LLM.
- That the results isolate architecture from model choice (see §7 confound).
- General AI safety, production readiness, or scientific conclusiveness.
- That all decisions are deterministic: the LLM performs real semantic interpretation in the router and composer, and Jev performs the bounded judgment. Determinism lives in fact establishment and contract enforcement, not in every decision.
src/
types.ts shared domain + trace types
data/mockData.ts local mock customers/subscriptions/invoices/payments/refunds
data/scenarios/*.json 11 live-app fixtures (user input + verified facts + faults)
lib/
deterministic/ pure, rule-based layer (retrieval, totals, duplicates,
eligibility, action that executes the refund — demo only)
adapters/
jev/types.ts JevAdapter contract (classify → distribution + confidence)
jev/mockJevAdapter.ts heuristic stand-in — clearly NOT the real Jev API
llm/types.ts LlmAdapter contract (generate → prose from structured facts)
llm/mockLlmAdapter.ts template stand-in
pipeline/
composed.ts the DEMO workflow (mock adapters, simulated execution)
llmOnly.ts the LLM-only comparison
runner.ts wires adapters + presets; reset demo state
live/
pipeline.ts the BENCHMARK pipeline (real LLM + real Jev, fail-closed)
gold.ts gold expectations + verdict logic
baseline.ts client for the live baseline panes
types.ts live run/stage/contract types
server/live.ts dev middleware: credentials + LLM/Jev/baseline backends
scripts/
eval.ts main benchmark harness
fault-eval.ts fault-injection matrix for baselines
lib/scenarios.ts 20-scenario set (fixtures + 9 edge variants)
lib/cot.ts lib/react.ts baseline implementations
lib/validator.ts deterministic structural validator arm
prompts/ versioned baseline prompts (v1 vs v2-told)
The demo pipeline never imports concrete adapters. lib/pipeline/runner.ts
is the only wiring point:
export const adapters = {
jev: realJevAdapter, // implement src/lib/adapters/jev/types.ts against real Jev
llm: realLlmAdapter, // implement src/lib/adapters/llm/types.ts against a real model
}The mock adapters are explicitly not the real Jev/LLM APIs:
The mock adapter exists to make the demo runnable without external Jev credentials. Benchmark claims about the Jev-backed architecture should not be inferred from mock-adapter behavior.
A real adapter maps the actual provider API onto the same narrow interface, so no pipeline or UI changes are needed.
- Semantic routing — LLM maps input to a route +
decision_space; no policy applied. decision_space— the bounded policy domain governing the request (e.g.plan_change_policy).- Verified facts — fixture/DB record; the epistemic basis. Never model memory.
- Bounded judgment — Jev applies the stated policy to the facts inside the decision space and returns a typed choice + confidence.
- Decision contract — the validated judgment (decision + reason code + confidence + evidence) that alone may reach the composer.
- Response composition — LLM renders the contracted decision into prose; zero decision authority.
- Fail-closed handling — halting (triage/error) instead of forcing a decision on corrupt/ambiguous input. Preferred over unqualified "safety".
- Deterministic execution — simulated refund/flag actions in the demo pipeline only; the benchmark pipeline executes nothing consequential.
Known code-level inconsistency (flagged, not changed): gold.ts expects
reason code regular_review for ambiguous-intent and missing-facts
escalations, while pipeline.ts (REASON_ALLOWED) permits only
irregular_review for request_review — the live pipeline can never emit
the gold reason code on those paths. The verdict logic treats a triage halt
on a request_review gold as a pass, so scoring is unaffected, but the
reason-code vocabulary is incoherent across the two files and should be
unified deliberately, not silently.