Skip to content

Repository files navigation

Kerness — Kernel for Harness

Kerness — Kernel for Harness.
The framework an AI harness sits on, assembled from plug-and-play components.
A Rust crate, with Python bindings over the same kernel.

CI License: MIT Rust 1.88+ Python 3.10+


What Kerness is

A harness is everything wrapped around a language model that turns it into a working system: who speaks and in what order, which tools are reachable, what counts as finished, what gets remembered, and what the run finally returns. A debate between three agents is a harness. A research pipeline is a harness. A code-review bot, a negotiation simulator, a poker table with three seats — all harnesses.

Building one from scratch means writing the same substrate every time. Provider transport and retries. Tool-call parsing across three incompatible dialects. Prompt assembly. A turn loop with phases and termination conditions. Access control on anything that touches the filesystem. Memory. Context compaction. Crash-resumable state. That substrate is where the weeks go, and none of it is the harness you actually wanted to build.

Kerness is that substrate — the kernel. It owns every piece listed above and exposes them as components you plug together. What is left for you is the part that is genuinely yours: a Markdown file declaring how your harness behaves, and whatever tools you want to hand the agents.

The name is the design: a kernel for a harness.

Two artifacts, one kernel

Kerness is a Rust crate. crates/kerness/ links no Python, spawns no threads, and runs a whole session on the calling thread — a stack trace from inside a tool handler reaches back to Session::run.

bindings/python/ is a binding: a thin PyO3 layer that decides nothing. It forwards into the same kernel, so a session driven from Python takes the same code path as one driven from main(). What Python adds is what Python callers expect — a Provider you subclass, a lambda as a tool handler, a pydantic model for structured output.

Neither surface is a wrapper around the other's use case, and neither is the "real" one. Pick the language; the framework is the same.

That last part is a rule and not an aspiration: a feature is written in Rust. The installed Python package holds the classes callers subclass, the few the extension cannot declare, and re-exports — never behaviour. Where a feature needs something only the interpreter has, the crate names the need as a trait and the binding installs it at import, so capsys captures a console channel, caplog sees a warning, and mock.patch intercepts a request, with the decision still made in one place.

The split

The kernel owns Your harness declares
Provider transport, retries, backoff Which models each agent uses
Tool dialects (OpenAI / Anthropic / text-fence) Which tools exist and what they do
Prompt assembly and ordering Roles, personas, system prompts, language
The orchestrator loop, phases, turn counting Phase names, round counts, instructions
Termination detection Which tokens end a run
Access decisions on commands and paths The policy those decisions are made against
Memory read/write, context compaction Whether the run may write memory
Session files and resume Where the state file lives
Skill discovery and progressive disclosure Which skills an agent may load

Nothing in the left column is something you should have to write again. Nothing in the right column is something the framework should decide for you.

Plug-and-play components

Every component is an interface with a working default. Use the default, or swap in your own — the rest of the kernel does not notice.

Component Ships with Swap it by
Provider OpenAiProvider, ClaudeProvider, OpenRouterProvider, CustomProvider; OAuth credentials where the vendor offers them implementing the Provider trait — one required method, chat
Channel ConsoleChannel, FileChannel, LogChannel, MultiChannel implementing Channel — one required method, send
Tools cmd, read_file, list_dir, write_memory add_tool for legacy handlers, add_tool_spec for complete specs, or add_contextual_tool for scoped capabilities
Skills challenge, fact-check, summarize, agent-browser dropping a SKILL.md directory on disk
Roles participant, orchestrator a .md file whose frontmatter declares a position:, or inline prose
Personas pragmatic_engineer, devils_advocate a .md file, or inline prose
Gameplans debate, discussion, research a new Markdown file — see below
Access closed by default; a workspace that grants its own contents, glob and regex command allow-lists, and path allow-lists that reach past the workspace an AccessPolicy, plus an approval callback
Memory FileMemory — a plain .md file per scope, read-only unless asked; SummarizingMemory — recent notes verbatim, the rest folded into a running summary at the end of the run implementing MemoryStore — two required methods, read and append — and passing it as memory_store
Session file Versioned JSON snapshots of turns, suspended tools, approvals, and usage session_file — absent means persist nothing

The names are the Rust ones. Python spells the two acronym providers the way Python callers expect — OpenAIProvider, plus OpenAIOAuthProvider and ClaudeOAuthProvider — and a trait to implement becomes a class to subclass; everything else carries the same name in both.

A CustomProvider pointed at any OpenAI-compatible endpoint covers most local inference servers without implementing anything at all.

Providers may speak native tool calling or fall back to text fences. The dialect is resolved per agent, so one session can mix an Anthropic model, an OpenAI model, and a local endpoint that supports no tool calling at all, and the orchestrator sees one normalized stream of calls.

A harness is a Markdown file

The YAML frontmatter is a machine-readable contract the framework validates and enforces. The body below it is the orchestrator's manual, in prose.

---
name: debate
description: Adversarial debate, then a revisit against a neutral summary.
agents:
  orchestrator: { required: true }
  participants: { min: 2, max: 6 }
loop:
  max_turns: 50
  max_rounds: 3
  terminate_on: [END_SESSION, CONSENSUS_REACHED]
  phases:
    - name: think
      rounds: 1
      instruction: Give your own independent opinion. Do not rebut anyone yet.
    - name: argue
      rounds: 2
      instruction: Choose a side and present a forceful argument for it.
    - name: rethink
      rounds: 1
      rethink: true
      instruction: Re-examine your opening position against the summary.
result:
  consensus: { type: bool, description: Whether participants converged. }
  summary:  { type: str,  description: The final neutral summary. }
---

# Debate

You are running an adversarial debate. Your job is to make the disagreement
productive, not to resolve it prematurely.

Everything in that contract is enforced, not advisory: a session with one participant is refused before the first API call, a tools: entry naming a tool nobody registered is an error rather than a silent drop, and every problem is reported at once instead of one per run. The declared result: fields come back as typed values on the session result.

A new harness is a new Markdown file. It is not a new runtime, not a subclass, and not a fork.

Install

Rust — MSRV 1.88, no build script, no system libraries:

[dependencies]
kerness = { git = "https://github.com/xwings/kerness" }

Pythonabi3 wheels from CPython 3.10 up, so one build covers every supported interpreter:

pip install kerness                  # runtime
pip install 'kerness[structured]'    # plus pydantic, for OpenAIProvider(output_type=...)

From a checkout, either side. The root is a Cargo workspace, so a Python build starts from the binding's own directory:

cargo test --workspace                           # the kernel

cd bindings/python && maturin develop            # the binding
python -m kerness.selfcheck                      # pass = "OK: all core checks passed"

Set it once on the session, override it per agent

The session is configured before any agent exists, and everything it carries — provider, model, reasoning effort, persona, language, system prompt, memory scope, workspace — is a default. An agent that names none of them inherits every one; an agent that names one overrides it, for itself alone. That is the whole rule, and it has two deliberate exceptions.

Provider and model inherit as a pair. A model name only means something on the backend it was written for, so an agent that brings its own provider must name its own model; inheriting the session's would silently ask one vendor for another's model. It is an error at run(), naming the agent.

The workspace only ever narrows. A session's workspace grants its own contents — every path under it is readable without an allow-list entry — and it is the working directory a command starts in. Unset, it is the directory the program was launched from. allowed_dirs and allowed_files reach past it, which is how a session confined to one project still reads /tmp; the two together are the whole of what a session can touch, and an approval callback cannot add to them. An agent may set a workspace of its own, and it is intersected with the session's rather than replacing it — otherwise an agent stanza would be a way to hand itself more of the filesystem than the session was given.

One further asymmetry, and it is not about inheritance: role has no session default at all. A session-wide role would make every agent the orchestrator at once. role is what an agent is in the session — its position and its job — and it is a built-in name, a path to a .md role file, or that job written out as prose. persona is a different question, who the agent is, and it reaches the prompt and nothing else. Unset, role seats a participant, and prose seats a participant too: only a role file declaring position: orchestrator in its frontmatter can seat the chair, so privilege comes from a declaration and never from a substring somebody wrote.

A run, in Rust

Two providers, four agents, a tool, and two skills — the whole configuration contract in one file.

use std::sync::Arc;

use kerness::access::AccessPolicy;
use kerness::provider::{
    ClaudeConfig, ClaudeCredential, ClaudeProvider, OpenAiConfig, OpenAiProvider,
};
use kerness::tooling::Arguments;
use kerness::{Agent, ConsoleChannel, Provider, ReasoningEffort, Session, SessionConfig};
use serde_json::{json, Value};

let openai: Arc<dyn Provider> = Arc::new(OpenAiProvider::new(OpenAiConfig {
    api_key: std::env::var("OPENAI_API_KEY")?,
    ..Default::default()
})?);
let claude: Arc<dyn Provider> = Arc::new(ClaudeProvider::new(ClaudeConfig {
    credential: ClaudeCredential::ApiKey(std::env::var("ANTHROPIC_API_KEY")?),
    ..Default::default()
}));

// Everything on the config is a default. Agents fill in from it at `run()`.
let mut session = Session::new(SessionConfig {
    gameplan: "research".to_string(),
    topic: "Should the cache be write-through?".to_string(),
    provider: Some(openai),
    model: Some("gpt-4o".to_string()),
    reasoning_effort: ReasoningEffort::High,
    channel: Some(Arc::new(ConsoleChannel::default())),
    memory: "/srv/work/notes.md".to_string(),     // the scope the store is asked for
    memory_write: true,                           // ...and, here, may append to
    memory_store: None,                           // None is FileMemory: a scope is a path
    session_file: Some("/srv/work/run.json".to_string()), // None persists nothing
    access_policy: Some(AccessPolicy {
        workspace: Some("/srv/work".to_string()),  // this tree, and nothing else
        allowed_dirs: vec!["/tmp".to_string()],    // ...except what is named here
        allowed_commands: vec!["rg *".to_string()],
        ..AccessPolicy::new()
    }),
    ..Default::default()
})?;

// A tool of your own. The handler is a closure; the kernel translates the
// schema into whichever dialect each agent's provider speaks, parses the call
// back out, feeds the result in, and stops a model that loops on bad calls.
session.add_tool(
    "lookup_price",
    "Look up the current price of a ticker.",
    json!({"type": "object",
           "properties": {"ticker": {"type": "string"}},
           "required": ["ticker"]}),
    Arc::new(|args: &Arguments, _actor: &str| {
        let ticker = args.get("ticker").and_then(Value::as_str).unwrap_or_default();
        Ok(format!("{ticker} is at 41.20"))
    }),
)?;
session.add_skill("fact-check")?;
session.add_skill("summarize")?;

// Alice names nothing but a persona, so she takes every default above.
session.add_agent(Agent {
    persona: Some("pragmatic_engineer.md".to_string()),
    ..Agent::new("Alice")
})?;

// Bob overrides one thing — a cheaper model on the same backend — and confines
// himself to a directory inside the session's workspace.
session.add_agent(Agent {
    persona: Some("devils_advocate.md".to_string()),
    workspace: Some("/srv/work/scratch".to_string()),
    ..Agent::new("Bob").with_model("gpt-4o-mini")
})?;

// Carol brings her own vendor, so she must name her own model: "gpt-4o" means
// nothing to Anthropic, and inheriting it silently would be the wrong answer.
// `with_provider` takes both for that reason. Effort is portable, so it
// inherits or overrides on its own.
session.add_agent(Agent {
    reasoning_effort: Some(ReasoningEffort::Medium),
    skills: Some(vec!["fact-check".to_string()]),   // Alice and Bob get both
    ..Agent::new("Carol").with_provider(claude, "claude-sonnet-4-5")
})?;

// The chair, seated by a role file that declares `position: orchestrator`.
session.add_agent(Agent::new("Mod").with_role("orchestrator"))?;

let result = session.run()?;
println!("{}", result.summary());
println!("{:?}", result.fields.get("findings"));   // the gameplan's declared fields
println!("{} rounds, ended on {}", result.rounds_run, result.end_reason);

Full file: crates/kerness/examples/debate.rscargo run -p kerness --example debate.

To watch a session run without an API key, use offline_debate, which drives the debate gameplan against a scripted provider:

cargo run -p kerness --example offline_debate    # no key, no network

Other examples cover per-agent providers, memory, structured output, custom tools and channels, host control, and durable approvals.

The same run, in Python

The binding mirrors the crate, in the shapes Python callers expect: keyword arguments instead of a config struct, a plain callable instead of a closure in an Arc, a dataclass instead of an AccessPolicy literal. Same kernel, same resolution rules, same errors.

import os

from kerness import (
    AccessPolicy, ClaudeProvider, ConsoleChannel, OpenAIProvider, Session,
)

openai = OpenAIProvider(api_key=os.environ["OPENAI_API_KEY"])
claude = ClaudeProvider(api_key=os.environ["ANTHROPIC_API_KEY"])

session = Session(
    gameplan="research",
    topic="Should the cache be write-through?",
    provider=openai,
    model="gpt-4o",
    reasoning_effort="high",
    channel=ConsoleChannel(),
    memory="/srv/work/notes.md",
    memory_write=True,
    session_file="/srv/work/run.json",  # resumable; omit to persist nothing
    access_policy=AccessPolicy(
        workspace="/srv/work",      # this tree, and nothing else
        allowed_dirs=["/tmp"],      # ...except what is named here
        allowed_commands=["rg *"],
    ),
)

prices = {"KRN": "41.20"}
session.add_tool(
    "lookup_price",
    "Look up the current price of a ticker.",
    {"type": "object",
     "properties": {"ticker": {"type": "string"}},
     "required": ["ticker"]},
    lambda args: prices.get(args["ticker"], "unknown"),
)
session.add_skill("fact-check")
session.add_skill("summarize")

session.add_agent("Alice", persona="pragmatic_engineer.md")
session.add_agent("Bob", persona="devils_advocate.md",
                  model="gpt-4o-mini", workspace="/srv/work/scratch")
session.add_agent("Carol", provider=claude, model="claude-sonnet-4-5",
                  reasoning_effort="medium", skills=["fact-check"])
session.add_agent("Mod", role="orchestrator")

result = session.run()
print(result.summary)
print(result.fields["findings"])   # the gameplan's declared fields
print(result.rounds_run, result.end_reason)

The add_* calls return the session, so registration chains if you would rather write it that way.

Host-controlled runs

Session::start(RunOptions) transfers configuration into an owned SessionRun. Choose RunMode::HostDriven to select participants yourself; an orchestrator is needed only when the gameplan explicitly requires one. step(RunInput) returns progress, an identified waiting state, or a typed terminal outcome with partial history, result diagnostics, usage, and any original framework error.

Each step dispatches at most one engine-selected logical provider operation, tool invocation, compaction, or maintenance scope, and may settle several local effects. Providers may retry synchronously, and callbacks can call provider APIs; supplied metering seams count those calls against the run. Events report progress; inputs select an agent, add user text at a turn boundary, answer an approval, reconcile an interrupted action, or finish with a host-supplied result. Finish validates the declared result without an implicit agent or judge call to generate it; configured memory maintenance may still call a provider. A separate RunControl requests cooperative cancellation.

Custom contextual handlers receive trusted actor/run/turn/call identity and file, command, and memory capabilities. Those handles enforce the actor's policy and expire when the invocation ends. A side-effect-free preflight can request confirmation before the handler runs. A saved pending approval resumes with its same request ID and completed tool results; an interrupted action with an unknown outcome requires reconciliation instead of automatic replay.

RunOptions also accepts operation/tool limits, host-supplied token pricing, and measured token/cost/time thresholds. Unsupported hard token/cost caps are rejected. Cancellation cannot forcibly interrupt an arbitrary blocking provider or user callback.

Run the complete offline examples:

cargo run -p kerness --example host_control
cargo run -p kerness --example resume_approval
python bindings/python/examples/host_control.py

host_control.rs selects one participant, observes events, and supplies a validated result. resume_approval.rs saves a pending confirmation, drops and rebinds the run, then completes two tools once each. Both create and clean their own temporary workspaces.

The Python example uses the same engine through session.start(mode="host_driven") and run.step(...). Inputs are dictionaries such as {"kind": "select_agent", "agent": "Advisor", "instruction": "Recommend a policy."}; outcomes carry status: "progress", "waiting", or "finished". ARCHITECTURE/bindings.md documents the thin API, callback signatures, and handle lifetimes.

What the kernel does while it runs

Neither of the above is a different runtime, so the following holds for both.

Skills use progressive disclosure: prompts carry only names and descriptions, and the full instructions load on demand through a turn-local Skill tool. A skill also says which tools the turn should hold: allowed-tools: narrows it, requires-tools: adds back what the skill cannot work without, out of whatever the session registered. A skill requiring a tool nobody registered is refused before the first call rather than quietly doing nothing. When the policy trusts bundles, loading a skill also grants read access to its own directory.

There is no daemon and no server. A run given a session file saves coherent execution boundaries and resumes from that file when the host reattaches its providers and handlers. Resume checks identity and the saved runtime contract before continuing; valid version 1 turn-boundary snapshots are migrated.

Layout

Cargo.toml       # the only manifest at the root
crates/kerness/  # the kernel, pure Rust — no PyO3, no Python
  src/           #   kernel implementation, with unit tests inline
  tests/         #   integration tests over the public API
  examples/      #   10 runnable Rust harnesses, three needing no key
  assets/        #   bundled gameplans, roles, personas, skills
bindings/python/ # everything the wheel is built from
  pyproject.toml #   the wheel's manifest — `pip install .` runs here
  src/           #   the PyO3 extension module, kerness._core
  kerness/       #   the Python package: the subclassable classes, shims, assets
  tests/         #   pytest suite, over the binding
  examples/      #   runnable Python harnesses
ARCHITECTURE/    # one document per subsystem

The bundled debate, discussion, and research gameplans are worked examples of the contract, not the product.

Testing

Each suite proves its own layer. The Rust integration tests compile against the crate's public API exactly as a dependent does, so a break in that surface fails a test here rather than a downstream build; the pytest suite proves the binding carries the kernel's behaviour across the FFI boundary intact.

cargo fmt --all -- --check
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspace                     # unit and public API integration suites
cargo build -p kerness --examples          # every example still compiles
cargo run -p kerness --example offline_debate   # a whole session, no key

python -m pytest bindings/python/tests -q      # binding suite
python -m kerness.selfcheck                    # exit 0
ruff check bindings/python

CI runs all of it on every push, on Rust stable and the 1.88 MSRV floor, and on Python 3.10 and 3.13.

Releasing

Pushing a v* tag runs Release: it builds the Python wheels and source distribution, installs the source distribution in a clean interpreter, then uploads all artifacts to PyPI and attaches them to a draft GitHub Release. PyPI publishing happens on the tag push; publishing the GitHub Release draft is a separate step.

Configure PyPI once after registering your account. At PyPI Publishing, add a pending publisher for a new project with these exact GitHub settings:

Field Value
PyPI project name kerness
Owner xwings
Repository name kerness
Workflow name release.yml
Environment name pypi

The first successful upload creates the project. If you already own the kerness project on PyPI, add the same Trusted Publisher under that project's Publishing settings instead. GitHub exchanges its identity for a short-lived PyPI credential, so no API token or repository secret is needed.

In GitHub's pypi environment, allow Tags matching v* under deployment branches and tags. A main branch rule alone does not permit tag runs. Leave required reviewers unset for automatic uploads on each release.

For each release:

  1. Set [workspace.package] version in the root Cargo.toml to the release version, for example 0.1.1. Run cargo check --workspace to update Cargo.lock, commit both files, and merge the release changes to main. The Python package derives its version from Cargo; the tag does not set it.

  2. Once CI passes, tag that release commit with the matching version and push:

    git switch main
    git pull --ff-only
    git tag v0.1.1
    git push origin v0.1.1
  3. Check the Release workflow for a successful Publish to PyPI job, then publish the GitHub Release draft.

Use a new version for each release: PyPI does not allow replacing an uploaded file. To check the build before tagging, run Release → Run workflow on a branch; branch runs build and verify artifacts without publishing.

Documentation

ARCHITECTURE.md is the entry point: mission, workspace layout, boot flow, well-known constants, and an index of one document per subsystem under ARCHITECTURE/. Each carries live file:line references and the commands that prove it works.

License

MIT. See LICENSE.

About

A harness is everything wrapped around a language model that turns it into a working system

Resources

Stars

9 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages