Read README.md for repo concepts and instructions for running evals locally.
First, determine the eval suite for your scenario:
- Regression evals are suitable for most scenarios. If we notice agents make a narrow mistake, we track it here to reproduce the issue, verify a fix, and monitor for regression. These scenarios are not included in the benchmark so they don't inflate scores.
- Benchmark evals are scenarios we've intentionally selected for the published benchmark report. These should be representative of the user journey on Supabase to cover a breadth of dimensions.
- Docs evals are owned by docs team for their own analysis of how agents interpret docs pages, refreshed as-needed.
- CLI evals are owned by the CLI team: one scenario run unchanged across forced CLI environments (Docker available / daemon unreachable / absent; pinned, stable and beta CLI) via the
cliexperiment suite.
Then add a folder under evals/<suite>/ containing:
PROMPT.mdwith frontmatter metadata and the task the agent sees.EVAL.tswith the scorer.- Optional
remote/data when the scenario needs to seed hosted project state, such as database, logs, or functions. - Optional
local/files when the scenario needs to seed a local filesystem, such as a localsupabase/project.
If your scenario contains anything not self-explanatory, consider adding a README.md to the folder with a brief explanation of how it's set up and what it's testing.
Every new scenario needs a motivation: defined in PROMPT.md frontmatter that cites some evidence for the scenario being a part of the Supabase user journey, ideally a pain point. Examples include support tickets, GitHub issues, Linear issues, or social media threads.
For new benchmark scenarios, we need to see at least one agent, ideally more, failing the new scenario to ensure we're getting signal from results. If agents are already acing your scenario, consider hardening it with a more ambiguous or misleading prompt, unusual seed data, or subtle footgun. Run locally and review agent failures to ensure they're legitimate reasoning mistakes, not eval framework limitations. We also want to keep benchmarks representative of the user journey. Review the Evals coverage table and make sure you're not over-indexing on a niche use case.
Prompts should reflect what a real user would send to an agent. Prompts should NOT reflect deep familiarity with Supabase nor specify every detail of a request, as users should expect agents to fill in the gaps themselves. They should be short and casual messages, not highly formatted specs.
Instead of spoonfeeding agents in the prompt, move details into seed data to let agents discover context and infer user intent. For example, a seeded database table can help agents resolve the true names of columns or preferred naming conventions for a project, seeded edge functions can provide a template for desired functionality, and inline comments can help explain a project's structure beyond what the code shows.
Prefer deterministic checks where possible for stability and efficiency. Avoid being overly prescriptive with the process an agent takes to reach a solution (unless critical to the scenario), prefer checking the end state by inspecting the project or filesystem.
If deterministic checks are too inflexible or convoluted, use an LLM-as-a-judge check via judge() to check semantic correctness.
Prefer building checks declaratively and returning the list in one place instead of accumulating checks within branching logic, so the list remains stable if one path fails.
Add a *.experiment.ts file under experiments/<owner>/ for the agent, model, and runtime setup you want to compare. Experiment discovery only scans this owner directory depth, so supporting files can live beside experiments or in nested directories. Reuse the base configs exported from experiments/presets.ts where they fit.
Select the experiment's suite: depending on your use case. If this experiment should be part of our published benchmark, assign suite: ["benchmark"] and include a corresponding *-no-skills variant to compare results with and without skills. You can also assign custom experiment suites for grouping related experiments for other head-to-head comparisons as desired.
Before submitting an eval for review, try running it locally to sanity check that it can complete without errors. It's okay if agents fail the eval, we just don't want them to be scored unfairly for framework limitations.
When you create a PR, use GitHub Actions to refresh the results in CI so we can verify the results in a trusted environment. Currently, results are tracked in Git and committed to the repo, so the refresh results workflow can either commit result changes directly to a branch or generate a PR to propose the change.
You have a few options to run evals in CI:
- Add the
run-evals-changedlabel to your PR to refresh only theevals/changed in that PR and commit merged results directly to your branch. - Add the
run-evalslabel to run every benchmark eval across thebenchmarkandno-skillsexperiment suites. Use this when a change can affect results broadly, such as framework changes. - Dispatch the Refresh eval results workflow manually to target any branch and choose specific evals, experiments, or other options. It can commit results directly to the selected branch or open a separate results PR.
Include refreshed results for PRs with new/changed evals so a reviewer can see results directly from your PR or Vercel preview build.
The docs team owns evals/docs/ and its results. Docs evals run without skills on a single experiment (codex-gpt-6-luna-no-skills).
Common workflows:
- Add or change a docs eval. Add the scenario under
evals/docs/<id>/(see Adding an eval), open a PR, and add therun-evals-changedlabel. Results for the changed evals are committed back to your branch and viewable in the Vercel preview. - Refresh every docs eval. Dispatch the Refresh eval results workflow on
mainwithsuite: docsandexperiment_suite: docs. It opens a draft PR with the updateddocs-eval-results.jsonfor you to review and merge. - Analyze results over time. Every merge that changes
docs-eval-results.jsonappends a snapshot todocs-results.jsonlon GitHub Pages, alongside the benchmark and regression histories.
The CLI team owns evals/cli/ and its results. CLI evals run on codex-gpt-6-luna-cli-{pinned,stable,beta,nodaemon,absent} under experiments/cli/: pinned runs the repo's pinned CLI version, stable/beta install the latest stable or beta CLI, and nodaemon/absent additionally force Docker-less sandboxes — comparing the same scenario across CLI environments.
Which evals each arm picks up:
- pinned, stable, beta run every
interface: clieval that isn'thostedProject: true. - nodaemon, absent additionally only run evals that also set
needsDocker: falseandprojectRunning: false— the harness cannot pre-start a stack, or link a hosted project, without Docker.
Common workflows:
- Add or change a CLI eval. Add the scenario under
evals/cli/<id>/(see Adding an eval); setneedsDocker: falsein itsPROMPT.mdfrontmatter if it can run without Docker, open a PR, and add therun-evals-changedlabel. Results for the changed evals are committed back to your branch and viewable in the Vercel preview. - Refresh every CLI eval. Dispatch the Refresh eval results workflow on
mainwithsuite: cliandexperiment_suite: cli. It opens a draft PR with the updatedcli-eval-results.jsonfor you to review and merge. Leavecli_stable_version/cli_beta_versionblank to resolve npm's latest dist-tags, or pin them to reproduce a specific run. - Analyze results over time. Every merge that changes
cli-eval-results.jsonappends a snapshot tocli-results.jsonlon GitHub Pages, alongside the benchmark, regression, and docs histories. - Run the unit tests.
pnpm --filter @supabase-evals/framework test:cli-lib(the CLI skip predicates inexperiments/cli/lib/plus every CLI eval's scorer tests) — also part ofpnpm check.