Skip to content

Latest commit

 

History

History
113 lines (87 loc) · 5.09 KB

File metadata and controls

113 lines (87 loc) · 5.09 KB

Judge scorecard

Internal assessment of the 2026-08-20 V8 candidate. These are not judge scores; they are deliberately strict estimates used to find what can still lose.

Honest verdict

Lacuna is credible finalist material and can win on technical depth, HydraDB usage and originality. It is not honestly winner-ready until the new candidate is deployed and retested, the exact voice is provider-proved, and a tight three-minute film demonstrates the real product. A judge can currently find unfinished cross-client promises faster than the architecture can explain its strongest idea.

Technical execution: 8.5/10

Strong evidence:

  • 79 unit files and 1,344 passing tests.
  • Deterministic resolver, evidence trail, history, contradiction and structural abstention.
  • Interactive overview graph and exact provenance DAG/table.
  • Two bounded agent roles with persisted lifecycle, Context Pack handoff, reviewer verdict and explicit no-write policy.
  • OAuth signature/claim verification, PKCE, nonce, provider binding and session revocation are covered at the HTTP boundary.
  • Private MCP has digest-only random capabilities, revoke, cross-workspace refusal, streaming body caps and rate limits.

Why it is not a 10:

  • HydraDB provides no CAS/transaction seam to this adapter. Scheduler claims, agent idempotency, quota records and registry writes are not atomic across concurrent serverless instances.
  • The candidate has not yet completed production/browser acceptance after the security changes.
  • Rate limits are process-local rather than distributed spend controls.

HydraDB and graph-native approach: 9/10

HydraDB is load-bearing: it stores the public corpus, private memory records, agent/schedule state and capability digests. Lacuna preserves source, evidence, claim, entity, supersession, contradiction, dependency and temporal edges, and can page both a readable overview and exact proof graph. The answer surface resolves current standing from graph-backed evidence rather than treating the latest message as truth.

The missing point is operational: document upsert without CAS is not enough for strong distributed coordination or identity uniqueness. The product now says so and disables the unsafe hosted password-creation path.

Product completeness and usability: 7.5/10

The product has a coherent landing journey, public proof workspace, dashboard, plain-English Ask with artifacts, dense table, graph, ingest, two agents, recommendations, schedules, tools, models and a voice surface. Signed-in users can mint and revoke a private MCP bearer from the UI.

The visible gaps are material:

  • Production voice is unavailable until the exact Vaibhav Lalwani Professional ElevenLabs clone is selected by id and STT, TTS, interruption and fallback pass in a real session.
  • ChatGPT and Claude are protocol targets, not completed client proofs.
  • No packaged SDK ships, and CLI/MCP do not control agent lifecycle.
  • Spotify and other native connectors are examples, not integrations.
  • Existing legacy Google accounts need a verified linking migration.

Quality of results: 8.5/10

The core evaluation is strong within its stated scope: 64/64 on the generated corpus, evidence-backed answers, revisions preserved, conflicts exposed and unsupported premises refused. The Context Pack makes agent output inspectable instead of asking a model to invent a rationale after the fact.

The main limitation is external validity. The corpus is generated by this project and the LongMemEval work is an ingestion check, not a full published benchmark score. The video must not blur that distinction.

Originality: 9/10

The differentiated thesis is clear: portable memory should govern what the next agent may believe, not merely retrieve more text. Current, historical, conflicting and missing are product states; provenance and abstention survive across web, CLI and MCP. Memory-derived agent suggestions with bounded, no-write execution make that thesis judge-visible.

What will decide the hackathon

The strongest three-minute sequence is:

  1. Show a correction and a contradiction entering one memory.
  2. Ask a natural-language question and reveal the evidence artifact.
  3. Switch between overview and exact proof graph/table.
  4. Show Lacuna recommending a bounded agent because of that memory, then run Researcher to Reviewer and show the no-write verdict.
  5. Show the same memory over CLI and private MCP.
  6. Explain precisely what HydraDB stores and where it is load-bearing.
  7. End with the portable-memory thesis and one honest future-scope sentence.

Do not spend the film on decorative animation, a connector catalogue, or unproved ChatGPT/Claude/Spotify claims. The product earns the animation after the proof is visible.

Release blockers

  1. Deploy this candidate and rerun live OAuth, MCP, graph, agent and schedule acceptance.
  2. Configure and prove the exact selected ElevenLabs clone.
  3. Capture clean real product/CLI/MCP footage from the accepted deployment.
  4. Build the sub-three-minute HyperFrames preview and obtain owner approval before rendering the final MP4.
  5. Keep YouTube upload and submission as explicit owner actions.