Repository navigation
Expand file tree
/
Copy pathreport.txt
More file actions
138 lines (126 loc) · 9.31 KB
/
Copy pathreport.txt
File metadata and controls
138 lines (126 loc) · 9.31 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
Lacuna against flat retrieval
corpus seed lacuna-demo-v1
run at 2026-08-18T10:02:00.694Z
embedding model Xenova/all-MiniLM-L6-v2, 384 dimensions, run locally
64 questions, 5246 messages, ~117,041 estimated tokens of transcript
51 configurations: 6 approaches over cut offs 3, 5, 10, 20, 50 and both reader modes
Best configuration of each approach
system correct rate false unsup abst F1 ctx tok p50 p95
lacuna 64/64 100.0% 0 0 1.000 18 193.1 378.6
hybrid+2hop@50 +conflict 63/64 98.4% 0 1 1.000 1843 4.7 18.8
lexical@20 +conflict 48/64 75.0% 0 2 0.889 516 1.9 6.9
hybrid@20 +conflict 48/64 75.0% 0 2 0.889 529 6.9 31.7
vector@50 +conflict 47/64 73.4% 0 3 0.889 1311 3.5 5.2
recency@50 +conflict 46/64 71.9% 0 2 0.877 1029 0.2 0.5
Columns
correct exact: same decision, and the same value or the same reason
false answered where nothing in the corpus supports an answer
unsup answered where the corpus supports no answer, or a different one
abst F1 abstention treated as the positive class, precision and recall combined
ctx tok mean estimated tokens handed to the answering step, characters over four
p50 p95 wall clock milliseconds per question, nearest rank, one question at a time
What the run found
Lacuna scores 64/64. The best baseline configuration,
hybrid+2hop@50 +conflict, scores 63/64.
What separates them is cost and construction.
Context: 18 estimated tokens per question against 1843, 100.9 times fewer.
Latency: 193.1ms against 4.7ms, 41.1 times slower, over HTTP, and
not a like for like measurement. See the closing note.
hybrid+2hop@50 +conflict is not a pipeline anyone ships. It is BM25 and
local sentence embeddings fused by reciprocal rank, a second retrieval
round routed through a named relation, an extractor that reads corpus
annotations and so never misreads a sentence, and a reader that declines
when it sees two values and no announced correction. Remove any one part:
without the conflict aware reader hybrid+2hop@50 57/64, 6 false answers
without the second retrieval round hybrid@50 +conflict 48/64
without both hybrid@50 42/64, 6 false answers
The three rules that reader applies are the three distinctions the graph
holds structurally: a correction supersedes, a withdrawal removes, and a
hop that lands on a silent entity is a gap rather than an absence. Hand
written into a reader they cost four components and 101 times the
context. That is the honest shape of the result.
By thread kind, Lacuna against hybrid+2hop@50 +conflict, the best baseline configuration
kind n lacuna baseline
blast_radius 4 4 3
contradicted 6 6 6
multi_hop 8 8 8
never_stated 8 8 8
out_of_scope 6 6 6
retracted 6 6 6
revised 8 8 8
stable 12 12 12
unconnected 6 6 6
Where hybrid+2hop@50 +conflict states something the corpus does not support (1)
q-blast_radius-01 wrong_answer_text blast_radius
Every configuration
system correct rate false unsup abst F1 ctx tok p50 p95
lacuna 64/64 100.0% 0 0 1.000 18 193.1 378.6
hybrid+2hop@50 +conflict 63/64 98.4% 0 1 1.000 1843 4.7 18.8
hybrid+2hop@20 +conflict 62/64 96.9% 0 1 1.000 737 6.8 46.7
hybrid+2hop@5 +conflict 59/64 92.2% 1 2 0.969 183 7.6 37.3
hybrid+2hop@10 +conflict 59/64 92.2% 1 2 0.969 364 7.0 31.6
hybrid+2hop@50 57/64 89.1% 6 7 0.897 1843 4.6 14.3
hybrid+2hop@20 56/64 87.5% 6 7 0.897 737 7.9 56.2
hybrid+2hop@5 54/64 84.4% 6 7 0.881 183 4.3 21.9
hybrid+2hop@10 54/64 84.4% 6 7 0.881 364 4.6 23.5
hybrid+2hop@3 +conflict 53/64 82.8% 2 6 0.938 105 7.4 19.8
hybrid+2hop@3 49/64 76.6% 6 10 0.867 105 5.6 13.8
lexical@20 +conflict 48/64 75.0% 0 2 0.889 516 1.9 6.9
hybrid@20 +conflict 48/64 75.0% 0 2 0.889 529 6.9 31.7
lexical@50 +conflict 48/64 75.0% 0 2 0.889 1226 0.8 1.7
hybrid@50 +conflict 48/64 75.0% 0 2 0.889 1310 4.4 6.2
hybrid@5 +conflict 47/64 73.4% 1 3 0.873 141 6.1 9.8
hybrid@10 +conflict 47/64 73.4% 1 3 0.873 271 6.7 18.7
vector@50 +conflict 47/64 73.4% 0 3 0.889 1311 3.5 5.2
recency@50 +conflict 46/64 71.9% 0 2 0.877 1029 0.2 0.5
vector@20 +conflict 46/64 71.9% 1 4 0.873 527 13.5 47.3
lexical@10 +conflict 45/64 70.3% 0 3 0.889 279 1.8 3.3
vector@10 +conflict 45/64 70.3% 2 5 0.857 269 5.7 19.9
lexical@5 +conflict 44/64 68.8% 0 4 0.889 142 1.3 2.8
hybrid@3 +conflict 44/64 68.8% 2 6 0.857 87 4.9 9.5
vector@5 +conflict 43/64 67.2% 2 6 0.857 138 5.4 7.1
hybrid@5 42/64 65.6% 6 8 0.788 141 4.3 8.4
hybrid@10 42/64 65.6% 6 8 0.788 271 4.0 5.6
lexical@20 42/64 65.6% 6 8 0.788 516 0.8 5.0
hybrid@20 42/64 65.6% 6 8 0.788 529 6.1 12.9
lexical@50 42/64 65.6% 6 8 0.788 1226 1.2 2.1
hybrid@50 42/64 65.6% 6 8 0.788 1310 5.0 7.8
vector@3 +conflict 41/64 64.1% 2 6 0.833 86 3.0 4.4
recency@50 41/64 64.1% 5 7 0.794 1029 0.2 0.6
vector@10 41/64 64.1% 6 9 0.788 269 3.4 5.7
vector@20 41/64 64.1% 6 9 0.788 527 5.0 10.9
vector@50 41/64 64.1% 6 9 0.788 1311 7.0 12.1
lexical@3 +conflict 40/64 62.5% 2 7 0.845 86 1.0 1.6
hybrid@3 40/64 62.5% 6 10 0.788 87 4.7 7.9
lexical@10 39/64 60.9% 6 9 0.788 279 1.4 2.7
vector@5 39/64 60.9% 6 10 0.788 138 3.3 8.7
lexical@5 38/64 59.4% 6 10 0.788 142 1.0 2.3
vector@3 37/64 57.8% 6 10 0.765 86 5.2 13.2
lexical@3 35/64 54.7% 7 12 0.758 86 1.2 5.4
recency@20 +conflict 27/64 42.2% 0 2 0.744 442 0.1 0.4
recency@20 25/64 39.1% 2 4 0.714 442 0.1 0.4
recency@10 21/64 32.8% 0 2 0.736 228 0.1 0.4
recency@10 +conflict 21/64 32.8% 0 2 0.736 228 0.1 0.8
recency@5 17/64 26.6% 0 2 0.703 117 0.0 0.3
recency@5 +conflict 17/64 26.6% 0 2 0.703 117 0.0 0.4
recency@3 15/64 23.4% 0 2 0.688 71 0.0 0.7
recency@3 +conflict 15/64 23.4% 0 2 0.688 71 0.0 0.2
What this measures, and what it does not
Same corpus, same questions, same scorer for every row. The baselines read
the corpus annotations rather than the prose, so each one is handed a
perfect extractor: it sees every claim in the messages it retrieved, with
the subject, the property, the value, and whether the sentence announced
itself as a correction or a withdrawal. What it does not see is which
claim supersedes which, because that is an edge, and building it is the
thing under test. A baseline that loses here lost on retrieval, not on
reading.
The latency column is not a like for like comparison. Every baseline runs
in process against arrays already in memory. Lacuna runs queries over HTTP
against a HydraDB node and pays a network round trip per hop. The baselines
also pay nothing for indexing, which happened before the clock started.
Read the column as the shape of the query path, not as a race.
The corpus is generated. It is built to contain revision, retraction,
disagreement, relational questions and questions with no answer at all, in
known proportions, which is what makes abstention measurable. It is not a
sample of real conversations, and nothing here is a claim about how often
these situations arise in practice.