Quality & measurement
How well does Brain actually answer?
A fixed set of 14 questions with known-good expectations, run against the live pipeline and scored 0 to 3. This is how we tell whether a change made answers better or worse, instead of guessing — and it's the Phase 0 baseline the 90-day plan calls for.
The rubric
0
Wrong / empty
No answer, or a confidently wrong one.
1
Weak
Retrieves something related but misses the core.
2
Good
Correct and useful, minor omissions.
3
Strong
Correct, specific, cites the right source.
Honesty guards work in reverse. Some cases are deliberately unanswerable from our
data. For those, a confident fabrication scores 0 and an honest "I don't have that"
scores 3. They guard against the pipeline drifting back toward confident-but-wrong answers.
Run a case
Score this run
Latest runs
| Case | Category | Score | Cost | Scored by | When |
|---|---|---|---|---|---|
No runs yet. Run a case above to start building the baseline. | |||||
Cases live in CURACITY_BRAIN_EVAL_SET.md (source of truth). Want to add one? Suggest a case and it gets reviewed before it joins the set.