Quality & measurement

How well does Brain actually answer?

A fixed set of 14 questions with known-good expectations, run against the live pipeline and scored 0 to 3. This is how we tell whether a change made answers better or worse, instead of guessing — and it's the Phase 0 baseline the 90-day plan calls for.

The rubric

0
Wrong / empty
No answer, or a confidently wrong one.
1
Weak
Retrieves something related but misses the core.
2
Good
Correct and useful, minor omissions.
3
Strong
Correct, specific, cites the right source.
Honesty guards work in reverse. Some cases are deliberately unanswerable from our data. For those, a confident fabrication scores 0 and an honest "I don't have that" scores 3. They guard against the pipeline drifting back toward confident-but-wrong answers.

Run a case


Score this run

Latest runs

CaseCategoryScoreCostScored byWhen
No runs yet. Run a case above to start building the baseline.

Cases live in CURACITY_BRAIN_EVAL_SET.md (source of truth). Want to add one? Suggest a case and it gets reviewed before it joins the set.