Quality & measurement

How well does Brain actually answer?

A fixed set of questions with known-good expectations, run against the live pipeline and scored on a shared rubric. This is how we tell whether a change made answers better or worse, instead of guessing. Anyone at Curacity can run a case, see the score, and suggest new ones.

The rubric

0
Wrong / empty
No answer, or a confidently wrong one.
1
Weak
Retrieves something related but misses the core.
2
Good
Correct and useful, minor omissions.
3
Strong
Correct, specific, cites the right source.
Honesty guards work in reverse. Some cases are deliberately unanswerable from our data. For those, a confident fabrication scores 0 and an honest "I don't have that" scores 3. They guard against the pipeline drifting back toward confident-but-wrong answers.

Run a case


Response will appear here after you run the case.
Score this run

Latest runs

Case Category Score Scored by When
No runs yet. Run a case above to start building the baseline.

Cases live in CURACITY_BRAIN_EVAL_SET.md (source of truth). Want to add one? Suggest a case and it gets reviewed before it joins the set.