Skip to main content

Evaluation

Evaluation answers one question: did this change to a prompt, a model or the code make the study material better or worse? Everything runs locally; no results leave the machine.

Suites​

CommandWhatTypical time
tests/evals/run.sh --smoke6 fixed cases: 2 questions, 2 notes, 2 card setsabout 12 minutes
tests/evals/run.shFull suitewell over an hour
tests/evals/run.sh --judge-onlyRe-score cached outputsjudge time only
tests/evals/run_judge_bench.shLabelled claim/context pairs: how often is the judge right?about 10 minutes
eval/run_eval.pySlide reading on synthetic slides with exact answersminutes

Each run writes eval/results/<timestamp>-<suite>/summary.json with the models, prompt versions, case counts and metric means. That folder is the run history; it is ignored by git.

How a run works​

The generator model produces answers, notes and cards. Then the runner swaps in a different model as judge, scores the outputs with DeepEval metrics (Faithfulness, Answer Relevancy, and G-Eval criteria for correctness, coverage and card quality), and swaps the generator back, even if the run fails.

A run refuses to start while capture is running in live mode.

Reading results​

  • Smoke runs have two cases per kind. Compare pass or fail per case, not means; a 0.07 difference on two cases is noise.
  • Check what a failure is before fixing the model. In the first full baseline, most failed question-answering cases were the judge penalising the answer's source list, which it could not see in the context. The harness now scores the answer body only.
  • Trust the judge only as far as the judge benchmark. A perfect score on easy pairs says little; add hard pairs (derived claims, partial support).

Adding cases from your own classes​

Cases built from a real class are course material. Keep them in files named tests/evals/session_*, which git ignores and the content guard blocks. They run with the full suite on your machine.