Skip to main content

Testing strategy

Cheapest first. The first three layers need no model and no running capture, so they run on every change and in CI.

LayerCoversRunsTime
1. Unit (pytest)Session splitting, time-window sessions, transcript parts, title rules, card JSON salvage, keyframe text comparison, word error rate, the content guardEvery change; CIseconds
2. Contract--json output of each tool and the Swift types that decode itEvery change; Swift tests in CI on macOSunder a minute
3. Replay (planned)Recorded health sequences and a small database fixture drive the watchdog and the buildEvery changeunder a minute
4a. Smoke eval6 fixed cases through generator and judgePrompt, model or build-logic change; not during captureabout 12 minutes
4b. Full evalAll question, notes and card cases, plus your private session casesOn requestover an hour
4c. Judge benchmarkLabelled pairs: is the judge right?When the judge or its prompt changesabout 10 minutes

Every fault becomes a named test​

TestAsserts
test_window_session_keeps_early_audioA window starting before the first frame includes the earlier audio
test_long_transcript_is_fully_coveredParts of a long transcript join back to the input exactly
test_two_slides_of_one_deck_differ_by_textTemplate-alike slides are kept apart by their text
test_ui_chrome_is_not_slide_textMeeting-app controls do not count as slide content
test_judge_scores_the_answer_not_its_source_listThe judge sees the answer body only
test_the_repository_itself_is_cleanThe tracked tree passes the content guard

Planned with the capture guard: a zero-writes replay must turn the Screen light red within 60 seconds, and a protected-video pause must not.

Not automated​

Starting real capture, permission prompts and protected-video behaviour need the signed app and a person. They are covered by the preflight in Capture a live class.

Test data​

Invented subjects only (biology, astronomy). Never real course content: the content guard fails the build if it appears.