Skip to main content

Lecture Lens — system design (v1)

1. Requirements​

Functional

  • F1 Capture: start/stop, portal/live mode, allowed domains; live health of screen frames, App-Tap audio, vision model, oMLX; permissions checklist.
  • F2 Lectures: list sessions/lectures, open transcript, slides (image + description), notes and cards; build or rebuild one.
  • F3 Ask: grounded Q&A over captures with citations; feedback on any AI output.
  • F4 Evals: run the vision eval and the DeepEval suite; show score history; turn feedback into test cases; benchmark judges (local oMLX vs Laya).
  • F5 Settings: models, thinking mode, capture tuning, domains, paths.

Non-functional

  • The app is the TCC-responsible process for screenpipe (it launches it).
  • Offline: only 127.0.0.1 traffic (screenpipe :3030, oMLX :8001). No telemetry.
  • One implementation of each capability. The app calls the tested toolkit and never re-implements it.
  • Never hides a failure. Every background failure becomes a visible state.
  • The UI never blocks. Heavy jobs run serially (one GPU), with visible progress.

Constraints: single user; Swift 6.4 / macOS 27 / Xcode 27; no Apple Development certificate yet (ad-hoc signing means permissions are re-granted after each rebuild until one exists).

2. High-level design​

┌──────────────────────── Lecture Lens.app (SwiftUI) ───────────────────────┐
│ MenuBarExtra ── status lights · start/stop · mode · open window │
│ Window: Capture │ Lectures │ Ask │ Evals │ Settings │
│ │
│ AppModel (@Observable, @MainActor) │
│ ├─ HealthMonitor ── polls: screenpipe /health (5 s), oMLX /api/status (10 s),│
│ │ `lecture doctor --json` (60 s, vision probe on demand) │
│ ├─ CaptureController ─ runs sp-start / sp-stop / sp-mode / sp-domains │
│ ├─ ToolRunner ─ Process + JSON contract (stdout = one JSON doc, stderr = log)│
│ ├─ JobQueue ─ one heavy job at a time (build, eval, benchmark); progress │
│ ├─ LectureStore ─ reads lectures/*/manifest.json + files (read-only) │
│ ├─ FeedbackStore ─ appends feedback/feedback.jsonl │
│ └─ EnvFile ─ line-preserving edits of .env │
└──────┬───────────────────┬──────────────────────────┬─────────────────────────┘
│ launches │ runs (JSON contract) │ HTTP (localhost)
▼ ▼ ▼
sp-start → screenpipe study / lecture / eval screenpipe :3030 oMLX :8001
(+ sp-autoswitch) (Python, stdlib) (health, search) (LLM+VLM)
│ │
▼ ▼
~/.screenpipe/db.sqlite lectures/… feedback/… eval/results/…
eval/deepeval/ (uv venv: deepeval, laya)

Process lifecycle and permissions​

The app launches sp-start with Process. sp-start runs screenpipe in the background and starts sp-autoswitch, and the autoswitcher restarts screenpipe on mode changes. Every process in that chain descends from the app, so macOS treats Lecture Lens as the responsible app: permission prompts name it, and the grants in System Settings belong to it. screenpipe keeps running if the window closes (menu-bar app); Quit asks whether to stop capture too.

3. Contracts​

3.1 Tool JSON contract​

Every tool the app calls gets a --json flag:

  • stdout: exactly one JSON document.
  • stderr: human progress, streamed into the job log.
  • exit code: 0 success, 1 domain failure (JSON still printed, with "ok": false), 2 usage error.
CommandJSON (abridged)
lecture doctor --json{ok, checks:[{id, ok, level, message, fix}]} — ids: database, frames, app_audio, ffmpeg, pillow, vision, vault
lecture list --json [--since 30d]{sessions:[{id, title, start, end, minutes, frames, built, folder}]}
lecture build <id> --json{ok, folder, manifest:{…}}
study ask "…" --json [--since]{answer, sources:[{n, ts, src, where, text}], queries, model, latency_s}
sp-mode --json{running, mode, pinned, autoswitch}

3.2 Feedback record (feedback/feedback.jsonl, append-only)​

{"id":"uuid","ts":"2026-09-29T18:02:11Z","target":"answer|note|slide|card",
"ref":{"lecture":"20260929-1830","slide":4,"question":"…"},
"input":"question or slide path","output":"what the model said",
"context":["cited source texts…"],"rating":1,"correction":"optional text",
"model":"Swift-Qwen3.8-27b-oQ4e-mtp","prompt_version":"slide-v2"}

Append-only, so nothing is ever rewritten. eval/deepeval/feedback_goldens.py turns it into DeepEval goldens: a correction becomes the expected output, and a thumbs-down without a correction becomes a known-bad example.

3.3 Eval results​

Every run writes eval/results/<ts>-<suite>/summary.json:

{"suite":"deepeval-ask|vision|judge-bench","model":"…","judge":"…",
"metrics":{"faithfulness":{"mean":0.91,"pass":18,"n":20}, "…":{}},
"latency":{"p50":4.7,"p95":8.4},"started":"…","git":"50624d0"}

The app plots history from these files. No database is needed.

4. Deep dive​

4.1 Health model​

Each signal is ok | degraded | down | unknown, with a reason and a fix action:

SignalSourceDown meansFix action
Screenframes in last 2 min while capturingpermission missing or display asleepopen Privacy → Screen & System Audio Recording
App audioApp-Tap transcripts in last 10 min while Chrome playstap refused by OSsame pane; restart capture
Visiondoctor probe ("PROBE 4721")model loaded text-only"Reload model in VLM mode" (oMLX admin API)
oMLX/api/statusserver downopen oMLX
Screenpipe/healthnot runningStart

4.2 Jobs​

JobQueue runs heavy work serially: lecture builds, evals and the judge benchmark all compete for the one GPU. Ask requests are interactive and never queue behind a build; they run immediately and may be slower while a build runs, and the UI says so. Every job keeps its stderr log for the log view.

4.3 Failure handling​

  • Tool exits non-zero → the job shows a red card with the last 20 stderr lines and the command to rerun.
  • JSON parse failure → treated as a tool bug; the raw stdout is kept for the log.
  • oMLX down during Ask → an inline error instead of a spinner; retry button.
  • .env edits are atomic (temp file + rename) and keep comments; the previous file is kept as .env.bak.
  • screenpipe crashes → HealthMonitor shows it within 5 s; one-click restart.

5. Evaluation architecture (DeepEval + judge benchmark)​

fixtures (synthetic lecture: transcript + slide descriptions + questions + expected answers)
+ feedback goldens (your corrections)
│
▼
deepeval test cases ── system under test = the real study/lecture_kit code paths
│ (retrieval, answer, notes, cards, slide description)
▼
metrics: Faithfulness · AnswerRelevancy · ContextualRelevancy · Hallucination · GEval(correctness)
judge: OmlxJudge (DeepEvalBaseLLM → 127.0.0.1:8001, thinking off, temperature 0)
│
└─ judge benchmark: claim-support pairs with known labels
System 2: OmlxJudge verdicts vs System 1: Laya `noul` P(supported)
→ accuracy, F1, ECE, latency; later: agreement with your labels (κ)

6. Trade-offs​

DecisionChosenCostRevisit when
SwiftPM package + bundling script vs Xcode projectSwiftPMno Interface Builder or asset catalogsUI needs heavy assets or App Store packaging
App calls tools vs reimplementing in Swifttools~100–300 ms process spawn per calla call becomes interactive-latency-critical (then add a daemon)
Polling health vs pushpolling (5–60 s)small CPU costscreenpipe exposes events we need
JSONL feedback vs SQLiteJSONLfull scan to read (fine below ~100k rows)feedback exceeds ~50k rows
Unsandboxed apprequired (launch processes, read ~/.screenpipe)no App Storenever, for this use
Ad-hoc signingonly option todaypermissions re-granted on each rebuildan Apple Development cert exists (sign into Xcode)