Assessment test bench

Belt Test — three spec versions, one page

Spec
Grader
LLM: … voice: … grader calls: 0

V1: email → self-select your belt → adaptive E→F→P verification with probes; retake-style endings. V2: email → hidden 10-question diagnostic (hard gates) → seamless verification, ±1 travel cap, profile + courses. V3.1: email → 2 routers + dynamic adaptive concept items (two-signal rule, second chances, skips, code/low-code wording) → capability-first verification, cross-track report, periphery. Typed answers are graded semantically against each spec’s rubrics by the selected grader — gpt-5.6-luna (default, ~1s/grade) or claude-sonnet-5. Own-level practical answers grade in the background so the flow never waits on them. Typed questions accept voice input via ElevenLabs scribe_v2 — the transcript lands in the text box for review and editing before you submit.

Inspector — internal state (never shown to real users)