HITL Monitor
Generated 2026-05-28T03:27:24.111178Z • 0 questions • 0 human labels
Calibration map (b̂ vs human label)
Y-axis: Easy = 0, Medium = 0.5, Hard = 1. Few labels = expected early on.
b̂ is the difficulty estimator's continuous score — roughly
−3 (very easy) to +3 (very hard). Each dot is one human-labelled
question at its b̂ (x) versus the human's label (y). Good
calibration shows the dots rising left-to-right: low b̂ gets
"Easy", high b̂ gets "Hard". Isotonic regression is the
monotonic (never-decreasing) curve fitted through these dots — the
function that turns a raw b̂ into a calibrated difficulty, re-fit
every 30 labels. A dot that breaks the rising trend (high b̂ but
labelled "Easy") flags a question the model mis-rated. The κ panel
below is the single-number summary of that agreement.
Cohen's κ — model vs human
Threshold for healthy: κ ≥ 0.5. With <30 labels, expect noisy/below-threshold values.
Cohen's κ measures how often the model's difficulty bucket
matches the human's, corrected for the agreement you would get by
random chance. κ = 1 is perfect agreement, κ = 0 is no better
than chance, negative κ is systematic disagreement. It is recomputed
every 30 ingested labels.
Recent warnings & errors (0)
Problems the run loop would otherwise only print to the terminal —
a generation slot dropped on a JSON parse failure, a rubric-mining
reply that would not parse, a generation/ingest that errored. They
pile up here (newest first) so you can spot and debug them async.
No warnings or errors logged.
Constitution — uninitialised
Core principles are hand-authored or curriculum-seeded rules,
always applied. Mined principles are learned from reviewer
notes; each carries a Beta(α, β) belief, where α = 1 +
supporting reviews and β = 1 + contradicting reviews (so α=3,
β=1 means 2 reviews backed it and 0 went against it). Its estimated
support rate is α/(α+β). A mined principle is promoted to
active — and only then fed into generation prompts — once it
has at least 5 supporting reviews and the Wilson 95% lower
bound on its support rate clears the promotion threshold (about 5
clean supports). An active principle is retired only if it
later collects 2 or more contradictions. "Inactive" below means
simply not-yet-promoted.
Core principles (0)
Active mined principles (0)
- (no active mined principles yet)
Inactive / retired mined principles (0)
Pending review queue (0)
High-score items are highest information-gain. Walk top-down.
Per-question log (0)
Click any row to expand the human feedback that was submitted for it.
H — the model's uncertainty about the question's
difficulty: the Shannon entropy (in bits, 0 to about 1.58) of its
Easy/Medium/Hard posterior. H near 0 means the model is confident
the question sits in one bucket; a high H means it is torn between
buckets.
Score — the active-learning priority: how much the
system expects to learn from a human review of this question. It
blends H, Δ (the gap between the model's estimate and the difficulty
the slot asked for) and novelty (how unlike the existing canonical
examples the question is). Higher = review sooner — the pending
queue is sorted by it.
| ID | Pattern / Section | Declared | Estimated | Reviewed | H | Score | Outcome | Feedback | Preview |
| (no questions generated yet) |