HITL Monitor

Generated 2026-05-28T03:27:24.111178Z • 0 questions • 0 human labels

Difficulty distribution

Declared
Estimated
Human

Calibration map (b̂ vs human label)

Y-axis: Easy = 0, Medium = 0.5, Hard = 1. Few labels = expected early on.
is the difficulty estimator's continuous score — roughly −3 (very easy) to +3 (very hard). Each dot is one human-labelled question at its b̂ (x) versus the human's label (y). Good calibration shows the dots rising left-to-right: low b̂ gets "Easy", high b̂ gets "Hard". Isotonic regression is the monotonic (never-decreasing) curve fitted through these dots — the function that turns a raw b̂ into a calibrated difficulty, re-fit every 30 labels. A dot that breaks the rising trend (high b̂ but labelled "Easy") flags a question the model mis-rated. The κ panel below is the single-number summary of that agreement.

Cohen's κ — model vs human

Threshold for healthy: κ ≥ 0.5. With <30 labels, expect noisy/below-threshold values.
Cohen's κ measures how often the model's difficulty bucket matches the human's, corrected for the agreement you would get by random chance. κ = 1 is perfect agreement, κ = 0 is no better than chance, negative κ is systematic disagreement. It is recomputed every 30 ingested labels.

Recent warnings & errors (0)

Problems the run loop would otherwise only print to the terminal — a generation slot dropped on a JSON parse failure, a rubric-mining reply that would not parse, a generation/ingest that errored. They pile up here (newest first) so you can spot and debug them async.
No warnings or errors logged.

Constitution — uninitialised

Core principles are hand-authored or curriculum-seeded rules, always applied. Mined principles are learned from reviewer notes; each carries a Beta(α, β) belief, where α = 1 + supporting reviews and β = 1 + contradicting reviews (so α=3, β=1 means 2 reviews backed it and 0 went against it). Its estimated support rate is α/(α+β). A mined principle is promoted to active — and only then fed into generation prompts — once it has at least 5 supporting reviews and the Wilson 95% lower bound on its support rate clears the promotion threshold (about 5 clean supports). An active principle is retired only if it later collects 2 or more contradictions. "Inactive" below means simply not-yet-promoted.
Core principles (0)
Active mined principles (0)
Inactive / retired mined principles (0)

Pending review queue (0)

High-score items are highest information-gain. Walk top-down.

Per-question log (0)

Click any row to expand the human feedback that was submitted for it.
H — the model's uncertainty about the question's difficulty: the Shannon entropy (in bits, 0 to about 1.58) of its Easy/Medium/Hard posterior. H near 0 means the model is confident the question sits in one bucket; a high H means it is torn between buckets.
Score — the active-learning priority: how much the system expects to learn from a human review of this question. It blends H, Δ (the gap between the model's estimate and the difficulty the slot asked for) and novelty (how unlike the existing canonical examples the question is). Higher = review sooner — the pending queue is sorted by it.
IDPattern / SectionDeclaredEstimatedReviewedHScoreOutcomeFeedbackPreview
(no questions generated yet)