HopperLabs.aiStudy Lab

Measurement

Confidence becomes useful when it meets the score

Pair a confidence judgment with the actual response outcome to find confident errors, fragile correct answers, and better next prompts.

Confidence is an internal judgment, not an answer key. It becomes actionable when the learner records it before feedback and the system pairs it with an independently evaluated response. The gap between confidence and outcome is calibration. That gap can reveal a different practice need than accuracy alone.

Record the judgment before the reveal

Use a small, clearly labeled scale such as low, medium, and high, or a probability range the learner understands. Capture it after the response and before showing the target. Do not repeatedly ask for elaborate ratings that overwhelm the practice itself. The scale must remain stable long enough to interpret and should be accessible without color.

OutcomeConfidenceInterpretation to investigate
CorrectHighPotentially secure; confirm after delay and another cue
CorrectLowFragile or lucky retrieval; strengthen the relationship
IncorrectLowRecognized gap; supply feedback and a tractable next cue
IncorrectHighConfident misconception, interference, bad key, or ambiguous prompt; inspect before repeating

These labels are routing hints, not diagnoses. A high-confidence error can come from a misleading cue or evaluator mistake. A low-confidence correct answer can reflect a cautious response style rather than weak knowledge. Preserve enough evidence to inspect the item before drawing a learner-level conclusion.

Measure calibration over eligible groups

For a simple categorical scale, compare accuracy within each confidence band and track how those rates change over time. With a probability judgment, a proper scoring rule can measure the distance between the forecast and outcome. Do not collapse all categories, cue types, and item difficulties into one flattering number. Report sample sizes and abstentions.

The objective is not maximum confidence. It is confidence that tracks performance closely enough to guide study choices and identify when a source, cue, or answer decision needs review.

Route confident errors to discrimination

A confident miss often benefits from a prompt that contrasts the chosen response with the target on one diagnostic relationship. Ask what feature separates the two, then retrieve each from a different cue later. If the source or accepted variant is contested, freeze the item rather than strengthening a possibly wrong distinction.

Keep the data private

Individual responses, confidence, timing, study schedules, and chat questions are private product state. They do not belong in a public contestant profile, SubjectGraph release, analytics export, or feeder article. Aggregate learning metrics require an explicit purpose, minimum cohort, retention rule, and privacy review before they leave the study system.

Use the product's private analytics surface only within the local proof; it is not public graph evidence.