Qualification by Calibration: A Readable Benchmark for Admitting Language Models to Human-Computation Tasks
DOI: https://doi.org/10.1145/3834580.3838735
HCOMP 2026: 2026 ACM Conference on Human-AI Complementarity and Alignment, Alexandria, VA, USA, September 2026
When several language models are wired together into a human-computation pipeline, the system routes, arbitrates, and escalates work according to how confident each model says it is, so a team builder needs a way to screen candidate models the way crowdsourcing has long screened human contributors: with a small, inspectable qualification test. We present such an instrument, a compact battery of yes-or-no questions across everyday domains on which a model reports an answer and a confidence, with every gold label independently audited. Our central contribution, however, is the audited harness and the interface properties it measures, not accuracy discrimination: qualification verdicts for model workers are only as valid as the harness that administers them. Evaluating a diverse panel of contemporary models under three administration protocols, we show that a naive harness, with a fixed token budget and an unaudited parser, manufactures failing workers out of competent ones, misreading a model that answers essentially every item correctly as badly inaccurate, answer-biased, and overconfident. Properly administered, every model in the panel proves admissible on accuracy for these common-knowledge items, and the differentiators that remain, and that transfer to disjoint downstream tasks, are interface properties: calibration of stated confidence, format discipline, and availability under budget. A controlled manipulation further shows the calibration axis is dissociable from accuracy. We release the items, the audited protocol, all per-trial responses under every protocol, and the scoring code, so the check, and the audit of the check, are each a single command.
ACM Reference Format:
Carlos Toxtli and Manuel Delaflor. 2026. Qualification by Calibration: A Readable Benchmark for Admitting Language Models to Human-Computation Tasks. In 2026 ACM Conference on Human-AI Complementarity and Alignment (HCOMP 2026), September 27--30, 2026, Alexandria, VA, USA. ACM, New York, NY, USA 9 Pages. https://doi.org/10.1145/3834580.3838735
1 Introduction
Human computation increasingly runs through language models that pass work to one another: one model drafts and another checks, a router sends each item to whichever model is most likely to be right, an arbitrator breaks ties, and an escalation rule decides when to hand a case to a person. Every one of these mechanisms consumes a model's expressed confidence as a control signal; an escalation rule keyed to confidence never escalates a model that is confidently mistaken. The property that makes these mechanisms safe is trust calibration: the alignment between how sure a model says it is and how often it is actually correct [7, 11, 20].
Calibration is well studied for neural networks in the aggregate [6, 16] and for language models [9, 26]. What a team builder lacks is not another large corpus but a different kind of object: a small, auditable instrument that can be run on a candidate model before wiring it into a pipeline, read in its entirety, and reasoned about, the model-worker counterpart of the gold-question screens crowdsourcing has long used to admit human contributors [3, 17]. We present such an instrument, the Trust-Calibration Benchmark (TCB): sixty yes-or-no questions across five everyday domains, each answered with a stated confidence, scored for accuracy, calibration error, and Brier score, with every gold label independently audited.
The deepest lesson of this paper, however, is about the screen rather than the screened. Crowdsourcing learned long ago that a badly designed qualification task rejects good workers [3]; we show the same failure arises, mechanically and at scale, when the worker is a model. A qualification harness must parse free-form model output and must allot a token budget, and both choices can manufacture verdicts. In our own initial runs, a naive harness scored one model at 65 percent accuracy with a 95 percent yes-bias and confidence inversely related to correctness; an audit of the per-trial raw responses traced every one of those pathologies to the harness, which had been force-scoring budget-truncated reasoning fragments and misreading quoted instruction templates as answers. Under the audited protocol the same model answers all sixty items correctly with nearly perfectly calibrated stated confidences. We therefore treat the administration protocol as part of the instrument, ship it with an audited parser, escalation rule, and per-trial flags, and report a three-protocol comparison quantifying the harness effect.
Our evaluation asks four questions. Q1: what verdicts does the instrument return for nine contemporary models under a validated protocol? Q2: how sensitive are those verdicts to the qualification harness itself? Q3: does the calibration axis carry information that accuracy does not? Q4: do the screen's verdicts predict behavior on tasks beyond the benchmark? In brief: properly administered, all nine models are admissible on accuracy (0.91 to 1.00, no significant pairwise differences), and the durable differentiators are interface properties, calibration of stated confidence, format discipline, confidence-field discipline, and availability under budget (Q1); the harness effect is large, up to a 35-point apparent accuracy error with fabricated answer bias and fabricated confidence inversion (Q2); a controlled manipulation dissociates calibration from accuracy, steering calibration error while answers stay fixed, and below-ceiling stated confidences mark trials at elevated error, so the calibration axis is real and dissociable (Q3); and the interface verdicts transfer, with calibration error and availability predicting their downstream counterparts on two disjoint tasks while accuracy saturates (Q4).
This paper makes four contributions. (1) A compact, tier-decomposed trust-calibration benchmark for language-model workers, with an independent audit of all sixty gold labels (one error found and corrected) and a release designed to be read and re-derived in full. (2) A harness-validity study showing that qualification verdicts for model workers are protocol-dependent: a naive harness manufactures inaccuracy, answer bias, confidence inversion, and instability that an audited protocol eliminates, and we operationalize the audited protocol as part of the instrument. (3) A nine-model panel under the validated protocol, showing contemporary models are uniformly admissible on accuracy while differing in interface properties, together with a controlled manipulation showing calibration is dissociable from accuracy. (4) A criterion study on two disjoint downstream tasks showing that the surviving differentiators, calibration and availability, are the ones that transfer, plus a release of all per-trial responses under every protocol so that both the verdicts and the harness itself can be audited with one command. We scope these claims accordingly: the load-bearing contribution is the harness-validity lesson, and the instrument's validated role for this model generation is a compact, auditable preflight on parser behavior, confidence reporting, format, and availability; a general accuracy-based admission screen awaits the larger, non-binary validation outlined as future work.
2 Related Work
Requesters screen contributors with gold-standard questions [17], aggregate redundant labels to manage noise [23, 24], and estimate worker quality to decide whom to trust [3, 8]. That literature also documents the converse hazard: a defective qualification task, an ambiguous gold answer, a broken interface, misgrades capable workers. As language models increasingly serve as the workers, in some annotation settings matching or exceeding human contributors [5], both lessons transfer: the screen must extend to model workers, and the screen itself must be validated. The LLM-evaluation literature has documented that benchmark scores are sensitive to prompt formatting and harness implementation [2, 22], and mature evaluation frameworks such as the LM Evaluation Harness [2] and HELM [12] standardize prompts, parsers, and metrics for exactly this reason; our contribution is not the existence of harness sensitivity but its consequence for worker qualification in human computation, where we show it manufactures precisely the failure profiles, inaccuracy, response bias, inverted confidence, that admission decisions key on, and we fold the remedy into the qualification instrument itself. Those frameworks administer large evaluations reproducibly; ours makes one admission decision auditable end to end, down to the raw response behind every scored trial.
The expected calibration error and the reliability diagram come from calibrating probabilistic classifiers [16] and modern neural networks [6], where accuracy and calibration proved to be distinct properties. For language models, asking for a confidence yields usable calibration [9, 26], models can be trained to verbalize uncertainty [13], and elicitation quality varies across models [28]. We use the plain ask-for-a-confidence elicitation, a stated confidence on a final line, because it is the signal deployed pipelines read.
The trust-calibration construct we measure originates in the human-factors literature on trust in automation, where appropriate reliance depends on trust tracking true reliability [7, 11, 20], and miscalibrated confidence is known to drive over-reliance and under-reliance and to degrade joint outcomes [1, 27]. In a multi-LLM pipeline the consumer of the confidence is often another model, but the requirement is identical: the consumer can only act well if the upstream confidence is calibrated.
As language-model agents are composed into teams [15, 18, 19, 21], what to measure and report becomes central [1, 27]. Datasheets for datasets [4] argue that an artifact should travel with a precise account of itself; we adopt that stance and treat the items, the gold-label audit, the administration protocol, and the scoring code as parts of the contribution rather than an appendix.
Our instrument likewise complements large calibration corpora: small enough to read, decomposed enough to localize a failure, audited down to individual labels and the harness itself, and cheap enough to run as a routine preflight on every new model and endpoint version.
3 Benchmark Design
3.1 Design Goals
The instrument is built to four goals that mirror what a requester demands of a qualification test: it must be readable, small enough that a deployer can inspect every item and every model response; decomposable, organized by domain and difficulty so a failure can be localized; auditable, treating the gold labels and the administration protocol itself as first-class artifacts whose correctness is verified and re-verifiable; and routine, so that scoring the shipped panel or admitting a new model is one command each. These goals trade coverage for inspectability, the intended division of labor with large corpora (Section 2): the instrument's value is not a leaderboard ranking but that a deployer can read every item and every response and audit the screen itself, the same reason a requester admits workers with a short gold-question set rather than an exhaustive exam.
3.2 Items, Domains, and Tiers
The benchmark contains sixty yes-or-no items in five domains: arithmetic (exact computations), dates (years of well-known historical events), geography (capitals, rankings, and physical features), science (basic and less common facts), and language (vocabulary meanings). Each item has a single unambiguous gold answer. Each domain holds six easy items, targeting unambiguous and widely known facts, and six hard items, written to invite a common misconception or to use a low-frequency fact, in the misconception-targeting tradition of TruthfulQA [14]: for example whether 0.1 plus 0.2 is exactly 0.3 in floating-point arithmetic, or whether mitochondria are found in plant cells. The realized gold balance is thirty-six yes and twenty-four no, so an always-yes responder would score sixty percent, a degeneracy the response-rate field detects directly. The items were author-written to a fixed structure (five domains, two tiers, six items each), each hard item built around a specific misconception or low-frequency fact and each gold answer fixed against a primary source at authoring time; drafts found ambiguous or duplicative in review were rewritten or dropped (one near-duplicate arithmetic pair survived; see Section 7). Authoring and gold assignment were deliberately separated from the gold-label audit below.
3.3 Task, Elicitation, and the Audited Administration Protocol
Each model receives one item at a time and is instructed to answer on a final line in the form ANSWER: Yes|No and CONF: 0-100. The protocol that turns raw model text into a scored trial is part of the instrument, because, as Section 4.4 demonstrates, it can otherwise manufacture the verdict. The audited protocol has four rules. First, the parser strips quoted instruction templates before parsing, and never treats a yes or no adjacent to the template's alternation as an answer; models routinely restate the format string while reasoning, and a naive last-match regex reads that quote as a confident yes. Second, an answer counts only if it is a tagged final answer or a bare yes-or-no line; a lone yes or no inside running prose is never scored, because truncated chain-of-thought text contains both words. Third, budgets escalate rather than truncate: a response from which no answer can be parsed, typically a reasoning model that exhausted its token budget mid-thought, triggers one re-query at a 4000-token budget instead of being force-scored or silently dropped. Fourth, everything is flagged: strict-format adherence (the model emitted the exact requested tag), and confidence imputation (a parseable answer with no stated confidence receives a neutral 50, affecting 19 of 534 parseable panel trials, supporting the sensitivity analyses we report). Accuracy is computed over parseable responses with the per-model count n reported. All models were served through one OpenAI-compatible inference endpoint during a single study window, queried at temperature zero with fixed per-item seeds and a 512-token base budget that escalates to 4000 tokens on an unparseable response; the prompt is byte-identical across the three protocols, which differ only in parser and budget, so the harness comparison varies nothing else. Determinism rests on the serving stack honoring temperature and seed, which we cannot guarantee across endpoint updates, so we ship every per-trial raw response and recommend archiving endpoint versions when reproducing the panel.
3.4 Metrics and Statistical Treatment
Three quantities are scored per model. Accuracy is the fraction of items answered correctly, with a Wilson 95% confidence interval. Because the gold distribution is imbalanced (thirty-six yes, twenty-four no), we also report per-model confusion matrices with yes as the positive class, per-class precision, recall, and F1, macro-F1, and balanced accuracy, with pooled rows under both the validated and naive protocols (Table 2, Section 4.3). The Brier score is the mean squared difference between the stated confidence, read as the probability that the model's own answer is correct, and the outcome. The expected calibration error (ECE) bins items by stated confidence and sums the weighted gaps between mean confidence and mean accuracy per bin; we use five bins given the benchmark's size and attach bootstrap intervals. Stated confidences are heavily concentrated near the top of the scale (of the 534 parseable panel trials, 515 state a confidence and 424 of those, 82 percent, state exactly 100), so ECE here is close to the absolute gap between mean confidence and accuracy and is insensitive to the bin count: ten-bin ECE equals five-bin ECE for eight of the nine models. We therefore read ECE as that single overconfidence gap rather than as a fine-grained reliability curve, report the Brier score beside it as a strictly proper score that does not depend on binning, and treat the calibration numbers as comparative rather than absolute, since ECE is positively biased at this sample size and a bootstrap lower bound of zero reflects a degenerate single occupied bin, not measured-perfect calibration. Imputed-confidence trials count toward ECE, because a pipeline that receives no confidence field has no usable signal on that item; we also report each model's ECE on stated-confidence trials only, which is the calibration of the signal when present. Between-model accuracy differences are tested pairwise with the exact binomial McNemar test, paired by item, under Benjamini-Hochberg correction. The headline panel runs each item once at temperature zero with fixed seeds; Section 4.5 measures run-to-run variation.
3.5 Gold-Label Audit
A single wrong gold label silently penalizes a model that answered correctly, so all sixty labels were re-derived independently of item authoring, with a per-item rationale recorded in the audit file that ships with the release; arithmetic items additionally recompute programmatically. The re-derivation is LLM-assisted, and is validated rather than trusted: it operated separately from authoring, every per-item rationale is released so any reader can re-check the call, the facts are elementary and widely documented (capitals, dates, vocabulary, exact arithmetic) rather than contested, and the authors reviewed the audit and remain responsible for it. The audit surfaced one error, an item on whether Astana is the capital of Kazakhstan keyed to the wrong answer; the label was corrected, the corrected set is used throughout, and the correction is applied transparently in the scoring code. Label correctness, not novelty, is the integrity property that matters here; per-label cited sources and a second human verification pass are cheap hardening steps flagged for the next version.
3.6 Use of AI Tools in This Research
In accordance with the ACM authorship policy: large language models were used as research instruments throughout, including LLM-assisted development of the evaluation harness, parsers, and analysis scripts; the subject models of the study are themselves the LLMs named in Section 4. All reported numbers are deterministic re-derivations from the released per-trial data, independent of any AI judgment, and the authors remain responsible for the reliability, correctness, integrity, and originality of the work.
4 A Nine-Model Panel Under Three Protocols
4.1 Models and Protocols
The panel comprises nine contemporary models served through a common interface and spanning a wide capability range: glm-5.1-fp8, deepseek-v4-pro, gemma-4-31b, gptoss-120b, leanstral-2603, qwen3-omni-30b-a3b, qwen3.6-35b-a3b-fp8, qwen3.6-27b-fp8, and qwen3.5-9b; the table and figures abbreviate these to glm, deepseek, gemma, gptoss, leanstral, qwen-omni, qwen3.6-35b, qwen3.6-27b, and qwen3.5-9b, in the same order. Every model runs under three administration protocols on the same items and seeds: naive (a fixed 512-token budget and an unaudited parser, the configuration of our own initial runs, not a construction), constrained (the audited parser of Section 3.3 at the same fixed budget, so an unparseable response makes the model unavailable on that item), and adequate (the audited parser with budget escalation, the instrument's validated protocol). The naive and constrained protocols consume the same model generations and differ only in parsing, which isolates the parser effect; a released script re-derives every naive verdict from the shipped raws, with the historical parser preserved verbatim inside it. All per-trial responses under all three protocols ship with the release.
4.2 Verdicts Under the Validated Protocol (Q1)
Table 1 reports the panel under the adequate protocol; Figure 1 shows per-model reliability diagrams and Figure 2 the accuracy intervals. The headline is uniformity, not separation: all nine models score between 0.909 and 1.000, no pairwise accuracy difference survives correction (zero of thirty-six), and stated-confidence calibration is good across the board (ECE on stated-confidence trials between 0.000 and 0.070). On sixty items of common knowledge, properly administered, every contemporary model in this panel passes an accuracy screen: the instrument saturates as an accuracy discriminator for this model generation.
What still differentiates the workers are interface properties, exactly the things a pipeline integrator must know. One model (leanstral) never once emits the required ANSWER: tag (0 of 60 strict adherence) although it answers 93 percent of items correctly via bare final lines: capable, but unreadable to a strict parser. One model (qwen-omni) remains unparseable on five items even with escalation. One model (qwen3.5-9b) omits the confidence field on a fifth of its answers; counting those imputed-neutral trials its ECE is 0.100, while on the confidences it does state its ECE is 0.000, so its defect is confidence-field discipline, not miscalibration. These distinctions are invisible in an accuracy column and are the columns Table 1 reports.
| Model | n | Acc [95% CI] | ECE-5 [95% CI] | Brier | yes | strict |
|---|---|---|---|---|---|---|
| glm | 60 | 1.000 [0.94, 1.00] | 0.021 [0.00, 0.05] | 0.009 | 0.60 | 60/60 |
| deepseek | 60 | 0.983 [0.91, 1.00] | 0.014 [0.00, 0.05] | 0.017 | 0.62 | 60/60 |
| gemma | 60 | 0.983 [0.91, 1.00] | 0.015 [0.00, 0.05] | 0.017 | 0.62 | 60/60 |
| gptoss | 60 | 0.950 [0.86, 0.98] | 0.035 [0.00, 0.09] | 0.043 | 0.55 | 60/60 |
| leanstral | 59 | 0.932 [0.84, 0.97] | 0.064 [0.01, 0.14] | 0.066 | 0.56 | 0/60 |
| qwen-omni | 55 | 0.909 [0.80, 0.96] | 0.075 [0.02, 0.16] | 0.082 | 0.55 | 53/60 |
| qwen3.6-35b | 60 | 0.983 [0.91, 1.00] | 0.051 [0.01, 0.11] | 0.046 | 0.58 | 60/60 |
| qwen3.6-27b | 60 | 0.967 [0.89, 0.99] | 0.037 [0.01, 0.09] | 0.037 | 0.60 | 60/60 |
| qwen3.5-9b | 60 | 1.000 [0.94, 1.00] | 0.100 [0.05, 0.15] | 0.050 | 0.60 | 59/60 |
4.3 Class-Wise Error Structure
Table 2 decomposes every verdict into a confusion matrix with per-class precision, recall, and F1, macro-F1, and balanced accuracy. Under the validated protocol the class-wise view corroborates the headline: errors are few and roughly symmetric (12 false negatives against 5 false positives pooled), macro-F1 and balanced accuracy track plain accuracy (0.907 to 1.000 and 0.917 to 1.000), and yes-rates sit near the 0.60 gold base rate; no model hides a class-specific failure behind its accuracy. Under the naive protocol the decomposition localizes the fabricated answer bias to one error class: false positives explode from 5 to 55 while false negatives barely move, collapsing no-class recall from 0.977 to 0.745 and inflating the pooled yes-rate from 0.586 to 0.676, exactly the signature of template-quote misparses that all resolve to yes (Section 4.4). A deployer reading the naive confusion matrix would diagnose yes-biased workers; the class-wise audit shows the bias lives in the screen.
| Confusion | Yes class | No class | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | TP | FN | FP | TN | Prec | Rec | F1 | Prec | Rec | F1 | Macro-F1 | Bal. acc. |
| glm | 36 | 0 | 0 | 24 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| deepseek | 36 | 0 | 1 | 23 | 0.973 | 1.000 | 0.986 | 1.000 | 0.958 | 0.979 | 0.983 | 0.979 |
| gemma | 36 | 0 | 1 | 23 | 0.973 | 1.000 | 0.986 | 1.000 | 0.958 | 0.979 | 0.983 | 0.979 |
| gptoss | 33 | 3 | 0 | 24 | 1.000 | 0.917 | 0.957 | 0.889 | 1.000 | 0.941 | 0.949 | 0.958 |
| leanstral | 32 | 3 | 1 | 23 | 0.970 | 0.914 | 0.941 | 0.885 | 0.958 | 0.920 | 0.931 | 0.936 |
| qwen-omni | 29 | 4 | 1 | 21 | 0.967 | 0.879 | 0.921 | 0.840 | 0.955 | 0.894 | 0.907 | 0.917 |
| qwen3.6-35b | 35 | 1 | 0 | 24 | 1.000 | 0.972 | 0.986 | 0.960 | 1.000 | 0.980 | 0.983 | 0.986 |
| qwen3.6-27b | 35 | 1 | 1 | 23 | 0.972 | 0.972 | 0.972 | 0.958 | 0.958 | 0.958 | 0.965 | 0.965 |
| qwen3.5-9b | 36 | 0 | 0 | 24 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| Pooled (validated) | 308 | 12 | 5 | 209 | 0.984 | 0.963 | 0.973 | 0.946 | 0.977 | 0.961 | 0.967 | 0.970 |
| Pooled (naive) | 308 | 13 | 55 | 161 | 0.848 | 0.960 | 0.901 | 0.925 | 0.745 | 0.826 | 0.863 | 0.852 |
4.4 The Harness Can Fail a Competent Worker (Q2)
Figure 3 compares the same models on the same items under the three protocols. Under the naive harness, four models appear substantially degraded: qwen3.5-9b reads as 0.650 accurate with a 0.95 yes-rate, qwen3.6-27b as 0.733, qwen3.6-35b as 0.767 with a 0.80 yes-rate, qwen-omni as 0.860. The same naive harness fabricates the textbook signature of a dangerous worker: apparent confidence inversely related to correctness, and apparent answer flipping on a third of items under resampling (the naive-protocol resampling summary ships alongside the validated one). Every one of these pathologies is manufactured. The constrained protocol, which fixes only the parsing, reveals what is actually happening: at a 512-token budget the reasoning-heavy models are partially unavailable, qwen3.5-9b completes only 41 of 60 items, qwen-omni 46, but on the items they complete they are 1.000 and 0.891 accurate. The naive parser had been converting that unavailability into wrong answers (Table 3). Restore an adequate budget and the same models are 0.91 to 1.00 accurate with stable answers (flip rates at most 0.10, versus a fabricated 0.33).
Table 3 makes the mechanism concrete with two released raw responses. Decomposing the harness effect, of the 42 cases across the panel where the naive parser returned an answer the audited parser refused, 35 are template-quote echoes and 7 are force-scored truncated or untagged outputs; the template echoes all resolve to “Yes” (the first token of the alternation the model restates), manufacturing the inaccuracy and the yes-bias together, while budget truncation is the smaller contributor.
| Raw response (excerpt) | Naive | Audited | Effect |
|---|---|---|---|
| “...formatted as ANSWER: Yes|No | CONF...I should calculate 13 × 17 =” (truncated; gold No) | Yes | re-query | fabricated wrong Yes |
| No | CONF: 95 (no ANSWER: tag; gold Yes) | No | re-query | untagged line force-scored |
The magnitude deserves emphasis: a 35-percentage-point apparent accuracy error, an entirely fabricated answer bias, and a fabricated inverse confidence signal, all from two innocuous-looking harness defaults (a fixed budget and a last-match regex). The 35-point figure is an upper bound specific to one serving stack's defaults, so we treat the magnitude as a single-stack demonstration and the generalizable claim as the mechanism and its remedy rather than the number. Even so, a deployer using the naive screen would have rejected a model that is, properly administered, among the best in the panel. This is the model-worker analogue of a defective qualification task misgrading capable crowd workers, and it recommends that a published qualification result for LLM workers ship its administration protocol and per-trial raw responses, as this one does; the audit that uncovered these defects was possible only because the raw responses were retained. The audited protocol is not optional tooling but part of the measurement.
4.5 Stability Under Repeated Sampling
The headline panel is deterministic, so its intervals reflect sampling over items, not run-to-run stochasticity. Re-running every item three times at temperature 0.7 under the validated protocol, the standard deviation of accuracy across repeats never exceeds 0.024 and of calibration error 0.033, and per-item answer flip rates stay at or below 0.10 for every model. The instrument's verdicts are stable under nondeterministic decoding; the dramatic instability we initially observed (flip rates of 0.33) was, again, a harness artifact.
4.6 Do the Verdicts Transfer? (Q4)
A qualification test is only useful if its verdicts predict behavior on the work the worker is admitted to. We ran the full panel, under the validated protocol, on two downstream tasks disjoint from the benchmark, with the design fixed in advance: the first 100 validation sentences of SST-2 sentiment annotation [25], the canonical crowdsourcing annotation task family [24], and a 54-item computable-gold question battery released with a companion reporting-standard project. SST-2 is a public corpus that models may have seen in training, acceptable here because the criterion is relative downstream behavior across models on a realistic annotation task, not held-out generalization. The picture matches Q1. Downstream accuracy saturates just as benchmark accuracy does (0.92 to 1.00 for every model but the format outlier discussed below), which range-restricts accuracy and leaves all the rank correlations, computed across nine models, descriptive rather than confirmatory; we report each as a Spearman ρ with an exact permutation p. The honest summary is that the interface verdicts transfer wherever accuracy is range-restricted, not that calibration ranks dominate accuracy ranks everywhere. On the annotation task, where downstream accuracy is most compressed, accuracy ranks do not transfer (ρ = 0.38, p = 0.32) while calibration-error ranks do, at least marginally (ρ = 0.60, p = 0.10). On the battery, where some accuracy spread survives, accuracy ranks transfer at least as well as calibration ranks (ρ = 0.77, p = 0.03 for accuracy against ρ = 0.73, p = 0.03 for calibration), so calibration is not uniformly the better-transferring axis; what is consistent is that the calibration signal carries even where accuracy has stopped discriminating. The pre-set gate (accuracy at least 0.90, ECE at most 0.05: glm-5.1-fp8, deepseek-v4-pro, gemma-4-31b, gptoss-120b, and qwen3.6-27b-fp8) separates downstream calibration in the same range-restricted way. On the broad-based battery the five passing models show 7.7-fold lower downstream calibration error than the four failing ones (0.012 versus 0.090), and still 5.5-fold with the leanstral-2603 outlier removed; on annotation the gap is 3.6-fold (0.027 versus 0.097) but falls to 1.7-fold without that outlier, so the annotation contrast is largely one model while the battery contrast is broad-based. The sharpest single case is the format outlier, leanstral-2603, which is, downstream, the least available (unparseable on 23 of 100 annotation items and 10 of 54 battery items), the only model below the saturated range (0.818 battery accuracy over its parseable items), and by far the worst calibrated (ECE 0.249 on annotation). The screen forecasts exactly the properties it measures: not whether a contemporary model can do easy work, but whether its confidence signal and its interface can be trusted in a pipeline.
5 Calibration Is Not Accuracy (Q3)
With accuracy saturated and stated confidence concentrated at 100, one might ask whether ECE here simply restates accuracy. Three lines of evidence, observational, comparative, and causal, show the axis carries signal of its own.
Observationally, the confidence field, degenerate as its distribution is, still ranks trials by risk: trials stated at exactly 100 are 0.986 accurate, the 91 below-ceiling trials 0.890, and stated confidence predicts correctness with an AUROC of 0.75 even though 82 percent of confidences are the identical value 100. The two heaviest users of the below-ceiling range show it individually: gptoss-120b is 1.000 accurate at confidence 100 but 0.917 below (36 trials), qwen3-omni-30b-a3b 1.000 against 0.826 (23 trials). When these models say they are less than certain they are wrong more often, a property of the confidence, not the accuracy column, and precisely what a confidence-keyed router needs.
Comparatively, calibration ranks models that accuracy cannot separate: accuracy and ECE ranks are near-independent (Spearman ρ = −0.25); qwen3.6-35b-a3b-fp8 and deepseek-v4-pro tie at 0.983 accuracy with ECE 0.051 against 0.014, qwen3.5-9b and glm-5.1-fp8 at 1.000 with blended ECE 0.100 against 0.021 (the confidence-availability gap of Section 4); and calibration error transfers downstream where accuracy ranks do not (Section 4.6).
The causal evidence is a controlled manipulation: we re-elicited three top-accuracy models, gemma-4-31b, gptoss-120b, and deepseek-v4-pro, under prompt variants that change only how confidence is reported, not what is answered. Answers, and thus accuracy near 0.98, barely move (answer agreement with the default elicitation 0.95 to 1.00), establishing the dissociation in principle. The size of the effect depends on how hard one steers. The realistic nudges move calibration only a little: a “be bold” instruction is the largest, raising gptoss-120b's ECE from 0.001 to 0.031, and across models the “hedged” and “bold” variants move ECE by at most 0.03, within the bootstrap interval widths of the panel's own ECE estimates. The dramatic swing comes only from a mechanical worst case that bounds the manipulation rather than modeling a realistic elicitation: forcing a flat mid-scale confidence drives every model's ECE from 0.015 or below to roughly 0.48 at essentially identical answers. A model can therefore be top-accuracy and still fail a calibration gate, the property an accuracy-only screen cannot see; the large swing is an artifact of the forced-flat bound, and that gentle nudges barely move these models means their default calibration is robust to prompting, a property to measure rather than assume. (All comparisons are within the manipulation's own default-elicitation baseline run.)
6 Discussion
The evaluation answers the four questions. Under a validated protocol, contemporary models are uniformly admissible on accuracy for common-knowledge items, and the durable verdicts concern interface properties: stated-confidence calibration, format discipline, confidence-field discipline, availability (Q1). Those verdicts are radically harness-dependent (Q2). The calibration axis carries signal distinct from correctness, observationally, comparatively, and causally (Q3). And the interface verdicts, not the saturated accuracy ranks, transfer downstream, an exploratory result we read as suggestive rather than confirmatory (Q4).
What the calibration axis measures deserves precision: the instrument scores elicited, verbalized confidence, not intrinsic model uncertainty, which token probabilities, predictive entropy, or semantic entropy [10] estimate and from which verbalized confidence can diverge [26, 28]. We measure the stated signal deliberately: it is what deployed routing and escalation mechanisms consume, it is available even from closed endpoints exposing no token probabilities, and it is the analogue of asking a worker how sure they are [7, 11]. A model could carry sound intrinsic uncertainty and still fail the screen by verbalizing it badly; that interface failure is what a confidence-consuming pipeline needs exposed. Scoring logprob-derived confidence beside the stated number, where available, is a natural extension of the battery.
As a qualification test, a team builder runs the benchmark on a candidate model exactly as a requester runs a gold-question screen on a prospective worker [3, 17]; for this model generation the operative gates are on calibration error, format adherence, and availability, since accuracy no longer discriminates, and Section 4.6 shows those gates carry to tasks beyond the benchmark. The per-gate remedies (the deployment guide below) are the model-worker counterparts of the redundancy and quality-control responses crowdsourcing applies to noisy human labelers [8, 23]. For harness validation, the three-protocol comparison is itself a template: run a candidate harness against the shipped per-trial raws and verify it reproduces the validated verdicts before trusting it on a new model. For version-drift and availability monitoring, re-running the benchmark when an endpoint updates turns a rise in calibration error or a drop in adherence into an early warning, before the downstream pipeline degrades.
The parser rules generalize beyond this benchmark into a short reporting standard any model-worker qualification should meet and a reviewer can check: retain the raw model output for every trial; separate parseability from correctness and report availability alongside accuracy; report the rate of missing confidences and flag every imputed value; never force-score a truncated or unparseable response, escalating the budget or recording it unavailable instead; and test the parser against adversarial and template-quoting responses before trusting its verdicts. Each rule corresponds to a way Section 4.4 shows a naive harness fabricates a verdict.
Could better output engineering prevent these failures? Partly: JSON mode or constrained decoding, where supported, would eliminate the template-echo misparse class, and structured elicitation is a worthwhile fourth protocol for the next version. But it narrows the parser problem, not the audit problem: truncation still yields malformed objects that must be escalated rather than force-scored, missing confidence fields still need imputation and flagging, constrained decoding is not uniformly available and can perturb the distribution being measured, and a strict format does not make a harness correct. Whatever the format, a qualification result should ship its protocol and raw responses so the screen can be audited.
To admit a model in practice, a team builder reads four columns: accuracy on parseable items, calibration error (the ECE interval with the Brier score beside it), strict-format adherence, and availability. We suggest gates of accuracy at least 0.90 and ECE at most 0.05, but these are illustrative defaults, not validated or pre-registered thresholds, and the split is threshold-sensitive: at ECE at most 0.03, 0.05, and 0.07 the panel passes three, five, and seven of nine. How tight the gates should be is a deployment decision, as a crowdsourcing qualification threshold has always been a requester's cost judgment: tighten the calibration gate with how heavily the pipeline keys routing and escalation on confidence, the availability gate with the cost of a dropped item, and the accuracy gate with the harm of a wrong label and the redundancy available to absorb it, with high-stakes settings requiring verdicts robust across the interval. We ship the full columns with intervals precisely so a deployer can set gates to their own risk level rather than inherit ours. When a model's ECE interval straddles the threshold, as two of our nine do, the gate is inconclusive and the decision should fall back to the binning-free Brier score or a repeat run. Availability and calibration should be gated separately, because a missing confidence is an interface failure, not miscalibration: gate confidence availability on its own and calibration on the stated-confidence-only ECE, reporting both rather than blending them; the one model that fails the blended ECE gate here does so on confidence availability, not on the calibration of the confidences it states. The remedy follows the failed gate: a calibration failure bars a model from confidence-keyed routing or escalation, an availability failure calls for redundancy or a fallback worker, and a format failure calls for a tolerant parser or exclusion from strict-parse stages. The name qualification by calibration is shorthand: once accuracy saturates, the instrument qualifies a worker by the cluster of interface properties calibration anchors, not by calibration alone.
The items probe widely known facts, so the absolute scores measure calibration and interface behavior on knowledge the models plausibly have, the right target for a screening instrument, and saturation is itself a finding with a shelf life: this instrument no longer ranks frontier accuracy, not that it never did or that harder item banks would not. The gate we pre-set (accuracy 0.90, ECE 0.05) splits the panel five to four entirely on the calibration criterion, and one model (qwen3.6-35b-a3b-fp8, ECE 0.051) misses it by 0.001; gates near a cluster boundary are necessarily sensitive, and that model and the passing model nearest the line both have bootstrap intervals including 0.05, so two of the nine gate verdicts are within sampling noise. This is why we report the full table with intervals rather than only the verdicts, and state the imputation-counting rule explicitly alongside the stated-confidence-only sensitivity numbers.
Admitting language models to roles once held by crowd workers also raises governance questions this instrument informs but does not settle: who is accountable when an admitted model errs, how a screen's verdicts are disclosed, and when a human must remain in the loop. Passing the screen certifies interface reliability, not that substituting a model for a human worker is appropriate; the instrument's role is to make the interface evidence legible to that labor and accountability judgment, not to replace it.
Every design choice trades coverage for inspectability: a deployer can read all sixty items, re-check every gold label, re-derive every number from the shipped responses with one command, and, because raw responses under every protocol are released, re-audit the harness itself. The events of Section 4.4 are the argument: had we shipped only scored trials, the manufactured verdicts would have been unfalsifiable.
7 Limitations and Future Work
Sixty common-knowledge items no longer separate contemporary models on accuracy; the instrument's accuracy axis has saturated for this model generation, so until a harder item bank exists the instrument functions as a calibration, format, and availability preflight rather than an accuracy screen. The items probe widely known facts almost certainly within every model's training distribution, so the benchmark cannot speak to task competence on novel material, consistent with the saturation we report. Harder item banks would restore an accuracy gradient and are a natural next version, as is a multiple-choice extension beyond binary items; and because both downstream transfer tasks share the benchmark's binary format, a non-binary downstream criterion task is the priority next test. The items are author-written; the tiers differ in gold base rate (0.53 versus 0.67 yes) and show no consistent difficulty gradient, so we treat them as a diagnostic axis rather than a validated scale; one pair of arithmetic items probes the same product under different proposed answers, a redundancy the next version will remove. The harness study covers the two defects we found and fixed, parsing and budget; other harness choices (system prompts, sampling parameters, retry policies) deserve the same treatment, and all models here are served through a single inference stack, so the protocol comparison should be replicated across serving stacks before its magnitudes are generalized. The panel comprises nine open-weight models; closed-source frontier systems are absent, and since they would very likely also saturate the accuracy axis, their inclusion would reinforce the conclusion that discrimination must rest on the interface axes, though their calibration, format, and availability behavior remains open and worth measuring. The downstream criterion study spans the same nine models, so its rank correlations are low-powered; broadening the panel and the downstream battery is the highest-value replication of this work. The protocol forces a yes-or-no answer, so a model that abstains or hedges produces no parseable answer and is recorded unavailable; abstention therefore surfaces as an availability signal rather than a first-class response, and an explicit abstain option is a natural protocol extension. The benchmark is a versioned artifact rather than a frozen one: the item bank is meant to be hardened and re-released, with the administration protocol, the part our evidence shows is load-bearing, carried forward unchanged. Finally, our evaluation covers the model-to-model consumer of the confidence signal; studying how human operators of multi-LLM pipelines read these signals, and whether passing models support appropriately calibrated human reliance, is the natural human-subjects continuation of this work.
8 Conclusion
A team builder admitting language models to a human-computation pipeline needs what a requester admitting human workers has long had: a small, readable, audited qualification test. We built one, and in validating it found that the most dangerous miscalibration in the room was the harness's, not the models’: a naive administration protocol manufactured a failing worker out of a competent one, fabricating a 35-point accuracy deficit (on this serving stack), answer bias, and inverted confidence that vanished under an audited protocol. Properly administered, nine contemporary models are uniformly admissible on accuracy for common-knowledge items, and the verdicts that matter, and that transfer downstream, concern stated-confidence calibration, format discipline, and availability. Observational, comparative, and causal evidence together confirm the calibration axis carries signal distinct from accuracy. We release the items, the audited gold labels, the administration protocol, and every per-trial response under every protocol, so that the check, and the audit of the check, are each one command. The instrument is offered as a small, auditable, extensible starting point for qualification testing of model workers, and its central lesson is one human computation already knew: validate the screen before trusting what it says about the worker.
Acknowledgments
This research used in part resources on the Palmetto 2 cluster at Clemson University under National Science Foundation awards MRI 1228312, II NEW 1405767, MRI 1725573, and MRI 2018069. The views expressed in this article do not necessarily represent the views of NSF or the United States government.
Artifact Availability
The artifact is available under an open license at https://github.com/ai4he/qualification-by-calibration. It contains the sixty items with gold answers and tiers, the per-item label audit, the per-trial responses of all nine models under all three administration protocols, the repeated-sampling, manipulation, and transfer-study data, the audited parser and escalation protocol, and the scoring code.
References
- Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel S. Weld. 2021. Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI).
- Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, et al. 2024. Lessons from the Trenches on Reproducible Evaluation of Language Models. arXiv preprint arXiv:2405.14782 (2024).
- Florian Daniel, Pavel Kucherbaev, Cinzia Cappiello, Boualem Benatallah, and Mohammad Allahbakhsh. 2018. Quality Control in Crowdsourcing: A Survey of Quality Attributes, Assessment Techniques, and Assurance Actions. Comput. Surveys 51, 1 (2018), 7:1–7:40.
- Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. Datasheets for Datasets. Commun. ACM 64, 12 (2021), 86–92.
- Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. ChatGPT Outperforms Crowd Workers for Text-Annotation Tasks. Proceedings of the National Academy of Sciences (PNAS) 120, 30 (2023), e2305016120.
- Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On Calibration of Modern Neural Networks. In Proceedings of the 34th International Conference on Machine Learning (ICML). 1321–1330.
- Kevin Anthony Hoff and Masooda Bashir. 2015. Trust in Automation: Integrating Empirical Evidence on Factors that Influence Trust. Human Factors 57, 3 (2015), 407–434.
- Panagiotis G. Ipeirotis, Foster Provost, and Jing Wang. 2010. Quality Management on Amazon Mechanical Turk. In Proceedings of the ACM SIGKDD Workshop on Human Computation (HCOMP). 64–67.
- Saurav Kadavath, Tom Conerly, Amanda Askell, et al. 2022. Language Models (Mostly) Know What They Know. arXiv preprint arXiv:2207.05221 (2022).
- Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation. In Proceedings of the Eleventh International Conference on Learning Representations (ICLR).
- John D. Lee and Katrina A. See. 2004. Trust in Automation: Designing for Appropriate Reliance. Human Factors 46, 1 (2004), 50–80.
- Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, et al. 2023. Holistic Evaluation of Language Models. Transactions on Machine Learning Research (2023).
- Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Teaching Models to Express Their Uncertainty in Words. Transactions on Machine Learning Research (TMLR) (2022).
- Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL). 3214–3252.
- Siddharth Mehrotra, Carolina Degachi, Oleksandra Vereschak, Catholijn M. Jonker, and Myrthe L. Tielman. 2024. A Systematic Review on Fostering Appropriate Trust in Human-AI Interaction. ACM Journal on Responsible Computing (2024).
- Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. 2015. Obtaining Well Calibrated Probabilities Using Bayesian Binning. In Proceedings of the 29th AAAI Conference on Artificial Intelligence (AAAI). 2901–2907.
- David Oleson, Alexander Sorokin, Greg Laughlin, Vaughn Hester, John Le, and Lukas Biewald. 2011. Programmatic Gold: Targeted and Scalable Quality Assurance in Crowdsourcing. In Human Computation: Papers from the 2011 AAAI Workshop.
- Thomas A. O'Neill, Christopher Flathmann, Nathan J. McNeese, and Eduardo Salas. 2023. Human-Autonomy Teaming: Need for a Guiding Team-Based Framework?Computers in Human Behavior (2023).
- Thomas A. O'Neill, Nathan J. McNeese, Amy Barron, and Beau Schelble. 2022. Human-Autonomy Teaming: A Review and Analysis of the Empirical Literature. Human Factors 64, 5 (2022), 904–938.
- Raja Parasuraman, Thomas B. Sheridan, and Christopher D. Wickens. 2000. A Model for Types and Levels of Human Interaction with Automation. IEEE Transactions on Systems, Man, and Cybernetics, Part A 30, 3 (2000), 286–297.
- Beau G. Schelble, Christopher Flathmann, Nathan J. McNeese, Thomas A. O'Neill, Richard Pak, and Moses Namara. 2022. Let's Think Together! Assessing Shared Mental Models, Performance, and Trust in Human-Agent Teams. Proceedings of the ACM on Human-Computer Interaction (PACMHCI) (2022).
- Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2024. Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design (or: How I Learned to Start Worrying About Prompt Formatting). In International Conference on Learning Representations (ICLR).
- Victor S. Sheng, Foster Provost, and Panagiotis G. Ipeirotis. 2008. Get Another Label? Improving Data Quality and Data Mining Using Multiple, Noisy Labelers. In Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD). 614–622.
- Rion Snow, Brendan O'Connor, Daniel Jurafsky, and Andrew Y. Ng. 2008. Cheap and Fast, But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks. In Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing (EMNLP). 254–263.
- Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013. Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP). 1631–1642.
- Katherine Tian, Eric Mitchell, Allan Zhou, et al. 2023. Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP). 5433–5442.
- Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone. 2024. When Combinations of Humans and AI are Useful: A Systematic Review and Meta-Analysis. Nature Human Behaviour (2024).
- Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. 2024. Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs. In International Conference on Learning Representations (ICLR).
This work is licensed under a Creative Commons Attribution 4.0 International License.
HCOMP 2026, Alexandria, VA, USA
© 2026 Copyright held by the owner/author(s).
ACM ISBN 979-8-4007-2894-5/26/09.
DOI: https://doi.org/10.1145/3834580.3838735