Author companion · Multi-agent AI

MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination

Preprint · arXiv:2605.22949 · Version 4, 8 October 2026 · First posted 21 May 2026

When several AI models answer the same question, should the most confident one win? Only if their confidence means the same thing. MARGIN learns what each model’s confidence actually means from the outcomes it observes.

Confidence does not mean the same thing across models

A coordinator that combines answers from several foundation models has to decide how much weight each answer gets. Self-reported confidence is the obvious signal. But the same number can mean different things for different models. A model that says 90% may be less reliable than one that says 70%.

On the hard subset of BigCodeBench, this goes beyond noise. Across an 18-model pool, models with higher average stated confidence had lower pass rates (correlation −0.69). Among pairs of answers where exactly one was correct, picking the more confident model chose the correct answer only 42.4% of the time, worse than a coin toss.

A second difficulty is that the workload changes. A correction learned on one kind of task may not suit the next one, even though the models themselves have not changed.

Correct each model’s confidence from its own record

  1. 01Answer

    Each model returns an answer and a stated confidence.

  2. 02Correct

    MARGIN rescales each confidence using that model’s record in the same confidence band.

  3. 03Weigh

    The corrected confidences weight a vote over the candidate answers.

  4. 04Learn

    When correctness is known, it updates the record for each model that answered.

The MARGIN loop. Each answer is scored before its own outcome updates any record.

MARGIN divides stated confidence into a few bands, three in the main experiments. For each model and band it keeps two running averages, one of observed accuracy and one of stated confidence. Their ratio gives the correction.

corrected confidence = (recent accuracy ÷ recent stated confidence) × stated confidence

A model that is right 60% of the time when it states about 90% gets a factor of about 0.67. A new answer at 90% is then counted at about 60%. A model whose stated confidence already matches its accuracy keeps a factor near 1.

A band with few observations borrows from the same model’s other bands, and relies on its own record as evidence accumulates. It never borrows another model’s history. The averages weight recent outcomes more heavily, so the correction can follow a change in workload.

MARGIN needs only the answers, their stated confidence and later correctness feedback. It does not retrain the models, read their internals or need a separate calibration dataset.

Successes and failures update the record at the same rate. The paper proves that, for outcomes with a fixed success probability, unequal rates bias the tracked accuracy, and equal rates are the only setting in that family that avoids it.

Watch the correction change the decision

Two models disagree. Model A states 90% confidence and Model B states 70%. A raw confidence vote picks Model A’s answer.

Constructed example · one confidence band per model
ModelStated confidenceRecent accuracy in bandRecent stated confidence in bandFactorCorrected confidence
A0.900.600.900.670.60
B0.700.700.701.000.70

After correction, Model B’s answer carries more weight and wins the vote. Nothing about either model changed. Only the coordinator’s reading of their confidence did.

Illustrative numbers. Model A’s figures are the illustration used in the paper’s method section. The example ignores the blending toward each model’s overall record, which matters only for sparsely observed bands.

What I evaluated

The experiments cover code generation, question answering and mathematics. The pool has 18 models, nine cloud-served and nine served locally, and a fixed nine-model cloud subset is used for the workload-change experiments. Calibration error is measured as expected calibration error (ECE) on a 0 to 100 scale, where lower is better.

Adapting when the workload changes

Each method first learns on one workload, keeps what it learned, and continues on a different one. All methods receive the same feedback. Online Platt scaling was the strongest of the five online alternatives.

Calibration error after the change (ECE ↓)
TransitionOnline PlattMARGINOutcome against the five alternatives
HumanEval → BigCodeBench15.3212.99Lower than all five
MBPP → BigCodeBench13.4211.89Lower than all five
MMLU STEM → MMLU-Pro business3.172.85Lower than four, with the Platt comparison inconclusive

Source: the paper’s matched-feedback comparison on the nine-model cloud subset. The histogram and windowed-accuracy alternatives were 13 to 31 points worse on the coding transitions and about 4 points worse on the question-answering transition.

Choosing better answers

A separate set of experiments starts without any history and asks whether corrected confidence improves the coordinator’s choice. Raw and corrected confidence use the same voting rule.

18-model pool, code generation
BenchmarkAnswer selection, raw (%)Answer selection, MARGIN (%)More confident is correct, raw (%)More confident is correct, MARGIN (%)
HumanEval94.5195.3346.5381.66
MBPP87.9492.2649.9481.25
BigCodeBench13.8927.8742.3763.36

Source: the paper’s coordination experiments. “More confident is correct” counts pairs in which exactly one of two answers was correct, so 50% is chance. Highlighted cells are differences with a 95% interval above zero. The HumanEval selection difference is inconclusive.

Where the claims stop

MARGIN needs correctness feedback. The main experiments provide it for every model that answered, which suits tasks with executable tests. If only the chosen model’s answer is ever checked, the rarely chosen models stop updating, and calibration error rose by roughly 30 points in that diagnostic.

When calibration starts with no history at all, MARGIN holds no general advantage. On BigCodeBench, four of the five online alternatives reached lower calibration error. The advantage reported above concerns adapting a correction that already exists.

The gains in answer selection are relative to uncalibrated confidence weighting. They do not show that combining several models beats deploying the single strongest model, and they say nothing about cost. The responders answer independently. Debate, planning and multi-step agents are outside the evaluated setting.

From one model’s record to collective decisions

I place MARGIN within my work on multi-agent coordination and institutional design, which asks how evidence about past performance should affect an actor’s influence over a decision.

TACIT applies the same principle to network applications that arbitrate configuration changes. Its reputation deliberately falls faster after a wrong prediction than it rises after a right one. MARGIN has a different job. It estimates how often a model is right at a given stated confidence, and for that estimate equal update rates are what keep it unbiased.

Paper and citation

J. Armstrong, “MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination,” arXiv:2605.22949, 2026. DOI 10.48550/arXiv.2605.22949.

Download BibTeX citation · Version described here: arXiv v4