Imagine evaluating a retrieval-augmented-generation system, where a user asks a question, the system retrieves a text passage, and an LLM judge decides whether it’s relevant. To reduce noise, you ask several judge models to evaluate the same passage. Eight say “relevant”; two say “not relevant”. Eight out of 10 feels convincing, but the important question is not only how many judges agreed but how independently they arrived at that agreement.

If the eight agreeing judges are genuinely different sources of evidence, then agreement is a strong signal. But if they share a prompt template, a training lineage, a model family, or a common blind spot, they may be repeating the same mistake. The vote count makes the evidence look stronger than it really is. Correlation between different judges' outputs limits the utility of multijudge panels.

A new paper, “Dependence-aware label aggregation for LLM-as-a-judge via Ising models,” coauthored with Shiva Kasiviswanathan and presented at this year’s International Conference on Machine Learning (ICML), addresses this problem. The authors present a method for assessing the correlations between judges’ outputs and adjusting the aggregate score accordingly, to ensure a diversity of opinion.

The method is designed for the unsupervised setting, learning from judge outputs without using human reference labels for training. It treats each item's true label as a latent variable to infer jointly with the parameters describing judge reliability and dependence. The algorithm combines each item's votes to estimate the probability that its true label is positive, then alternates between updating those probabilities and re-estimating judge reliability and pairwise dependence from them.

The authors evaluated their approach on three binary tasks: relevance classification for retrieved information, toxicity classification, and summarization assessment. The judge panel contained 10 judge models, all run at temperature zero. The dependence-aware models outperformed the best-performing baseline — a panel of judges weighted according to historical accuracy — by 9% to 14% on standard metrics.

The learned relationships among judges can be used during audits to help identify redundant judges and task-specific shared blind spots. The same learned network can help answer practical questions, such as whether similar models are adding independent evidence or simply reinforcing each other.

The Importance of Dependence-Aware Aggregation

Dependence-aware aggregation suggests a few useful habits for teams using LLM-as-a-judge pipelines. First, evaluate the judge panel, not just the individual judges. A set of individually strong judges can still be redundant if they fail in the same way. Second, treat model diversity as statistical diversity. Mixing model families or architectures is helpful only to the extent that it changes the error patterns that matter for the task.

Finally, report uncertainty with dependence in mind. Ten correlated votes should not always produce the same confidence as 10 independent votes. When LLM judges agree, we should ask why. Sometimes agreement is independent evidence. Sometimes it is a shared blind spot. A good aggregation method should be able to tell the difference.

Dieser Artikel wurde mit Unterstützung von KI verfasst.
News Factory APP - agentische News für besseres SEO & AEO.