Research ·

Multi-Agent Code Judge Reliability: Label-Free Measures and Decline Mechanism

66Developing1 reportarXiv cs.AI

AI brief

AI-written

Why it mattersIt improves AI code review credibility and reduces risks from incorrect automated judgments.

Newly proposed label-free measurement method lets multi-agent code review systems proactively decline judgments when insufficient evidence exists

What happened

arXiv researchers addressed the common problem of multi-agent code review systems outputting conclusive judgments without sufficient evidence. They ran 80 conditional measurement tests on the publicly available MARCH review framework across two code evaluation benchmarks, finding the framework judged two code samples as equivalent in 78% to 95% of comparisons, with only 4.4% accuracy—far lower than the 43.7% accuracy of direct single-model code review. The team extracted two label-free measurement dimensions from process logs, setting filtering thresholds for comparisons with insufficient judgment evidence to let the system proactively decline these unsubstantiated comparisons. After filtering, system accuracy improved from 20.7% to 36.9%, while still covering half of all comparison tasks.

Key facts

Research subject
MARCH multi-agent code review framework
Observed core problem behavior
78% to 95% of comparisons judged two code samples as equivalent, with only 4.4% accuracy
Baseline performance
Direct single-model code review achieves 43.7% accuracy
Post-optimization performance
After filtering unsubstantiated comparisons, accuracy improved from 20.7% to 36.9%, while still completing half of all comparison tasks
Core innovation
Label-free method to identify when code review judgments lack sufficient supporting evidence

Background

When running code reviews, current large models often output conclusive, reasoning-backed judgments even without sufficient supporting evidence, making these unsubstantiated outputs indistinguishable from evidence-based judgments. Multi-agent verification frameworks work well for document retrieval evidence scenarios, but their original applicable conditions break down in code review use cases.

Why it matters

For the code review industry, this label-free evidence identification method reduces invalid AI review judgments, improving review output trustworthiness. For developers, this filtering mechanism cuts meaningless code comparisons to boost development efficiency. For end users, code tools built on this technology deliver more reliable outputs, reducing usage risks from unaddressed code vulnerabilities.

Written by AI from the original article. It may contain mistakes; the original is the source of truth.

Source

  1. arXiv cs.AI ↗Multi-Agent Code Judge Reliability: Label-Free Measures and Decline MechanismThe paper proposes label-free reliability measures and a code judge that avoids unfounded guesses.
Back to AI News