Research ·
Source-Position Coherence Bias in AI Evaluation
AI brief
AI-writtenWhy it mattersHighlights evaluation bias to improve AI assessment fairness.
Research confirms AI evaluations of text are biased by alignment between the speaker’s identity and stated stance, rather than a simple preference for specific sources.
What happened
The research team ran two pre-registered descriptive studies plus a follow-up supplemental experiment with a European sample, using 6 fixed texts covering U.S. AI policy, Germany’s debt brake, and Swiss nuclear energy. They collected 2,976 valid AI scores via different source attribution pairings. For example, on a national security-themed text, the average score was 0.359 under the CODEPINK attribution and 0.639 under the Republican University attribution. When the same topic was rephrased as a civil rights framing, the score gap between the two sources narrowed to 0.021, ruling out an explanation of constant preference for a single source.
Key facts
- Valid study sample size
- 2,976 responses
- Number of test texts covered
- 6 texts
- Topic coverage
- U.S. AI policy, Germany’s debt brake, Swiss nuclear energy
- Representative score gap
- Up to 0.28 between different attributions for the same text
Background
Most existing AI evaluation focuses on factual accuracy of model outputs, with little attention to whether AI judgments of text can be induced by the speaker’s identity. Prior related research has mostly focused on human subjective evaluation.
Why it matters
The finding means AI-generated evaluations are not fully neutral. When judging public policy or stance-based text, assigning different source identities may produce contradictory scores. For evaluation developers, this requires additional controls for identity variables in stance-consistency scenarios. For general users, AI judgments of public viewpoints should not be treated as objective conclusions.
What to watch
Follow-up work will further verify whether this bias weakens as large model reasoning capabilities improve, and whether prompt engineering can eliminate the source-stance induction effect.
Written by AI from the original article. It may contain mistakes; the original is the source of truth.