Paper: arXiv 2609.19866

Authors: Veronika Batzdorfer, Carlo Romano Marcello Alessandro Santagiustina

Abstract

High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission’s AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran’s I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.

Complexity vs Empirical Score

  • Math Complexity: 3.0/10
  • Empirical Rigor: 7.0/10
  • Quadrant: Street Traders — practical and empirical, lighter on theory

Why this score: This paper presents a novel and empirically rigorous analysis of LLM measurement validity, distinguishing it from reproducibility. While the mathematical complexity is moderate, the robust statistical methodology and clear findings contribute to a high overall score. The insights into institutional communication are particularly valuable.

Research Flowchart

  flowchart TD
    A[Research Question: Reproducibility vs. Construct Validity in LLM Measurement?] --> B{Key Methodology: Comparing LLM annotations vs. Survey data};
    B --> C1[Data Input: EC AI Act Consultation Survey Responses];
    B --> C2[Data Input: EC AI Act Consultation Free-Text Submissions];
    C1 & C2 --> D[Computational Process: LLM Annotation & Statistical Analysis];
    D --> E{Key Finding 1: High LLM Reproducibility (>0.99 ICC) but Poor Construct Convergence};
    D --> F{Key Finding 2: Divergence Varies by Stakeholder (e.g., Business Associations: LLM > Survey AI Risk Concern, g=+1.0)};
    D --> G{Key Finding 3: Spatial Autocorrelation of Divergence (Moran's I = 0.347) & Survey Concerns vs. Explainability Support};