AMALIA matches larger models on agreement with human coders for moral-authority annotation, yet its recovery gap shows most of that performance is not reproduced by the theory-defined route—revealing a validity shortfall invisible to agreement metrics.
FotiosFitsilis, MariaKamilaki, BasilisGatos, VassilisKatsouros, andGeorgeMikros
12 Pith papers cite this work, alongside 10,874 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Grain calibration decomposes theoretical constructs into clause-level components, tests each with extractive evidence, and combines results through explicit theory-derived rules to validate LLM coding beyond agreement with human annotators.
Larger LLMs acquire basic situation modeling before mentalizing on false-belief tasks, with performance depending on size, training volume, and post-training, yet remaining sensitive to non-factive verbs and agent knowledge states.
Simulations show standard CI methods underperform for classifier metrics in small and nested datasets, while Agresti-Coull, Wilson, Clopper-Pearson, and a new pseudo-count regularized bootstrap perform better, with specific adjustments needed for nested structures.
Semantic mapping of 8,954 definitions and 2,700 scales from 14,000+ papers shows learner agency and autonomy span task regulation, personal motivation, and sociocultural dimensions, with existing scales and generative AI research underrepresenting the sociocultural dimension.
CFA and Generalizability Theory applied to LLM leaderboards show latent general-factor slopes are stable (R_g=0.97) while manifest scaling-law slopes are unreliable (R_β=0.53), with contributor metadata explaining more rank variance than architecture.
Semantic projection of Sentence-BERT embeddings onto axes from validated clinical scales yields continuous scores for depression, anxiety, and worry that correlate with standard measures, especially in structured text formats.
ValueBlindBench is a preregistered agreement-gated stress-test protocol for deciding when LLM-judged investment-rationale claims are stable enough to report, using 1,100 trajectories and 5,500 judge calls to gate claims by weighted kappa agreement.
The Stakeholder Grounding Exercise shows neural text embeddings are 19-26pp less reliable than human experts at capturing semantic distinctions, with misalignment strongly correlated to poorer clustering performance (ρ=0.9), replicated across Danish policy and US AI domains.
Introduces the Mechanism Plausibility Scale, a four-level framework separating generative sufficiency from mechanistic plausibility in LLM-based agent-based models.
AI benchmark evaluations require standardized item-level data releases as core infrastructure to support validity assessment, demonstrated via the OpenEval archive of 10M responses across 155k items.
Multimodal LLMs in robots develop self-identification and predictive awareness through sensorimotor loops, with structural equation modeling linking sensory integration to dimensions of the minimal self.
citing papers explorer
-
Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA
AMALIA matches larger models on agreement with human coders for moral-authority annotation, yet its recovery gap shows most of that performance is not reproduced by the theory-defined route—revealing a validity shortfall invisible to agreement metrics.
-
Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs
Grain calibration decomposes theoretical constructs into clause-level components, tests each with extractive evidence, and combines results through explicit theory-derived rules to validate LLM coding beyond agreement with human annotators.
-
Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models
Larger LLMs acquire basic situation modeling before mentalizing on false-belief tasks, with performance depending on size, training volume, and post-training, yet remaining sensitive to non-factive verbs and agent knowledge states.
-
Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data
Simulations show standard CI methods underperform for classifier metrics in small and nested datasets, while Agresti-Coull, Wilson, Clopper-Pearson, and a new pseudo-count regularized bootstrap perform better, with specific adjustments needed for nested structures.
-
Large-scale semantic mapping of learner agency and autonomy reveals what measurement and generative AI research overlook
Semantic mapping of 8,954 definitions and 2,700 scales from 14,000+ papers shows learner agency and autonomy span task regulation, personal motivation, and sociocultural dimensions, with existing scales and generative AI research underrepresenting the sociocultural dimension.
-
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
CFA and Generalizability Theory applied to LLM leaderboards show latent general-factor slopes are stable (R_g=0.97) while manifest scaling-law slopes are unreliable (R_β=0.53), with contributor metadata explaining more rank variance than architecture.
-
Measuring Psychological States Through Semantic Projection: A Theory-Driven Approach to Language-Based Assessment
Semantic projection of Sentence-BERT embeddings onto axes from validated clinical scales yields continuous scores for depression, anxiety, and worry that correlate with standard measures, especially in structured text formats.
-
ValueBlindBench: Agreement-Gated Stress Testing of LLM-Judged Investment Rationales Before Returns Are Observable
ValueBlindBench is a preregistered agreement-gated stress-test protocol for deciding when LLM-judged investment-rationale claims are stable enough to report, using 1,100 trajectories and 5,500 judge calls to gate claims by weighted kappa agreement.
-
Grounding Text Embeddings in Stakeholder Associations
The Stakeholder Grounding Exercise shows neural text embeddings are 19-26pp less reliable than human experts at capturing semantic distinctions, with misalignment strongly correlated to poorer clustering performance (ρ=0.9), replicated across Danish policy and US AI domains.
-
Mechanism Plausibility in Generative Agent-Based Modeling
Introduces the Mechanism Plausibility Scale, a four-level framework separating generative sufficiency from mechanistic plausibility in LLM-based agent-based models.
-
AI Evaluation Should Require Standardized Item-Level Data Releases
AI benchmark evaluations require standardized item-level data releases as core infrastructure to support validity assessment, demonstrated via the OpenEval archive of 10M responses across 155k items.
-
Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots
Multimodal LLMs in robots develop self-identification and predictive awareness through sensorimotor loops, with structural equation modeling linking sensory integration to dimensions of the minimal self.