A four-stage human-subject study finds that agent-generated feedback on test reports leads to better revisions, improved first submissions on new tasks, and partial transfer of reporting practices across three real-world applications.
arXiv preprint arXiv:2411.07407 , year=
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
Survey of 112 agentic AI for social good papers reveals moral-geographic asymmetry with 73% lacking geographic context (lowest for SDG 16) and only 25% reporting deployments.
Introduces the TCR framework to evaluate educational LLM assistants on transparency, consistency, and refinement in multi-turn interactions, complementing aggregate metrics.
citing papers explorer
-
More than a Judge: An Empirical Study of Agent-Human Interaction in Crowdsourced Testing Assessment
A four-stage human-subject study finds that agent-generated feedback on test reports leads to better revisions, improved first submissions on new tasks, and partial transfer of reporting practices across three real-world applications.
-
Whose Good, Whose Place? The Moral Geography of Agentic AI for Social Good
Survey of 112 agentic AI for social good papers reveals moral-geographic asymmetry with 73% lacking geographic context (lowest for SDG 16) and only 25% reporting deployments.
-
Evaluating Multi-turn Human-AI Interaction
Introduces the TCR framework to evaluate educational LLM assistants on transparency, consistency, and refinement in multi-turn interactions, complementing aggregate metrics.