Pith. sign in

REVIEW 4 major objections 5 minor 27 references

The role of joint utility and pragmatic reasoning in cooperative communication

T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read People choose cooperative signals by joint utility, not by the receiver's cost alone.

desk verdict A genuinely interesting behavioral result, but the joint-utility conclusion is not supported by the design; the same data can be explained by a cost-sensitive receiver-centered pragmatic inference. read the letter →

arxiv 2504.21224 v1 pith:TJCX7HHM submitted 2025-04-29 stat.AP

classification stat.AP
keywords cooperativecommunicationjointutilitypragmaticreasoningRationalSpeechActsreferencegamesignalingbehaviorgridworldtasksharedagency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tests whether human speakers in a cooperative task choose ambiguous signals by reasoning about joint utility, meaning what is best for the team, rather than only the receiver's individual utility. In a grid-world game where a signaler describes a target with a single feature, moving a barrier so that one candidate item becomes costly for the receiver made participants more likely to send the optimal signal, at 91.4% versus 77.1%. The authors argue that because the barrier does not change the receiver's own path to the target, an individual-utility account predicts no such difference. They also compare human behavior with an established computational pragmatics model extended with individual utility, and that model predicts the opposite pattern. If the interpretation is right, cooperative communication can be powered by shared-agency utility reasoning rather than by increasingly deep pragmatic inference alone.

What carries the argument

The load-bearing mechanism is a two-step decision process: first, joint utility filters the candidate items to those the receiver can reasonably reach, excluding the critical item behind the barrier; second, within that filtered set a single feature uniquely identifies the target. The comparison machinery is a soft-max expected-utility speaker with rationality coefficient $\lambda = 4$, whose listener distribution is a literal, non-pragmatic interpretation of the signal. The direction of the disagreement between this model and human data is what carries the argument: the model's equal weighting of all literal referents makes the optimal signal look bad when it can refer to a costly item, whereas the joint-utility filter makes it look good.

What would settle it

Running a version of the Simple trials in which the barrier moves near the receiver but the set of candidate items near the receiver is held fixed would settle it: if the optimal-signal advantage disappears, the effect is a change in the referent set, not joint utility.

Watch

Extended reading notes

Core claim

The central claim is that, before interpreting a signal, cooperative speakers perform a joint-utility step: they restrict the set of candidate referents to the items that are the receiver's responsibility, and only then does a single feature become unambiguous. The behavioral signature is that in Simple trials, where a barrier near the receiver makes the critical competitor expensive, participants send the optimal feature signal 91.4% of the time versus 77.1% when the barrier is near the signaler, with faster reaction times. A Rational Speech Acts speaker extended with individual utility predicts the reverse pattern, because it weights all literally true referents equally and therefore avoids a signal that could refer to the costly item. The paper concludes that human behavior supports joint-utility reasoning, while Difficult trials, which require both pragmatics and joint utility, did not show a barrier effect, leaving the integration of the two mechanisms unresolved.

Load-bearing premise

The conclusion rests on the claim that moving the barrier does not change the receiver's own cost of reaching the target; if the barrier instead changes which items are near or relevant to the receiver, then a purely receiver-centered pragmatic account could produce the same 91.4% versus 77.1% pattern without assuming joint utility.

Editorial extensions

If this is right

  • If humans genuinely use a joint-utility filter first, then ambiguity in cooperative communication can be reduced before any pragmatic inference starts, so deep recursive reasoning is not always required.
  • A utility-extended Rational Speech Acts model, as implemented here, is incomplete for cooperative signaling: it predicts that speakers should avoid the optimal signal when a barrier raises a referent's cost, opposite to human behavior.
  • Faster reaction times with the barrier near the receiver suggest that joint-utility reasoning makes the signaling problem easier, not harder.
  • Because the Difficult trials showed no significant barrier effect, the conditions under which joint utility and pragmatic reasoning combine remain unresolved; the paper suggests that scenes with too many alternatives may exceed the pragmatic process.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: The argument could be sharpened by building a joint-utility-first speaker that applies a pragmatic listener only to the filtered referent set; such a model should reproduce the Simple-trial effect and may also explain the Difficult-trial null result.
  • My inference: Treating the receiver's path cost as a continuous variable, rather than a binary near/far distinction, would let the same design measure how finely people tune their signals to the receiver's effort.
  • My inference: The lack of feedback after experimental trials removes learning and reputation effects, so the observed joint-utility behavior is likely a default cooperative stance; whether it persists when the receiver's competence is in doubt is a testable extension.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports a behavioral experiment (N=28, 21 after exclusions) in a grid-world cooperative communication game. Participants (signalers) chose to either navigate to a target or send a single feature-based signal to a simulated receiver, under conditions that varied barrier location (near receiver vs. signaler) and trial difficulty (Simple vs. Difficult). The main finding is that in Simple trials participants selected the 'optimal' signal more often when the barrier was near the receiver (91.4% vs. 77.1%, p=4.45e-3); reaction times were also faster. The authors compare human behavior with RSA simulations that incorporate a literal listener and an action-utility-aware speaker (λ=4) and find the RSA speaker shows the opposite pattern. The paper concludes that humans use joint utility reasoning, and that RSA's individual-utility extension is insufficient.

Significance. If the conclusion were established, the paper would contribute a behavioral demonstration that human cooperative signaling can go beyond RSA-style pragmatic inference, in a physically grounded task. The within-subject design and the use of a shared grid-world environment are strengths, and the paper is candid that its proposed joint-utility model is not implemented. However, the central theoretical claim is not yet supported because the reported effect is also predictable from a receiver-centered, cost-sensitive pragmatic account, and the RSA baseline does not incorporate such cost sensitivity at the inference stage. The manuscript also lacks a fitted joint-utility model and a principled treatment of the λ parameter. With additional modeling and analysis, the study could become a valuable empirical contribution; in its current form the evidence is consistent with joint utility but does not uniquely implicate it.

major comments (4)
  1. [Section 4] The inference in Section 4 that 'Since the receiver's path to the target was not affected by the barrier, the utility for the receiver to reach the target remained unchanged. Therefore, if human participants only considered the individual utility of the receiver, we would not observe this difference' is not secure. Although the target's reachability cost is unchanged, the barrier also changes the receiver's cost to reach the critical distractor, and hence alters which items are cheap literal referents of a feature signal. A pragmatic listener that uses a naïve utility calculus based on individual receiver costs would update its interpretation of 'circle' differently across barrier locations, making the target more probable when the distractor is behind the barrier. The observed Simple-trial difference (91.4% vs. 77.1%) is therefore compatible with an individual-utility mechanism, and the Section 4 claim that only joint utility can explain it is unsupported.
  2. [Section 3, Eq. (1)-(3)] The rationality parameter λ is selected post hoc by grid search ('We tried values from 1 to 10 in increments of 1 and used λ = 4'), not estimated from data or justified by a preregistered criterion. Because λ controls the speaker's sensitivity to utility differences, the qualitative contrast between human and RSA behavior may depend on this choice. Please report a sensitivity analysis across λ and, if possible, fit λ to the human data (e.g., by maximum likelihood) rather than selecting it to make the simulation behave a certain way.
  3. [Section 3, statistical tests] The reported p-values for percentages (e.g., p = 4.45×10−3) appear to be computed at the trial level, ignoring participant clustering; with 21 participants and repeated trials, this can inflate significance. Report mixed-effects logistic regressions with random intercepts for participants (and items), and provide test statistics. In addition, the multiple comparisons across Simple/Difficult, barrier location, and reaction time are not corrected; a correction or a pre-specified analysis plan would strengthen the claims.
  4. [Section 4 and Abstract] The Abstract and Discussion state that the 'results provide support for a joint utility reasoning mechanism,' but the paper does not implement or fit a joint-utility model; it explicitly defers this to future work ('Our future research aims to build a model based on joint utility calculus'). Without a quantitative joint-utility model that makes distinct predictions from an individual-utility pragmatic model, the behavioral data can at most be said to be consistent with joint utility reasoning, not to provide support for it. Either reframe the conclusions to this more modest claim or add a model comparison.
minor comments (5)
  1. [Section 2, Appendix A.4] The manuscript reports 650 trials from participants; with 28 participants × 36 trials = 1008, it is unclear how many trials remain after exclusions and removal of '56 paired trials.' Clarify the accounting and define what a 'paired trial' is.
  2. [Figures 2 and 3] The captions for Figure 2 and Figure 3 should define what error bars (if any) represent and state the statistical test used; currently they only describe the plotted variables.
  3. [Section 4] The sentence '...weighted equally regardless of the cost to reach of them' contains a typo: 'reach of them' should be 'reaching them.'
  4. [Section 3, reaction time] The paper does not state a hypothesis or effect-size expectation for the reaction-time measure; please specify what a shorter reaction time is predicted to indicate.
  5. [Throughout] The manuscript should state explicitly whether the experiment and analysis were preregistered; if not, this should be acknowledged.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the behavioral inference is an empirical claim, and the RSA contrast model is not fitted to the human data.

full rationale

The paper's central inference does not reduce to its own inputs. The 'optimal signal' is defined by a joint-utility account, but participants were free to choose other signals, so observing that they preferentially choose it is a testable empirical outcome rather than a definitional equivalence. The RSA simulation is presented as a comparison model, not as a fitted predictor of human behavior; the rationality parameter lambda=4 is calibrated by grid search, but the model's simulated barrier effect is opposite to the human effect, so it cannot be a post-hoc fit to the target result. The Section 4 assertion that an individual-utility receiver account would predict no barrier effect is an empirical/modeling claim. A cost-sensitive individual-utility listener could in principle reproduce the effect, but this is an alternative-explanation or confound concern, not circularity under the requested taxonomy: no equation is defined in terms of the conclusion, and no self-citation is load-bearing. The cited frameworks (RSA, naive utility calculus, joint action) are external to the authors, and no uniqueness theorem or prior author result is invoked to forbid alternatives. The paper even states that building a joint-utility model is future work, which undercuts any suggestion that the conclusion is derived from a self-constructed model. Therefore, the manuscript exhibits no significant circularity.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The model comparison rests on a small number of hand-set design choices; the most consequential is lambda=4. No new theoretical entities are introduced.

free parameters (1)
  • Rationality parameter lambda = 4
    In Section 3, the authors state 'We tried values from 1 to 10 in increments of 1 and used lambda = 4,' with no stated selection criterion. The RSA simulation results depend on this value.
assumptions (3)
  • domain assumption The receiver's goal inference follows Bayes rule with a literal signaler and a uniform prior over goals.
    Assumed in the RSA simulation (Section 3: 'PR(goal|signal) proportional to PlitS(signal|goal)P(goal)'). Not tested against other priors.
  • domain assumption Utility is modeled as additive step costs and rewards shared between agents.
    The RSA extension assigns utilities based on walking costs and reward (Section 3, equations 1-3). This is a modeling choice not independently validated.
  • domain assumption Participants believed the receiver was intelligent and motivated to help.
    The procedure tells participants the receiver is 'intelligent and motivated to help' (Section 2), but the receiver was pre-programmed. The validity of the cooperative communication measure depends on this belief.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The role of joint utility and pragmatic reasoning in cooperative communication." pith.science (2026). https://pith.science/paper/TJCX7HHM

@misc{pith2026250421224,
  author       = {Pith},
  title        = {Pith review of: The role of joint utility and pragmatic reasoning in cooperative communication},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TJCX7HHM}},
  note         = {Machine review of arXiv:2504.21224}
}
read the original abstract

Humans are able to communicate in sophisticated ways with only sparse signals, especially when cooperating. Two parallel theoretical perspectives on cooperative communication emphasize pragmatic reasoning and joint utility mechanisms to help solve ambiguity. For the current study, we collected behavioral data which tested how humans select ambiguous signals in a cooperative grid world task. The results provide support for a joint utility reasoning mechanism. We then compared human strategies to predictions from Rational Speech Acts (RSA), an established model of language pragmatics.

Figures

Figures reproduced from arXiv: 2504.21224 by the authors.

Figure 1
Figure 1. Paired example trials. the Difficult trials, either feature of the target unambiguously referred to the target using pragmatic reasoning. When the barrier shifted towards R, the critical item became farther from R. In this case, one feature of the target was less ambiguous because it had less literal alternatives among the items that were still near R. In [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Signaling behavior. "Optimal Feature" was the optimal signal when the barrier was near R. "Suboptimal Feature" was the other feature of the target. "Do" was when the signaler walked by itself. "Irrational" was when the signal was not a feature of the target. Furthermore, the reaction time was significantly shorter when the barrier was near R both in the Simple trials (µR = 4.85, µS = 7.01, p = 5.63 × 10−7 ) and in t… view at source ↗
Figure 3
Figure 3. Reaction time. Measured from participants seeing the trial to making a decision. 4 Discussion This task provides evidence that humans do consider joint utility as a heuristic for how to disambiguate signals. In the grid world environment, when the barrier is near the receiver, joint utility reasoning allows agents to exclude some items from consideration. Participants were sensitive to this change as evidenced by th… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

27 extracted references · 27 canonical work pages

  1. [1]

    L., Saxe, R., and Tenenbaum, J

    Baker, C. L., Saxe, R., and Tenenbaum, J. B. (2009). Action understanding as inverse planning. Cognition, 113(3):329–349

  2. [2]

    Chun, M. M. and Jiang, Y . (1998). Contextual cueing: Implicit learning and memory of visual context guides spatial attention. Cognitive psychology, 36(1):28–71

  3. [3]

    Clark, H. H. (1996). Using language. Cambridge university press

  4. [4]

    and Franke, M

    Degen, J. and Franke, M. (2012). Optimal reasoning about referential expressions. Proceedings of SemDIAL, pages 2–11

  5. [5]

    E., Warneken, F., and Tomasello, M

    Fletcher, G. E., Warneken, F., and Tomasello, M. (2012). Differences in cognitive processes underlying the collaborative activities of children and chimpanzees. Cognitive Development, 27(2):136–153

  6. [6]

    Frank, M. C. and Goodman, N. D. (2012). Predicting pragmatic reasoning in language games. Science, 336(6084):998–998

  7. [7]

    Gilbert, M. (2013). Joint commitment: How we make the social world . Oxford University Press

  8. [8]

    Goodman, N. D. and Frank, M. C. (2016). Pragmatic language interpretation as probabilistic inference. Trends in Cognitive Sciences, 20(11):818–829

Show all 27 references
  1. [9]

    Grice, H. P. (1975). Logic and conversation. In Speech acts, pages 41–58. Brill

  2. [10]

    Grosse, G., Moll, H., and Tomasello, M. (2010). 21-month-olds understand the cooperative logic of requests. Journal of Pragmatics, 42(12):3377–3383

  3. [11]

    E., and Tenenbaum, J

    Jara-Ettinger, J., Gweon, H., Schulz, L. E., and Tenenbaum, J. B. (2016). The naïve utility calculus: Computational principles underlying commonsense psychology. Trends in cognitive sciences, 20(8):589–604. 5

  4. [12]

    Kao, J., Bergen, L., and Goodman, N. (2014). Formalizing the pragmatics of metaphor understanding. In Proceedings of the 36th Annual Meeting of the Cognitive Science Society , volume 36

  5. [13]

    K., Austerweil, J

    Kleiman-Weiner, M., Ho, M. K., Austerweil, J. L., Littman, M. L., and Tenenbaum, J. B. (2016). Coordinate to cooperate or compete: abstract goals and joint intentions in social interaction. In Proceedings of the 38th Annual Conference of the Cognitive Science Society

  6. [14]

    D., Tenenbaum, J

    Liu, S., Ullman, T. D., Tenenbaum, J. B., and Spelke, E. S. (2017). Ten-month-old infants infer the value of goals from the costs of actions. Science, 358(6366):1038–1041

  7. [15]

    and Potts, C

    Monroe, W. and Potts, C. (2015). Learning in the rational speech acts model. CoRR

  8. [16]

    Nagel, T. (1989). The view from nowhere. oxford university press

  9. [17]

    and Franke, M

    Qing, C. and Franke, M. (2015). Variations on a bayesian theme: Comparing bayesian models of referential reasoning. In Bayesian natural language semantics and pragmatics , pages 201–220. Springer

  10. [18]

    Sacks, H. (1985). The inference-making machine: Notes on observability. Handbook of discourse analysis, 3:13–23

  11. [19]

    Sugden, R. (1993). Thinking as a team: Towards an explanation of nonselfish behavior. Social philosophy and policy, 10(1):69–89

  12. [20]

    Tomasello, M. (2010). Origins of human communication . MIT press

  13. [21]

    Tomasello, M. (2014). A natural history of human thinking . Harvard University Press

  14. [22]

    Török, G., Pomiechowska, B., Csibra, G., and Sebanz, N. (2019). Rationality in joint action: maximizing coefficiency in coordination. Psychological science, 30(6):930–941

  15. [23]

    V ogel, A., Potts, C., and Jurafsky, D. (2013). Implicatures and nested beliefs in approxi- mate decentralized-pomdps. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (V olume 2: Short Papers), pages 74–80

  16. [24]

    Vygotsky, L. S. (2012). Thought and language. MIT press

  17. [25]

    Warneken, F., Chen, F., and Tomasello, M. (2006). Cooperative activities in young children and chimpanzees. Child development, 77(3):640–663

  18. [26]

    and Tomasello, M

    Warneken, F. and Tomasello, M. (2006). Altruistic helping in human infants and young chimpanzees. science, 311(5765):1301–1303

  19. [27]

    i won’t lie, it wasn’t amazing

    Yoon, E. J., Tessler, M. H., Goodman, N. D., and Frank, M. C. (2017). "i won’t lie, it wasn’t amazing": Modeling polite indirect speech. In Proceedings of the 39th Annual Meeting of the Cognitive Science Society. A Appendix A.1 Participants Twenty-eight University of Californi...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.