Pith. sign in

REVIEW 6 cited by

Generative Emergent Communication: Large Language Model is a Collective World Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.00226 v2 pith:IXJWLJA7 submitted 2024-12-31 cs.AI cs.CL

Generative Emergent Communication: Large Language Model is a Collective World Model

classification cs.AI cs.CL
keywords collectivelanguageworldmodelgenerativeprocessframeworkllms
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large Language Models (LLMs) have demonstrated a remarkable ability to capture extensive world knowledge, yet how this is achieved without direct sensorimotor experience remains a fundamental puzzle. This study proposes a novel theoretical solution by introducing the Collective World Model hypothesis. We argue that an LLM does not learn a world model from scratch; instead, it learns a statistical approximation of a collective world model that is already implicitly encoded in human language through a society-wide process of embodied, interactive sense-making. To formalize this process, we introduce generative emergent communication (Generative EmCom), a framework built on the Collective Predictive Coding (CPC). This framework models the emergence of language as a process of decentralized Bayesian inference over the internal states of multiple agents. We argue that this process effectively creates an encoder-decoder structure at a societal scale: human society collectively encodes its grounded, internal representations into language, and an LLM subsequently decodes these symbols to reconstruct a latent space that mirrors the structure of the original collective representations. This perspective provides a principled, mathematical explanation for how LLMs acquire their capabilities. The main contributions of this paper are: 1) the formalization of the Generative EmCom framework, clarifying its connection to world models and multi-agent reinforcement learning, and 2) its application to interpret LLMs, explaining phenomena such as distributional semantics as a natural consequence of representation reconstruction. This work provides a unified theory that bridges individual cognitive development, collective language evolution, and the foundations of large-scale AI.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Specification-Guided Synthesis of Deadlock-Free Communication Protocol Refinements with Large Language Models

    cs.SE 2026-07 conditional novelty 6.5

    Syntropy synthesises asynchronous multiparty session-type subtypes with 95.6–99.5% checker-accepted validity via LoRA fine-tuning and two-level constrained decoding.

  2. Emergent Communication between Heterogeneous Visual Agents through Decentralized Learning

    cs.CV 2026-05 unverdicted novelty 6.0

    Heterogeneous visual agents form shared symbols via decentralized Metropolis-Hastings captioning, where encoder similarity shapes the content and symmetry of the resulting language.

  3. Decentralized Collective World Model for Emergent Communication and Coordination

    cs.MA 2025-04 unverdicted novelty 6.0

    A decentralized collective world model integrates predictive coding with bidirectional communication to achieve simultaneous symbol emergence and coordination, outperforming non-communicative baselines in a two-agent ...

  4. On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective

    cs.AI 2026-05 unverdicted novelty 5.0

    Post-training reweights a pretrained model's behavior distribution either within its existing accessible support (elicitation) or by expanding that support (creation), with both SFT and RL acting as free-energy minimi...

  5. SANEmerg: An Emergent Communication Framework for Semantic-aware Agentic AI Networking

    cs.AI 2026-05 unverdicted novelty 5.0

    SANEmerg enables emergent communication among bounded-intelligence AI agents for semantic-aware task fulfillment in AgentNet systems via a bandwidth-adaptable importance filter and MDL-based complexity regularizer.

  6. Interoceptive Divergence in Aesthetic Evaluation and Implications for Human-AI Alignment

    cs.HC 2026-04 unverdicted novelty 5.0

    LLMs approximate human patterns in beauty-emotion links and image feature priorities but diverge in emotional response distributions and especially in beauty-bodily sensation relationships.