{"id":"30e290db-cfe5-461f-9f0d-afc939633b8c","arxiv_id":"2505.03798","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A position paper proposes replacing token-based representations in foundation models with outcome-driven digital twin representations that explicitly encode physical and semantic structure.","lead":"This position paper argues that foundation models should use digital twin representations instead of discrete tokens. The authors say this would let models preserve physical structure and support causal reasoning, not just statistical pattern matching.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's core impossibility premise—that token limitations cannot be overcome by scaling—is asserted, not demonstrated; the cited token-count scaling law does not test model/data scaling and controls neither architecture nor training objective.","rationale":"The reader's weakest assumption—that token limitations are not shown to be inherent to the representation rather than to training objectives, data scale, or architecture—is precisely the load-bearing concern. The paper's evidence for impossibility is a single token-count scaling law, which is not the same as model/data scaling and does not include controlled ablations. I also checked the paper's own Alternative Views section, which explicitly concedes that token representations may overcome limitations through architectural and training improvements, further undermining the asserted impossibility. Given that this is a position paper and the DT advantages are framed as assumptions, the appropriate verdict is CONDITIONAL, not REJECT: the agenda is reasonable, but the central claim is not established. The reader's conditional verdict already captures this, so no verdict adjustment is needed.","tokens_in":14947,"tokens_out":1893,"duration_ms":21453,"concrete_test":"Replicate the scaling experiment behind [49] while separately varying model parameters and training data at fixed token count: train a token-based vision-language model at 0.5B, 2B, and 8B parameters on the same data and measure performance on fine-grained spatio-temporal reasoning and causal counterfactual tasks. If performance improves substantially with scale at fixed token count, the 'cannot be overcome by scaling' premise is falsified; if it plateaus across all scales, the premise gains support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central motivating claim is that token representations have inherent limitations that 'cannot be overcome by simply scaling up model size or expanding datasets' (Abstract, Section 1). This claim is load-bearing because, without it, the case for replacing tokens with DT representations reduces to a preference rather than a necessity. The only supporting evidence is in Section 4.1, which cites [49] for a weak scaling law S(N) ≈ (c/N)^alpha, where N is the number of vision tokens. That evidence addresses scaling in token count, not scaling in model parameters or dataset size. More importantly, it is a single empirical trend from a specific model family; it does not control for architecture choices, training objectives, or data quality. The paper's own Alternative Views section acknowledges that token models might overcome these limitations through better architectures and training methods. If the impossibility claim fails, the paper's central motivation for the DT paradigm collapses, even though DT representations might still be a useful research direction. The five listed advantages of DT representations are explicitly labeled assumptions, but the impossibility premise is asserted as fact without a controlled experiment or formal argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that token representations in foundation models have inherent limitations that cannot be overcome by scaling, and proposes digital twin (DT) representations as an alternative. It defines DT representations, sketches a formalization, and lists five assumptions about how DT representations would improve knowledge encoding, data synthesis, causal reasoning, semantic unification, and interpretability. The paper is an explicit position paper, with the five advantages labeled as assumptions and an 'Alternative Views' section acknowledging counterarguments.","tokens_in":15152,"tokens_out":4708,"duration_ms":45138,"significance":"If the central claim were established, the paper would motivate a significant reorientation of multimodal foundation model research. The paper is honest in labeling its five advantages as assumptions, and it provides concrete examples from robotics, surgery, and autonomous driving. However, the key premise that token limitations are insurmountable by scaling is asserted rather than demonstrated, and the formalization is too loose to carry the argument. As a position paper, it has the merit of opening a debate, but it currently overstates its case.","major_comments":[{"comment":"The central claim that token limitations 'cannot be overcome by simply scaling up model size or expanding datasets' is load-bearing but unsupported. The only cited evidence is the token-count scaling law S(N) ≈ (c/N)^α in §4.1, which concerns the number of vision tokens N, not model parameters or dataset size. This does not test the scaling claim, and no experiment controls for architecture or training objective. The Alternative Views section even concedes that token models might overcome these limitations through better architectures and training methods. The paper should either provide direct evidence against model/data scaling or reframe the claim as a hypothesis.","section":"Abstract; §1; §4.1"},{"comment":"There is a circularity between the definition of DT representations and the claimed advantages. Definition 3.2 defines DT representations as 'outcome-driven digital representations' that capture task-specific properties such as geometric and physical constraints, and Assumptions 1 and 3 then attribute to DT representations the ability to encode exactly these physical and causal properties. The advantages are thus partly built into the definition. The paper should define DT representations in a way that does not presuppose the benefits, or the assumptions should be stated independently of the definition.","section":"§3.2 and §4.1/§4.3"},{"comment":"The attempted formalization is not actually used. The equation R = Φ(P, S) = Φe(S, Φd(P)) is never employed to derive any of the five assumptions, and the functions Φd and Φe are not specified with any formal semantics. This makes the formalization decorative rather than load-bearing. Either the formalization should be developed into a framework with stated properties, or it should be removed and the position stated directly.","section":"§3.2"}],"minor_comments":[{"comment":"Typo: 'Charecter-level' should be 'Character-level'; also 'WorkPiece' should be 'WordPiece'.","section":"§2, 'Token Representation for Text'"},{"comment":"The figure caption contains 'Phyical Process' instead of 'Physical Process', and §3.1 refers to 'Fig. 3.1' although the figure is numbered Figure 1.","section":"Figure 1 and §3.1"},{"comment":"The variables m and n are not defined; the paper should specify that m and n are the numbers of representations and raw data modalities, respectively.","section":"§3.2"},{"comment":"The paper alternates between 'FMs' and 'LLMs' without a clear distinction; both terms should be defined at first use and used consistently.","section":"Throughout"},{"comment":"The conclusion lists future research directions but gives no concrete evaluation benchmarks or success criteria for DT representations, which would strengthen the argument as a position paper.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"This is a position paper, and the criteria for acceptance should account for genre. The paper is honest and well-structured, but the central impossibility claim needs to be either supported or softened. Given that the authors explicitly concede counterarguments, a revise-and-resubmit is appropriate. I would not reject, since the topic is timely and the paper could be repairable with a more modest framing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid position paper, not a research result. It argues that foundation models should be built on digital twin (DT) representations rather than tokens, and it does the genre justice: the writing is clear, the advantages are explicitly labeled as assumptions, and the Alternative Views section engages real counterarguments. If you want a citable statement of the DT-for-ML agenda with a decent reference list, this is useful.\n\nWhat is new is confined to the framing: Definitions 3.1 and 3.2 and the attempted formalization in Section 3.2 give the idea a name and a shape, but the underlying proposal — structured, physically grounded, object-centric representations instead of raw tokens — is not new. The paper cites its own earlier work on surgical digital twins and reasoning segmentation, plus related lines like NeRF and neuro-symbolic approaches. That is fine; the contribution is the synthesis and the call to take the paradigm seriously, not the invention of the concept.\n\nThe soft spots are real but proportionate to the genre. The biggest one is the abstract's claim that token limitations 'cannot be overcome by simply scaling up model size or expanding datasets.' That is load-bearing: without it, the paper is arguing a preference, not a necessity. The only evidence offered is a token-count scaling law S(N) ~ (c/N)^alpha from a single paper, which says nothing about scaling model parameters or data, and the paper's own Alternative Views section concedes that architectural and training improvements might overcome the limitations. So the central premise is asserted, not demonstrated. In a position paper that can be acceptable, but it should be flagged as a conjecture, not a fact.\n\nThe other soft spot is a mild circularity: Definition 3.2 defines DT representations as outcome-driven representations that capture task-specific properties like geometry and physical constraints, and Assumptions 1 and 3 then credit them with encoding exactly those properties. To some degree the advantages are baked into the definition. That doesn't invalidate the research direction, but it means the benefits are partly a matter of stipulation until someone builds a DT-based FM that beats a strong token-based baseline on the same benchmarks.\n\nRecommended treatment: send it to peer review if the venue handles position papers. It will provoke useful discussion. A serious referee should ask for a sharper statement of what would falsify the impossibility claim, and ideally a small proof-of-concept comparison on one controlled task. Not a desk reject.","headline":"A clear, honest position paper whose agenda is reasonable but whose load-bearing claim that token limitations cannot be scaled away is asserted, not shown.","tokens_in":15656,"tokens_out":2379,"would_cite":false,"duration_ms":23063,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that foundation models should abandon token representations for digital twin representations—outcome-driven encodings of physical entities and interactions—because tokenization's limits cannot be fixed by scaling alone.","keywords":["foundation models","token representations","digital twin representations","multimodal learning","physical grounding","causal reasoning","interpretability","representation learning"],"falsifier":"A matched experiment on the same benchmark suite—fine-grained spatio-temporal video reasoning, counterfactual physical prediction, and cross-modal semantic alignment—with a token-based model and a DT-representation model trained on identical data and compute would settle the claim; if the token model performs comparably or better, the assertion that tokenization is the bottleneck is falsified.","tokens_in":14767,"feed_emoji":"🧩","tokens_out":7933,"duration_ms":75477,"temperature":0.7,"pith_summary":"This is a position paper, not an experimental study. It argues that the discrete tokens used by current multimodal foundation models—image patches, video frames, subword units—fragment continuous physical processes and force models to learn world knowledge, causality, and cross-modal semantics purely from statistical correlation. The paper's central claim is that these limitations are inherent to tokenization and cannot be overcome by scaling model size or data, so the field should adopt digital twin (DT) representations: outcome-driven digital encodings of task-specific entities, properties, and interactions, grounded in physical and domain constraints. If the claim is right, foundation models built on DT representations would be more data-efficient, physically consistent, causally reliable, and interpretable. The paper supports the position with five assumptions and existing examples, but it does not run a head-to-head comparison between token and DT representations.","feed_headline":"Replace tokens with digital twin representations in foundation models","feed_subtitle":"A position paper argues discrete tokens cap what models can learn from the physical world.","key_machinery":"The carrying mechanism is the digital twin representation, formalized as a set $R=\\{r_0,\\dots,r_m\\}$ of outcome-driven digital representations derived from raw sensor data $S=\\{s_0,\\dots,s_n\\}$ through a pipeline $R=\\Phi(P,S)=\\Phi_e(S,\\Phi_d(P))$, where $\\Phi_d$ is outcome-driven representation design and $\\Phi_e$ is representation extraction; the foundation model $\\Theta$ then produces target outputs $T=\\Theta(R)$. This formalization does the argument's work by contrasting explicit, physically grounded construction with token representations' 'direct vectorization' of raw data into uniform patches, frames, or subwords. The paper also leans on a cited scaling law $S(N)\\approx (c/N)^\\alpha$ for vision-token counts to argue that adding more tokens cannot recover fine-grained world knowledge.","core_discovery":"The central claim is that token representations have inherent limitations that scaling cannot repair, and that digital twin representations offer a paradigm shift. A DT representation is defined as an outcome-driven digital representation extracted from raw data that serves as a building block of a digital twin, capturing task-specific entities, their geometric and physical properties, and their interactions under domain constraints. The paper proposes that foundation models should consume DT representations rather than tokens, which would preserve the continuous nature of physical processes, encode explicit causal mechanisms and physical laws, provide a unified semantic space across modalities, and make reasoning chains transparent. It organizes this into five assumptions—better real-world knowledge encoding, practical data synthesis, causal reasoning, unified semantics, and interpretability—each illustrated with existing systems, but the paper is explicit that these are assumptions forming a position rather than demonstrated results.","pith_inferences":["Editorial inference: the strongest test of the paper's position is a hybrid architecture that keeps tokens but injects structured physical or causal constraints, because if that matches DT representations, the bottleneck is the objective rather than the representation.","Editorial inference: DT representations are defined as outcome-driven and task-specific, which may tension with foundation-model generality; a key open question is whether one DT representation can serve many downstream tasks without being redesigned per task.","Editorial inference: the paper's examples mostly encode the answer into the representation (e.g., object masks plus geometry), so a fair benchmark must separate the value of the representation from the value of the task-specific information already contained in it.","Editorial inference: a concrete next step is to measure whether DT representations reduce the data needed to reach a given level of physical or causal reasoning, since the paper claims data efficiency but offers no quantitative comparison."],"forward_implications":["If the claim is right, multimodal foundation models could be built to use physical laws and domain constraints directly, so they would not need to rediscover geometry, kinematics, or conservation laws from data.","DT-based synthetic data generation could produce physically valid long-tail scenarios—like adverse weather driving conditions—and close part of the sim-to-real gap for robotics and autonomous driving.","Encoding causal structure into representations could improve interventional, attributional, and counterfactual reasoning, reducing hallucination and spurious correlations.","A unified semantic space grounded in shared physical and geometric properties could make cross-modal alignment more robust than aligning visual patches with text tokens.","Physically grounded internal states would make model decisions traceable, improving verification and safety in medical, robotic, and industrial applications."],"supporting_citations":[{"why":"Supplies the scaling law for vision-token counts that the paper uses to argue token scaling cannot fix world-knowledge limitations.","marker":"[49]"},{"why":"Provides the review of digital twin definitions that the paper builds on to define DT representations as outcome-driven and cyclic.","marker":"[28]"},{"why":"Example of a digital twin representation encoding task-specific status for an LLM-based agent, supporting the knowledge-encoding assumption.","marker":"[23]"},{"why":"Example of unified digital twin scene representation improving out-of-distribution generalization in surgical phase recognition.","marker":"[24]"},{"why":"Example of just-in-time digital twin representations preserving semantic, spatial, and temporal relationships in video reasoning.","marker":"[72]"},{"why":"Provides the digital cousins concept used as evidence that DT representations enable sim-to-real transfer and counterfactual reasoning.","marker":"[18]"},{"why":"Provides the TWICE dataset as an example of physics-aware synthetic data generated with digital twin representations.","marker":"[62]"},{"why":"Shows neural scene representations as a precedent for a unified continuous 3D semantic space across views and modalities.","marker":"[59]"},{"why":"Documents that large language models perform near chance on pure causal inference, grounding the paper's causal-reasoning limitation.","marker":"[36]"}],"fun_headline_variants":["Swap tokens for digital twins in foundation models","Digital twin representations: the next AI paradigm?","Why tokens fail: digital twins for causal AI","Rethink the token: digital twin representations","Foundation models need digital twin reps, argue researchers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes the failures it blames on tokens are caused by the representation itself and would not disappear with better training objectives, larger data, or new architectures; if any token-based model overcomes them, the central motivation for switching to digital twins collapses.","fun_headline_variants_meta":{"raw":{"variants":["Swap tokens for digital twins in foundation models","Digital twin representations: the next AI paradigm?","Why tokens fail: digital twins for causal AI","Rethink the token: digital twin representations","Foundation models need digital twin reps, argue researchers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000281,"raw_usage":{"total_tokens":1605,"prompt_tokens":828,"completion_tokens":777,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":444,"completion_tokens_details":{"reasoning_tokens":707}},"tokens_in":444,"tokens_out":777,"duration_ms":8519,"temperature":1.0,"reasoning_tokens":707,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:32:29.646094+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A matched experiment on the same benchmark suite—fine-grained spatio-temporal video reasoning, counterfactual physical prediction, and cross-modal semantic alignment—with a token-based model and a DT-representation model trained on identical data and compute would settle the claim; if the token model performs comparably or better, the assertion that tokenization is the bottleneck is falsified.","supporting_citations":[{"cited_title":"Scaling capability in token space: An analysis of large vision language model","cited_arxiv_id":null,"evidence_quote":"Supplies the scaling law for vision-token counts that the paper uses to argue token scaling cannot fix world-knowledge limitations."},{"cited_title":"TWICE Dataset: Digital Twin of Test Scenarios in a Controlled Environment","cited_arxiv_id":"2310.03895","evidence_quote":"Provides the TWICE dataset as an example of physics-aware synthetic data generated with digital twin representations."}],"review_version":1}