{"id":"42c3762b-1547-434a-9e18-23f8d07d4fc9","arxiv_id":"2608.10886","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A 4D scene graph augmented with atomic human-object interactions and goal-driven events improves retrospective question answering about human activities in dynamic scenes.","lead":"GESTO builds a robot memory that records not just what objects were present, but what people did with them and what larger activities those actions formed. The paper shows that this activity-aware memory improves a robot's ability to answer retrospective questions about past human behavior.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline comparison rests on an unreported LLM-judged semantic score over small per-category query counts; without per-query scores and variance, the 0.01 text lead and the 0.08 time gap may be within the noise floor.","rationale":"The reader's weakest assumption correctly identifies the evaluation protocol as the least secure link in the paper's strongest claim. I considered whether the more serious problem is the fully automatic EGG baseline described in Section V-A(d), which is constructed by exposing EGG's tool-calling agent to GESTO's predicted activity information and may not represent a fair native baseline; that would undermine the 'substantially outperforming' sub-claim. However, the primary claim as stated is the competitive claim against the ground-truth-grounded EGG(auto-caption), and that claim depends directly on the semantic score and on the small per-category query counts. The paper reports no error bars, no per-query data, no judge identity, and no scoring prompt, so the 0.01 text lead is smaller than the resolution of a single query in a 20-item category. The extended queries are author-authored and scored with the same unreleased LLM metric, adding another layer of unverifiability. This is not an accusation of unfairness; it is a statement that the evidence, as currently presented, cannot distinguish a real advantage from scoring noise or judge bias. The proposed bootstrap re-scoring with two independent judges is a concrete way to settle the question. If the gaps survive, the central claim is strengthened; if they collapse, the conclusion should be weakened to 'comparable within noise.' Since the reader's verdict is already CONDITIONAL and this concern supports that verdict, no adjustment is needed.","tokens_in":12239,"tokens_out":6575,"duration_ms":64295,"concrete_test":"Release per-query scores for all methods on the standard and extended query sets, together with the exact semantic-scoring prompt and judge settings. Then re-score every text answer with two independent LLM judges (e.g., GPT-4o and Claude) at temperature 0 using the same rubric, and compute bootstrap 95% confidence intervals over queries with category sizes reported. If the GESTO-vs-EGG(auto-caption) text difference (0.71 vs 0.70) flips sign under either judge, or if the time gap (0.70 vs 0.78) falls within the bootstrap CI, then the 'approaching ground truth' claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that GESTO's automatically built hierarchy approaches a method given ground-truth event and object grounding—rests on three scores (0.71 text, 0.75 binary, 0.70 time) from the 80-query benchmark of [8]. In Section V-A(b), node queries are excluded because ground-truth point clouds are unavailable, so the reported categories likely contain only about 15–30 queries each. Section V-A(d) defines the text metric only as 'an LLM-based semantic score between 0 and 1' and time correctness as within two minutes, without specifying the judge LLM, the scoring prompt, the temperature, or the per-query scores. A single answer in a 20-question category changes a score by 0.05, so the 0.01 text advantage over EGG(auto-caption) and the 0.08 time deficit are both within the resolution of one or two questions. The 40 extended Space2Event and Event2Space queries were written by the authors, scored with the same LLM semantic metric, and are not released. Because GESTO returns answers 'together with the supporting evidence' (Section IV-D), an LLM judge that rewards verbosity or evidence citation could systematically inflate GESTO's text scores relative to a flat baseline. As reported, the evaluation has an unknown noise floor and an unexamined judge bias, so the headline 'approaching' and 'substantially outperforming' claims are not yet supported by the evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"GESTO couples a persistent 4D scene graph with a two-level activity hierarchy of atomic human–object interactions and goal-driven events, both grounded to persistent scene entities. The pipeline extracts interactions from RGB-D clips with a vision-language model, groups them into events via an LLM, refines uncertain object links using event context, and exposes the memory through a relation-aware tool-calling agent. The paper evaluates on the text, binary, and time QA categories from the EGG benchmark plus 40 new Space2Event/Event2Space queries, reporting that GESTO achieves 0.71/0.75/0.70 on the standard categories, 0.73/0.75 on the new ones, and that ablations show complementary benefits of video fragmentation, grounding refinement, and the event hierarchy.","tokens_in":12601,"tokens_out":5127,"duration_ms":47718,"significance":"If reliable, the central claim is that an automatically constructed, grounded two-level activity hierarchy is sufficient to bring a robot memory close to methods that receive externally provided event and object grounding — a meaningful step beyond object-centric 4D scene graphs. The paper's strengths include a clear system design, an explicit problem formulation (G+ = (G_S, G_A, L)), an ablation that isolates the event hierarchy and grounding refinement, and honest reporting of benchmark limitations such as the exclusion of node queries and known object-description inconsistencies. The promised release of prompts, queries, and code would support reproducibility. However, the headline comparisons currently rest on small, unreleased query sets and an underspecified LLM judge, so the strength of the contribution cannot yet be assessed.","major_comments":[{"comment":"The central comparison rests on aggregate scores over small per-category query counts with no error bars, per-query scores, or significance tests. Because node queries are excluded from the 80-query benchmark, the text, binary, and time categories likely contain only about 15–30 queries each; one answer changes a score by roughly 0.03–0.07, which covers the 0.01 text advantage over EGG(auto-caption) and part of the 0.08 time gap. Please report category sizes, per-query results, and confidence intervals, and temper the comparative claims accordingly.","section":"V-A(b), Table I"},{"comment":"The text metric is defined only as an LLM-based semantic score between 0 and 1 without specifying the judge model, prompt, temperature, or aggregation rule. Since GESTO returns answers together with supporting evidence (Section IV-D), an LLM judge that rewards verbosity or evidence citation could systematically inflate GESTO's text scores relative to flat baselines. Please specify the judge and scoring prompt, validate against human ratings on a subset, and check for judge bias toward evidence-rich answers.","section":"V-A(d), IV-D"},{"comment":"The 40 extended queries were authored by the authors before running any method, but they are not released, and the ground-truth answers are scored with the same unspecified LLM semantic metric. These queries are used to support the cross-domain reasoning claims (0.73 Space2Event, 0.75 Event2Space). Without releasing the queries and answers and providing human-rated or otherwise validated scores, these numbers cannot be independently reproduced or fairly compared. Please include the query set and human agreement.","section":"V-A(c)"},{"comment":"The EGG (w/o GT input) baseline is not the released EGG system; it is EGG's tool-calling agent exposed to the same automatically constructed scene and activity information used by GESTO. The phrase including its predicted activity information is ambiguous: if it includes GESTO's event structure and refined links, the baseline is not independent; if it includes only raw predicted interactions, the comparison conflates representation structure with tool-calling differences. Please specify precisely which inputs each baseline receives and how the agent was adapted, or run the original EGG with automatic event segmentation.","section":"V-B, Table I"}],"minor_comments":[{"comment":"Binary are evaluated using F1 scores should read Binary answers are evaluated using F1 scores.","section":"V-A(b)"},{"comment":"The definition of an atomic interaction uses o_action in the tuple but later text refers to the interaction object label of o action; please standardize the notation.","section":"III"},{"comment":"The free parameters, including the 15-second clip length, tau_mIoU = 0.3, tau_sem = 0.7, and one refinement round, are fixed without sensitivity analysis; a small sweep would strengthen the claim that the components transfer across settings.","section":"V-A(a)"},{"comment":"The statement that GESTO transfers across sensor configurations without retraining is based on a single qualitative example; please soften the wording or add a small quantitative egocentric evaluation.","section":"V-D"},{"comment":"The availability statement says code, prompts, and queries will be open-sourced upon acceptance; for a reproducibility-focused evaluation, releasing these artifacts at submission time or in a supplementary package would allow reviewers to verify the numbers.","section":"Availability statement"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a computer-vision or robotics venue, and the system is well engineered. My main concern is that the headline comparison is not yet verifiable: the author-written extended queries, the unreleased code, and the underspecified LLM judge make it impossible to separate method quality from evaluation noise. I would encourage the editor to require the release of per-query scores and the judge prompt before final acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"GESTO is a real step forward in representation: a persistent 4D scene graph coupled with a two-level activity hierarchy—atomic interactions and goal-driven events—built fully automatically from RGB-D and grounded to persistent objects and places. No cited prior work does that end-to-end, so the novelty claim holds. The ablations are the strongest part: removing the event hierarchy collapses temporal accuracy (0.70→0.40) and Space2Event (0.73→0.53), and context-aware refinement recovers about 30% of otherwise ungrounded interactions. Those internal comparisons are instructive and internally consistent.\n\nThe soft spot is the evaluation against external scores. The standard benchmark has 80 queries split across three reported categories, so each has roughly 15–30 questions; one or two answers change a score by 0.03–0.05. The 0.01 text lead over EGG(auto-caption) and the 0.08 time deficit are within that resolution. The text metric is described only as an LLM-based semantic score between 0 and 1, with no judge model, scoring prompt, temperature, or per-query scores. The 40 extended queries were authored by the same team and are not released; writing them before running methods is honest, but it is not an independent test. The EGG(w/o GT input) baseline is also nonstandard: it feeds EGG's agent GESTO's automatically constructed scene graph and activity information, so the 'substantially outperforming' claim conflates representation structure with a pipeline you built. The egocentric transfer is only qualitative, which is fine for illustration but adds no quantitative evidence.\n\nI do not think the central claim is circular. No fitted parameters define the answers, and event summaries are inferred from observation, not from benchmark answers. The method description is clear and the limitations section is candid—it flags the inconsistent object descriptions in the benchmark and the missing node-query ground truth, which is more than many papers do.\n\nWho is this for: robotics and scene-graph researchers working on long-horizon memory. The representation idea is worth citing now; the benchmark numbers are not. I would send this to peer review with a requirement to release code, prompts, the 40 queries, and per-query scores with variance. The paper deserves referee time—the idea is new and the ablation design is thoughtful—but the headline comparisons should not be taken at face value until the artifacts are out.","headline":"Genuinely new representation idea with solid ablations, but the evaluation is under-specified and the headline numbers are within the noise floor—worth peer review, not worth believing yet.","tokens_in":13098,"tokens_out":2911,"would_cite":true,"duration_ms":28083,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GESTO claims that a robot can answer retrospective questions about human activities by building its own two-level activity hierarchy over a persistent 4D scene graph, so that it matches methods supplied with human-annotated event and…","keywords":["4D scene graphs","activity hierarchy","human-object interaction","spatio-temporal memory","grounded reasoning","robot memory","question answering","dynamic scenes"],"falsifier":"A decisive check would be re-scoring GESTO and the fully automatic EGG baseline with an independent LLM judge plus a human-rated sample: if the text gap from 0.71 versus 0.33 collapses or reverses, the advantage is an artifact of grading, not of the activity hierarchy.","tokens_in":12070,"feed_emoji":"🤖","tokens_out":8584,"duration_ms":80586,"temperature":0.7,"pith_summary":"GESTO claims that a robot can keep a useful memory of what people did and with which objects, without any human-supplied event boundaries or object correspondences, by constructing the activity structure itself from an RGB-D observation stream. The proposed representation couples a persistent 4D scene graph, a record of objects and places over time, with a two-level activity hierarchy in which timestamped atomic human-object interactions are grouped under goal-driven events such as 'prepares coffee,' and both levels are grounded to persistent scene entities. On the text, binary, and time query categories of the benchmark, GESTO reports scores of 0.71, 0.75, and 0.70, close to a method given ground-truth event and object grounding and well above the same reasoning framework when those inputs are removed, with 0.73 and 0.75 on new Space2Event and Event2Space query sets. If these results hold, a grounded activity hierarchy built automatically would give robots a practical memory for answering retrospective questions about human behavior in dynamic scenes.","feed_headline":"Automatic activity memory rivals hand-annotated grounding in robot Q&A","feed_subtitle":"A two-level event hierarchy built from raw RGB-D answers which objects, places, and times an activity involved.","key_machinery":"The central object is the coupled pair $\\mathcal{G}^+ = (\\mathcal{G}_S, \\mathcal{G}_A, L)$: the spatial layer $\\mathcal{G}_S$ organizes persistent object and place nodes as in existing 4D scene graphs, while the activity layer $\\mathcal{G}_A$ organizes timestamped atomic human-object interactions $I$ under goal-driven events $E$, and the grounding links $L$ attach each interaction to the object node it acts on. This linkage is what lets one query move from an object to the events it participated in, or from an event to the objects and places that support it. The construction pipeline carries the argument: atomic interactions are extracted from short clips by a vision-language model, grounded by mean intersection-over-union between interaction boxes and tracked object boxes plus text-embedding similarity between labels and object descriptions, grouped into events by an LLM, and then refined through one alternating pass $L^{(1)} = R_L(L^{(0)}; E^{(0)})$, $E^{(1)} = R_E(E^{(0)}; L^{(1)})$ in which event context infers links for unlinked interactions and corrected links re-check event membership. A relation-aware tool-calling agent exposes these structures through event, interaction, object, place, and time-interval tools.","core_discovery":"GESTO's central discovery is that the missing ingredient for activity-centric robot memory is not more perception or larger language models, but a structured coupling between persistent spatial records and a two-level activity graph. The representation $\\mathcal{G}^+ = (\\mathcal{G}_S, \\mathcal{G}_A, L)$ pairs a spatio-temporal scene graph $\\mathcal{G}_S$ of objects and places with an activity graph $\\mathcal{G}_A$ in which atomic interactions $I$ are grouped into goal-driven events $E$, connected by grounding links $L$ that attach each interaction to the persistent object it acts on. The pipeline constructs everything automatically: a vision-language model reads short clips to extract timestamped interactions, geometric and semantic matching grounds them to tracked objects, an LLM groups them into coherent events, and one alternating refinement pass uses event context to infer links for ungrounded interactions and re-evaluate event membership. The result, as the paper reports, is performance on text, binary, and time queries that approaches a method provided with externally annotated events and object associations, and ablation evidence that the event level itself, not just the grounded interactions, is what carries temporal and activity-space reasoning.","pith_inferences":["The single alternating refinement pass is a deliberate stopping point; iterating link and event refinement to a fixed point could plausibly improve consistency, so the reported gains may be a lower bound on what the hierarchy can support.","The paper notes the vision-language model is weaker at localizing the handled object than at recognizing the action, an appearance-based limitation; adding affordance or hand-contact cues to the grounding score is a concrete way to test whether the time-accuracy gap to ground-truth-grounded EGG can be closed.","Since memory grows with session length, detecting and consolidating recurrent activities into routines would be the natural next step for long-term deployments, a capability the single-session benchmark cannot assess.","The two-minute tolerance on time answers means 'when' questions are coarse; re-scoring with a stricter tolerance would test whether the hierarchy's temporal precision is real or just enough to pass a generous window."],"forward_implications":["A robot equipped with GESTO could answer queries such as 'Which mug was used to drink coffee this morning?' from its own recordings, with no human annotation of event boundaries or object correspondences.","Removing the event hierarchy drops time accuracy from 0.70 to 0.40 and Space2Event from 0.73 to 0.53, which indicates that the goal-level organization, not the stored interactions alone, is what makes temporal and activity-space queries answerable.","Context-aware grounding refinement recovers interactions that initially lack a geometric match; without it, roughly 30% of interactions stay ungrounded and time accuracy falls from 0.70 to 0.59.","Because the pipeline treats the 4D scene graph as a swappable backbone and needs only a prompt change for egocentric video, the same memory construction can in principle be reused across robot-mounted and wearable camera configurations."],"supporting_citations":[{"why":"Supplies the benchmark protocol, the RGB-D dataset with human activities, and the closest activity-grounded baseline EGG that GESTO compares against under both ground-truth and fully automatic settings.","marker":"[8]"},{"why":"Provides the DAAAM 4D scene graph used as the object- and place-level backbone for GS and as the object-centric baseline that GESTO extends.","marker":"[1]"},{"why":"ReMEmbR is a long-horizon spatio-temporal memory baseline that GESTO outperforms, establishing the comparison for text, binary, and time queries.","marker":"[25]"},{"why":"Cosmos-Reason2 is the vision-language model used to extract atomic interactions, localize manipulated objects, group events, and judge borderline links, so it carries most of the perception load.","marker":"[37]"},{"why":"FastSAM produces the object masks used in the geometric IoU matching that grounds interactions to persistent scene-graph objects.","marker":"[38]"},{"why":"Sentence-T5-large computes the text-embedding similarity between interaction labels and object descriptions, the semantic half of the grounding score.","marker":"[39]"},{"why":"Hydra defines the persistent object/place graph formulation that GS follows, making the spatial backbone modular and swappable.","marker":"[14]"}],"fun_headline_variants":["GESTO auto-builds hierarchical activity memory from raw video","Robot memory rivals manual grounding using auto event hierarchy","Two-level event memory answers what, where, and when in scenes","GESTO: spatio-temporal memory that groups actions into goals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the evaluation protocol—an LLM-based semantic score for text answers plus a two-minute tolerance for time answers—is a fair, stable measure of reasoning ability, because if that judge favors GESTO's verbose evidence-rich responses or is noisy at the current query counts, the reported margins over the fully automatic baseline could shrink or disappear.","fun_headline_variants_meta":{"raw":{"variants":["GESTO auto-builds hierarchical activity memory from raw video","Robot memory rivals manual grounding using auto event hierarchy","Two-level event memory answers what, where, and when in scenes","GESTO: spatio-temporal memory that groups actions into goals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000566,"raw_usage":{"total_tokens":2750,"prompt_tokens":1079,"completion_tokens":1671,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":695,"completion_tokens_details":{"reasoning_tokens":1601}},"tokens_in":695,"tokens_out":1671,"duration_ms":12101,"temperature":1.0,"reasoning_tokens":1601,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:58:08.863216+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check would be re-scoring GESTO and the fully automatic EGG baseline with an independent LLM judge plus a human-rated sample: if the text gap from 0.71 versus 0.33 collapses or reverses, the advantage is an artifact of grading, not of the activity hierarchy.","supporting_citations":[{"cited_title":"Event-grounding graph: Uni- fied spatio-temporal scene graph from robotic observations,","cited_arxiv_id":null,"evidence_quote":"Supplies the benchmark protocol, the RGB-D dataset with human activities, and the closest activity-grounded baseline EGG that GESTO compares against under both ground-truth and fully automatic settings."},{"cited_title":"Cosmos-Reason2-8B: Physical ai common sense and embodied reasoning model,","cited_arxiv_id":null,"evidence_quote":"Cosmos-Reason2 is the vision-language model used to extract atomic interactions, localize manipulated objects, group events, and judge borderline links, so it carries most of the perception load."}],"review_version":1}