{"id":"56990e9e-b38e-461d-82ff-1c9d64ee640b","arxiv_id":"2607.05716","paper_version":3,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Scene-graph-aligned SFT plus node-as-proxy GRPO rewards let small MLLMs outperform larger baselines on fine-grained visual reasoning tasks.","lead":"SaGe turns flat image-text data into hierarchical scene graphs and trains MLLMs to reason by traversing nodes with attributes, boxes, and depth. This yields large gains on fine-grained perception and spatial benchmarks for 3B/7B models.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified that overturns the central empirical claim.","rationale":"The paper is a clean systems contribution whose central claim is the measured multi-benchmark improvement under a reproducible post-training recipe, not a claim of perfect scene graphs. The hybrid construction (Qwen2.5-VL-72B + compositional crop + Depth-Anything + SAM + LLM edges) is imperfect, and the Gemini audit quantifies that; however, the ablations, teacher-beating results, and multi-benchmark consistency already supply independent evidence that the structured signal is useful enough for the claimed gains. Dependence on large teachers for data/rewards and incomplete public release of the full 120K corpus are practical limitations, not soundness failures. No mathematical circularity or unsupported leap is present. Therefore the reader's ACCEPT / high-confidence verdict stands; the concrete noise-injection check would only further quantify robustness.","tokens_in":27773,"tokens_out":545,"duration_ms":5888,"concrete_test":"Independently re-run the full two-stage recipe on a held-out 5k SA-1B subset after deliberately injecting the observed error rates (≈5.5% node attribute/bbox noise, ≈7% edge noise) into the graphs used for CoT sampling and reward judging; if V* / CVBench lifts remain within 2–3 points of the published deltas, residual graph noise is non-load-bearing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest_assumption correctly flags residual graph noise (94.5% node / 92.9% edge precision on a 1k Gemini audit; incomplete recall; attribute/relation errors). That noise is real, yet it is not load-bearing against the strongest claim. The claim is empirical: after SFT on 120K node-articulated traces plus node-as-proxy GRPO, SaGe-3B/7B produce large, consistent lifts on V*, HRBench, CVBench-2D/3D, etc. Ablations (Tables 3–5) show cold-start alone already moves V* 75.4→83.2; full NPR reaches 89.0; CoT tags and dual rewards are complementary; SaGe-7B beats the 72B teacher on several axes. Two-round filtering (GME + MLLM) plus human spot-check (<1% failure) further dampens transfer of residual errors. Thus graph imperfections do not falsify the reported gains or the practical value of the pipeline.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes Scene Graph Thinking (SaGe), a paradigm that equips MLLMs with structured visual reasoning via explicit hierarchical scene graphs. An automated engine converts flat image–text pairs into graphs with compositional nodes (attributes, bboxes, depth ranges from Depth-Anything + SAM) and multi-type edges built from node priors. From these graphs the authors sample 120K node-/edge-centric QA pairs and multi-hop traces with node-articulated CoT (<entity>, <bbox>, <depth>). Two-stage post-training follows: SFT cold-start on the 120K corpus, then GRPO with complementary node-as-proxy rewards (node-grounded visual–text alignment + node-relevance to query anchors) plus accuracy and format terms. On Qwen2.5-VL-3B/7B backbones, SaGe reports large gains on VStarBench (75.4→89.0 / 76.4→89.0), HRBench-4K/8K, CVBench-2D/3D, and modest lifts on MMStar, RefCOCO, ChartQA, often matching or exceeding larger open models and some proprietary systems. Ablations isolate SFT, GRPO, CoT tags, and each reward; a Gemini audit of 1k graphs reports 94.5% node / 92.9% edge precision; code is released.","tokens_in":28176,"tokens_out":1066,"duration_ms":9401,"significance":"If the empirical gains hold under independent re-implementation, the work supplies a practical, scalable recipe for injecting hierarchical relational structure into MLLMs without requiring online tool use or multi-turn search at inference. The combination of an automated hierarchical graph engine, node-articulated CoT, and node-as-proxy GRPO is a concrete advance over pure crop/zoom “think-with-image” methods. Strengths include consistent multi-benchmark lifts, teacher-beating results on several axes, complementary ablations (Tables 3–5), data-scale curves, and public code. Residual graph noise is acknowledged and partially quantified; the central claim remains empirical and falsifiable.","major_comments":[{"comment":"§3.2–3.3 and Appendix B.1: residual graph noise (94.5% node / 92.9% edge precision on a 1k Gemini audit; incomplete recall; attribute/relation error modes) is real. While two-round filtering and human spot-check (<1%) mitigate transfer, the paper should quantify how residual errors affect final student accuracy—e.g., by injecting controlled noise into a subset of the 120K traces or reporting performance stratified by graph-quality bins. Without this, the load-bearing assumption that node-articulated CoT and node-as-proxy rewards transfer robustly remains only partially stress-tested.","section":null},{"comment":"§3.4 Eq. (5)–(7) and Appendix D: both node-grounded and node-relevance rewards rely on external LLM/MLLM judges (Qwen3-VL-30B / Qwen3-30B). The manuscript should report inter-judge agreement or a small human validation of the reward labels, and ideally an ablation that replaces the learned judges with cheaper heuristics, to show that the reported GRPO gains are not artifacts of judge–student co-training.","section":null}],"minor_comments":[{"comment":"Table 1 vs. Table 2: GPT-4o numbers appear inconsistent across tables (e.g., CVBench-2D overall); please verify and unify.","section":null},{"comment":"Figure 1 and §1: the “heuristic search” caricature of prior work is slightly overstated; several cited methods already use multi-step visual search. Soften the contrast.","section":null},{"comment":"§4.1 Implementation: list the exact KL coefficient β, clip ε, and rollout temperature used for GRPO so that the RL stage is fully reproducible from the main text.","section":null},{"comment":"Appendix C data mixture: the 25K multi-hop + 20K caption split is useful; a short table summarizing exact counts per subtype would help readers re-create the mixture.","section":null},{"comment":"Typos / notation: “SaGe” vs. “SaGe” capitalization is consistent, but “hieratically” (§3.2) should be “hierarchically”; “proxy-as-node” in Fig. 4 caption should match “node-as-proxy” used elsewhere.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The central empirical claim is solid and the ablations are among the cleaner ones in recent MLLM-RL papers. The two major comments are addressable with modest additional experiments or analysis and do not require redesign of the method. Fit for a top ML venue is good once residual-noise sensitivity and judge reliability are clarified."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean empirical methods paper that actually moves the needle on fine-grained and spatial MLLM tasks. The core idea is simple: stop treating visual search as isolated crops and instead pre-build hierarchical scene graphs (compositional nodes with attributes/bbox/depth, prior-based edges), sample 120K node-articulated CoT traces from them, SFT, then GRPO with two complementary node-as-proxy rewards (grounded visual match + relevance to query anchors). That combination is the real contribution; the ingredients are known, but the pipeline and reward design are coherent and new enough to matter.\n\nWhat works: the numbers are large and consistent. SaGe-3B lifts V* 75.4→89.0 and CVBench-2D 67.0→77.8; 7B is competitive with or better than much larger open models and several tool/RL baselines. Ablations isolate cold-start, GRPO, each CoT tag (<entity>/<bbox>/<depth>), and each reward; NRR + NGR are complementary (NGR alone can hurt). Teacher-beating on several axes and data-scale curves help. Code, prompts, and hyper-parameters are public; the Gemini audit (94.5% node / 92.9% edge precision) plus two-round filtering is more transparency than most data-engine papers give.\n\nSoft spots are real but proportionate. Graphs still depend on large teachers (Qwen2.5-VL-72B, Depth-Anything, SAM, LLM judges) and residual attribute/relation noise plus incomplete recall remain; the paper does not claim perfect graphs. Full 120K corpus release is not guaranteed. Free parameters (reward tiers, mixture sizes, LR/batch) are standard for this genre. None of that overturns the empirical claim that the pipeline produces the reported lifts.\n\nThis is for people working on multimodal post-training, visual search, or agentic MLLMs who need practical gains on high-res and spatial suites. It deserves a serious referee. I would engage with it and expect to cite the data engine and node-proxy rewards if I am in that space.","headline":"Solid systems paper: hierarchical scene-graph data + node-articulated CoT + dual node-as-proxy GRPO rewards give large, ablated gains on fine-grained and spatial MLLM benchmarks.","tokens_in":28779,"tokens_out":549,"would_cite":true,"duration_ms":6104,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Training multimodal models on hierarchical scene graphs yields large gains in fine-grained perception and spatial reasoning.","keywords":["scene graph","multimodal large language models","fine-grained visual perception","structured visual reasoning","node-as-proxy rewards","chain-of-thought","spatial understanding","GRPO"],"falsifier":"Train an otherwise identical model on the same volume of data but with the graph structure ablated (flat captions or random crops only); if the large gains on VStarBench, HRBench and CVBench disappear, the claim that the graphs themselves are doing the work is falsified.","tokens_in":28672,"feed_emoji":"🕸️","tokens_out":783,"duration_ms":19684,"temperature":0.7,"pith_summary":"Current multimodal large language models still treat scenes as collections of isolated objects and rely on heuristic crops or zooms, which makes target navigation inefficient and hurts fine-grained tasks. This paper shows that converting ordinary image–text pairs into hierarchical scene graphs—nodes holding entities, attributes, boxes and depth, edges holding spatial and semantic relations—supplies the missing structure. From those graphs the authors sample 120K node- and edge-centric reasoning traces, then run a two-stage post-training pipeline: supervised fine-tuning internalizes the graph-aligned chain-of-thought, and reinforcement learning with node-as-proxy rewards consolidates efficient, query-relevant traversal. The resulting models (SaGe-3B/7B) deliver consistent lifts across eight visual benchmarks, often surpassing larger open models and matching or beating proprietary systems on high-resolution and spatial tests.","feed_headline":"Scene graphs lift 3B models past larger ones on visual tasks","feed_subtitle":"Node-articulated traces and proxy rewards teach efficient relational navigation in images.","key_machinery":"Node-as-proxy graph rewards (node-grounded visual alignment plus node-relevance to the query) inside GRPO, together with the automated hierarchical scene-graph construction and the 120K sampled reasoning traces that train the model to follow those graphs.","core_discovery":"Explicit scene-graph representations, when used both to generate node-articulated chain-of-thought data and to define node-as-proxy rewards for reinforcement learning, equip multimodal models with fine-grained, relation-aware visual reasoning that isolated-object or pure zoom-in methods lack.","pith_inferences":["Residual 5–7 % node/edge error rates imply that stronger foundation detectors would further amplify the observed gains.","The same node-proxy idea could regularize video or multi-view reasoning by treating temporal or cross-view links as additional edges.","Forcing every reasoning step to name a grounded node may reduce visual hallucinations even on tasks that do not explicitly require localization."],"forward_implications":["High-resolution visual search becomes a guided walk over relational pathways instead of repeated heuristic zooms.","Explicit depth ranges on nodes give smaller models competitive 3D spatial reasoning without specialized 3D architectures.","Node-articulated intermediate steps supply verifiable evidence that can be audited or used for further self-correction.","The same data engine can be reapplied to new image corpora, turning any flat multimodal collection into structured training fuel."],"fun_headline_variants":["Scene graphs equip MLLMs with relation-aware visual reasoning","Node-as-proxy rewards teach efficient graph exploration in images","Explicit scene graphs lift fine-grained perception in multimodal models","Graph-aligned training beats isolated-object methods on visual tasks","SaGe converts flat corpora into structured reasoning for MLLMs"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the automatically constructed scene graphs are accurate and complete enough for the sampled traces and node rewards to teach genuine structured reasoning rather than merely imitate residual annotation noise.","fun_headline_variants_meta":{"raw":{"variants":["Scene graphs equip MLLMs with relation-aware visual reasoning","Node-as-proxy rewards teach efficient graph exploration in images","Explicit scene graphs lift fine-grained perception in multimodal models","Graph-aligned training beats isolated-object methods on visual tasks","SaGe converts flat corpora into structured reasoning for MLLMs"]},"model":"grok-4.5","effort":"low","cost_usd":0.00478,"raw_usage":{"total_tokens":1318,"prompt_tokens":734,"num_sources_used":0,"completion_tokens":87,"cost_in_usd_ticks":47800000,"prompt_tokens_details":{"text_tokens":734,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":497,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":734,"tokens_out":87,"duration_ms":4624,"temperature":1.0,"reasoning_tokens":497,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T16:12:01.345835+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train an otherwise identical model on the same volume of data but with the graph structure ablated (flat captions or random crops only); if the large gains on VStarBench, HRBench and CVBench disappear, the claim that the graphs themselves are doing the work is falsified.","supporting_citations":[],"review_version":2}