{"id":"b94acc95-f243-4da6-bee4-8658b0a71771","arxiv_id":"2607.23588","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"JarvisHub open-sources a three-layer canvas-state, protocol-bridge, and agent-runtime harness so multimodal creative agents can inspect and update a shared editable project graph over long workflows.","lead":"JarvisHub is an open canvas-native harness that treats an editable multimodal canvas as the shared project state, memory, and action space for long-horizon creative agents. It matters because closed commercial creative agents are hard to study, and open tooling could standardize how agents plan, revise, and recover across images, video, web, and decks.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The central claim is comparative — the canvas harness does what chat/prompt/node systems \"only partially support\" — yet the paper runs no baseline of any kind, so nothing distinguishes canvas-state benefits from simply giving GPT-5.5 good tools.","rationale":"The reader's weakest_assumption already identifies attribution — harness vs. underlying generators — as the soft spot, and flags the qualitative-only, three-task evidence conceded in §5. My read confirms that concern and sharpens it: the paper's own §1 framing makes the claim comparative against chat/prompt/node systems, so the missing piece is not just \"more evaluation\" but specifically the canvas-vs-chat ablation under identical tools, plus perturbations that exercise the protocol bridge and recovery tools the paper advertises. This does not change the verdict: CONDITIONAL is correct — the work is defensible as architecture + open software + illustrative workflows, and the limitations section is honest. The contribution stands if positioned as a harness, and the proposed ablation is the cheapest decisive check. I do not find internal inconsistency in the formalism (Eqs. 1-5 are descriptive contracts, not derivations with hidden assumptions), and the open repo plus named backends give middling-but-real reproducibility. No stronger downgrade is warranted; no upgrade is warranted until the ablation exists.","tokens_in":12431,"tokens_out":1301,"duration_ms":70474,"concrete_test":"Run all three §3.2 tasks under two conditions with identical backends (GPT-5.5, GPT Image 2, Seedance 2.0) and identical tool families: (a) full JarvisHub canvas + protocol bridge, (b) chat-only baseline where C_t is replaced by a linear transcript and mutations are unchecked. N≥20 seeds each, with 2-3 scripted perturbations per run (a failed generation, a user rejection of an accepted candidate, a late style change). Score repair success, retention of previously accepted artifacts, and locality of revision (fraction of unaffected nodes preserved). If (b) matches (a) on these process metrics, the canvas/protocol claim reduces to tool access; if (a) wins, the central claim gains its first real support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is not merely \"JarvisHub works on three tasks\"; it is that the canvas-as-shared-state architecture is what enables sustained consistency, recovery, and human-steerable revision, in contrast to prompt-, chat-, and node-based systems that \"discard intermediate context\" or \"require manually specified workflows\" (§1, Figure 1). For this to hold, the canvas state layer (Eq. 1) and protocol bridge (Eq. 3-4) must contribute something beyond the agent backend and tool families. But §3 contains zero controlled comparisons: no chat-only ablation (same GPT-5.5, same generation/native tools, linear transcript instead of C_t), no unbridged ablation (canvas present but mutations unchecked), and no process metrics (repair success rate, reference retention, revision locality, intervention count) even for the single JarvisHub condition. Each demo shows one curated successful trajectory per task; Figure 4-9 captions assert continuity and inspectability without any measurement, and §5 concedes the benchmark/leaderboard is unfinished. Worse, §2.3's protocol bridge and §2.5's feedback-repair loop are exactly the components the paper claims differentiate it, yet they are never exercised adversarially — no injected tool failure, no contradictory user feedback, no competing-version conflict appears in any trace. The concern is not that the system is broken; it is that the paper's own framing (\"moves creative agents beyond isolated tool use\") makes the comparative mechanism the load-bearing element, and that element is untested. The attribution problem the reader flagged is real, but the sharper issue is that the distinguishing ablation is cheap to run and absent.","agreement_with_reader":"agree"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a real open systems artifact—canvas as shared project graph, protocol bridge, runtime with tool families/skills/memory/subagents—and the writeup is clearer than most agent-harness papers. What it does not yet show is that the canvas layer, rather than GPT-5.5 plus good tools, is what buys long-horizon consistency.\n\nWhat is actually new is the integrated formalization and shipped harness, not any single piece. Canvas UIs, Comfy-style graphs, prompt chains, MCP, and skills/memory/subagents are all prior art they cite. Treating the editable multimodal graph as external memory, action space, and inspectable state under checked grants (Ct, Γt, Ωt, τ) is a useful research object. Equations 1–5 are bookkeeping, not deep theory, but they make the contract legible. The limitations section is honest: qualitative demos only, quality still rides on external models, trajectories need filtering before reuse. Repo is public. That is real infrastructure credit.\n\nThe soft spot is the one the stress-test names, and it lands. Figure 1 and the intro frame prompt/chat/node systems as only partially supporting project state; §3 then shows three curated successful traces (drama, photo site, ML deck) with no chat-only ablation, no unchecked-mutation ablation, no process metrics (reference retention, local repair rate, intervention count), and no injected failures. So you cannot yet attribute continuity to the canvas versus a strong backend with tools. That is not fatal for a systems paper if they position it as architecture + software + illustrative workflows; it is load-bearing if they sell “moves agents beyond isolated tool use” as demonstrated fact.\n\nCitations look appropriate; self-cites sit in a line of related agent work rather than padding. No circular math games—mild systems tautology (the loop does what the loop defines).\n\nWho it is for: people building creative-agent runtimes, project-state benchmarks, or trajectory datasets. Not for someone hunting a new generation model or a scored leaderboard. I would bring it to a systems/HCI-agents reading group as “here is the open harness shape,” not as settled behavioral science. Send it to referees; ask for baselines or a hard repositioning of claims. Worth engaging if you care about open creative-agent infrastructure.","headline":"Open canvas-as-state harness with a clean three-layer design; the comparative claim outruns the three qualitative demos.","tokens_in":13963,"tokens_out":580,"would_cite":true,"duration_ms":15867,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Creative agents need a shared editable canvas as project state, not chat logs or fixed pipelines.","keywords":["creative agents","canvas-native interface","multimodal generation","long-horizon workflows","project state","agent harness","human-in-the-loop","trajectory recording"],"falsifier":"Build a controlled project-state benchmark with fixed initial canvases, tools, and feedback events; compare JarvisHub against chat-only and fixed-pipeline agents on process metrics (context preservation, dependency correctness, local repair success) and final quality; if the canvas harness does not improve process metrics or recoverability when generators are held fixed, the central claim fails.","tokens_in":13534,"feed_emoji":"🎨","tokens_out":964,"duration_ms":22778,"temperature":0.7,"pith_summary":"Real creative work is long-horizon: references, drafts, alternatives, failures, versions, and feedback accumulate into a project state that single prompts and linear chats cannot hold. This paper argues that the missing piece is not another generator but an open harness in which an editable multimodal canvas is the workspace, external memory, action space, and shared state for both user and agent. JarvisHub represents artifacts, dependencies, versions, and feedback as typed nodes and links, and routes agent behavior through a three-layer design—canvas state, a protocol bridge that checks reads and writes, and a runtime that plans, calls tools, and records trajectories. On narrative media, interactive web, and presentation tasks, the system shows agents planning, generating, revising, and organizing while humans can inspect and intervene. If the claim holds, creative AI research can study sustained, steerable production instead of isolated prompt–output steps.","feed_headline":"Creative agents get a shared canvas, not just chat","feed_subtitle":"An open harness treats the editable board as memory, action space, and project state for long multimodal work.","key_machinery":"The canvas-native three-layer harness: the canvas as typed artifact graph Ct = (Gt, Xt, Mt, Ut, Lt); a protocol bridge that issues capability manifests and execution grants and commits only checked mutations; and an agent runtime that selects granted actions, invokes tool families, and records full trajectories including feedback and repairs.","core_discovery":"JarvisHub establishes that long-horizon multimodal creation can be formalized as an agent process over an editable project graph on a canvas, and that a three-layer harness—canvas state, protocol-constrained interaction, and an agent runtime with tool families, skills, memory, and subagents—lets agents plan, generate, revise, and organize multimodal projects in a way that remains inspectable and human-steerable, unlike prompt-only, chat-only, or fixed node-pipeline systems.","pith_inferences":["Closed commercial creative agents may already use similar internal state, so open canvas harnesses could become the main way the field measures and compares long-horizon creative behavior.","Once trajectory datasets exist at scale, progress may shift from better single-shot generators toward models trained specifically on canvas mutations and local repair policies.","The same graph-plus-grant pattern could transfer to other long-horizon multimodal domains (scientific figure pipelines, game content, multi-page design systems) that also outgrow linear chat.","Without standardized process metrics, harness papers risk being judged only on final visuals that mostly reflect the underlying image/video models."],"forward_implications":["Creative agents can keep prompts, references, drafts, versions, and feedback as addressable canvas state instead of discarding them in chat history.","Researchers can define project-state benchmarks with initial canvases, tools, constraints, and checkpoints rather than only prompt–answer pairs.","Evaluation can combine final artifact quality with process measures: context preservation, tool appropriateness, dependency correctness, feedback adherence, and repair success.","Recorded trajectories of states, actions, observations, feedback, and repairs become training data for planning, tool choice, state tracking, and local repair.","Users can inspect, edit, and steer the same project graph the agent reads and writes throughout long workflows."],"fun_headline_variants":["JarvisHub: canvas as memory and action space for creative agents","Open harness makes multimodal agents work on an editable project graph","Long-horizon creation formalized as agents over a shared canvas state","Three-layer harness keeps creative agents inspectable and steerable","Beyond chat: typed canvas nodes link drafts, versions, and feedback"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That qualitative workspace traces and final artifacts on three hand-chosen tasks with strong external generators are enough to show the harness itself—not the backends—delivers long-horizon consistency, recovery, and competence.","fun_headline_variants_meta":{"raw":{"variants":["JarvisHub: canvas as memory and action space for creative agents","Open harness makes multimodal agents work on an editable project graph","Long-horizon creation formalized as agents over a shared canvas state","Three-layer harness keeps creative agents inspectable and steerable","Beyond chat: typed canvas nodes link drafts, versions, and feedback"]},"model":"grok-4.5","effort":"low","cost_usd":0.002068,"raw_usage":{"total_tokens":961,"prompt_tokens":870,"num_sources_used":0,"completion_tokens":71,"cost_in_usd_ticks":20684000,"prompt_tokens_details":{"text_tokens":870,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":20,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":870,"tokens_out":71,"duration_ms":2531,"temperature":1.0,"reasoning_tokens":20,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T18:09:13.245170+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Build a controlled project-state benchmark with fixed initial canvases, tools, and feedback events; compare JarvisHub against chat-only and fixed-pipeline agents on process metrics (context preservation, dependency correctness, local repair success) and final quality; if the canvas harness does not improve process metrics or recoverability when generators are held fixed, the central claim fails.","supporting_citations":[],"review_version":1}