Pith. sign in

Paper Citation Record · LEDGER

Qwen3-VL Technical Report

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2511.21631.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.21631 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 100 of 1869 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:17:53.872394Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fe585dc7-9c0b-4d92-9e1e-a4bd2b0b919e · inbound

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO cites this paper.

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO Qwen3-VL Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T00:22:18.787000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T00:21:51.621582Z digest=sha256:367ca56f8ceafb31b773140985c224b3144a287c5ddd674f3054f6ad1ccb5240

Observation ab1d400d-5dc2-4de9-9379-10855b88bb59 · inbound

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence cites this paper.

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence Qwen3-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:11:35.674800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:07:11.548885Z digest=sha256:3a748d0c2bf19984a27e984ec225d20ed2ff8d3d3b65275e6cefbd5dc15ac61d

Observation bbd98ed3-fe46-4315-93c2-261e898b1235 · inbound

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks cites this paper.

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks Qwen3-VL Technical Report

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T14:37:35.792455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T14:36:30.827705Z digest=sha256:538fb2dada2657f4cccabea24d16e7d884377fea4a7e0cf44455171ee4f8768d

Observation bd703d6e-3bec-4583-9014-7a79ec11f3e2 · inbound

Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification? cites this paper.

Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification? Qwen3-VL Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:53.872394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:53.872394Z digest=sha256:ff3a30cab5b4779e7b588ce57fad62e0889a269b4e41b1ca86e9da6cf279d30c

Observation e3e94efc-dcf7-41d6-a470-493466e02992 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Qwen3-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.126418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.126418Z digest=sha256:7c8bb189a02aceaaa99d0ee7f5c42b2fcda780bb1e684301c925297b79af3754

Observation 1cab8f1d-c5f0-4b28-9d7b-118e6deff237 · inbound

VideoGuard: Protecting Video Content from Unauthorized Editing cites this paper.

VideoGuard: Protecting Video Content from Unauthorized Editing Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T04:27:04.607470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:27:04.607470Z digest=sha256:e692dc6f7af44284d38c549623ac475f2798f669374553d671d02be374ada40a

Observation cfae5e5e-4b17-4f63-898a-c6606e079d79 · inbound

SkillWrapper: Generative Predicate Invention for Task-level Robot Planning cites this paper.

SkillWrapper: Generative Predicate Invention for Task-level Robot Planning Qwen3-VL Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:44:07.201529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T05:43:56.774051Z digest=sha256:3e5a83ad6c68b2a8c6dc9b1540e12b26f733cf46795e93fae999b2cd5c1d60a2

Observation 05b71cb1-b3c3-4e6b-8538-f15b051ac6e4 · inbound

SkillWrapper: Generative Predicate Invention for Task-level Robot Planning cites this paper.

SkillWrapper: Generative Predicate Invention for Task-level Robot Planning Qwen3-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:53:33.614695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:53:33.614695Z digest=sha256:62ac9cc766dce3bf464f9452f9005729b1bf48cc342df19bd5c2ab5e6028097a

Observation 39d336c0-c180-46fc-b34b-2885586ebb50 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Qwen3-VL Technical Report

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:11:26.492650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:d8fa0b6f1727acb345ce97c0a0628b915ca1b8c5e55529f89c254db35ea30a6c

Observation 13e34751-035f-4fb0-9e60-2f3c880588ff · inbound

ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos cites this paper.

ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:51:28.038076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:50:17.955287Z digest=sha256:c05c512fa53e75bf7f44ffe23bb7ed9635f3ec7cf9eb04d3ada8bd67e720ba3a

Observation 3f6a00a5-566b-4481-bf30-9f350e0fc805 · inbound

SAM3-I: Segment Anything with Instructions cites this paper.

SAM3-I: Segment Anything with Instructions Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:08:51.399219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:06:35.864921Z digest=sha256:665f48b72ab95a2c70e2aed98f355af2cac0d3860338999ad2b86fe1b6c80602

Observation 967822dd-61a2-4198-9923-dc5f5a705185 · inbound

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension cites this paper.

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T18:15:04.635320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:15:04.635320Z digest=sha256:2025a59d7589d9186a66d6ecf5b12ed9f67bb5876a6636e48b10a1b93b4f4a0c

Observation 03c856b0-d022-46bb-9ac0-9359badea614 · inbound

Mull-Tokens: Modality-Agnostic Latent Thinking cites this paper.

Mull-Tokens: Modality-Agnostic Latent Thinking Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:58:38.598410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:57:04.802871Z digest=sha256:fb714f7c4d4f764ae432590b1c3622d41b8f37337380d060e7b577b62da37df2

Observation 1c1e00cf-e3de-43f1-b1ac-6a77a874954d · inbound

Are vision-language models ready to zero-shot replace supervised classification models in agriculture? cites this paper.

Are vision-language models ready to zero-shot replace supervised classification models in agriculture? Qwen3-VL Technical Report

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:18:32.173903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T21:15:20.717705Z digest=sha256:a632bc572b871930aba93e3c6783d610af9193403fc0dbfac408b1e2b9dfe1dc

Observation 16b730f1-8fb9-4e40-9ae2-c3da6938de0b · inbound

Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding cites this paper.

Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding Qwen3-VL Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:31:13.416234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T19:28:35.576661Z digest=sha256:4a30a56ad87defd71ab636b01143dd0ffd3ca8b42a1f6cbb00d18c3556aad0fb

Observation 2543ab7f-c882-41ff-a549-10703700eb72 · inbound

VPTracker: Global Vision-Language Tracking via Visual Prompt cites this paper.

VPTracker: Global Vision-Language Tracking via Visual Prompt Qwen3-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:03:18.733192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T19:02:27.777162Z digest=sha256:c7d9ef4a373a03473d032551714e3f340d800ea1ce33df111816990485e67fcc

Observation 3bcc6ae7-329c-4345-af8c-2fd4eca81b51 · inbound

S1-MMAlign: A Large-Scale, Multi-Disciplinary Dataset for Scientific Figure-Text Understanding cites this paper.

S1-MMAlign: A Large-Scale, Multi-Disciplinary Dataset for Scientific Figure-Text Understanding Qwen3-VL Technical Report

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:18:14.439263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T18:16:45.825451Z digest=sha256:ff9f4a84bacc1efe6014caf81a60853dba97b8ad9846354a6a5fc15e2f3d0e28

Observation e4b9d45e-f926-478d-a5e4-eac5d5258435 · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.737643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.737643Z digest=sha256:3466d13501172f7db2f2ca8d38e0e01362a1d433f0223a5a8da3da2034f130d5

Observation 768ad5cf-84db-4e70-a90b-7928f392fe03 · inbound

BabyVision: Visual Reasoning Beyond Language cites this paper.

BabyVision: Visual Reasoning Beyond Language Qwen3-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:36.210179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:36.210179Z digest=sha256:eeb5c35adb31c63b5b44d9c76e481ed6a2ea1b333fa48d328699608dd09e7ec9

Observation 357b5972-8863-4e53-870b-80c51b8ac1b2 · inbound

Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos cites this paper.

Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos Qwen3-VL Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T14:58:01.128329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T14:56:15.554554Z digest=sha256:827a87197734ab1c3c239aa48d017cb9adfa09d1f912055d8ed62b1a9070225a

Observation 1142b4bd-0af4-493e-b6ee-7302f29b8ab7 · inbound

Ministral 3 cites this paper.

Ministral 3 Qwen3-VL Technical Report

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T19:12:24.675756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:104aaf98aef79c4a443867a74a459d07e8100d17fb612d42fbcf3bd13fc18060

Observation 223de39f-6967-4380-9bab-bf6031f185b8 · inbound

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding cites this paper.

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding Qwen3-VL Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T10:49:58.837335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:49:58.837335Z digest=sha256:c596838aa04ae9cdd9bade35138c91c5d7e2da1ec65c64410b1702a5fe481f71

Observation 233866a4-1b1a-4365-a3c8-527afeb6e726 · inbound

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding cites this paper.

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding Qwen3-VL Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T04:21:29.723925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:21:29.526008Z digest=sha256:33aa184f391e64354dd997e1a3ed4f3c2fa35b0d0591bbe3a73b7850f640a307

Observation 36d7964a-70ea-467d-975f-6f9dc0137e63 · inbound

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch cites this paper.

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch Qwen3-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:12:54.854639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:12:01.889341Z digest=sha256:1465f0fa6cb8226acef6a8efcce0734a90645c69455c090f19adfb52ad6a3da6

Observation b9b5dc81-54ec-4d9e-8b4a-050b10c77247 · inbound

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs cites this paper.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:58.420522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:58.420522Z digest=sha256:ae47484964712715b0f934ca9669ac6af456078fb762c7078ae8d3ea9a333f19

Observation 70df9ede-de91-42c0-a414-7f8b96845227 · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Qwen3-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:57:53.908356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:a04f67b5c00b78df2fee547b4da29cf9ba09c043137c736eeef82edaafa82574

Observation 80dce210-42de-4b05-a844-e28ab1cab5b6 · inbound

Common to Whom? Regional Cultural Commonsense and LLM Bias in India cites this paper.

Common to Whom? Regional Cultural Commonsense and LLM Bias in India Qwen3-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:37:52.899451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T12:37:25.252161Z digest=sha256:386f226929ffdd4f4baf09d409e63f56ca528a5e10a3fdd86afd69df0a0d8c53

Observation d27e6d48-5024-4215-9ff5-7be6ff71a249 · inbound

Scaling medical imaging report generation with multimodal reinforcement learning cites this paper.

Scaling medical imaging report generation with multimodal reinforcement learning Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T08:29:47.265364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:29:47.265364Z digest=sha256:598073c500815cd36f283a7b9b487b5343ed86254a75025bbb18c4d1e7fc090c

Observation 03c9571d-0916-4217-badf-cdd4652565bc · inbound

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents cites this paper.

GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents Qwen3-VL Technical Report

Reference 106

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:32:48.303676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T11:31:16.087579Z digest=sha256:ea01248cee40a567a98876e190bf629957201e1ac6557ef9cbc8af9d0fff0c9d

Observation 73410ca9-f773-4739-bc2d-a4c91508ac2f · inbound

Focus on What Really Matters in Low-Altitude Governance: A Management-Centric Multi-Modal Benchmark with Implicitly Coordinated Vision-Language Reasoning Framework cites this paper.

Focus on What Really Matters in Low-Altitude Governance: A Management-Centric Multi-Modal Benchmark with Implicitly Coordinated Vision-Language Reasoning Framework Qwen3-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:07:48.358867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T11:03:37.176829Z digest=sha256:8e2dfe13f726f16966d4ea81602d37c7f935facef6759984625a5e0fcb7c73ed

Observation b53df7da-44a1-4b47-b9e5-bb76d9e6d7ea · inbound

Advancing Open-source World Models cites this paper.

Advancing Open-source World Models Qwen3-VL Technical Report

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:07:01.026530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:07:00.904794Z digest=sha256:6743a519e1e2302a4f9428ec20b24752b77510810b527fb1d9b347cc887b74da

Observation cc8683e6-e8bd-4cbc-8c80-c31113ec594e · inbound

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding cites this paper.

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:17:44.704604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:14:15.589472Z digest=sha256:721ada4ee10fbda02869f723ebb5601533d3841e869059562906257d02dac6ff

Observation 79ac3723-2a3c-4df7-80b3-aeb7e722c2a1 · inbound

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation cites this paper.

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T07:03:11.740274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:03:11.740274Z digest=sha256:880a0e838689f552487ee0bcc89990e3bd02cabd07f71c3ce99eb4ab980fa73d

Observation 09b98cdc-1397-4b82-adef-3997f6b400fc · inbound

TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning cites this paper.

TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:35:28.715407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T07:33:15.865358Z digest=sha256:4e8da68662be82524c07bca5b9add2aff18fc4dd7880f12af3c7f3a3794d8799

Observation 9f433c2c-a656-463c-9866-010787960c87 · inbound

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models cites this paper.

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models Qwen3-VL Technical Report

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T14:50:14.733038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T14:48:21.787919Z digest=sha256:530ebdfd57e65815da383b6925c5b1e039443c00f697a39aca7d35b10fe68059

Observation 11aa3de7-df12-4be4-9338-347592404301 · inbound

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability cites this paper.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Qwen3-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:31.646881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:31.646881Z digest=sha256:a3b3c2709739dbc1254c0d81790702299c113720b3d2c3a1742da70454eed3cb

Observation 0c31f05b-b385-42da-a432-c1498d167a03 · inbound

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning cites this paper.

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:02:42.459666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:02:20.477517Z digest=sha256:b3882d92f496d66407708029823878db00bd130c0bd8e413044a65a2cf7deef9

Observation fa75d537-19ca-46fb-bce3-4b07da7aca58 · inbound

Dual Latent Memory for Visual Multi-agent System cites this paper.

Dual Latent Memory for Visual Multi-agent System Qwen3-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T06:04:50.189057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:04:50.189057Z digest=sha256:f86d3359237f3aef1f39ee45694896e44a5983ac021705a2a4d830823c167a9f

Observation 5943ccf5-4882-4409-bf94-b3a63ff65977 · inbound

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding cites this paper.

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding Qwen3-VL Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:32:36.542385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:30:50.984873Z digest=sha256:11ef42549fe91bac5e9bf38f4676c65a6d85fb26f85083e6d7c55210220450c4

Observation da3fd933-8924-4d36-93e0-bd608cfddeef · inbound

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning cites this paper.

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Qwen3-VL Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:27.189984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:27.189984Z digest=sha256:132a19a73b67960b313196b70318a75ca429e02a50ef99aee3f9e06039ffa9a5

Observation aefbea46-0aa4-40a4-84b7-404cf41b371f · inbound

LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation cites this paper.

LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation Qwen3-VL Technical Report

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T05:29:03.479558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:29:03.479558Z digest=sha256:3add31753b55e9313be6cbf99f842214af59a05ae700befd52db7645a5648109

Observation 1b798d46-87d2-4ba7-9c74-07b0059f3204 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence Qwen3-VL Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-10T16:09:05.268836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:87ba67147ba9fea6293f26948cca55c588f0a42f45593e5e42404b7728470223

Observation 2d0bec1c-6591-4868-9f89-16e07bbc61a3 · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:40:46.405043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:38:49.075457Z digest=sha256:84ed6a6453e79ab790337098a312faa30ba0253fd868d9c17a99e641dd69045a

Observation f4288644-0504-42b7-9442-5a9859579005 · inbound

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation cites this paper.

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation Qwen3-VL Technical Report

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:22.508983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:14:22.508983Z digest=sha256:5bd2a998d1dff7fdd6c0595c5706d59d1768d6b8bc12443673dafd6d3fec7382

Observation eb666695-a8e2-49f6-a2db-7e3bfd543aac · inbound

Decoupling Skeleton and Flesh: Efficient Multimodal Table Reasoning with Disentangled Alignment and Structure-aware Guidance cites this paper.

Decoupling Skeleton and Flesh: Efficient Multimodal Table Reasoning with Disentangled Alignment and Structure-aware Guidance Qwen3-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T05:02:40.578669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:02:40.578669Z digest=sha256:92c7a08e2420125b9f6312b878fb0fd35ce8b1647629109844a47403350aaba2

Observation 19b4b2cb-a4cc-4f11-b1b9-f826892306b0 · inbound

Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data cites this paper.

Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T04:37:17.655839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:37:17.655839Z digest=sha256:d65bf70660e4cc1ee5f92ab0c745f5c660505b3d065108366d47669accd7cdc2

Observation 166a2c47-85ae-4c04-906c-9eee29fe86a9 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:50.269672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:50.269672Z digest=sha256:31a73ed5b7de6e712f13018fe2f07d31725ee9e0bf03c10636380affe3ebd03f

Observation 4b9e2f50-5aff-4032-a80d-7fb7772c10d2 · inbound

MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation cites this paper.

MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation Qwen3-VL Technical Report

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T07:27:31.674384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T07:24:11.276711Z digest=sha256:095b6ae015f77d819a1f70546d359c2177953cd7d873c2bcf7487eac55a4b330

Observation f6b0e7ea-2f1a-4be9-9163-f1c090d9150d · inbound

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy cites this paper.

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:47.134762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:56:47.134762Z digest=sha256:514b7232541e741dcb5394a37bfb036195516ccace1fd970548505b552d7d137

Observation 990ac962-f728-465b-a11b-9560849c980d · inbound

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization cites this paper.

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization Qwen3-VL Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:10:43.127503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T07:09:46.254851Z digest=sha256:7cd6490053d970208cb3d58bd4245b30785dc67d980c03682c177255d0245fb9

Observation 469223d4-0eed-45ea-a6ec-f883a1df4cc4 · inbound

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models cites this paper.

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models Qwen3-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:37:50.165649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:37:50.165649Z digest=sha256:c9a956ed2cc2430e5b6cedafff8f0b03d6a0aa27a2827e4001184471e9067ce0

Observation 495d751b-5049-4590-a6ed-885cc8edab04 · inbound

Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning cites this paper.

Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning Qwen3-VL Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:04:11.944041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T14:03:48.795572Z digest=sha256:8b49f624191ebd7b45085b3cbda25e51b274b5dad853ce97c1e8e57785731313

Observation a2fd6f15-800d-4fc6-b488-6859a4873b9e · inbound

Prism: Spectral-Aware Block-Sparse Attention cites this paper.

Prism: Spectral-Aware Block-Sparse Attention Qwen3-VL Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T03:22:55.371524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:22:55.371524Z digest=sha256:75ae193a3cdaf6820f2fecf726ce1f0f33baec7c3ae6972c3a82288bdd38100f

Observation c8e65144-6508-45a3-8d93-4f6280a36fd6 · inbound

Multimodal Learning for Arcing Detection in Pantograph-Catenary Systems cites this paper.

Multimodal Learning for Arcing Detection in Pantograph-Catenary Systems Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:16:09.825495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:16:09.825495Z digest=sha256:ca17389a33c9c18280f4da46e655a77d0c9eb7ad0ca45cb5385ddb172672748a

Observation f615c7e5-4e8c-4516-9c83-bc1fd9a19778 · inbound

Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense cites this paper.

Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense Qwen3-VL Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:08:29.381598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:08:29.381598Z digest=sha256:97b67533aad74fc72fd73c7fde4a396f91f13198a817d178c2d94490e9e91429

Observation 11cf1db0-9ef5-4f1f-bdd8-01bb1507c1c8 · inbound

Towards Explainable Industrial Anomaly Detection via Knowledge-Guided Latent Reasoning cites this paper.

Towards Explainable Industrial Anomaly Detection via Knowledge-Guided Latent Reasoning Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:10:31.981569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T03:08:58.617137Z digest=sha256:a4ca7469462122edb6a21829e9068eadd4bf2c2b9c43632bd3112ab89d2d03cb

Observation b349e52a-5994-4ac3-97de-6beed00b7849 · inbound

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories cites this paper.

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T01:02:32.574737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:02:32.574737Z digest=sha256:124459fe224c66d335fac87c230ffc97d339461ff038a85d28415baf7f6b95e1

Observation 04a2e567-54f6-4f53-8fbd-1f3c19c7cee5 · inbound

Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling cites this paper.

Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:45:26.198129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T06:41:33.493927Z digest=sha256:ed970af36b90044b5a1c0b232b8296910ff04d93f22db44358a002baa9680bfd

Observation 2388e3c1-f114-4da2-bc4e-ac6bc6ea63b6 · inbound

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning cites this paper.

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:12:11.842647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T03:11:52.645633Z digest=sha256:6908ff59bcc357c421c4cb72ad7abc92dbf600933e3f65d52b78b4642cddf77d

Observation 3ed0c257-6fec-4467-b33e-a268a69c9105 · inbound

FAIL: Flow Matching Adversarial Imitation Learning for Image Generation cites this paper.

FAIL: Flow Matching Adversarial Imitation Learning for Image Generation Qwen3-VL Technical Report

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-02T23:59:12.875366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:59:12.875366Z digest=sha256:8bc66cd5704deb834024873976f1630862f7803f4098c72a163ea31a4c8794f4

Observation 4a0c099a-dbb6-4ba8-b96f-722ebc087593 · inbound

CAPTS: Channel-Aware, Preference-Aligned Trigger Selection for Multi-Channel Item-to-Item Retrieval cites this paper.

CAPTS: Channel-Aware, Preference-Aligned Trigger Selection for Multi-Channel Item-to-Item Retrieval Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T23:49:54.782056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:49:54.782056Z digest=sha256:39a897063d16a9d282353c4d62980e9fb04e8ebbb2746a412b2e7fd30cbce476

Observation 4a9581ac-5a48-441e-a833-0c63b98d7c91 · inbound

JARVIS: An Evidence-Grounded Retrieval System for Interpretable Deceptive Reviews Adjudication cites this paper.

JARVIS: An Evidence-Grounded Retrieval System for Interpretable Deceptive Reviews Adjudication Qwen3-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:31:43.758096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T22:31:09.746554Z digest=sha256:263af18caf644150b205f11d7bbb519190d323ef4d1b5c1c3ebeb52d881ae809

Observation f705114d-687f-4cf4-96ac-5af4d77ad5a5 · inbound

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding cites this paper.

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T23:44:34.956040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:44:34.956040Z digest=sha256:1258fc949f4ddf3b0a428d56511c386aedf135ceed9e1b255eed8e26d5156b7e

Observation 792637b8-b22c-4e0e-93d8-fe3dff08b3d3 · inbound

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension cites this paper.

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension Qwen3-VL Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T03:30:33.088129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T03:27:53.694506Z digest=sha256:e3ebe73874728e919dd45b2565ccf28df7dcd4bea1d1d11d06710d9495bc75ce

Observation e4382789-4f06-4374-8035-9c43e06223f3 · inbound

SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification cites this paper.

SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T23:33:14.940131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:33:14.940131Z digest=sha256:7fbf59b0f328bc479b19a3bbd891142b0982073934f1d541baca3895edc9bc83

Observation 37eea6c8-c8d5-484a-95ef-ad0a1a6522ae · inbound

Exploring a Multimodal Chatbot as a Facilitator in Therapeutic Art Activity cites this paper.

Exploring a Multimodal Chatbot as a Facilitator in Therapeutic Art Activity Qwen3-VL Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:11:42.404880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T22:10:37.970689Z digest=sha256:98a23138fd6ffdb05a860cdd59d3cb8a8b2c43eacefda2d2d59fa78ff70ad979

Observation f52e3b97-deae-46e5-a889-34adbc6bf5cf · inbound

ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision cites this paper.

ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision Qwen3-VL Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T21:36:40.000612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:ccfe41a2570c6ca4bf73934346324bf69dc2d67a3aeff0aa1ca1ea8be9208ee9

Observation aa6a7a59-a574-48e1-acbc-c67ed29039cc · inbound

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems cites this paper.

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems Qwen3-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T22:58:04.174651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:58:04.174651Z digest=sha256:51e1e6459fa2cbf70646ce53460a16fc1e117d8633d7c484772b377e25e7b8c6

Observation 69d5d5c3-7316-4d86-b69b-d048f96e6ada · inbound

DODO: Discrete OCR Diffusion Models cites this paper.

DODO: Discrete OCR Diffusion Models Qwen3-VL Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T22:27:40.232320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:27:40.232320Z digest=sha256:c23eaf52d74e7ac6183656323aaa0ee4ff1b294ff404a498954ef2dc1de3e6b5

Observation 66a0e07c-dd83-4b94-9929-b405c103bbd6 · inbound

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments cites this paper.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Qwen3-VL Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:39.292349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:39.292349Z digest=sha256:56198488e956c69a6f797ec4ca00e9131d019135200dcc3db6922805e14be6ce

Observation b2529d4c-c5b3-4386-af60-45f20360ac71 · inbound

VLANeXt: Recipes for Building Strong VLA Models cites this paper.

VLANeXt: Recipes for Building Strong VLA Models Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:00:09.843841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T12:58:30.777235Z digest=sha256:3572d4a2f37cc56af1df385ef01892dc2be14b901229273a28a09c53adb40354

Observation 4a1d9cd5-0a6d-4593-9242-8ac4b7c4089f · inbound

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery cites this paper.

FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T21:47:00.159441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:47:00.159441Z digest=sha256:13d3c091330b77a0b4e8c9008cf3024b078b640584c63d7dc568cd42bf1e421a

Observation 06d9cbd1-2ac9-4d59-a461-e906f25cf304 · inbound

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics cites this paper.

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T21:43:06.921012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:43:06.921012Z digest=sha256:676a2b105812647ff70bba948359b5618754b19da804719dc5917cca9d8f100f

Observation 43ab2a23-547a-4dd6-b6d7-5893326c027a · inbound

Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework cites this paper.

Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework Qwen3-VL Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:56:37.029054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:56:06.005832Z digest=sha256:8e5b6062020d30da2d3ccfc6a0c9b51f985320dd6c2745c3bb5b3f86a13dc735

Observation 3c42f98b-5a6c-4c41-b741-1c377628af9f · inbound

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning cites this paper.

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning Qwen3-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:06:33.818788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:05:30.324309Z digest=sha256:29c92e3b9267d411e2198474345a8d37961036587a63a077ce3bf7286f1f7995

Observation a2e33ca8-ac2a-465d-8ccf-3a79d0f16854 · inbound

Efficient Scaling of LLM Training with Flexible Context Parallelism cites this paper.

Efficient Scaling of LLM Training with Flexible Context Parallelism Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T20:59:24.170775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:59:24.170775Z digest=sha256:c2b9f4f93369b3d01b68a5528203277339209c2fd3ff98a95b873fdbbbf560ce

Observation d7567fb1-0730-4be7-9d34-20a5cd28eb33 · inbound

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL cites this paper.

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL Qwen3-VL Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T20:51:40.372101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T20:51:40.372101Z digest=sha256:a02af45cc024baf82a480ca8199f81586ced2422341ed063ae61f7b313d71526

Observation 2d4e5937-16ab-44ae-b78b-56bf087e50c2 · inbound

Imagination Helps Visual Reasoning, But Not Yet in Latent Space cites this paper.

Imagination Helps Visual Reasoning, But Not Yet in Latent Space Qwen3-VL Technical Report

Reference 1971

Resolution
unresolved
no resolver link, observed 2026-08-02T20:40:29.696203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:40:29.696203Z digest=sha256:fe1ca142198dacbecaf4c98e0db000d402f9da33f92571835d075246d7704135

Observation 20d0d4a4-2c64-4f7d-8f5d-5fd0ad7a898a · inbound

PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models cites this paper.

PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T20:38:23.446746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:38:23.446746Z digest=sha256:10f47bc8e131f64de6d5ad3f921e948cf7d9dab397aec0d7e27f6cb93b9bd17b

Observation 788c4ca7-7e55-40db-bbe8-21b1414254f0 · inbound

Can Agents Distinguish Visually Hard-to-Separate Diseases in a Zero-Shot Setting? A Pilot Study cites this paper.

Can Agents Distinguish Visually Hard-to-Separate Diseases in a Zero-Shot Setting? A Pilot Study Qwen3-VL Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T20:37:20.863237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:37:20.863237Z digest=sha256:6e9a1d030d6107aa20bdd94d386a658b09d17cb2e33c8833c20f89e4531e2a9d

Observation 1fc72b8e-0fd1-4fe6-a8f1-f87eb540b08b · inbound

CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era cites this paper.

CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era Qwen3-VL Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:51:30.676601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:50:13.036260Z digest=sha256:b2a18f1dfc93851d4b060b3030dafc9338522aab1487240c41611ac2f7362eaf

Observation 8376e032-b09c-4640-abc4-20f5a17c8be1 · inbound

DUCX: Decomposing Unfairness in Tool-Using Chest X-ray Agents cites this paper.

DUCX: Decomposing Unfairness in Tool-Using Chest X-ray Agents Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T19:51:05.718724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:51:05.718724Z digest=sha256:fcb07946c0368cafedb9155acf914718988652d3755367180abeec547fc10953

Observation f8d0bc55-c0f5-4958-9db3-140b6651204d · inbound

HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents cites this paper.

HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents Qwen3-VL Technical Report

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T18:36:28.415729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:35:45.900606Z digest=sha256:009e8b092561d64f8e826423f33f6dcece9e4c010277b25b6b9d813d0210394a

Observation 9c1d15f3-2039-474a-8bf9-59d136cf6d44 · inbound

HVR-Met: A Hypothesis-Verification-Replanning Agentic System for Extreme Weather Diagnosis cites this paper.

HVR-Met: A Hypothesis-Verification-Replanning Agentic System for Extreme Weather Diagnosis Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T19:46:43.537373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T19:46:43.537373Z digest=sha256:3e2811e1ee25b9aa392cab3ba050c3fc1bd018a4408e0511bdafd58584a972c3

Observation 63e9e5d5-987d-4a92-98b3-2e9a28d28df9 · inbound

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models cites this paper.

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T19:35:37.630960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:35:37.630960Z digest=sha256:3d500e372243bba0c2f053ccda9839a99fb1ca0ed064142b39c3b1df42846d2b

Observation 8bbd1663-c092-4c95-98bc-67fea50770e0 · inbound

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons cites this paper.

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons Qwen3-VL Technical Report

Reference 124

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:00:12.635655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:59:52.365630Z digest=sha256:7210b3be629b75c9e426363966c98126e0e57dbf904f2b8319a8085871e0b50b

Observation 9c45cfd5-90ea-4e28-aeab-7beb28bdf80a · inbound

Structure-Aware Text Recognition for Ancient Greek Critical Editions cites this paper.

Structure-Aware Text Recognition for Ancient Greek Critical Editions Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T19:18:51.740315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:18:51.740315Z digest=sha256:17df798d870862d2c5c4c4c43d95949f46d341b82625515df22192d185d38fc8

Observation 3138430f-9e50-4162-bc2c-2a31fc4172ee · inbound

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild cites this paper.

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild Qwen3-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T18:54:55.012353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:54:55.012353Z digest=sha256:3d5ac1d2f9de1fb837859f77b901e54f2a9d033f6a6d7569e11e77124e3dadd8

Observation ae58ab0a-0b3f-43bc-aa63-f1328bccf7e7 · inbound

Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training cites this paper.

Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training Qwen3-VL Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:12:35.442378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:12:25.063204Z digest=sha256:fb15abc7975dcdcaf88610a9d4459e50a3f39a3bb28e0e7697a613b98ec66275

Observation 7d0447e6-8406-4d11-b36f-1a22ab8430f8 · inbound

Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks cites this paper.

Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T16:06:14.495880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T16:04:58.851378Z digest=sha256:c54c68658fe9c9a0887a95071ec35c3e18dc1eecb5605b23c76baa05e42d9f97

Observation d676fc0c-0864-4950-91a9-99bdc5bda74a · inbound

InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context cites this paper.

InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T18:47:14.582142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:47:14.582142Z digest=sha256:e3271880c985a7f552fcbb0aebc4eca45b817cfd4776ae1b342ee947e04bc73d

Observation 8a6e3125-d124-4aa2-a667-d273fd0c59b7 · inbound

Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine cites this paper.

Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:36:29.064109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:32:04.537458Z digest=sha256:68f5e0b50ef6a90ab0ff1dc0b35deeaeeb90e123e9eee6767e46285d0a4be6ef

Observation 820e71b2-b28c-409d-ae56-37d90b4b6b16 · inbound

TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images cites this paper.

TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images Qwen3-VL Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:35:55.884850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T14:32:59.720416Z digest=sha256:dc009c2d7e7850c4b9cdf639fed4fb6a69934c04984d4e93770bb20c4e869552

Observation 9933e431-8add-448b-b2d0-e385865d2774 · inbound

Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge cites this paper.

Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T18:40:08.348967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:40:08.348967Z digest=sha256:4fbdcb23869b69fe5bbd4b9e878007c16bbe18c7d9f1c53d140e1b45590f4d8d

Observation 2a75cda4-351e-4679-bacf-f1388c6c5b10 · inbound

ICLR: In-Context Imitation Learning with Visual Reasoning cites this paper.

ICLR: In-Context Imitation Learning with Visual Reasoning Qwen3-VL Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T18:38:30.711742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:38:30.711742Z digest=sha256:8d480363c177164face4f42e0ef2a68ef8fdc421175b3b6566ba8d3675d8b4a9

Observation 5b6eefd0-b323-4a6c-8750-cca73fe3d73b · inbound

SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents cites this paper.

SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T11:16:27.898882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T11:14:47.943242Z digest=sha256:62879777a68a3378ae55d9596c94973552366e759f9a373963dd2cdb659bbca3

Observation a55b1774-8239-4a4d-b69a-581af614b838 · inbound

Logics-Parsing-Omni Technical Report cites this paper.

Logics-Parsing-Omni Technical Report Qwen3-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T13:40:01.762091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T13:37:44.189839Z digest=sha256:a9f3af5d4747b509db48119f8efca461a0cc0171900130b677c107c471652492

Observation 997b43c7-f9b3-4b20-a73a-413775b14ed4 · inbound

TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans cites this paper.

TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans Qwen3-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T18:31:40.087308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:31:40.087308Z digest=sha256:563b0cdd71c41fb53a6fd68b11a3506c5cc7ee3933dc9cb25383b3c0a8626a9f

Observation c835e926-ff75-476b-bb2c-d70d00eff150 · inbound

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing cites this paper.

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing Qwen3-VL Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T18:25:51.149194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:25:51.149194Z digest=sha256:a86e99d3d97533860910108c1916ec0c39af99f887fbd7f44754bf324e42c494

Observation 0409a229-3d47-4f6a-a19b-0ea9b7642dbe · inbound

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next cites this paper.

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T05:50:27.140276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:50:27.140276Z digest=sha256:0f549a2cf310b70a3fb4e92d001a8e24fc3637875a2600ccff8a84a24f6bb7cc