Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2511.21631.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:17:53.872394Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
3
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation fe585dc7-9c0b-4d92-9e1e-a4bd2b0b919e · inbound
FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab1d400d-5dc2-4de9-9379-10855b88bb59 · inbound
SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence Qwen3-VL Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bbd98ed3-fe46-4315-93c2-261e898b1235 · inbound
FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd703d6e-3bec-4583-9014-7a79ec11f3e2 · inbound
Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification? Qwen3-VL Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3e94efc-dcf7-41d6-a470-493466e02992 · inbound
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Qwen3-VL Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cab8f1d-c5f0-4b28-9d7b-118e6deff237 · inbound
VideoGuard: Protecting Video Content from Unauthorized Editing Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfae5e5e-4b17-4f63-898a-c6606e079d79 · inbound
SkillWrapper: Generative Predicate Invention for Task-level Robot Planning Qwen3-VL Technical Report
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05b71cb1-b3c3-4e6b-8538-f15b051ac6e4 · inbound
SkillWrapper: Generative Predicate Invention for Task-level Robot Planning Qwen3-VL Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39d336c0-c180-46fc-b34b-2885586ebb50 · inbound
OneThinker: All-in-one Reasoning Model for Image and Video Qwen3-VL Technical Report
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13e34751-035f-4fb0-9e60-2f3c880588ff · inbound
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f6a00a5-566b-4481-bf30-9f350e0fc805 · inbound
SAM3-I: Segment Anything with Instructions Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 967822dd-61a2-4198-9923-dc5f5a705185 · inbound
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c856b0-d022-46bb-9ac0-9359badea614 · inbound
Mull-Tokens: Modality-Agnostic Latent Thinking Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c1e00cf-e3de-43f1-b1ac-6a77a874954d · inbound
Are vision-language models ready to zero-shot replace supervised classification models in agriculture? Qwen3-VL Technical Report
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16b730f1-8fb9-4e40-9ae2-c3da6938de0b · inbound
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding Qwen3-VL Technical Report
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2543ab7f-c882-41ff-a549-10703700eb72 · inbound
VPTracker: Global Vision-Language Tracking via Visual Prompt Qwen3-VL Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3bcc6ae7-329c-4345-af8c-2fd4eca81b51 · inbound
S1-MMAlign: A Large-Scale, Multi-Disciplinary Dataset for Scientific Figure-Text Understanding Qwen3-VL Technical Report
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4b9d45e-f926-478d-a5e4-eac5d5258435 · inbound
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768ad5cf-84db-4e70-a90b-7928f392fe03 · inbound
BabyVision: Visual Reasoning Beyond Language Qwen3-VL Technical Report
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 357b5972-8863-4e53-870b-80c51b8ac1b2 · inbound
Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1142b4bd-0af4-493e-b6ee-7302f29b8ab7 · inbound
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 223de39f-6967-4380-9bab-bf6031f185b8 · inbound
M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding Qwen3-VL Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 233866a4-1b1a-4365-a3c8-527afeb6e726 · inbound
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding Qwen3-VL Technical Report
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36d7964a-70ea-467d-975f-6f9dc0137e63 · inbound
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b9b5dc81-54ec-4d9e-8b4a-050b10c77247 · inbound
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70df9ede-de91-42c0-a414-7f8b96845227 · inbound
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Qwen3-VL Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80dce210-42de-4b05-a844-e28ab1cab5b6 · inbound
Common to Whom? Regional Cultural Commonsense and LLM Bias in India Qwen3-VL Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d27e6d48-5024-4215-9ff5-7be6ff71a249 · inbound
Scaling medical imaging report generation with multimodal reinforcement learning Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c9571d-0916-4217-badf-cdd4652565bc · inbound
GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents Qwen3-VL Technical Report
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73410ca9-f773-4739-bc2d-a4c91508ac2f · inbound
Focus on What Really Matters in Low-Altitude Governance: A Management-Centric Multi-Modal Benchmark with Implicitly Coordinated Vision-Language Reasoning Framework Qwen3-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b53df7da-44a1-4b47-b9e5-bb76d9e6d7ea · inbound
Advancing Open-source World Models Qwen3-VL Technical Report
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc8683e6-e8bd-4cbc-8c80-c31113ec594e · inbound
CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79ac3723-2a3c-4df7-80b3-aeb7e722c2a1 · inbound
Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09b98cdc-1397-4b82-adef-3997f6b400fc · inbound
TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-Tuning Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f433c2c-a656-463c-9866-010787960c87 · inbound
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models Qwen3-VL Technical Report
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11aa3de7-df12-4be4-9338-347592404301 · inbound
Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Qwen3-VL Technical Report
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c31f05b-b385-42da-a432-c1498d167a03 · inbound
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa75d537-19ca-46fb-bce3-4b07da7aca58 · inbound
Dual Latent Memory for Visual Multi-agent System Qwen3-VL Technical Report
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5943ccf5-4882-4409-bf94-b3a63ff65977 · inbound
CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding Qwen3-VL Technical Report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da3fd933-8924-4d36-93e0-bd608cfddeef · inbound
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning Qwen3-VL Technical Report
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aefbea46-0aa4-40a4-84b7-404cf41b371f · inbound
LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation Qwen3-VL Technical Report
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b798d46-87d2-4ba7-9c74-07b0059f3204 · inbound
Kimi K2.5: Visual Agentic Intelligence Qwen3-VL Technical Report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2d0bec1c-6591-4868-9f89-16e07bbc61a3 · inbound
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4288644-0504-42b7-9442-5a9859579005 · inbound
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation Qwen3-VL Technical Report
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb666695-a8e2-49f6-a2db-7e3bfd543aac · inbound
Decoupling Skeleton and Flesh: Efficient Multimodal Table Reasoning with Disentangled Alignment and Structure-aware Guidance Qwen3-VL Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19b4b2cb-a4cc-4f11-b1b9-f826892306b0 · inbound
Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 166a2c47-85ae-4c04-906c-9eee29fe86a9 · inbound
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b9e2f50-5aff-4032-a80d-7fb7772c10d2 · inbound
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation Qwen3-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6b0e7ea-2f1a-4be9-9163-f1c090d9150d · inbound
World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 990ac962-f728-465b-a11b-9560849c980d · inbound
OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization Qwen3-VL Technical Report
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 469223d4-0eed-45ea-a6ec-f883a1df4cc4 · inbound
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models Qwen3-VL Technical Report
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 495d751b-5049-4590-a6ed-885cc8edab04 · inbound
Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning Qwen3-VL Technical Report
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2fd6f15-800d-4fc6-b488-6859a4873b9e · inbound
Prism: Spectral-Aware Block-Sparse Attention Qwen3-VL Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8e65144-6508-45a3-8d93-4f6280a36fd6 · inbound
Multimodal Learning for Arcing Detection in Pantograph-Catenary Systems Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f615c7e5-4e8c-4516-9c83-bc1fd9a19778 · inbound
Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense Qwen3-VL Technical Report
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11cf1db0-9ef5-4f1f-bdd8-01bb1507c1c8 · inbound
Towards Explainable Industrial Anomaly Detection via Knowledge-Guided Latent Reasoning Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b349e52a-5994-4ac3-97de-6beed00b7849 · inbound
DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04a2e567-54f6-4f53-8fbd-1f3c19c7cee5 · inbound
Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2388e3c1-f114-4da2-bc4e-ac6bc6ea63b6 · inbound
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ed0c257-6fec-4467-b33e-a268a69c9105 · inbound
FAIL: Flow Matching Adversarial Imitation Learning for Image Generation Qwen3-VL Technical Report
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a0c099a-dbb6-4ba8-b96f-722ebc087593 · inbound
CAPTS: Channel-Aware, Preference-Aligned Trigger Selection for Multi-Channel Item-to-Item Retrieval Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a9581ac-5a48-441e-a833-0c63b98d7c91 · inbound
JARVIS: An Evidence-Grounded Retrieval System for Interpretable Deceptive Reviews Adjudication Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f705114d-687f-4cf4-96ac-5af4d77ad5a5 · inbound
HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 792637b8-b22c-4e0e-93d8-fe3dff08b3d3 · inbound
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4382789-4f06-4374-8035-9c43e06223f3 · inbound
SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37eea6c8-c8d5-484a-95ef-ad0a1a6522ae · inbound
Exploring a Multimodal Chatbot as a Facilitator in Therapeutic Art Activity Qwen3-VL Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f52e3b97-deae-46e5-a889-34adbc6bf5cf · inbound
ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa6a7a59-a574-48e1-acbc-c67ed29039cc · inbound
The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems Qwen3-VL Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69d5d5c3-7316-4d86-b69b-d048f96e6ada · inbound
DODO: Discrete OCR Diffusion Models Qwen3-VL Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66a0e07c-dd83-4b94-9929-b405c103bbd6 · inbound
JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Qwen3-VL Technical Report
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2529d4c-c5b3-4386-af60-45f20360ac71 · inbound
VLANeXt: Recipes for Building Strong VLA Models Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a1d9cd5-0a6d-4593-9242-8ac4b7c4089f · inbound
FUSAR-GPT : A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06d9cbd1-2ac9-4d59-a461-e906f25cf304 · inbound
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43ab2a23-547a-4dd6-b6d7-5893326c027a · inbound
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c42f98b-5a6c-4c41-b741-1c377628af9f · inbound
PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning Qwen3-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2e33ca8-ac2a-465d-8ccf-3a79d0f16854 · inbound
Efficient Scaling of LLM Training with Flexible Context Parallelism Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7567fb1-0730-4be7-9d34-20a5cd28eb33 · inbound
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL Qwen3-VL Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d4e5937-16ab-44ae-b78b-56bf087e50c2 · inbound
Imagination Helps Visual Reasoning, But Not Yet in Latent Space Qwen3-VL Technical Report
Reference 1971
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20d0d4a4-2c64-4f7d-8f5d-5fd0ad7a898a · inbound
PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 788c4ca7-7e55-40db-bbe8-21b1414254f0 · inbound
Can Agents Distinguish Visually Hard-to-Separate Diseases in a Zero-Shot Setting? A Pilot Study Qwen3-VL Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fc72b8e-0fd1-4fe6-a8f1-f87eb540b08b · inbound
CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era Qwen3-VL Technical Report
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8376e032-b09c-4640-abc4-20f5a17c8be1 · inbound
DUCX: Decomposing Unfairness in Tool-Using Chest X-ray Agents Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8d0bc55-c0f5-4958-9db3-140b6651204d · inbound
HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents Qwen3-VL Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c1d15f3-2039-474a-8bf9-59d136cf6d44 · inbound
HVR-Met: A Hypothesis-Verification-Replanning Agentic System for Extreme Weather Diagnosis Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e9e5d5-987d-4a92-98b3-2e9a28d28df9 · inbound
Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bbd1663-c092-4c95-98bc-67fea50770e0 · inbound
Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons Qwen3-VL Technical Report
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c45cfd5-90ea-4e28-aeab-7beb28bdf80a · inbound
Structure-Aware Text Recognition for Ancient Greek Critical Editions Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3138430f-9e50-4162-bc2c-2a31fc4172ee · inbound
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild Qwen3-VL Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae58ab0a-0b3f-43bc-aa63-f1328bccf7e7 · inbound
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training Qwen3-VL Technical Report
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d0447e6-8406-4d11-b36f-1a22ab8430f8 · inbound
Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d676fc0c-0864-4950-91a9-99bdc5bda74a · inbound
InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a6e3125-d124-4aa2-a667-d273fd0c59b7 · inbound
Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 820e71b2-b28c-409d-ae56-37d90b4b6b16 · inbound
TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images Qwen3-VL Technical Report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9933e431-8add-448b-b2d0-e385865d2774 · inbound
Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a75cda4-351e-4679-bacf-f1388c6c5b10 · inbound
ICLR: In-Context Imitation Learning with Visual Reasoning Qwen3-VL Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b6eefd0-b323-4a6c-8750-cca73fe3d73b · inbound
SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a55b1774-8239-4a4d-b69a-581af614b838 · inbound
Logics-Parsing-Omni Technical Report Qwen3-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 997b43c7-f9b3-4b20-a73a-413775b14ed4 · inbound
TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans Qwen3-VL Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c835e926-ff75-476b-bb2c-d70d00eff150 · inbound
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing Qwen3-VL Technical Report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0409a229-3d47-4f6a-a19b-0ea9b7642dbe · inbound
EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.