Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:17:33.970322Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2608.12209.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:17:33.970322Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
83 of 83 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e27eee67-71f9-4100-8bcb-157b5ef51e24 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LLaVA-OneVision: Easy Visual Task Transfer
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73daf702-45dc-4ba5-9b8c-c00c0f426859 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 581d4fef-c86b-4679-892c-e199b0988514 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 690036a4-7e49-4628-8a91-54aefacf29ac · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Visual Instruction Tuning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 073cf827-4cbb-4837-ab8f-b899bd3e0298 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Learning transferable visual models from natural language supervision
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66e1153a-e3fe-4ab7-a84d-c0a1ec0e8aba · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48223edc-b902-4f01-9d60-bb73da0595d2 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction An empirical anal- ysis on spatial reasoning capabilities of large multimodal models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dcd2796-7c6f-470c-b83f-a84dce81d3b9 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Why are visually-grounded language models bad at image classification? Advances in Neural Information Processing Systems, 37:51727–51753, 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 56769107-c26a-4d69-8ea9-0a62b49efece · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Question aware vision transformer for multimodal reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d29ae9be-8ff3-4050-9b9a-60a890622529 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8111b51c-97ea-491e-bf8a-5ac7e5ac0c30 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Show-o: One single transformer to unify multimodal understanding and generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1b064072-c6b8-46a6-a37c-2e788fcd5d72 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Transfusion: Predict the next token and diffuse images with one multi-modal model
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5357a682-5d79-41d5-9b67-aa6be7d681e1 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Emerging Properties in Unified Multimodal Pretraining
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f0b369f-d2e4-4afc-9563-01b08ee21b37 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mammothmoda2: A unified ar-diffusion framework for multimodal understanding and generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 569f7263-26fc-45b3-8025-46d8129b94ac · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Lance: Unified Multimodal Modeling by Multi-Task Synergy
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63003f27-2aa1-4654-9c49-363a64cad2c1 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Multimodal learning with next-token prediction for large multimodal models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84770594-92dd-4d11-8867-6f6eccdd453c · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Janus: Decoupling visual encoding for unified multimodal understanding and generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da4b6b3-7fb0-4f5a-befe-3a99babd2c83 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mmada: Multimodal large diffusion language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 75ccea02-af2c-4f5b-a6b5-5904a0f2c0b2 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Longcat-next: Lexicalizing modalities as discrete tokens
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64c0448f-797f-402e-b382-d4c47a23a458 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Next-embedding prediction makes strong vision learners
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c21e1c-d592-4b6d-ab4d-9da81a50ac3f · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Unihetero: Could generation enhance understanding for vision-language-model at large data scale? arXiv preprint arXiv:2512.23512, 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 942b4361-6ee5-44ce-ba8e-a87358e57908 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0c30f67c-517b-4927-81ba-45a7ce43b7b8 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Uni-x: Mitigating modality conflict with a two-end-separated architecture for unified multimodal models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6985abff-97f7-4ef9-9bb6-0f31241ac41d · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1dfc456c-42b0-4591-ae33-cf1c8703e850 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Qwen3 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ef9f9ac-56f1-4129-9351-ddfa7291dc70 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Adam: A Method for Stochastic Optimization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d32e69c6-17cd-44c4-abf9-8a2e7f63c81b · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Zero: Memory optimizations toward training trillion parameter models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 41ae2848-0ec8-47b1-90a8-5f6a04382004 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mme: A comprehensive evaluation benchmark for multimodal large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 183cc7c9-3aab-49f4-9be4-b059d9ce5cfd · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f66b54ef-2f8e-4fa1-b366-dae59ce38d49 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Blink: Multimodal large language models can see but not perceive
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1930387-009c-43db-90c4-b86a88be5003 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Charxiv: Charting gaps in realistic chart understanding in multimodal llms
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39e60b28-eb04-4718-9a89-3abff8768d58 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a74c4a37-3aca-4d5d-94d0-28c8a50ffbdb · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Measuring multimodal mathematical reasoning with math-vision dataset
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320d399e-1b4b-4016-ac2b-82b8f1efb5be · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f06d5710-f9c5-4b5f-a98d-e20707970c57 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02f09eb4-7996-4dfb-af6e-2b535747e291 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Visulogic: A benchmark for evalu- ating visual reasoning in multi-modal large language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 26c84fe8-d53e-442a-a4e8-85a94ad09106 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Teaching CLIP to Count to Ten
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1804ee86-6b67-4b94-947f-7787318f7452 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction PaliGemma: A versatile 3B VLM for transfer
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation def81ece-2147-47a0-bae5-2907a8f1f1a4 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 40f3269a-697b-44c9-a810-275fa14c1ae3 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 314f5c6e-7cb5-418f-b2df-3c51bcb1741d · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cbd7d36-a227-4539-97e0-6a8033974b9a · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd0019bf-cc06-4fde-a865-57ac6f7ea711 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Thyme: Think beyond images
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 523eabbf-f825-472c-b160-5d5f89a617c0 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Improved baselines with visual instruction tuning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83854f49-131d-44c5-bbf6-29f72bce05df · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Qwen2.5-VL Technical Report, February 2025
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e64aa5f8-a4cc-45e2-8d69-c4de94a0fee5 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Metamorph: Multimodal understanding and generation via instruction tuning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6a154796-6cdf-4c11-a87a-e2d00e90aabb · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c858fa4b-9750-475f-82ff-832482d94d18 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Show-o2: Improved native unified multimodal models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation afd85e1c-5dd4-4d8f-8ef3-6c2d91e10f02 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Cheers: Decoupling patch details from semantic representations enables unified multimodal comprehension and generation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50dfeed4-2f57-42eb-a738-4a3adbb969c6 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48670c1e-7796-4337-84f5-53a46388c1c1 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Imagenet: A large-scale hierarchical image database
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ecf47df-c7d9-4053-b9f6-444cd7a69bcc · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Gen- eration and comprehension of unambiguous object descriptions
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3091c69c-ee9c-4e96-a108-d3f6d8b20ad8 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Referitgame: Referring to objects 23 in photographs of natural scenes
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 26584f7a-e5c5-430d-8e1e-9bb1d93aa0af · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c70cbff7-5db0-4c6d-b223-3480478740f4 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 905e03c1-049e-4f64-a7bc-241d4ddf8505 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Flamingo: a visual language model for few-shot learning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9268cac0-a2b7-4868-b137-d12d81434bcb · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 84d506e9-63f8-4616-8f32-61a9a384bdfe · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 806d6f5a-3a59-438d-b1ec-c1e99e70adca · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd887cbb-d5da-4d59-9661-9de455213db5 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction gpt-5-system-card, 2025
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aacebb07-2ed8-446c-91cb-8ecf55737892 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction URL https://blog
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 53b7ce97-9450-476f-be32-e5a9b92933bf · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Gemini 3 flash: frontier intelligence built for speed, 2025
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fd929ecf-1804-4e14-a388-44982ba0013b · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction VGR: Visual grounded reasoning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e58da99f-15e6-4f75-ae40-28dbe6609a75 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Explain Before You Answer: A Survey on Compositional Visual Reasoning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f707088-def0-4a1a-9541-a31fde45abf9 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac6385c7-46e4-4ff2-ae91-b502cbe09e08 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Video-xl: Extra-long vision language model for hour-scale video understanding
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a929b0c3-f6b3-4cd3-b19b-0091995455d4 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Videochat-flash: Hierarchical compression for long-context video modeling
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a4a286f7-2de5-4dce-8344-2dd178faa383 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Reconstructive visual instruction tuning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e779dc83-cfa0-4390-b4fb-f2934a771cda · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Autoregressive semantic visual reconstruction helps vlms understand better
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0dae9be8-18b4-462f-aedf-d247e85fa217 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Generation enhances understanding in unified multimodal models via multi-representation generation
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b8ab97a0-4659-420f-bbd2-40479f921270 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f820c2a0-3031-4f0e-87b8-b377bd0734ae · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LMFusion: Adapting pretrained language models for multimodal generation
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 95c572ab-905d-4dd6-bc6f-629fc15d2719 · outbound
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3b16e616-1bc7-4869-aa16-07fc00ab01c8 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction All figures share the same two-shape overlapping composition with consistent contour and inner-line connectivity
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2c55ad56-d9a3-48ba-9d36-a7298d51ff94 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Option A is a triangle pair with a mismatched shape; Option B is a triangle pair with inconsistent line count; Option C is a quadrilateral but with a different stacking style
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 64359b35-023e-4c33-8ad6-d23ce8ab9a83 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction A bipartite graph is a graph whose vertices can be divided into two disjoint sets such that every edge connects a vertex in one set to a vertex in the other set
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ace91903-90da-48b4-8a18-a0dfcaaadbf3 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction If there are no odd-length cycles, the graph is bipartite
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fed92e39-9f71-4eb9-b408-a516c7389120 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction This is because we can color one set of vertices with one color and the other set with a different color, ensuring that no two adjacent vertices share the same color
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7a933e72-25a3-440a-84da-7f5ef6be6931 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction </think> The final answer is2
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1cb9d5ad-aa2a-48ce-adbc-690eaec86790 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction There are two visible wooden poles supporting the tree
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 349c59f0-13ea-4c7c-ab7d-5e1d5209b7af · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction There are no additional wooden poles visible in the background or elsewhere in the image
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation efe688c9-dee5-47cf-8bef-abf21d9a764e · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Given this analysis, the correct answer is: **C
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ff1816c0-e15b-4c2c-99e6-b95510212187 · outbound
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Unresolved cited work
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
No inbound Pith citation observations are available.