Pith. sign in

Paper Citation Record · LEDGER

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

As of 16 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2608.12209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.12209 v1

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:17:33.970322Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

83 of 83 outbound references displayed

  • verified exact2
  • verified fuzzy39
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e27eee67-71f9-4100-8bcb-157b5ef51e24 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LLaVA-OneVision: Easy Visual Task Transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.494370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.494370Z digest=sha256:1928b6426781a0d4ad040f0bfdc1f9b70b4f6ec128414243567dd89ff768ceba

Observation 73daf702-45dc-4ba5-9b8c-c00c0f426859 · outbound

This paper cites Qwen3-VL Technical Report.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.499723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.499723Z digest=sha256:fdbef77b7d3d6657b99e82d1912d62f6a1d20612ac9ca72f60798d729f917292

Observation 581d4fef-c86b-4679-892c-e199b0988514 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.503957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.503957Z digest=sha256:bd2967194ca44028366877d26f2e270dbf2e6188c7862930900a229035848e38

Observation 690036a4-7e49-4628-8a91-54aefacf29ac · outbound

This paper cites Visual Instruction Tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Visual Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.508431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.508431Z digest=sha256:929eed1abb6e1817e34b73a8a95cf33ddfa5befb5c86b7efa4e780e14d3caad8

Observation 073cf827-4cbb-4837-ab8f-b899bd3e0298 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Learning transferable visual models from natural language supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.513081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.513081Z digest=sha256:9be462ff458f9f5d0520c1aa0d85a014f379efb8c1ba7e5ea9411b0c8211f020

Observation 66e1153a-e3fe-4ab7-a84d-c0a1ec0e8aba · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.517459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.517459Z digest=sha256:68deb72684992e160c2dbdd722b5dd9b1fa0639324452482f8783faf8b87fd77

Observation 48223edc-b902-4f01-9d60-bb73da0595d2 · outbound

This paper cites An empirical anal- ysis on spatial reasoning capabilities of large multimodal models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction An empirical anal- ysis on spatial reasoning capabilities of large multimodal models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.522320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.522320Z digest=sha256:ae82db2e7d557d66ae048dcdbe89b2bf922302b61903416c6b9753a9e7e7fdd8

Observation 9dcd2796-7c6f-470c-b83f-a84dce81d3b9 · outbound

This paper cites Why are visually-grounded language models bad at image classification? Advances in Neural Information Processing Systems, 37:51727–51753, 2024.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Why are visually-grounded language models bad at image classification? Advances in Neural Information Processing Systems, 37:51727–51753, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.537516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.526218Z digest=sha256:06fb713b6ba834dafd6f15e528978791e39ab163a5bf4a7f037e93ee49c2ff45

Observation 56769107-c26a-4d69-8ea9-0a62b49efece · outbound

This paper cites Question aware vision transformer for multimodal reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Question aware vision transformer for multimodal reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.521970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.529994Z digest=sha256:38f3e78a1a81c3e38f809f7b4995d1e19770f8bb7cb9e082cf2d77ed4af61292

Observation d29ae9be-8ff3-4050-9b9a-60a890622529 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.534077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.534077Z digest=sha256:2b772c6a2b93120e6b168f94d7b98ae433d01b9627519a4a0e489e73407f7fde

Observation 8111b51c-97ea-491e-bf8a-5ac7e5ac0c30 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Show-o: One single transformer to unify multimodal understanding and generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.505930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.539092Z digest=sha256:3b2c3d1948f96a97d73f04b10d1df60fd8e988def83c8a42451988bf527e1d55

Observation 1b064072-c6b8-46a6-a37c-2e788fcd5d72 · outbound

This paper cites Transfusion: Predict the next token and diffuse images with one multi-modal model.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Transfusion: Predict the next token and diffuse images with one multi-modal model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.490249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.544262Z digest=sha256:b10b0fdd73519ccdc563086cd3121d77cc55e32800992e7819cbab9d4a7aa448

Observation 5357a682-5d79-41d5-9b67-aa6be7d681e1 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Emerging Properties in Unified Multimodal Pretraining

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.553581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.553581Z digest=sha256:0935d1b98d9067dc1e7f99c3c13239648a051ffabe0e10e402b625eb94260381

Observation 3f0b369f-d2e4-4afc-9563-01b08ee21b37 · outbound

This paper cites Mammothmoda2: A unified ar-diffusion framework for multimodal understanding and generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mammothmoda2: A unified ar-diffusion framework for multimodal understanding and generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.558435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.558435Z digest=sha256:4d2c80f01c78a6216b548b5bf5e4bda03c5f026917db678f587f76383de24652

Observation 569f7263-26fc-45b3-8025-46d8129b94ac · outbound

This paper cites Lance: Unified Multimodal Modeling by Multi-Task Synergy.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Lance: Unified Multimodal Modeling by Multi-Task Synergy

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.562911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.562911Z digest=sha256:9d9241da4187a8f9b19b569796b68e7deb57f00123ec8ce95e32fe31b1b04de6

Observation 63003f27-2aa1-4654-9c49-363a64cad2c1 · outbound

This paper cites Multimodal learning with next-token prediction for large multimodal models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Multimodal learning with next-token prediction for large multimodal models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.567839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.567839Z digest=sha256:b5c88b3b0f2121c164b2b42dd5782f8f6f1d0c1cd657f4606a9bc4e2b84eaa93

Observation 84770594-92dd-4d11-8867-6f6eccdd453c · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.572476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.572476Z digest=sha256:8ca7f4dc898e203833925ecafb283921a47196e8cbd998ab4e468b3291b2b3b3

Observation 5da4b6b3-7fb0-4f5a-befe-3a99babd2c83 · outbound

This paper cites Mmada: Multimodal large diffusion language models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mmada: Multimodal large diffusion language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.447949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.577175Z digest=sha256:ed82a351785fab3b6a6608957e64fa0335741e6cb2376795c8ed6296fc952b6b

Observation 75ccea02-af2c-4f5b-a6b5-5904a0f2c0b2 · outbound

This paper cites Longcat-next: Lexicalizing modalities as discrete tokens.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Longcat-next: Lexicalizing modalities as discrete tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.581554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.581554Z digest=sha256:e273dd951a487d9b83ccad67acfc4bc911b4f85fc0f2f4ce964e5c428d4d6118

Observation 64c0448f-797f-402e-b382-d4c47a23a458 · outbound

This paper cites Next-embedding prediction makes strong vision learners.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Next-embedding prediction makes strong vision learners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.585968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.585968Z digest=sha256:a84132d600c01790873a80116096f2b4fecd27369edf9d614b3c77b34f9a80f3

Observation 37c21e1c-d592-4b6d-ab4d-9da81a50ac3f · outbound

This paper cites Unihetero: Could generation enhance understanding for vision-language-model at large data scale? arXiv preprint arXiv:2512.23512, 2025.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Unihetero: Could generation enhance understanding for vision-language-model at large data scale? arXiv preprint arXiv:2512.23512, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.590538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.590538Z digest=sha256:493d28edc5c00ba4eae2c804e36bfb3e20ec95b1bc36fad3f27a2d3f9e000798

Observation 942b4361-6ee5-44ce-ba8e-a87358e57908 · outbound

This paper cites UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:17:34.447253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.594876Z digest=sha256:a82347ae50a07c2cc7624386e8c2626a9c887bd1fdf24a3034fb6a21bb263c44

Observation 0c30f67c-517b-4927-81ba-45a7ce43b7b8 · outbound

This paper cites Uni-x: Mitigating modality conflict with a two-end-separated architecture for unified multimodal models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Uni-x: Mitigating modality conflict with a two-end-separated architecture for unified multimodal models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.599748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.599748Z digest=sha256:2d01beb0228f8b2ce8a7261f59b52f40127f8d9b753e521319423e17b36c06e4

Observation 6985abff-97f7-4ef9-9bb6-0f31241ac41d · outbound

This paper cites Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.431635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.604319Z digest=sha256:dabb7f7fa9d5fd67fb6541c53443cb35d84a3e9c91e217e7b6f17bd9f118b130

Observation 1dfc456c-42b0-4591-ae33-cf1c8703e850 · outbound

This paper cites Qwen3 Technical Report.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Qwen3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.608967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.608967Z digest=sha256:089bd673304d58792e0888523b2032f14ac918c43aa97a3dc1f7e17afd617bc6

Observation 7ef9f9ac-56f1-4129-9351-ddfa7291dc70 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Adam: A Method for Stochastic Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.613869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.613869Z digest=sha256:77dd3f551a3290bab96225d4ab1019fcd03ac1930e8a63804e9b54d727e0a314

Observation d32e69c6-17cd-44c4-abf9-8a2e7f63c81b · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Zero: Memory optimizations toward training trillion parameter models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.415500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.618475Z digest=sha256:e433e71b6bc4d09db744b01ed6ee8fc2439e61321e2433dadf9ec12a4914e289

Observation 41ae2848-0ec8-47b1-90a8-5f6a04382004 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.399745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.622968Z digest=sha256:26c0db3ba8710052838cf2fa7d87d4ce2b988f6422d01b8e4b0b520ce370286c

Observation 183cc7c9-3aab-49f4-9be4-b059d9ce5cfd · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.384039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.627474Z digest=sha256:9d8be9c0d39b2561e078e981ca9e637987bf5a7309540bd43aa786dd2608108b

Observation f66b54ef-2f8e-4fa1-b366-dae59ce38d49 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Blink: Multimodal large language models can see but not perceive

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.632306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.632306Z digest=sha256:2854cecae3bb04f1e541aadaf4d5ec8855a6fa7127c00c27e556072f93f2d63b

Observation d1930387-009c-43db-90c4-b86a88be5003 · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.636732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.636732Z digest=sha256:80a34f7801061bfd2c5372daf85f4963493ce3cb61c6e99e497a319b992dd3dd

Observation 39e60b28-eb04-4718-9a89-3abff8768d58 · outbound

This paper cites Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.641183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.641183Z digest=sha256:b2bc8e5dbc2a9b55f653cd4207d5a99238e49376235fbceb4b35601b15c9d09e

Observation a74c4a37-3aca-4d5d-94d0-28c8a50ffbdb · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Measuring multimodal mathematical reasoning with math-vision dataset

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.645477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.645477Z digest=sha256:d872fa61b16db84e8017f018875c1a3027f6c96da40b17a7ba3486d8f6c2f87b

Observation 320d399e-1b4b-4016-ac2b-82b8f1efb5be · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.326990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.650019Z digest=sha256:a73baaed5d7df9f15a4d2f3feecc6a433a1fc175bd9b84728c28c63102d8c0ca

Observation f06d5710-f9c5-4b5f-a98d-e20707970c57 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.654244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.654244Z digest=sha256:b863b3302020555d9750af39f1faf78583b1f0071c1e5047cf6eb4a037b49d5b

Observation 02f09eb4-7996-4dfb-af6e-2b535747e291 · outbound

This paper cites Visulogic: A benchmark for evalu- ating visual reasoning in multi-modal large language models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Visulogic: A benchmark for evalu- ating visual reasoning in multi-modal large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.311942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.658494Z digest=sha256:49a74030910516a87250706a1313229f752ada282e8f88ac27c46ddc503d36ed

Observation 26c84fe8-d53e-442a-a4e8-85a94ad09106 · outbound

This paper cites Teaching CLIP to Count to Ten.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Teaching CLIP to Count to Ten

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.662073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.662073Z digest=sha256:8c202e741c34d189d7c377c1f63c0582aa9835ef7bde6b638daf27fb427ef947

Observation 1804ee86-6b67-4b94-947f-7787318f7452 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction PaliGemma: A versatile 3B VLM for transfer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.665902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.665902Z digest=sha256:fed5a259165c719f3654d74fd07355d4bffa6f8e554c56a85ae1354eb5547050

Observation def81ece-2147-47a0-bae5-2907a8f1f1a4 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.296526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.670137Z digest=sha256:da8e9a2c9688ef15fc5ca406dcfd73b1226f93d8a2183a564cbae6a180dab6db

Observation 40f3269a-697b-44c9-a810-275fa14c1ae3 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.280336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.674031Z digest=sha256:3102f7f583b4b187d39a015ac533b4d49da7f3267813a26bc61f1756d56a8f51

Observation 314f5c6e-7cb5-418f-b2df-3c51bcb1741d · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.677995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.677995Z digest=sha256:d3cb2c8dfb90b3c8429bee1fe2ddd147d277f7907e7bea591173a0c61c2d4454

Observation 2cbd7d36-a227-4539-97e0-6a8033974b9a · outbound

This paper cites VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.682022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.682022Z digest=sha256:6d812ccfc4d9c011acbefcfc5686d5addae2d187ae322190d03f0d7441ce26c0

Observation cd0019bf-cc06-4fde-a865-57ac6f7ea711 · outbound

This paper cites Thyme: Think beyond images.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Thyme: Think beyond images

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.254442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.686813Z digest=sha256:ff53bbb84846d88ec3985ca3a203f14a6820886af2b2ddfa9e415d4d67d9ff59

Observation 523eabbf-f825-472c-b160-5d5f89a617c0 · outbound

This paper cites Improved baselines with visual instruction tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Improved baselines with visual instruction tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.691642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.691642Z digest=sha256:fa32463d801dc17a12c48d7bb64658f3a424f0b75274897b48cb07648b6e4cac

Observation 83854f49-131d-44c5-bbf6-29f72bce05df · outbound

This paper cites Qwen2.5-VL Technical Report, February 2025.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Qwen2.5-VL Technical Report, February 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.229128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.695959Z digest=sha256:5b09a87494fa01aa0b53daa114031c1a0fb91d6a644903fc3bf7278f5eb2939c

Observation e64aa5f8-a4cc-45e2-8d69-c4de94a0fee5 · outbound

This paper cites Metamorph: Multimodal understanding and generation via instruction tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Metamorph: Multimodal understanding and generation via instruction tuning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.212857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.700366Z digest=sha256:2c2af19e66e13f88308c0124a5c7acec0a481e67c6e46d5ac2264319e1caff4c

Observation 6a154796-6cdf-4c11-a87a-e2d00e90aabb · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.704691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.704691Z digest=sha256:d5fba76e15c7bb15be7e65a1b0966e0984a13bebffaf7c48adefd7878bfd5e2a

Observation c858fa4b-9750-475f-82ff-832482d94d18 · outbound

This paper cites Show-o2: Improved native unified multimodal models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Show-o2: Improved native unified multimodal models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.196673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.709277Z digest=sha256:e77d278204fe5c6e15a93e9a90ef0898732e0cd89bd7926c6642a9eb3efd7b97

Observation afd85e1c-5dd4-4d8f-8ef3-6c2d91e10f02 · outbound

This paper cites Cheers: Decoupling patch details from semantic representations enables unified multimodal comprehension and generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Cheers: Decoupling patch details from semantic representations enables unified multimodal comprehension and generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.713410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.713410Z digest=sha256:fa5169d69e95e0a552e4675c7fc4739290187262897131db14bb3aab9e555b8a

Observation 50dfeed4-2f57-42eb-a738-4a3adbb969c6 · outbound

This paper cites Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.718062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.718062Z digest=sha256:71a86ac8abe7c6b1140cf3dcebbb088e91d48df2647009641d6137e56c5f002c

Observation 48670c1e-7796-4337-84f5-53a46388c1c1 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Imagenet: A large-scale hierarchical image database

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.722979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.722979Z digest=sha256:3ef8644fef80290ddad587a11c9f804d04b3aa47e803ad321df09c3e914bd004

Observation 5ecf47df-c7d9-4053-b9f6-444cd7a69bcc · outbound

This paper cites Gen- eration and comprehension of unambiguous object descriptions.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Gen- eration and comprehension of unambiguous object descriptions

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.161417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.727985Z digest=sha256:2f5e9a9bf4b3f20790fa8354c5934ccfa0f4b58492f47e3ad13acac72a37ddf8

Observation 3091c69c-ee9c-4e96-a108-d3f6d8b20ad8 · outbound

This paper cites Referitgame: Referring to objects 23 in photographs of natural scenes.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Referitgame: Referring to objects 23 in photographs of natural scenes

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.145538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.732783Z digest=sha256:2f7ee47a0b0fccf4cc6e224f6fc21159474245067b9bcfebc67fbab464c169cc

Observation 26584f7a-e5c5-430d-8e1e-9bb1d93aa0af · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.737344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.737344Z digest=sha256:39b0ceb6407307a19bc4824e75fe0013ce36ebbda53427560ccca0c284f95ef4

Observation c70cbff7-5db0-4c6d-b223-3480478740f4 · outbound

This paper cites an unresolved cited work.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.742241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.742241Z digest=sha256:b3c4faa8a679569e3fc9e7c924a2d63f4a6ae1c021e9532bd7c179f1cb1f05d7

Observation 905e03c1-049e-4f64-a7bc-241d4ddf8505 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Flamingo: a visual language model for few-shot learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.747386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.747386Z digest=sha256:d3907ab575136cb48e27d1d0e4570698d4437c3e9fb7d391de156cebfadfa170

Observation 9268cac0-a2b7-4868-b137-d12d81434bcb · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.111944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.752571Z digest=sha256:ba5750208474b3a3577146aae7da16b192765a46dd8bf3dc8b56cd03ab7cd57b

Observation 84d506e9-63f8-4616-8f32-61a9a384bdfe · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.758009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.758009Z digest=sha256:83e794885119e1b495ba55c49b63b2bbf0f95a747371884bbec5712f57e2297c

Observation 806d6f5a-3a59-438d-b1ec-c1e99e70adca · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.762555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.762555Z digest=sha256:76367c58d63b098940a2e9e7ef5b6c3c3ddf621c26837443e3bc2a82f86a9f84

Observation bd887cbb-d5da-4d59-9661-9de455213db5 · outbound

This paper cites gpt-5-system-card, 2025.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction gpt-5-system-card, 2025

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.097932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.767368Z digest=sha256:c9474740e6bc16a6394beb060c051b2bf5a6c715549552d3c36cc411db5d4ae2

Observation aacebb07-2ed8-446c-91cb-8ecf55737892 · outbound

This paper cites URL https://blog.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction URL https://blog

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.082321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.771718Z digest=sha256:f8d1e3a5f37aa37a79993f1828675a550532d4b251ebee536884e2642dbd3983

Observation 53b7ce97-9450-476f-be32-e5a9b92933bf · outbound

This paper cites Gemini 3 flash: frontier intelligence built for speed, 2025.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Gemini 3 flash: frontier intelligence built for speed, 2025

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.067638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.776922Z digest=sha256:b1885a8c8d08082b36475edd51c89896ccec02ea092fd3026a15673341f48224

Observation fd929ecf-1804-4e14-a388-44982ba0013b · outbound

This paper cites VGR: Visual grounded reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction VGR: Visual grounded reasoning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.052990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.781312Z digest=sha256:dbbee36c39ee8a5a7470aad2c23a4f54c97ea1492cca7ae1b8bc73c05bb26a52

Observation e58da99f-15e6-4f75-ae40-28dbe6609a75 · outbound

This paper cites Explain Before You Answer: A Survey on Compositional Visual Reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.785540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.785540Z digest=sha256:aab0532014c413b0145f8ba7828fa5d1411556da6b3a2238d04b1c1a3aad0cd9

Observation 8f707088-def0-4a1a-9541-a31fde45abf9 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.789852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.789852Z digest=sha256:353c3e4b0fc19bf9fa12c1e17333c5db5c40afc750a3471e5be6132a8cb15884

Observation ac6385c7-46e4-4ff2-ae91-b502cbe09e08 · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Video-xl: Extra-long vision language model for hour-scale video understanding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.036455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.892662Z digest=sha256:c59566d7e1b35fbe28fd10405d45c3e13af3cdc432907ff86e61e71bf2a2ef96

Observation a929b0c3-f6b3-4cd3-b19b-0091995455d4 · outbound

This paper cites Videochat-flash: Hierarchical compression for long-context video modeling.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Videochat-flash: Hierarchical compression for long-context video modeling

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.020711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.897426Z digest=sha256:6bd0bb6d3b3a3edc20b6d4001324449829cf01a0e019d5ed8523a0836bfc7143

Observation a4a286f7-2de5-4dce-8344-2dd178faa383 · outbound

This paper cites Reconstructive visual instruction tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Reconstructive visual instruction tuning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.004875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.902010Z digest=sha256:899b286e7df74dca0db118356fcce00aaef82f10e156359aaa46ef2d3254a77d

Observation e779dc83-cfa0-4390-b4fb-f2934a771cda · outbound

This paper cites Autoregressive semantic visual reconstruction helps vlms understand better.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Autoregressive semantic visual reconstruction helps vlms understand better

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.989707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.906676Z digest=sha256:1d1f114edb5b5c20206304d2cc8759b83fa798b13f0165c04a93a86e335d42b0

Observation 0dae9be8-18b4-462f-aedf-d247e85fa217 · outbound

This paper cites Generation enhances understanding in unified multimodal models via multi-representation generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Generation enhances understanding in unified multimodal models via multi-representation generation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.975035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.911508Z digest=sha256:b21bb498fac0eaa7c8fc59e072456eee78f84fc130280010071b4829fa6c0cae

Observation b8ab97a0-4659-420f-bbd2-40479f921270 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.916989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.916989Z digest=sha256:ed249933d6af8579b83e7722cec4a526eaa04cf5388c64c6020e2658ac5c0b43

Observation f820c2a0-3031-4f0e-87b8-b377bd0734ae · outbound

This paper cites LMFusion: Adapting pretrained language models for multimodal generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LMFusion: Adapting pretrained language models for multimodal generation

Reference 72

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:17:34.087176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.921910Z digest=sha256:e4e45f13ad5a3efdcf0b6e18eb9b7bf807b9d4202bba302ac48b61c910d3ead2

Observation 95c572ab-905d-4dd6-bc6f-629fc15d2719 · outbound

This paper cites - Row 1, Fig.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction - Row 1, Fig

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.960789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.926801Z digest=sha256:69b9edf6a191d0fd54541e53a55b92809cd9595137fa1bbaf15677a1dbcb869a

Observation 3b16e616-1bc7-4869-aa16-07fc00ab01c8 · outbound

This paper cites All figures share the same two-shape overlapping composition with consistent contour and inner-line connectivity.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction All figures share the same two-shape overlapping composition with consistent contour and inner-line connectivity

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.947379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.931751Z digest=sha256:d033dd36b237a6a423e41802ff5dd73b9e7f7165edc38cf2ba7b248f7cbb9163

Observation 2c55ad56-d9a3-48ba-9d36-a7298d51ff94 · outbound

This paper cites Option A is a triangle pair with a mismatched shape; Option B is a triangle pair with inconsistent line count; Option C is a quadrilateral but with a different stacking style.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Option A is a triangle pair with a mismatched shape; Option B is a triangle pair with inconsistent line count; Option C is a quadrilateral but with a different stacking style

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.931691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.935730Z digest=sha256:75c0c868be4618fa1d689fb99ee1fc3166fd111a1d52586bb87fe832ae7d90d5

Observation 64359b35-023e-4c33-8ad6-d23ce8ab9a83 · outbound

This paper cites A bipartite graph is a graph whose vertices can be divided into two disjoint sets such that every edge connects a vertex in one set to a vertex in the other set.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction A bipartite graph is a graph whose vertices can be divided into two disjoint sets such that every edge connects a vertex in one set to a vertex in the other set

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.915799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.940312Z digest=sha256:900350965612194cbb0336a11e1a92ac0e6f060fd83e0b561ac81fbabe4d448d

Observation ace91903-90da-48b4-8a18-a0dfcaaadbf3 · outbound

This paper cites If there are no odd-length cycles, the graph is bipartite.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction If there are no odd-length cycles, the graph is bipartite

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.899750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.946065Z digest=sha256:afbd1278c600a690701ca945a007e2fd50f8819d9af382c41118f9d92456a20d

Observation fed92e39-9f71-4eb9-b408-a516c7389120 · outbound

This paper cites This is because we can color one set of vertices with one color and the other set with a different color, ensuring that no two adjacent vertices share the same color.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction This is because we can color one set of vertices with one color and the other set with a different color, ensuring that no two adjacent vertices share the same color

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.883082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.950502Z digest=sha256:701f7d19eb1b1ef1bd22c211360b1f35716150323fb597641622b2172d8a5a1c

Observation 7a933e72-25a3-440a-84da-7f5ef6be6931 · outbound

This paper cites </think> The final answer is2.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction </think> The final answer is2

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.866845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.955074Z digest=sha256:77b67372a4657364e76a06d747b780cb55102beb0dae8f5d368f538f554653dc

Observation 1cb9d5ad-aa2a-48ce-adbc-690eaec86790 · outbound

This paper cites There are two visible wooden poles supporting the tree.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction There are two visible wooden poles supporting the tree

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.850621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.959994Z digest=sha256:c4993e450516936c2c4617e4b27a21e7a6b6af3870e94230422c874afd727ef0

Observation 349c59f0-13ea-4c7c-ab7d-5e1d5209b7af · outbound

This paper cites There are no additional wooden poles visible in the background or elsewhere in the image.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction There are no additional wooden poles visible in the background or elsewhere in the image

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.833597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.964690Z digest=sha256:06c15218150d7d1f3a357a8986eaeecd5d60b7c7cfbd47dffcaf1a2b45951587

Observation efe688c9-dee5-47cf-8bef-abf21d9a764e · outbound

This paper cites Given this analysis, the correct answer is: **C.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Given this analysis, the correct answer is: **C

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.817360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.970322Z digest=sha256:be50910b97fe94ba76d05c63ddd6702280eab49b81886ca702cde360f67ab082

Observation ff1816c0-e15b-4c2c-99e6-b95510212187 · outbound

This paper cites an unresolved cited work.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:17:35.474150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.548992Z digest=sha256:22ae94ba30eaefb06834524cb91f192a9885db9ff02c6ca8690f9c5d2356238b

Pith citing papers

No inbound Pith citation observations are available.