Pith. sign in

Paper Citation Record · LEDGER

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

As of 23 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2608.12209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.12209 v1

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:17:33.970322Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

83 of 83 outbound references displayed

  • verified exact2
  • verified fuzzy39
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e27eee67-71f9-4100-8bcb-157b5ef51e24 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LLaVA-OneVision: Easy Visual Task Transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.494370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.494370Z digest=sha256:46930216b2e6134d22e5c1318b4c3f959ac982e8c5701a06ed4c0919be359de1

Observation 73daf702-45dc-4ba5-9b8c-c00c0f426859 · outbound

This paper cites Qwen3-VL Technical Report.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.499723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.499723Z digest=sha256:4420192868591aa4091e8650a5f5882384c3e4c6171522558a6a8d8356ae66eb

Observation 581d4fef-c86b-4679-892c-e199b0988514 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.503957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.503957Z digest=sha256:dfea871553f3e848e53c78b008b7def9ec9fae3f14ccf0f2f89b016498d49965

Observation 690036a4-7e49-4628-8a91-54aefacf29ac · outbound

This paper cites Visual Instruction Tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Visual Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.508431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.508431Z digest=sha256:58c488b8489dc8bf4a31d5dab8d8a3017072f83718bbac2503d2ea3524cef8b1

Observation 073cf827-4cbb-4837-ab8f-b899bd3e0298 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Learning transferable visual models from natural language supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.513081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.513081Z digest=sha256:186e8d978ddefae9e29d7d70491eb5462a0e4fc29f1b3167dd6dae6cf6505bac

Observation 66e1153a-e3fe-4ab7-a84d-c0a1ec0e8aba · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.517459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.517459Z digest=sha256:4abe1bbc9427f1a7a743bc2306c462bd762d82a31f4efa88d161f8e07503133a

Observation 48223edc-b902-4f01-9d60-bb73da0595d2 · outbound

This paper cites An empirical anal- ysis on spatial reasoning capabilities of large multimodal models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction An empirical anal- ysis on spatial reasoning capabilities of large multimodal models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.522320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.522320Z digest=sha256:eaaf39ece98a6381b5651c73c9c9d977392e8fe33e18f6ac222511fb1423bb90

Observation 9dcd2796-7c6f-470c-b83f-a84dce81d3b9 · outbound

This paper cites Why are visually-grounded language models bad at image classification? Advances in Neural Information Processing Systems, 37:51727–51753, 2024.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Why are visually-grounded language models bad at image classification? Advances in Neural Information Processing Systems, 37:51727–51753, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.537516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.526218Z digest=sha256:64f558b03832e5cf92adb4caa7b59645ae20a94d1c0c56a27d6095ef74c4a664

Observation 56769107-c26a-4d69-8ea9-0a62b49efece · outbound

This paper cites Question aware vision transformer for multimodal reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Question aware vision transformer for multimodal reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.521970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.529994Z digest=sha256:fb6f33073d14994c01f9a722166ed6a778c85e56debaf5f907b08ab75d50ead6

Observation d29ae9be-8ff3-4050-9b9a-60a890622529 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.534077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.534077Z digest=sha256:ab3e9584eba3a73f3a48b83940781463727009c168cbc8cb26f75b0f6e53927f

Observation 8111b51c-97ea-491e-bf8a-5ac7e5ac0c30 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Show-o: One single transformer to unify multimodal understanding and generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.505930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.539092Z digest=sha256:9995a1cbcdac358d6807e26812e0b9b891c85c2b5b19c22691bf22db237c4699

Observation 1b064072-c6b8-46a6-a37c-2e788fcd5d72 · outbound

This paper cites Transfusion: Predict the next token and diffuse images with one multi-modal model.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Transfusion: Predict the next token and diffuse images with one multi-modal model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.490249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.544262Z digest=sha256:c0f9b1a7bbf725e7ed96f8750da191dc7f4d5fc99003404cc3e49899fb2d9e9b

Observation 5357a682-5d79-41d5-9b67-aa6be7d681e1 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Emerging Properties in Unified Multimodal Pretraining

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.553581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.553581Z digest=sha256:dd71f2f79a6498c6f79ad4dcc2f7151d1db27ec6d5f666dfc7bc4d82b5612cc3

Observation 3f0b369f-d2e4-4afc-9563-01b08ee21b37 · outbound

This paper cites Mammothmoda2: A unified ar-diffusion framework for multimodal understanding and generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mammothmoda2: A unified ar-diffusion framework for multimodal understanding and generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.558435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.558435Z digest=sha256:3b9e8879c8e6a7adeb6149dd62a633b6ab00ace194825d91a875afa16993be1d

Observation 569f7263-26fc-45b3-8025-46d8129b94ac · outbound

This paper cites Lance: Unified Multimodal Modeling by Multi-Task Synergy.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Lance: Unified Multimodal Modeling by Multi-Task Synergy

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.562911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.562911Z digest=sha256:4b1b3abebd465737fa36bc0934a3a67e39ac6b9dc930545a6724b7b0a55b6f88

Observation 63003f27-2aa1-4654-9c49-363a64cad2c1 · outbound

This paper cites Multimodal learning with next-token prediction for large multimodal models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Multimodal learning with next-token prediction for large multimodal models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.567839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.567839Z digest=sha256:010fb34a932e1578fee42e005814579d89b669060320969cc764f349f7d82abc

Observation 84770594-92dd-4d11-8867-6f6eccdd453c · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.572476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.572476Z digest=sha256:a5959e8e645b93c7e91b9abbf85e2e4acd7d77b591ec7f6f4cb80a6673a9126f

Observation 5da4b6b3-7fb0-4f5a-befe-3a99babd2c83 · outbound

This paper cites Mmada: Multimodal large diffusion language models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mmada: Multimodal large diffusion language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.447949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.577175Z digest=sha256:99c0d5da2b8a374acb04e227f76e1a2b053a058a099c77d0fbb5b249afb099f4

Observation 75ccea02-af2c-4f5b-a6b5-5904a0f2c0b2 · outbound

This paper cites Longcat-next: Lexicalizing modalities as discrete tokens.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Longcat-next: Lexicalizing modalities as discrete tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.581554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.581554Z digest=sha256:7c5ffc07807e2ec73a3fa00406725bd2d56ee56d20f9775fd7614b052d352431

Observation 64c0448f-797f-402e-b382-d4c47a23a458 · outbound

This paper cites Next-embedding prediction makes strong vision learners.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Next-embedding prediction makes strong vision learners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.585968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.585968Z digest=sha256:5e4444b33a74575b696eb3681fd3f7c58fd1bca7fafe567ea1b219f54ecc7b6e

Observation 37c21e1c-d592-4b6d-ab4d-9da81a50ac3f · outbound

This paper cites Unihetero: Could generation enhance understanding for vision-language-model at large data scale? arXiv preprint arXiv:2512.23512, 2025.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Unihetero: Could generation enhance understanding for vision-language-model at large data scale? arXiv preprint arXiv:2512.23512, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.590538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.590538Z digest=sha256:f95f556956a51df460cea8e19e1f60f5fbb04df14061cc0c465a0aad4f1ade2f

Observation 942b4361-6ee5-44ce-ba8e-a87358e57908 · outbound

This paper cites UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:17:34.447253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.594876Z digest=sha256:f5cfb6c71e92217a46b246ecaa76d4bd516ab0892908daae57b0ad4d2af414c6

Observation 0c30f67c-517b-4927-81ba-45a7ce43b7b8 · outbound

This paper cites Uni-x: Mitigating modality conflict with a two-end-separated architecture for unified multimodal models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Uni-x: Mitigating modality conflict with a two-end-separated architecture for unified multimodal models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.599748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.599748Z digest=sha256:eb6a2b6f00e285f27a894c39f5965c83a00219e78042f0332cad8b42c81ee43e

Observation 6985abff-97f7-4ef9-9bb6-0f31241ac41d · outbound

This paper cites Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.431635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.604319Z digest=sha256:60106de97f4f09b11ba0bc69dacc9d51727ad141930eb6ed8c006b42e78f68e4

Observation 1dfc456c-42b0-4591-ae33-cf1c8703e850 · outbound

This paper cites Qwen3 Technical Report.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Qwen3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.608967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.608967Z digest=sha256:8abf3efa30435666b96dd0645f93fe5983fb1d0ea7bd944b5ba88ab5483fc6da

Observation 7ef9f9ac-56f1-4129-9351-ddfa7291dc70 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Adam: A Method for Stochastic Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.613869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.613869Z digest=sha256:9099360ecb001f659afe726682e71102501e56831e7737791aa10c9609988a32

Observation d32e69c6-17cd-44c4-abf9-8a2e7f63c81b · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Zero: Memory optimizations toward training trillion parameter models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.415500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.618475Z digest=sha256:1229aa28ae498b96b906bd9ee6e1c49528dc0efc54a78c2892c592e9d4c2834d

Observation 41ae2848-0ec8-47b1-90a8-5f6a04382004 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.399745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.622968Z digest=sha256:94ff8f1a24328436544d4216b9ff97b376cb8e9ea7d16108679d4b251d10b20a

Observation 183cc7c9-3aab-49f4-9be4-b059d9ce5cfd · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.384039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.627474Z digest=sha256:a1a7258d112ded01b7bccf8707de9ffc7f71156a233e58c74b72959739351564

Observation f66b54ef-2f8e-4fa1-b366-dae59ce38d49 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Blink: Multimodal large language models can see but not perceive

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.632306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.632306Z digest=sha256:e0432a0b7f81803881644e959f05315aa0f1df5f8fab254694e2fc3166c5f7ea

Observation d1930387-009c-43db-90c4-b86a88be5003 · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.636732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.636732Z digest=sha256:ce83fb964645adcdd81c3c765bc450867539f7067374dbbbb7411e2c2bbcf0af

Observation 39e60b28-eb04-4718-9a89-3abff8768d58 · outbound

This paper cites Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.641183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.641183Z digest=sha256:25aead9d020320a6a180c8f5916d55fd570fc02b3fe033fc10e16615aec423a3

Observation a74c4a37-3aca-4d5d-94d0-28c8a50ffbdb · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Measuring multimodal mathematical reasoning with math-vision dataset

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.645477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.645477Z digest=sha256:580274eb967fc6b87e9924d842a891a6961900a22395949946ed8a1e7502892b

Observation 320d399e-1b4b-4016-ac2b-82b8f1efb5be · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.326990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.650019Z digest=sha256:43b1738e1f232986736bb6153d6e6d281d07dc96aba614bae952cc9c1909994a

Observation f06d5710-f9c5-4b5f-a98d-e20707970c57 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.654244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.654244Z digest=sha256:fe415190f4cde13778720a5c6ec82850269c967d616d50f0c129042ae11dfea3

Observation 02f09eb4-7996-4dfb-af6e-2b535747e291 · outbound

This paper cites Visulogic: A benchmark for evalu- ating visual reasoning in multi-modal large language models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Visulogic: A benchmark for evalu- ating visual reasoning in multi-modal large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.311942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.658494Z digest=sha256:6572f4ac4bd46eb266e7efddfc4b4eefb82eff24f197e6085b457f3aa4e663b8

Observation 26c84fe8-d53e-442a-a4e8-85a94ad09106 · outbound

This paper cites Teaching CLIP to Count to Ten.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Teaching CLIP to Count to Ten

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.662073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.662073Z digest=sha256:d0ceda7b29a8678ad225d0c7100023e9998fcca3627cb11675efb3a4ae63b011

Observation 1804ee86-6b67-4b94-947f-7787318f7452 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction PaliGemma: A versatile 3B VLM for transfer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.665902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.665902Z digest=sha256:6368d7580b01482d4c7e9d8b20d7ffc1236b9c99569fc0dc6705ddcb33d10afd

Observation def81ece-2147-47a0-bae5-2907a8f1f1a4 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.296526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.670137Z digest=sha256:d4ed948732772cacb0177720c0dc4b9431fefcdaf1d3ad1a8575bd88451cf5a8

Observation 40f3269a-697b-44c9-a810-275fa14c1ae3 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.280336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.674031Z digest=sha256:06c9e67296e2ef16090ea4e14c734fa665e001a4aecc2425b5fdc2369aa8a55a

Observation 314f5c6e-7cb5-418f-b2df-3c51bcb1741d · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.677995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.677995Z digest=sha256:f3097ce0fe1b0941c81a45985e1e6f660747f87a51080f31af1fc6d9b3b11d85

Observation 2cbd7d36-a227-4539-97e0-6a8033974b9a · outbound

This paper cites VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.682022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.682022Z digest=sha256:785bdeba8e308bfa48c3a8537bf638148ef440a8d84fa85d1d490348b0460053

Observation cd0019bf-cc06-4fde-a865-57ac6f7ea711 · outbound

This paper cites Thyme: Think beyond images.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Thyme: Think beyond images

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.254442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.686813Z digest=sha256:8208c01f03e6e7049ae481b81a90b4c74408a4fff5494eae39e9e2f6fe2819aa

Observation 523eabbf-f825-472c-b160-5d5f89a617c0 · outbound

This paper cites Improved baselines with visual instruction tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Improved baselines with visual instruction tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.691642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.691642Z digest=sha256:14770fbba9f0fc10d63e1f107b82ff4e3aeddac1c3d8db13bb5b41a7bf353610

Observation 83854f49-131d-44c5-bbf6-29f72bce05df · outbound

This paper cites Qwen2.5-VL Technical Report, February 2025.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Qwen2.5-VL Technical Report, February 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.229128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.695959Z digest=sha256:623dcf3a9af829b1225719d9626af5087607929e5b37c653a9fe2a5acca127b2

Observation e64aa5f8-a4cc-45e2-8d69-c4de94a0fee5 · outbound

This paper cites Metamorph: Multimodal understanding and generation via instruction tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Metamorph: Multimodal understanding and generation via instruction tuning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.212857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.700366Z digest=sha256:562cec6af9de5d046ef59f36d4bceccd1a0d317d7e7c3697c87687613f966844

Observation 6a154796-6cdf-4c11-a87a-e2d00e90aabb · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.704691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.704691Z digest=sha256:a31a7a752ff8813d9af8487ebd1edfd68c204c0fdc67789701667ac4add7538a

Observation c858fa4b-9750-475f-82ff-832482d94d18 · outbound

This paper cites Show-o2: Improved native unified multimodal models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Show-o2: Improved native unified multimodal models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.196673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.709277Z digest=sha256:0ea81a3b80a37795ab1fe08845557d475c0e285f53475dbaec1c61ca3e0fa90d

Observation afd85e1c-5dd4-4d8f-8ef3-6c2d91e10f02 · outbound

This paper cites Cheers: Decoupling patch details from semantic representations enables unified multimodal comprehension and generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Cheers: Decoupling patch details from semantic representations enables unified multimodal comprehension and generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.713410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.713410Z digest=sha256:979716f0136af81c60c6c20e22b012125c95057f39c5ce37558dd36b2a9e77b8

Observation 50dfeed4-2f57-42eb-a738-4a3adbb969c6 · outbound

This paper cites Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.718062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.718062Z digest=sha256:53826e6d3a22ffbee45286c7bec2722b8a070dff08076b1899341e7cfc0eb529

Observation 48670c1e-7796-4337-84f5-53a46388c1c1 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Imagenet: A large-scale hierarchical image database

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.722979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.722979Z digest=sha256:b7092d179b2daa9c93b6e2d07188a63365e1f8dacfc11a48fa3a045ffbb80a7f

Observation 5ecf47df-c7d9-4053-b9f6-444cd7a69bcc · outbound

This paper cites Gen- eration and comprehension of unambiguous object descriptions.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Gen- eration and comprehension of unambiguous object descriptions

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.161417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.727985Z digest=sha256:00eca98e28be493d74d5e09dd8dab5d37e46461a4965e498d87735280969a84d

Observation 3091c69c-ee9c-4e96-a108-d3f6d8b20ad8 · outbound

This paper cites Referitgame: Referring to objects 23 in photographs of natural scenes.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Referitgame: Referring to objects 23 in photographs of natural scenes

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.145538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.732783Z digest=sha256:1d19398d36f93d23dc067fa5234ede81caf73eb038bf587d50bcaaefb4e05b85

Observation 26584f7a-e5c5-430d-8e1e-9bb1d93aa0af · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.737344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.737344Z digest=sha256:45fc05fb21c1dd6342498ad381c7967c41b120657804910134f552b602748af4

Observation c70cbff7-5db0-4c6d-b223-3480478740f4 · outbound

This paper cites an unresolved cited work.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.742241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.742241Z digest=sha256:219993aef6ed284672548d648075b4e888ebddbc58e29b08dc3331e9b56c2e34

Observation 905e03c1-049e-4f64-a7bc-241d4ddf8505 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Flamingo: a visual language model for few-shot learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.747386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.747386Z digest=sha256:a558f0c4066bd368dd4f60eba1743f359c7c2f64e7a6e8a4fdb06fa2737bdcfa

Observation 9268cac0-a2b7-4868-b137-d12d81434bcb · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.111944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.752571Z digest=sha256:e68876cc6cebda4088a871ab43a80c27f507c15e3abaec8cf01ee8d6d213d407

Observation 84d506e9-63f8-4616-8f32-61a9a384bdfe · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.758009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.758009Z digest=sha256:6c5c2bd64d66e4ff72eb03b248852b4faffc95aa2d846a4943a3614acc3f35ea

Observation 806d6f5a-3a59-438d-b1ec-c1e99e70adca · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.762555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.762555Z digest=sha256:f78aca1c60db11a31ec97e9de41de5e3d1f450856b6f7c8b8df170e5c1ed2215

Observation bd887cbb-d5da-4d59-9661-9de455213db5 · outbound

This paper cites gpt-5-system-card, 2025.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction gpt-5-system-card, 2025

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.097932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.767368Z digest=sha256:c03d5a6e84dd8ce2460c26804a81d9ff47bb56ac87fc0abb250f7ff1797e5f9e

Observation aacebb07-2ed8-446c-91cb-8ecf55737892 · outbound

This paper cites URL https://blog.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction URL https://blog

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.082321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.771718Z digest=sha256:e99b325c833beeabbd71955e8fa50e9f83636848ca82f0bccdb83b292570d5e5

Observation 53b7ce97-9450-476f-be32-e5a9b92933bf · outbound

This paper cites Gemini 3 flash: frontier intelligence built for speed, 2025.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Gemini 3 flash: frontier intelligence built for speed, 2025

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.067638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.776922Z digest=sha256:ae50d7664539a8424e02e1c7038c92b353e0da485ed02bd3b21e78fffcdf9d96

Observation fd929ecf-1804-4e14-a388-44982ba0013b · outbound

This paper cites VGR: Visual grounded reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction VGR: Visual grounded reasoning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.052990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.781312Z digest=sha256:11f24403b80e3580be059fa1a10f2b7ea893ffaa4501605420c35053f470205b

Observation e58da99f-15e6-4f75-ae40-28dbe6609a75 · outbound

This paper cites Explain Before You Answer: A Survey on Compositional Visual Reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.785540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.785540Z digest=sha256:cae46eb1e25abbd1235579237d7e4e8778448bb13da7e3693b7eb6f3408b2852

Observation 8f707088-def0-4a1a-9541-a31fde45abf9 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.789852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.789852Z digest=sha256:82b152de83dd125b4efcae88543bb1b7dd2033469f9c5cc671d4aa75ef97119a

Observation ac6385c7-46e4-4ff2-ae91-b502cbe09e08 · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Video-xl: Extra-long vision language model for hour-scale video understanding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.036455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.892662Z digest=sha256:6b531c86d3d525490ec79b8ff69692e31c00750a03a586893aa9dd93aacb8e74

Observation a929b0c3-f6b3-4cd3-b19b-0091995455d4 · outbound

This paper cites Videochat-flash: Hierarchical compression for long-context video modeling.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Videochat-flash: Hierarchical compression for long-context video modeling

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.020711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.897426Z digest=sha256:53a5e37e128a5e85566ae8a79df631a92671e371694e755867a9b2fabe9e847a

Observation a4a286f7-2de5-4dce-8344-2dd178faa383 · outbound

This paper cites Reconstructive visual instruction tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Reconstructive visual instruction tuning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.004875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.902010Z digest=sha256:759a57c1e0a7e5106c46241194fb71a357b66eb6a3ffaff715e0c5d3c596db27

Observation e779dc83-cfa0-4390-b4fb-f2934a771cda · outbound

This paper cites Autoregressive semantic visual reconstruction helps vlms understand better.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Autoregressive semantic visual reconstruction helps vlms understand better

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.989707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.906676Z digest=sha256:7b2ad69b6ba8723ceb2dfdb993b0b02fe401f8518c26ba0fb6886e4bed119edd

Observation 0dae9be8-18b4-462f-aedf-d247e85fa217 · outbound

This paper cites Generation enhances understanding in unified multimodal models via multi-representation generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Generation enhances understanding in unified multimodal models via multi-representation generation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.975035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.911508Z digest=sha256:2fb4a48c6c7c01a45afdb8cb24c3e86f2f273193293fbd8c71ce6d2288f4db91

Observation b8ab97a0-4659-420f-bbd2-40479f921270 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.916989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.916989Z digest=sha256:582e5e4f2a8854e8f2ced276244e370bdc91eb5dec27f94e92f71ef12a20d225

Observation f820c2a0-3031-4f0e-87b8-b377bd0734ae · outbound

This paper cites LMFusion: Adapting pretrained language models for multimodal generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LMFusion: Adapting pretrained language models for multimodal generation

Reference 72

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:17:34.087176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.921910Z digest=sha256:04b78ad96497fb299e23be6464a80546b0d2095f146b52c93f00793727f796dc

Observation 95c572ab-905d-4dd6-bc6f-629fc15d2719 · outbound

This paper cites - Row 1, Fig.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction - Row 1, Fig

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.960789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.926801Z digest=sha256:9255b80dd1aecf3ab8eae71e96b138675ecba03c4fa245b36a55eb54cf684e75

Observation 3b16e616-1bc7-4869-aa16-07fc00ab01c8 · outbound

This paper cites All figures share the same two-shape overlapping composition with consistent contour and inner-line connectivity.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction All figures share the same two-shape overlapping composition with consistent contour and inner-line connectivity

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.947379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.931751Z digest=sha256:dc7756ff7ec44ca7af028c1aa99c340e61e1dd443a1cda45858a97ea3c6e5eec

Observation 2c55ad56-d9a3-48ba-9d36-a7298d51ff94 · outbound

This paper cites Option A is a triangle pair with a mismatched shape; Option B is a triangle pair with inconsistent line count; Option C is a quadrilateral but with a different stacking style.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Option A is a triangle pair with a mismatched shape; Option B is a triangle pair with inconsistent line count; Option C is a quadrilateral but with a different stacking style

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.931691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.935730Z digest=sha256:c1bae7f3b44ce12479864dd0c679a89d025fc84057323c76a57a5b7ecd434533

Observation 64359b35-023e-4c33-8ad6-d23ce8ab9a83 · outbound

This paper cites A bipartite graph is a graph whose vertices can be divided into two disjoint sets such that every edge connects a vertex in one set to a vertex in the other set.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction A bipartite graph is a graph whose vertices can be divided into two disjoint sets such that every edge connects a vertex in one set to a vertex in the other set

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.915799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.940312Z digest=sha256:99a9218b705a33b53aab4be9848e124ba218c6e32abce8ca1a3733cbd3a9f1dc

Observation ace91903-90da-48b4-8a18-a0dfcaaadbf3 · outbound

This paper cites If there are no odd-length cycles, the graph is bipartite.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction If there are no odd-length cycles, the graph is bipartite

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.899750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.946065Z digest=sha256:644ce4e9371bf5a9244d6e109dc8f1b12caf66652be7d03d6eb6a302e08f9fac

Observation fed92e39-9f71-4eb9-b408-a516c7389120 · outbound

This paper cites This is because we can color one set of vertices with one color and the other set with a different color, ensuring that no two adjacent vertices share the same color.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction This is because we can color one set of vertices with one color and the other set with a different color, ensuring that no two adjacent vertices share the same color

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.883082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.950502Z digest=sha256:16c7575f4ca24f8aad4572a5f5ddf7ce38f4644a6698639a9eb4afa4d8841b91

Observation 7a933e72-25a3-440a-84da-7f5ef6be6931 · outbound

This paper cites </think> The final answer is2.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction </think> The final answer is2

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.866845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.955074Z digest=sha256:ad79161316bd68d3a37a62dc468e88bb258d4a44486d056a4deef9ca25ce4caf

Observation 1cb9d5ad-aa2a-48ce-adbc-690eaec86790 · outbound

This paper cites There are two visible wooden poles supporting the tree.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction There are two visible wooden poles supporting the tree

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.850621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.959994Z digest=sha256:97f270e562168cccaea516c82d760d96bed23874b551dd9acf5f2ed2db56e760

Observation 349c59f0-13ea-4c7c-ab7d-5e1d5209b7af · outbound

This paper cites There are no additional wooden poles visible in the background or elsewhere in the image.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction There are no additional wooden poles visible in the background or elsewhere in the image

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.833597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.964690Z digest=sha256:ba8aad6c5f977671754d50742c1de65dd0bf95d1b68a83ab037ec5f7e7e7d0dd

Observation efe688c9-dee5-47cf-8bef-abf21d9a764e · outbound

This paper cites Given this analysis, the correct answer is: **C.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Given this analysis, the correct answer is: **C

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.817360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.970322Z digest=sha256:da983f30ab00eb91f505f5f0cd57a666c901f842cef7d1b33f8554adff49d5a4

Observation ff1816c0-e15b-4c2c-99e6-b95510212187 · outbound

This paper cites an unresolved cited work.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:17:35.474150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:17:33.548992Z digest=sha256:7b5c015894c3a7cf7bf4d77a85d96eff3b3f477509dbaf393ee585ff7cefbf6a

Pith citing papers

No inbound Pith citation observations are available.