Pith. sign in

Paper Citation Record · LEDGER

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models

As of 12 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2606.05758.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.05758 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T03:03:23.071178Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact15
  • verified fuzzy0
  • unresolved62
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ccf5a602-8584-4e0b-baeb-56156e970724 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:e049d172cc67cc5f3ab04da19455938449fb41a018b1957582415f781836c838

Observation e2f5f7b5-a506-4685-b2a3-2377c7674897 · outbound

This paper cites Qwen3-VL Technical Report.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.645434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:0ab581b869363356dd612b8c63b869ef08274f9283ab91dd1706889fc63a1b12

Observation abd0245b-9c31-4e3f-a3f4-e997f35acf6e · outbound

This paper cites Qwen2.5-VL Technical Report.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.648478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:263de3b760ab4f4972abc595d15e705b4379057030ea97ce473a95f393526908

Observation bbbd61e1-c8f3-420c-a162-e50a31b2b5cb · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.633298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:99bf5db6e294c0f0421ab67b3e64bdd8ca3867e01cb8c05b54ec1ecc677e53a6

Observation 49ad2854-97ea-48a3-ad4a-d0b4ee09e0b1 · outbound

This paper cites an unresolved cited work.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:21f47455ecad2e0c60ddaa23d44a089aa7248a985bb0fc5456459bae57e6f404

Observation 369a943b-0c55-4b83-845c-c7f3658c6899 · outbound

This paper cites π0: A vision-language-action flow model for general robot control.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models π0: A vision-language-action flow model for general robot control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:3408287278806dbbbfb4662c3a005cbe1ba6e60841f1b9542e8e0926239cc7a8

Observation a47bdf3e-2509-4aa9-9c72-5f31048fb6c8 · outbound

This paper cites InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.638894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:ae06b2f9c0477f7d4b6577a17f00ff7c6684b44646b9cbf02207e78b370c900b

Observation 0495db46-511b-4d82-9032-17f7a70d1d70 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Diffusion policy: Visuomotor policy learning via action diffusion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:4f70088f32cef3846740749258f22343c590d724c9a6da96593b477922bf16a6

Observation f13f6095-bf0e-4f98-9dd9-28d7ac9ae1f4 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.636251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:e1923ac159c9d65d00432fa9363e1426f1d2aca260bdbeae32e6ba138ad371c5

Observation 20b7a120-a796-479a-8510-223bb3597978 · outbound

This paper cites Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:19def89716a09e404c2a64551ed2227408c46f19c4b33a351f2e8d28e2d68d54

Observation 71c96f1d-84ed-4942-bf21-aa3d4d9f4c95 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:2a36082b9a9ffff1b0ce4c772b11a049819b2a97b464ba68b0dd13fa9c5d025d

Observation 2dd58261-06c7-43f7-8ef2-c927534e9c04 · outbound

This paper cites TALL: Temporal activity localization via language query.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models TALL: Temporal activity localization via language query

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:c5017fea2d3856b85d0517c4a55f31c5edf9c5aa4244ff0a67045168a9899cf8

Observation 37f4d66c-d488-4476-98a5-0350c8a2f7ca · outbound

This paper cites Localizing moments in video with natural language.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Localizing moments in video with natural language

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:d10a50f165027d943c785d017c5c70c63572c66a29bb872a112fd17fd57d1930

Observation 574c89ba-5672-4fff-9c3d-3b004d70f5f5 · outbound

This paper cites Denoising diffusion probabilistic models.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Denoising diffusion probabilistic models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:c6f626fb18c06ea8d3c4a695b47aa2441940e6ff5fc8ce2fc5829dc0b9896289

Observation 51e06bdb-d2c9-447d-8612-64500553faeb · outbound

This paper cites Lora: Low-rank adaptation of large language models.Iclr, 1(2):3, 2022.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Lora: Low-rank adaptation of large language models.Iclr, 1(2):3, 2022

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:5f25c4ca518c2d76cbaf23b77180a69db5e663921da60d19482bd7b1fc5d2f68

Observation b2a3ab53-3dfc-45a8-ad00-215c2cf06940 · outbound

This paper cites VTimeLLM: Empower LLM to grasp video moments.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models VTimeLLM: Empower LLM to grasp video moments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:00e1d4558087a9e8972ba33e311a2aa9bab9b5a8c43865035cfc4d05f38e8b89

Observation b1ca992f-2d90-4deb-bce3-73f44f9ed9e4 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:0c083960e9de9620b1e548ddcab8e1d8e61fd2b31be127df714cab577b7516e9

Observation cfee4243-3740-4c21-ba9e-0c2c1f25203c · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Referitgame: Referring to objects in photographs of natural scenes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:ffaebd8308fc7bd88d2b5284b96d71b98c4d898e3b3dbec4ead14309f61971fc

Observation 7dae655b-af99-4981-8e5e-be67ccb24be7 · outbound

This paper cites Language-free training for zero-shot video grounding.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Language-free training for zero-shot video grounding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:955f6b362b3bf043c11326c17223c113bbdd1a34eac3d63c7ddaaf5dea107140

Observation a42c444d-2c8b-4c6f-b415-0e74a11f4f7b · outbound

This paper cites OpenVLA: An open-source vision-language-action model.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models OpenVLA: An open-source vision-language-action model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:41d05dd791a321c240de7ae6079d1bafe717410ab0dabb643de8163695506c34

Observation 0dfa550c-87e6-4a06-b4f2-fa3e4e705ec5 · outbound

This paper cites Dense- captioning events in videos.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Dense- captioning events in videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:306f2fc29ff96962e8301eda5a5821a99af72b785debadfa4050da2382240fed

Observation 0df7b436-0101-47c3-8394-4f12fb37bf14 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:0bdac87bec639e9cbb84248cf1acc2633d6ddcc464564a67e911f0238648bacc

Observation 22c7016f-6232-4ada-a241-944be645b0ef · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:307a29fc64a57e35fa84580d0503e581b18d5fee7a3febce50c4a374f8d193e0

Observation 791f67d9-2455-404c-a022-f47cc7db2edc · outbound

This paper cites Back to Basics: Let Denoising Generative Models Denoise.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Back to Basics: Let Denoising Generative Models Denoise

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.642747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:5ec2e48b0a12951d34f41b05fcb43286a036152614f0e35ad8950d27396a7a93

Observation 8f659ae9-0f3c-4e94-ab1e-499bcadc089f · outbound

This paper cites Evaluating real-world robot manipulation policies in simulation.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Evaluating real-world robot manipulation policies in simulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:d05ed61b639f0a13c9b251040ba077bb51ecc2c92d180ba88dd6c3b51a563174

Observation f411fe3c-e914-4462-91ca-471357330ada · outbound

This paper cites UniVTG: Towards unified video-language temporal grounding.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models UniVTG: Towards unified video-language temporal grounding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:9f1c7a5a3e8f5264f2287c8ca80734476e25ca2b30e9a9dbf85e2be4f0c31fbf

Observation 5cbceac8-27eb-4801-827c-92fcd496913a · outbound

This paper cites an unresolved cited work.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:d3fbbd43c9c775c6d06977109e5fdcdca3fa8007192c1acbdf398005b4609f81

Observation 17628147-70c9-47f9-a981-2470f4ee4642 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:0e55830bdbba93d93c1095688975c688c95f3268397a1e2a7f23f89bd12019fc

Observation 3e6e92c4-34f9-4f93-9cde-913d6c952b3f · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:dd139d36457769fd8607dc974a6a4c0e4ca66668839ac687dadb715b79733098

Observation ca1229c5-4588-4b09-bace-39d432ba5e56 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:009dd18ddc41959f9ca580f8a4528c0c0382356081d8d12ce3e29003d67a251c

Observation 32c98f78-c653-43c9-80e5-7a6b4cd2a1f3 · outbound

This paper cites Unitime: A language-empowered unified model for cross-domain time series forecasting.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Unitime: A language-empowered unified model for cross-domain time series forecasting

Reference 31

Resolution
malformed identifier
arxiv_id, observed 2026-07-02T11:46:55.651527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:73654f31aeee73e3271786547c4bf8c4cebcfc1128bd6eb5288e7a3a5d567973

Observation daf17f2c-62f8-4cae-b998-d210a1bc9d71 · outbound

This paper cites an unresolved cited work.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:81342269958b201c1009eb70725e9135e67e2ecfb176c8e409c113d68ea4dc78

Observation 798a8ea7-e220-42e8-85a8-2d3428905f9f · outbound

This paper cites Decoupled weight decay regularization.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Decoupled weight decay regularization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:4f9796701d977343368d0a34aba03bf820879b01e7cfb3bd307d6dd93cfcce96

Observation 144ab426-3db2-4f17-9c66-88a0b4c00fa4 · outbound

This paper cites Valley: Video assistant with large language model enhanced ability.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Valley: Video assistant with large language model enhanced ability

Reference 34

Resolution
verified exact
doi, observed 2026-06-28T03:11:30.318499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:9c22082830cb0ead18668b264cb06cd657f236b82f07176d41278afe6102da77

Observation cd70a692-d87c-447b-bc15-900ed1a3a118 · outbound

This paper cites Video-ChatGPT: Towards detailed video understanding via large vision and language models.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Video-ChatGPT: Towards detailed video understanding via large vision and language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:d21aa86d1c3c3d597b3b162b074e0adc373c52a4cb97aa996fbee8547eb7ad8f

Observation f39d479a-2aec-48df-9392-2455b6a3d960 · outbound

This paper cites Yuille, and Kevin Murphy.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Yuille, and Kevin Murphy

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:4aa100be1a797a2737a7e17a872ae6ae59ee638572b1040d24b64075f0825bf6

Observation 41069113-3eaf-46a5-b9d6-c279929d4e8e · outbound

This paper cites Trespassing the boundaries: Labeling temporal bounds for object interactions in egocentric video.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Trespassing the boundaries: Labeling temporal bounds for object interactions in egocentric video

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:5da74b9c635c736ab138e56a7ff86281ec04a8fbe4a6425a5bf225e9e27d496a

Observation 97aee435-53ad-4367-8f5a-8f8ab25b32e5 · outbound

This paper cites Snag: Scalable and accurate video grounding.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Snag: Scalable and accurate video grounding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:ea8d364983c46966a9be4634856e502539730a13d42d8bbed421264012445c0a

Observation fde93205-503a-4f94-8d15-1eac75398095 · outbound

This paper cites Zero-shot natural language video localization.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Zero-shot natural language video localization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:e83cd80720f40abb1ef67a6aa0a48d48bcf69f518cf8cca2577fbfaec9476ae4

Observation c1ddca02-e7af-4671-aa2f-b63e6d1b100d · outbound

This paper cites Henriques, Yang Liu, Andrew Zisserman, and Samuel Albanie.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Henriques, Yang Liu, Andrew Zisserman, and Samuel Albanie

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:93d324fcb4152ecca1a567ed779e794e03244bd5b7c7ede988c76918d92f338a

Observation c03464e7-fb9a-4817-9621-e3248787f36b · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:722b26cf41aaf3bf4cd3eeab0c2f6c0c93373645729b13db7bab7ae5b9adcc66

Observation 3ca90936-b3f9-47a2-a756-0b2f8201f8e7 · outbound

This paper cites Scalable diffusion models with transformers.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Scalable diffusion models with transformers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:1d11b707a0cd296d8fec3f241d64c6e1cb2d087784ce717a72873a0ef35e757c

Observation b808d883-fc05-4375-960c-89a75d39fd09 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Movie Gen: A Cast of Media Foundation Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.631004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:baae8fd544c87604598e34a99ac050506e8d88e4c53c52a83e7aa41f5a8094e7

Observation 95c57808-42c6-4aab-b54d-fcae42e2414c · outbound

This paper cites Enrich and detect: Video temporal grounding with multimodal LLMs.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Enrich and detect: Video temporal grounding with multimodal LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:e890c54135e04ed799bfe6c34aca15b764d708d1488896f7ac1d2d10e06e47dd

Observation 97057a09-647f-4a5a-8e9b-d28f0394a72a · outbound

This paper cites Momentor: Advancing video large language model with fine-grained temporal reasoning.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Momentor: Advancing video large language model with fine-grained temporal reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:c641be054e85ccce1a549bddf56c45eaf44efe1c7ccbb2ea452724cc01044a5e

Observation f7215a9a-bee1-4df8-a38e-97cbf0ada2f2 · outbound

This paper cites Chatvtg: Video temporal grounding via chat with video dialogue large language models.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Chatvtg: Video temporal grounding via chat with video dialogue large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:187fd4454bb3628dc0b0d4dddbe21edf5c126a687c29fcfc3e2e72440bb411aa

Observation 6cc68753-7076-4d98-b28c-40f8ed9f5351 · outbound

This paper cites TimeChat: A time-sensitive multimodal large language model for long video understanding.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models TimeChat: A time-sensitive multimodal large language model for long video understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:faabe66edb87e6b3dcc28664de96e2dc221bfc3a56d7665f621bf3ad45e2bd40

Observation 1caf53bf-d45b-45cd-ad5a-d2ab1025957c · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models High- resolution image synthesis with latent diffusion models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:d431e436a1a4ced7b447977c3af97791786e7a7e032545649ae4fc971c97c0fe

Observation 012bd99f-ca27-4811-a4da-e8f0aa6c523f · outbound

This paper cites Hollywood in homes: Crowdsourcing data collection for activity understanding.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Hollywood in homes: Crowdsourcing data collection for activity understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:7167f0ac3b17ecda9435138d2868a3b4d7ab7d18ebe37b93e037917f78f9fc45

Observation e2297b42-08f2-4f5c-9911-38ce6827710f · outbound

This paper cites Score-based generative modeling through stochastic differential equations.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Score-based generative modeling through stochastic differential equations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:b4fe78c295aead196d7aca257af609e5dcc21acfefe0622557fab6b0cf27c8ee

Observation 2b9516eb-9718-4fec-85e3-347a492617d4 · outbound

This paper cites Moment quantization for video temporal grounding.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Moment quantization for video temporal grounding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:5c188d475ac85e5c4f5bbfd5a1886a7c95076e56b94da060b5dff29367433bc6

Observation 61d34f0c-a5b5-47d7-9b1a-49af527c4358 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:fb8982faed2c361155ada8bac127dd32a93851ebb59c5250d0f6e63cdd84dd5b

Observation 5c26f7c0-931a-4256-8799-d8d0f9e8eb9a · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Wan: Open and Advanced Large-Scale Video Generative Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.663465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:5ccd958b5c0fe1acf7e462e42e96e9e0175444b780ea23f451730c763e52135e

Observation 499609a7-c8ef-4e56-817a-0679e2ae38f6 · outbound

This paper cites COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.666603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:76fc911c09ad12169c4a72dcf9edd24f4043691bd9e73f8b79cf0f04d93d7a86

Observation 2b4ef27c-3f60-4165-ae80-ced4a3e76709 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.669598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:9b23214b5d88999a3043aeb385e4cc1710df0b2c8b7346e93f8eb89616016766

Observation 815b454e-b518-40fb-b2f0-9ca95622d5ac · outbound

This paper cites InternVid: A large-scale video-text dataset for multimodal understanding and generation.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models InternVid: A large-scale video-text dataset for multimodal understanding and generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:c2df1ecb1a3a5ac283240b8647784334044c1cce744fe901c1ead748b9cf54ea

Observation 737b8a5d-8c30-4bc3-9818-ba8fe654a9b6 · outbound

This paper cites Vla-adapter: An effective paradigm for tiny-scale vision-language-action model.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Vla-adapter: An effective paradigm for tiny-scale vision-language-action model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:74fba0251bf622e07361eaa6b4b56da04aecf064c974f2fcb050ad84909aab84

Observation b4438bce-39e3-4a81-88d8-8fcb2c7a5e47 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.672672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:545dae5c0b266132fa5c166b9a16489d96d324290acf6c2a9a197c5c5f1d62ce

Observation f614ba8c-3bb3-4bad-befb-d39b9cd4c338 · outbound

This paper cites Towards visual grounding: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Towards visual grounding: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:34f29b98f4f97032706145c0a1446a78adb73d1566e6230dd488a02c48efa94e

Observation fd1a0dff-cd6c-42ed-9153-a053061bc930 · outbound

This paper cites an unresolved cited work.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:c3362a5d3e4a30afb4f071b5075052fbcb8e7955ba7dc6961137a9b893e3af12

Observation 8993a40a-b4e5-44ee-b1d1-47d6c70e7627 · outbound

This paper cites Vlaser: Vision-language-action model with synergistic embodied reasoning.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Vlaser: Vision-language-action model with synergistic embodied reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:1e57d646dc6e10d212652bb61539224295d8f86274b10f5bd1f812ceb7c80821

Observation e391880d-6be2-4892-b3f3-568f793cf3f4 · outbound

This paper cites World Action Models are Zero-shot Policies.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models World Action Models are Zero-shot Policies

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.654698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:1c99e48fde6e846ab44e4230affcd686460f2256cdfbfaed4f56437c259381ef

Observation 6837e88c-06f0-44dd-942d-02228bcc87f7 · outbound

This paper cites Modeling context in referring expressions.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Modeling context in referring expressions

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:ab372235b1632e4716bc423ca24fb2ead4070109d3c8e58ad21cc757aa77df5c

Observation 6d08aa57-362b-4f3e-8a93-5c3cf1e774d8 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Self-chained image-language model for video localization and question answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:c7d9f60d7313a3a3ccd62049c720c42c59d9462399660f007b3c11f1fc69c3b6

Observation e7367b85-f001-45c6-9f7e-acf84312fb1f · outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:46:55.660321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:4e52d2906104fd76bec016c24a231cd28f016245533d7b1f1b867917e93afa13

Observation e034433b-9690-4454-93e8-c8ec5b1434de · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:14b31622b055bb677cccf94425a1d2fd4284ce616175fbdb0e6250593b0030f8

Observation 2b277255-8762-4f73-a6bd-91c2fae93455 · outbound

This paper cites Hierarchical video-moment retrieval and step-captioning.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Hierarchical video-moment retrieval and step-captioning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:6c45d831b41bc56e0205c8eb3a70f706f6dfd2b71af88c3589d6d2c88b5ed02a

Observation 5408c1d3-ca35-48bc-b506-83cf769e3b7a · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video understanding.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Video-LLaMA: An instruction-tuned audio-visual language model for video understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:68172fba94b15557b5e6a9dce832da45f4266c28f7b83c44cf868e8b4b87eb41

Observation 8e3979c6-1517-4f9f-9dfb-ed7d09fa703d · outbound

This paper cites VLM4VLA: Revis- iting vision-language-models in vision-language-action models.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models VLM4VLA: Revis- iting vision-language-models in vision-language-action models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:5f3882b8da85726d4de2b51efbe16025f0b41ccae8c8a7d6aa5c886a250fbe6e

Observation 772d6add-ffbc-4ca8-8df8-ff7cedae1b29 · outbound

This paper cites arXiv preprint arXiv:2512.14698 , year=.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models arXiv preprint arXiv:2512.14698 , year=

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.657654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:029f4e4966303d75a9da2119221eafa048894f6b98a7d523c88f972b7a51fc37

Observation 290e7c1c-aa05-40ec-9f58-290336248b9a · outbound

This paper cites (Br +B h)Rn(H) + (Br +B h)2 r log(2/δ) n # . (13) Furthermore, sincew(t)≥1, R(0)− R( ˆh)≥E ∥m(X)∥ 2 2 −App H −Cτ −2.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models (Br +B h)Rn(H) + (Br +B h)2 r log(2/δ) n # . (13) Furthermore, sincew(t)≥1, R(0)− R( ˆh)≥E ∥m(X)∥ 2 2 −App H −Cτ −2

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:2c2c47418b25b3b7c7ea9dd1f5a1c28585f9cb6b9adf0ddd536aa2881b4c28d9

Observation c77c1107-59b9-4782-8bd7-99799bcbe3a0 · outbound

This paper cites 16 For fixed (r, t), the map u7→ϕ r,t(u) =w(t)∥u−r∥ 2 2 is Lipschitz in u with constant at most 2τ −2(Bh +B r).

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models 16 For fixed (r, t), the map u7→ϕ r,t(u) =w(t)∥u−r∥ 2 2 is Lipschitz in u with constant at most 2τ −2(Bh +B r)

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:3bb5038f750f0a1f11274b2513a0a36ab4e129230a46b80684c5c2fe8ee9ecc3

Observation cec65d27-60e5-46d8-af3b-1d290597ae96 · outbound

This paper cites Decomposing the right-hand side: L(ˆh)− L(m) = L(ˆh)− L(h ⋆ H) + L(h⋆ H)− L(m) = L(ˆh)− L(h ⋆ H) + AppH.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Decomposing the right-hand side: L(ˆh)− L(m) = L(ˆh)− L(h ⋆ H) + L(h⋆ H)− L(m) = L(ˆh)− L(h ⋆ H) + AppH

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:9085e584c3b69b7078f14b6682575235688be23e4c4506652ab0dc8fd055b381

Observation 455d2768-f2aa-4e32-bd24-b910a975abfb · outbound

This paper cites 17 Conversely, for any full-target predictor H, define hH(X)≜H( ˜X)−g(z) , where ˜X= (y t −(1− t)g(z), t, z).

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models 17 Conversely, for any full-target predictor H, define hH(X)≜H( ˜X)−g(z) , where ˜X= (y t −(1− t)g(z), t, z)

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:e9cf9dff2666cd6144bc553ba423f7542648a06c9096ecd0013ec6250526b3b5

Observation e83daf3e-1dd1-469e-9fd0-68cc1af1a474 · outbound

This paper cites an unresolved cited work.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:3c5286865fe6ce49c3ab14e99ec06587ff2b8f40cdb448a50a8b114e8d6b78bc

Observation 2eb9235e-6bc1-4216-b3ff-b7c09a34c83a · outbound

This paper cites an unresolved cited work.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:83a5224d09e05bce8029639044d5cc9ced706e57b4840740405f65024628bb16

Observation 0fc261e5-7b80-41a7-9437-cac0c5df0471 · outbound

This paper cites an unresolved cited work.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:1df9702c1caa104b0d8ceb359901995f49901074d3704c8af3f0633db1370ba7

Observation fcbb69e3-ce2e-4d11-a366-db4e54a5d4ba · outbound

This paper cites an unresolved cited work.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:f36c231f984f48ad328b8f2559c01babfa6ec38ad9fde7fbf0cd95d94cfd22b5

Observation f9215499-1e1d-4692-8c83-66965cac3320 · outbound

This paper cites LinearLinearLinear noise𝑡z! Self-Attention+𝑦Linear Sampler 𝑦!mix Base Predictor MLPTokenizerOr z! z! 𝑦.

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models LinearLinearLinear noise𝑡z! Self-Attention+𝑦Linear Sampler 𝑦!mix Base Predictor MLPTokenizerOr z! z! 𝑦

Reference 79

Resolution
malformed identifier
no resolver link, observed 2026-06-28T03:03:23.071178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T03:03:23.071178Z digest=sha256:66679218b21d7982f3de1d9a453c2a81697ca2eaf85ebac88994b1c24a3cf330

Pith citing papers

No inbound Pith citation observations are available.