Pith. sign in

Paper Citation Record · LEDGER

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor

As of 7 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2507.07106.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07106 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:52:37.501612Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T19:30:30.517564Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact2
  • verified fuzzy37
  • unresolved29
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8b87fd0-240e-418d-b6cf-8ec87ed4eb41 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Eyes wide shut? exploring the visual shortcomings of multimodal llms,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.207480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:31.474283Z digest=sha256:7e1d2f62aa4116e047188f50ddc19c297be83f1fdb8c76ea09284bfdbb714efa

Observation a12d77e5-20f2-4573-90e9-81ca44e1ab33 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Learning transferable visual models from natural language supervision,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.197431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:31.523702Z digest=sha256:7a3174b3100d39d4f3dfb9a6854a4e8e168bd68da56c925a2ee20f469e857022

Observation b602ef68-7beb-436a-b7af-92b84f6a69a2 · outbound

This paper cites DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:31.581690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:31.581690Z digest=sha256:6c996cf67b5857c691488ecdf63bb1fa26faad2c1998a76365df2860226836d1

Observation b3734393-a69e-4769-a1e9-4e7e7c1d021c · outbound

This paper cites Is clip the main roadblock for fine-grained open-world perception?,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Is clip the main roadblock for fine-grained open-world perception?,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.187174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:31.646858Z digest=sha256:33b261b620457b461ddb69b21227cad42eaef6b1731261ac2838f0b845b8610a

Observation 625bdc2b-6844-4ae1-ac8f-0e68ed309af3 · outbound

This paper cites Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:52:38.207894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:31.697090Z digest=sha256:2fcc2fc2a107e201eed6df7548a8dc87577fe30f42b5992188283a328d896d0a

Observation bf309773-992a-4ce3-a6c3-4332c53bc974 · outbound

This paper cites BRAVE: Broadening the visual encoding of vision-language models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor BRAVE: Broadening the visual encoding of vision-language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:31.767820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:31.767820Z digest=sha256:39060ac8b6ee9bd4ecdbc67c935b5656fc5ab4a51ee88596f4bbb596d882622a

Observation 67cf8678-d817-4b9f-bb5c-00e6edf6170c · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:31.825326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:31.825326Z digest=sha256:551288773829d908278a823860284469ff0f7dc3d9f5ada42b5bd06a9650e441

Observation ea4cd2ef-2034-4a1b-8b7a-f9b908f0b8ea · outbound

This paper cites From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:31.897432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:31.897432Z digest=sha256:5fea67f92b548e9366b5ef0e3d299efbea99f1c3c75a996352b016b9991975e3

Observation a6fd8af5-b0b1-4324-8852-e6e4af2443fc · outbound

This paper cites Mini-gemini: Mining the potential of multi-modality vision language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Mini-gemini: Mining the potential of multi-modality vision language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.176426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:31.974928Z digest=sha256:56da89c3a22da43588669b9621c0cd7bf8595e106b542e31c5425965fae6fe1b

Observation e9b9b89d-4508-462d-80ed-2c3b2ab5c80e · outbound

This paper cites Prismer: A Vision-Language Model with Multi-Task Experts.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Prismer: A Vision-Language Model with Multi-Task Experts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:32.060515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:32.060515Z digest=sha256:8ee735390968e46bd26a1c93283d748ab6541ad6e6e1faa2e0bb2b3ca45454c6

Observation 7cb52f24-dbbb-4ab1-81ec-a3d5c5cbafc0 · outbound

This paper cites Vcoder: Versatile vision encoders for multimodal large language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Vcoder: Versatile vision encoders for multimodal large language models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.166549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:32.175329Z digest=sha256:f3fbb17a2fa0b9decf75d01239fd9173bbcddc0286999399f9b4c53b372ced13

Observation b7103154-94f7-4435-bfa8-ac2c332fb38a · outbound

This paper cites Question aware vision transformer for multimodal reasoning,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Question aware vision transformer for multimodal reasoning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.156651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:32.267593Z digest=sha256:e14e6c6c2b89f4ccb1ec16d19218e178bb36b0b7ae240fd3efc31715b409c151

Observation 7283de56-cd47-43d2-a110-a861513486d7 · outbound

This paper cites Api: Attention prompting on image for large vision-language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Api: Attention prompting on image for large vision-language models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.146498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:32.368424Z digest=sha256:7377e0662cd12f024a949fa35bfa94d0c87f411b8f48fd9fc2c430e7c6fc1345

Observation cefb533f-5313-443f-b4f5-cca3a74cf1db · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Instructblip: Towards general-purpose vision-language models with instruction tuning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:32.475667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:32.475667Z digest=sha256:20513408290e8c1426d31c78467462c80ca8e48f1971118485d459e38cec1cf9

Observation 51279896-4373-4237-84bd-1fdf372892e6 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:32.491396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:32.491396Z digest=sha256:980804b2e6e86b6578c886133a078c15d1ef2aa067845d4a3ee4e33679685f23

Observation df85e6de-4f47-4c97-afdc-2213358707fb · outbound

This paper cites High-resolution image synthesis with latent diffusion models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor High-resolution image synthesis with latent diffusion models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:32.533667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:32.533667Z digest=sha256:323b999013fdd7388594f04ad8ba30320f3c64b71c680657b80ae7103358d48d

Observation bbcd5294-5e78-4b1c-bd97-3c5db57e9a4a · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Photorealistic text-to-image diffusion models with deep language understanding,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.116420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:32.611237Z digest=sha256:615f4e0ac7799393a1088745a29a5bd8a7181e5d46f20bf647d2b39d152755d9

Observation 296a3972-4983-498b-ad94-39a6c8bc9d1a · outbound

This paper cites Hierarchical text-conditional image generation with clip latents,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Hierarchical text-conditional image generation with clip latents,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:32.727639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:32.727639Z digest=sha256:a48525312d7c18ad4f99992508932a7271eee4e7af832351816b28bc43d19f62

Observation c0901327-ae10-4002-9c21-bba7c7ce2154 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Sdxl: Improving latent diffusion models for high-resolution image synthesis,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.097243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:32.845664Z digest=sha256:f99c1c624fa63f4378d401368a989a94ca5ccf8b84ac6a0cb29c9cf1ea00139f

Observation 55412a86-d16b-48bd-a6f9-528b92f8f1d8 · outbound

This paper cites Prompt-to-prompt image editing with cross attention control,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Prompt-to-prompt image editing with cross attention control,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.086788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:32.958030Z digest=sha256:d70e3a2b55168313bfa24a1ccca6fb8306b347b3ad589f87cea7c618f254e213

Observation b99865be-9966-49d5-ae52-cd4fc760b7dc · outbound

This paper cites What the daam: Interpreting stable diffusion using cross attention,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor What the daam: Interpreting stable diffusion using cross attention,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.077053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:33.072184Z digest=sha256:00cdae231283bf981077387a360db5d9cfbbef9fa3ba0fdc08d67227a88d4fb2

Observation 5715627b-f108-4911-a7ea-6493cff780c5 · outbound

This paper cites Towards understanding cross and self-attention in stable diffusion for text-guided image editing,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Towards understanding cross and self-attention in stable diffusion for text-guided image editing,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.067728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:33.171759Z digest=sha256:77a614239f5aa29f75749a915770480bf41c88afe824dd2c1f7323d9e9060551

Observation f68a552d-440a-4c04-83f3-8c6fe427aa1e · outbound

This paper cites Plug-and-play diffusion features for text-driven image- to-image translation,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Plug-and-play diffusion features for text-driven image- to-image translation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.058139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:33.269016Z digest=sha256:c1c3e1c93de24f3be0cf008c070a15a15dcefc4f2672b3a17d05b28493094ada

Observation 3de6ec26-b7fb-41ef-9c92-b1afb1298272 · outbound

This paper cites Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:33.421637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:33.421637Z digest=sha256:dc3c18b2f6163547d1df5eec4b7ed4177cd25ad636af51e58c9f6f09fa68de8c

Observation 8e32935a-f2fa-4f14-879f-b4239b68da79 · outbound

This paper cites Repurposing diffusion-based image generators for monocular depth estimation,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Repurposing diffusion-based image generators for monocular depth estimation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.048007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:33.537106Z digest=sha256:fa1338f351b1e5b882192cc30493474ebb5682d988ccb13b2b0938e9c8ec67a8

Observation 4c4dce4f-505e-4580-bd26-cbd6e386ba1f · outbound

This paper cites Do text-free diffusion models learn discriminative visual representations?.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Do text-free diffusion models learn discriminative visual representations?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:33.646987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:33.646987Z digest=sha256:bc1e289ea69ac949a0db352bd13bb2054386718b78d37b75f895e12d33874bd4

Observation 44a895a7-2e78-42cf-bce6-5f52c4eb49a3 · outbound

This paper cites Deconstructing denoising diffusion models for self-supervised learning,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Deconstructing denoising diffusion models for self-supervised learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.038820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:33.726982Z digest=sha256:b7affa7c7a1461a9060c35aae4192967d297ee3ecdd9cf1b681cc56addc65851

Observation 98a009da-c65d-4f22-9c99-2857b6e2c9a8 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Coca: Contrastive captioners are image-text foundation models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.028681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:33.824328Z digest=sha256:304128e03612a66e67b2931a4bc8923063ff09afa16dd01f49a0c70eca1f087f

Observation 790d6f1e-5c0d-4a5c-9047-606095c7cbd9 · outbound

This paper cites Multimodal few-shot learning with frozen language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Multimodal few-shot learning with frozen language models,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.017462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:33.897344Z digest=sha256:e93f851453a9c47748607f7e85d6718f9cfb90a55f93c9957998530a8b475d8e

Observation e5a0be6c-107b-4935-9338-c4a9398dd94c · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Flamingo: a visual language model for few-shot learning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.007156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:33.976289Z digest=sha256:c0908c58083f018760c56883b01e13570c7bc32e683dac1190ff4fadf379ff4b

Observation f20a7514-0702-4d53-ad7d-f58aaafc9dfa · outbound

This paper cites Visual instruction tuning,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Visual instruction tuning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.996235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:34.146798Z digest=sha256:a1f50502b3219532a8020908392a03cf938e2a3e51de681f79f9ebe5ebefe46b

Observation d0517812-1c5b-4029-9d48-4d901626aca3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.227206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.227206Z digest=sha256:3b45b2d33923faa0a390a2050196a7188be7b3368abb500f67e76c82bb047690

Observation 5cc71a27-f8e9-43c7-b0f5-e29d26238011 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.315551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.315551Z digest=sha256:6e22594d8166b7148ec0a8bd2f64c3f8f2a8feb872763ccfa01b57debbbe14f6

Observation 556d3ef5-1f4b-4c51-aee0-77d103f946d8 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.409335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.409335Z digest=sha256:bd35951ce8b460a4123f84b6e24471a2d0b0b73ff3e60abf219ec9c677845031

Observation 848a5aaa-b4e9-4115-82e6-decb4d529af6 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Evaluating Object Hallucination in Large Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.500395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.500395Z digest=sha256:8a12a9d489afc653851500247881958f38f2d53c046b91ecbbca5af5ae2e5f02

Observation 54725c8f-4675-4ac8-abfb-c2e9f1e25b01 · outbound

This paper cites Multi-modal hallucination control by visual information grounding,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Multi-modal hallucination control by visual information grounding,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.986332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:34.591011Z digest=sha256:6b9f2eebdceab9c16efe98b182c3619798bcbbd445e714ec9abd4c1190ba9dcd

Observation 5f489965-c74f-4fb8-aa60-306084b23fae · outbound

This paper cites Detecting and preventing hallucinations in large vision language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Detecting and preventing hallucinations in large vision language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.976233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:34.723073Z digest=sha256:066f07ead6d5e8db704d539ba2ec56f3ba0d29a19f90997e23d89181e83df0c0

Observation 5b55157d-f698-4213-833c-1e1eefac95f1 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.830166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.830166Z digest=sha256:fc1ed24d90610dea3155b72fb2b4fd7f195e2abc38f560a2506191f74553ee8a

Observation 36fe1c95-1620-4561-a3cd-5e6037917e1a · outbound

This paper cites HallE-Control: Controlling Object Hallucination in Large Multimodal Models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor HallE-Control: Controlling Object Hallucination in Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.896934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.896934Z digest=sha256:3599d5ab95cc83c0814fd65096f30e06ec8e247644f1c4ff511434940d566758

Observation f9bcfcae-e0d3-4335-bc9a-5553952852ae · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.983806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.983806Z digest=sha256:fcf4930fb367cd97212e083fa8c590fd0762bbbca58ca4c7ad47a6fcf8bd3d10

Observation ab5b5abd-9b09-4f0f-bccc-629838e11b29 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.095890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:35.095890Z digest=sha256:3bdc660fcb44ce34649eb70bf276dc06194c3c482bdfbe496e8f29e292871de2

Observation 201ad2db-d2af-4eb2-b99f-c19dc5e6cf4f · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.206253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:35.206253Z digest=sha256:ce960fe94e82ba8d41d9c64b0f296e431824252d1f079eb31df7eb7604375184

Observation 26b9cebe-850f-4bff-b75f-82e4460c81a5 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.255074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:35.255074Z digest=sha256:2ba0b04e1649e7ee204b4f0b6d7b862775eb1e59f59f13571d8623df96c9ef8f

Observation 1a775c40-908f-4be4-97fc-d183c552ff0e · outbound

This paper cites Improved baselines with visual instruction tuning,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Improved baselines with visual instruction tuning,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.415756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:35.415756Z digest=sha256:c9009bb26938a4ce2022a749bc9e2153dabc9e5b70f0a659f358c23810d6f6fd

Observation 4aa7ebc1-2d2f-4923-9cc2-1c696551b0bf · outbound

This paper cites 4m: Massively multimodal masked modeling,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor 4m: Massively multimodal masked modeling,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.952048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:35.541595Z digest=sha256:b7212e6b171f12652e1702498f7b8872f674e82ab720658755d89a499699c73c

Observation e08a5649-69ea-45a7-8806-6160ae111cfd · outbound

This paper cites 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.658391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:35.658391Z digest=sha256:d7a70ec322237fb6047d96b5e767042bf07ce8e9a925b4df7229d2f694fa2783

Observation 958be85f-58a5-495e-b952-93836b10ad0a · outbound

This paper cites Llava-plus: Learning to use tools for creating multimodal agents,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Llava-plus: Learning to use tools for creating multimodal agents,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.941704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:35.742782Z digest=sha256:c2652e9b467b52ead8afc2eab6a6de57ffe64af18fb6621096787a6be7b2129c

Observation 524a1413-4b8c-4d16-94c5-8e3364a96190 · outbound

This paper cites Visual programming: Compositional visual reasoning without training,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Visual programming: Compositional visual reasoning without training,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.931881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:35.829893Z digest=sha256:249c5ecd903022785762ebde2afa1e5039c4a3a69588bbc5110dcf10ce200a3c

Observation cd605569-c973-412d-a5ba-9693f9cea517 · outbound

This paper cites Spatialbot: Precise spatial understanding with vision language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Spatialbot: Precise spatial understanding with vision language models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.921200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:35.917707Z digest=sha256:35e69a985917d6149c38b790dec9c30abec05f648f9b8b444cfd597bfb4070a0

Observation adba8552-2fb7-450e-9f5f-74e68331758f · outbound

This paper cites Your diffusion model is secretly a zero- shot classifier,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Your diffusion model is secretly a zero- shot classifier,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.911539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:35.983026Z digest=sha256:3b6587071ae7319f59d67bc2002700dd11762d5e2fb2c28bce6bc618bf1f6661

Observation 6958a577-10b3-40c1-bfe4-a581be05c5bb · outbound

This paper cites Diffusion Models Beat GANs on Image Classification.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Diffusion Models Beat GANs on Image Classification

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:36.050247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:36.050247Z digest=sha256:a6a30d91ab504687679df406a0aee74ad39d304a9d36c535d2b158b5a1111b29

Observation 396ca2fd-a5fa-4ad7-9dc9-4f33babf60e7 · outbound

This paper cites Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:36.136857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:36.136857Z digest=sha256:f7019606058d5ffcd9171e27c6f95272892e0ff5f8ba45323ee29d50c9c375f1

Observation 57b935b2-94bf-41df-af2b-48c18dd48d23 · outbound

This paper cites Diffusion Models for Open-Vocabulary Segmentation.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Diffusion Models for Open-Vocabulary Segmentation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:52:37.892673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:36.201294Z digest=sha256:5ec033b5f7bf9c996a2304edd8d2778480ec6b5da4a2654967336cf364cdba45

Observation 2801e47e-4379-472f-b257-71088ab13e4b · outbound

This paper cites Not all diffusion model activations have been evaluated as discriminative features,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Not all diffusion model activations have been evaluated as discriminative features,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.901068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:36.301593Z digest=sha256:6a395e7ca7f4287a2698e91a16a2408f007007ace297099de80279fd07cb446a

Observation 7d74e8ff-50e6-4bf4-b698-08dacdd59fbb · outbound

This paper cites Dinov2: Learning robust visual features without supervision,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Dinov2: Learning robust visual features without supervision,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:36.377242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:36.377242Z digest=sha256:93dca6844c03173b4f5431c4a259415e902fe191f28848f4beff00db8e94ee65

Observation 00a9423d-718c-4cc2-8596-7644d85d2322 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.883626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:36.434918Z digest=sha256:3dbf869ac5173f60361df63755ef16d38fe63bdff2f992352eae8ef7fcf3fb5e

Observation 3ea49173-02b8-4e6d-9f24-224ee2b21304 · outbound

This paper cites Naturalbench: Evaluating vision-language models on natural adversarial samples,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Naturalbench: Evaluating vision-language models on natural adversarial samples,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.873717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:36.531122Z digest=sha256:31b663786681809a92f17ebd5fdd64823ed827046f56b0e4092ab01da8c9f2b3

Observation 2f3f7607-0885-427b-932b-be059f4fe581 · outbound

This paper cites Openclip,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Openclip,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.863477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:36.592231Z digest=sha256:47062143d8d6a80647eb3a97666effb8a9cfba59d05ca3701beab00e9b801595

Observation 6ffaac83-c65c-4172-b043-419a46750ca0 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Sigmoid loss for language image pre-training,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.853467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:36.661149Z digest=sha256:6f9d05e48c1b89f5be0897677e7a2f585169598397f378522531c299e0cdaa97

Observation 7e2d64dd-5c70-499f-b96d-4a57f5104283 · outbound

This paper cites Data filtering networks,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Data filtering networks,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.843141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:36.753898Z digest=sha256:4b585214b06c110479d4891f4b4b3bc09ab1a38b522d188bd529aa72a102fc76

Observation 298bcc47-91a2-48c3-a73b-f196b9054cd7 · outbound

This paper cites Demystifying clip data,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Demystifying clip data,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.833840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:36.822538Z digest=sha256:060de17837fda8cb0f668c7bd1f0bdb007b8737bb6e8a7b04c4527da73d8c21d

Observation f2e55f68-91ff-47ae-b704-e53d070df640 · outbound

This paper cites Eva-clip: Improved training techniques for clip at scale,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Eva-clip: Improved training techniques for clip at scale,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:36.885560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:36.885560Z digest=sha256:a668e016ed499ad5477bdee012967772f563337ab159564f459222d8e05237b6

Observation e10e12ee-90e0-4599-8d23-15f3ac9f1acf · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:36.954496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:36.954496Z digest=sha256:e2ebf8fdaf34d3fa774beb53bedfb91382a990777d5af6cff4fa979303fa462d

Observation 7d47a46a-4416-4d94-b899-9dc929b9eccc · outbound

This paper cites Accurate computation of the log-sum-exp and softmax functions,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Accurate computation of the log-sum-exp and softmax functions,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.817386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:37.033529Z digest=sha256:1fc04314a1f9e03e0484af138fd7f599b645b1c2cdef4b594b10332810bc903a

Observation f2f69d76-9387-498b-b3cd-f189c898635f · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Winoground: Probing vision and language models for visio-linguistic compositionality,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.806943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:37.111149Z digest=sha256:2910863163c5587a623ecdd508b89bdab82124cbe873bc096b53bd6c191203d7

Observation e1eba4e4-6e52-40d7-b4ef-7a107f9ca734 · outbound

This paper cites Similarity of neural network representations revisited,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Similarity of neural network representations revisited,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.796071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:37.189998Z digest=sha256:62e87dd3b2c54fa96bfc4bbecef3ea32029e1dd28879f78bbb8f9f3fd516dba7

Observation a693805c-afd9-4fbe-825d-cd9103233a47 · outbound

This paper cites Cider: Consensus-based image description evaluation,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Cider: Consensus-based image description evaluation,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.250094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:37.250094Z digest=sha256:9fe615f0b6ec9a263bd37084b11fa81dd94ec4da63e21094d43c926b233c0da7

Observation 9c03cdc7-06dd-4d99-8b0c-27ce0bf0c6ab · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Spice: Semantic propositional image caption evaluation,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.741941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:37.333853Z digest=sha256:93ff5d190ab6ab36686c725fdfcfa62289fc32e9b551f11d6d56480938e16946

Observation 95b8348a-10c1-416b-a666-9ee0fe2a9ff0 · outbound

This paper cites Decoupled weight decay regularization,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Decoupled weight decay regularization,

Reference 69

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:52:37.725447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:37.421346Z digest=sha256:c4f9407e1d69ce9231b0b157f6b3b969889a8cce64bb499add4f959ab208196a

Observation fdb9678a-2a92-4434-9134-38e0b3a60d74 · outbound

This paper cites Frisbees.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Frisbees

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:52:38.503754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:52:37.501612Z digest=sha256:b75760f51ebc2deebfa8bd9dba7e499b5a0a7eb9d70c777f126791dfada49386

Pith citing papers

Observation 92a96c14-d83f-43c4-b8d8-62cbd1f2168b · inbound

LAP: Fast LAtent Diffusion Planner for Autonomous Driving cites this paper.

LAP: Fast LAtent Diffusion Planner for Autonomous Driving Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T19:30:30.517564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:30:30.517564Z digest=sha256:ca6f8f25c4eb214e0f33d3fceac236c94af2a3d714a98ac0d026efa4336755c8