Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:52:37.501612Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2507.07106.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:52:37.501612Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T19:30:30.517564Z
A source-named dated measurement, never combined with another source.
Source: cited_works
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e8b87fd0-240e-418d-b6cf-8ec87ed4eb41 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Eyes wide shut? exploring the visual shortcomings of multimodal llms,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a12d77e5-20f2-4573-90e9-81ca44e1ab33 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Learning transferable visual models from natural language supervision,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b602ef68-7beb-436a-b7af-92b84f6a69a2 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3734393-a69e-4769-a1e9-4e7e7c1d021c · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Is clip the main roadblock for fine-grained open-world perception?,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 625bdc2b-6844-4ae1-ac8f-0e68ed309af3 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf309773-992a-4ce3-a6c3-4332c53bc974 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor BRAVE: Broadening the visual encoding of vision-language models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67cf8678-d817-4b9f-bb5c-00e6edf6170c · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea4cd2ef-2034-4a1b-8b7a-f9b908f0b8ea · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6fd8af5-b0b1-4324-8852-e6e4af2443fc · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Mini-gemini: Mining the potential of multi-modality vision language models,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9b9b89d-4508-462d-80ed-2c3b2ab5c80e · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Prismer: A Vision-Language Model with Multi-Task Experts
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb52f24-dbbb-4ab1-81ec-a3d5c5cbafc0 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Vcoder: Versatile vision encoders for multimodal large language models,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7103154-94f7-4435-bfa8-ac2c332fb38a · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Question aware vision transformer for multimodal reasoning,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7283de56-cd47-43d2-a110-a861513486d7 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Api: Attention prompting on image for large vision-language models,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cefb533f-5313-443f-b4f5-cca3a74cf1db · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Instructblip: Towards general-purpose vision-language models with instruction tuning,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51279896-4373-4237-84bd-1fdf372892e6 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df85e6de-4f47-4c97-afdc-2213358707fb · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor High-resolution image synthesis with latent diffusion models,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbcd5294-5e78-4b1c-bd97-3c5db57e9a4a · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Photorealistic text-to-image diffusion models with deep language understanding,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 296a3972-4983-498b-ad94-39a6c8bc9d1a · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Hierarchical text-conditional image generation with clip latents,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0901327-ae10-4002-9c21-bba7c7ce2154 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Sdxl: Improving latent diffusion models for high-resolution image synthesis,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55412a86-d16b-48bd-a6f9-528b92f8f1d8 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Prompt-to-prompt image editing with cross attention control,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b99865be-9966-49d5-ae52-cd4fc760b7dc · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor What the daam: Interpreting stable diffusion using cross attention,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5715627b-f108-4911-a7ea-6493cff780c5 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Towards understanding cross and self-attention in stable diffusion for text-guided image editing,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f68a552d-440a-4c04-83f3-8c6fe427aa1e · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Plug-and-play diffusion features for text-driven image- to-image translation,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3de6ec26-b7fb-41ef-9c92-b1afb1298272 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e32935a-f2fa-4f14-879f-b4239b68da79 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Repurposing diffusion-based image generators for monocular depth estimation,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c4dce4f-505e-4580-bd26-cbd6e386ba1f · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Do text-free diffusion models learn discriminative visual representations?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a895a7-2e78-42cf-bce6-5f52c4eb49a3 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Deconstructing denoising diffusion models for self-supervised learning,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98a009da-c65d-4f22-9c99-2857b6e2c9a8 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Coca: Contrastive captioners are image-text foundation models,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 790d6f1e-5c0d-4a5c-9047-606095c7cbd9 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Multimodal few-shot learning with frozen language models,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5a0be6c-107b-4935-9338-c4a9398dd94c · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Flamingo: a visual language model for few-shot learning,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f20a7514-0702-4d53-ad7d-f58aaafc9dfa · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Visual instruction tuning,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0517812-1c5b-4029-9d48-4d901626aca3 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cc71a27-f8e9-43c7-b0f5-e29d26238011 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 556d3ef5-1f4b-4c51-aee0-77d103f946d8 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 848a5aaa-b4e9-4115-82e6-decb4d529af6 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Evaluating Object Hallucination in Large Vision-Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54725c8f-4675-4ac8-abfb-c2e9f1e25b01 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Multi-modal hallucination control by visual information grounding,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f489965-c74f-4fb8-aa60-306084b23fae · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Detecting and preventing hallucinations in large vision language models,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b55157d-f698-4213-833c-1e1eefac95f1 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36fe1c95-1620-4561-a3cd-5e6037917e1a · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor HallE-Control: Controlling Object Hallucination in Large Multimodal Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9bcfcae-e0d3-4335-bc9a-5553952852ae · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab5b5abd-9b09-4f0f-bccc-629838e11b29 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 201ad2db-d2af-4eb2-b99f-c19dc5e6cf4f · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Llava-next: Improved reasoning, ocr, and world knowledge,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26b9cebe-850f-4bff-b75f-82e4460c81a5 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a775c40-908f-4be4-97fc-d183c552ff0e · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Improved baselines with visual instruction tuning,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aa7ebc1-2d2f-4923-9cc2-1c696551b0bf · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor 4m: Massively multimodal masked modeling,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e08a5649-69ea-45a7-8806-6160ae111cfd · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 958be85f-58a5-495e-b952-93836b10ad0a · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Llava-plus: Learning to use tools for creating multimodal agents,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 524a1413-4b8c-4d16-94c5-8e3364a96190 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Visual programming: Compositional visual reasoning without training,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd605569-c973-412d-a5ba-9693f9cea517 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Spatialbot: Precise spatial understanding with vision language models,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adba8552-2fb7-450e-9f5f-74e68331758f · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Your diffusion model is secretly a zero- shot classifier,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6958a577-10b3-40c1-bfe4-a581be05c5bb · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Diffusion Models Beat GANs on Image Classification
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 396ca2fd-a5fa-4ad7-9dc9-4f33babf60e7 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57b935b2-94bf-41df-af2b-48c18dd48d23 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Diffusion Models for Open-Vocabulary Segmentation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2801e47e-4379-472f-b257-71088ab13e4b · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Not all diffusion model activations have been evaluated as discriminative features,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d74e8ff-50e6-4bf4-b698-08dacdd59fbb · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Dinov2: Learning robust visual features without supervision,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00a9423d-718c-4cc2-8596-7644d85d2322 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ea49173-02b8-4e6d-9f24-224ee2b21304 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Naturalbench: Evaluating vision-language models on natural adversarial samples,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f3f7607-0885-427b-932b-be059f4fe581 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Openclip,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ffaac83-c65c-4172-b043-419a46750ca0 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Sigmoid loss for language image pre-training,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e2d64dd-5c70-499f-b96d-4a57f5104283 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Data filtering networks,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 298bcc47-91a2-48c3-a73b-f196b9054cd7 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Demystifying clip data,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2e55f68-91ff-47ae-b704-e53d070df640 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Eva-clip: Improved training techniques for clip at scale,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e10e12ee-90e0-4599-8d23-15f3ac9f1acf · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d47a46a-4416-4d94-b899-9dc929b9eccc · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Accurate computation of the log-sum-exp and softmax functions,
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2f69d76-9387-498b-b3cd-f189c898635f · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Winoground: Probing vision and language models for visio-linguistic compositionality,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1eba4e4-6e52-40d7-b4ef-7a107f9ca734 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Similarity of neural network representations revisited,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a693805c-afd9-4fbe-825d-cd9103233a47 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Cider: Consensus-based image description evaluation,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c03cdc7-06dd-4d99-8b0c-27ce0bf0c6ab · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Spice: Semantic propositional image caption evaluation,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95b8348a-10c1-416b-a666-9ee0fe2a9ff0 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Decoupled weight decay regularization,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fdb9678a-2a92-4434-9134-38e0b3a60d74 · outbound
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Frisbees
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92a96c14-d83f-43c4-b8d8-62cbd1f2168b · inbound
LAP: Fast LAtent Diffusion Planner for Autonomous Driving Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.