Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:24:55.227295Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2507.08513.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:24:55.227295Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
80 of 80 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 22aedcb8-1c69-4983-b2da-fb13d0a3b8c5 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Claude v3.0, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23ebc90b-877b-4010-9729-4f7bafea106f · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation ShapeNet: An Information-Rich 3D Model Repository
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee36dacf-f13c-487a-9253-a849d2a1a051 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c74c8ca2-29e0-459b-8ba7-1b55b8ab2e6c · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d46bf874-ec93-4731-aaae-f46aba4d136a · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a214032c-b48a-448e-b668-586e250128f4 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e78d2225-ca41-4d40-9ac2-964e0bcdc054 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Gonzalez, Ion Stoica, and Eric P
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0a5888e-9cac-4437-b9ef-a5f1d2f6639a · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Instructblip: Towards general- purpose vision-language models with instruction tuning,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d840d744-f67b-413a-a7ef-e6193407eb5d · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Objaverse: A universe of annotated 3d objects
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ace1d4d-271b-4046-a81c-b2968a3a5186 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Objaverse-xl: A universe of 10m+ 3d objects
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67c8ed4e-ed69-437f-8081-08a895706f20 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e86d4e51-015f-42ad-bb9f-5d43029e79b4 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 336ad5c2-4776-4d06-ae92-5c7c304cd5a7 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Training on Synthetic Data Beats Real Data in Multimodal Relation Extraction
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e757f10-5bb8-4a03-bdcd-822a6fc97c6c · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c730f22a-0a89-4e9e-8226-a4143ac99b18 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5dd674ed-5472-40dd-b3fa-5fe03143854e · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Kubric: A scalable dataset generator
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9356ccb3-194f-4f3a-b89f-5c7733380849 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Regiongpt: Towards region understanding vision lan- guage model
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 953075ef-eae8-48c0-b500-13bd6b9b84be · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Lvis: A dataset for large vocabulary instance segmentation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4cacf1a3-b313-423c-a996-4dda7392e2a0 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Vizwiz grand challenge: Answering visual questions from blind people
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9577be88-486f-4125-b157-813f5b0e6346 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Prompt-to-Prompt Image Editing with Cross Attention Control
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bf75061-9736-4928-9110-4fe24c39642c · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75415d62-cf08-44fe-a973-1581acc6461e · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04d06e66-6699-4701-9838-9557421ee249 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Denoising dif- fusion probabilistic models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31ab4b62-2a33-42f6-898d-1e9c7a8e186b · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 089ba357-1387-443c-8879-f522d36d1506 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 095a08c1-e4e2-4643-8893-59ab52b7ae9d · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8614b7fd-4116-423d-9fa5-881825ac47d2 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Learning Action and Reasoning-Centric Image Editing from Videos and Simulations
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b730cbda-5bef-43f6-8798-4771a4d6016f · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b07225c-fcd3-491d-a3f7-858971ef0997 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a10dd858-9346-4875-933a-7566f1fd5aac · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation What If We Recaption Billions of Web Images with LLaMA-3?
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e92f5ebe-1327-4337-bc68-0d2a152b7e54 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9b3ccb0-a351-44e7-9cea-820c995e8ae4 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Vila: On pre-training for vi- sual language models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f35e224e-dd1e-4a27-b02c-93a3aee0d20e · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Microsoft coco: Common objects in context
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 814c3bba-122f-4c28-9e23-cae8c89f935f · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Visual spa- tial reasoning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42b1530c-475b-427e-a737-908dacbf949c · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Improved baselines with visual instruction tuning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48815cf8-cfef-4993-a696-fbfce8c671f3 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Visual instruction tuning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d370933-8331-4859-a900-570bb89c1077 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Clevr-ref+: Diagnosing visual reasoning with referring ex- pressions
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0607d074-6d96-4b08-9d9c-b9ba782c61a4 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed303b27-b88d-462f-a43b-73855064240f · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab1081b8-88ad-4e2c-802a-313c49b41b59 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Generating images with 3d annotations using diffusion models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6cd09cf-de71-4d48-9097-e578dad25472 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4dd85980-213a-4cb5-871d-50e3964e5b8b · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Llama 3.2 vision, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c94a93ae-c02f-4e07-a0b7-2b79fc18a6e4 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Ocr-vqa: Visual question answering by reading text in images
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae0ea9c-8307-4059-b057-82dd1cdba90a · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation GPT-4 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1661cc17-f57a-44b8-b047-43babe783fc7 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation DINOv2: Learning Robust Visual Features without Supervision
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca9c2755-f9d7-435a-8a1f-717b50699f4e · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4052a144-3fa7-4a01-a69c-8fa3c4b4411a · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation High-resolution image synthesis with latent diffusion models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd505345-76be-404b-9b10-275b36819337 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Imagenet large scale visual recognition challenge
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9875cc28-2f98-468c-a8d9-b47623ed61cb · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d002c5df-03fe-4763-a28f-23eee2621d6b · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4797100-dcb6-4087-bc71-4461a0599388 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Towards vqa models that can read
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8aa1ea9-f8e2-4318-9320-7e0eb3c59cde · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Denoising Diffusion Implicit Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d583e9-9e00-4bba-96a4-659e5838cd43 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b332fb1-d509-4478-b476-b450b80f5090 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 57934c97-273e-42da-9534-3393f50de742 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaMA: Open and Efficient Foundation Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e37eb4b-4166-45b0-b6df-e75df5fd21a9 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8423d23e-82ee-47e6-9527-49274baf82e9 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Imagen editor and editbench: Advancing and evaluating text-guided im- age inpainting
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd5b5f6f-9e11-4231-831c-620c733cfe23 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Mebow: Monocular estima- tion of body orientation in the wild
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03e92d3b-bf6b-4034-992f-3cc3dd659b52 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Beyond pascal: A benchmark for 3d object detection in the wild
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 117e2623-f045-42d3-91da-bc5c6bedd6b6 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Imagere- ward: Learning and evaluating human preferences for text- to-image generation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a3ab2d6-e4ce-4703-be7f-c99d03379062 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a79d9d7-e38a-465d-85c8-7c341d6176ed · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ae6438-1357-4865-a2aa-6fea4b191f89 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Modeling context in referring expres- sions
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 720d82eb-b1da-4161-9cf6-1fba16e76b26 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation When and why vision-language models behave like bags-of-words, and what to do about it?
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63f700ad-48dc-4077-98d8-270a195668c4 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Magicbrush: A manually annotated dataset for instruction- guided image editing
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f26794e-2a13-4c39-9435-d62985f49828 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Adding conditional control to text-to-image diffusion models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40b253b0-067d-4f80-b54f-90b17167e97a · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c59fd0ac-0659-49c7-88c2-f56a3efe1277 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f97af725-1353-4477-a6e1-299e62af70b7 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Semantic under- standing of scenes through the ade20k dataset
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7931a257-a464-4c0c-b3df-4fbbc23f9357 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Corresponding section in main paper is Sec
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4260af33-793b-4be7-b525-314b47663afb · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation front" also means
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 662fd4c4-96c5-4aa1-a11b-8950e0b57382 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SYSTEM_PROMPT_FOR_GRADING_MLLM_RESPONSE =’You are a helpful and precise assistant for checking the quality of the answer
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11274d10-18db-4291-b450-0af94b271366 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 6, we show an example of using ImageReward [60] for the dataset curation by evaluating the alignment be- tween generated image and text prompts
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a833ed89-9c6d-4823-a8ec-80624261cff8 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 3, we perform quantitative comparisons on general image visual quality between different DM backbones: SD V1.5 and SDXL
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8e97d91-b1ca-4777-87d7-625d4fce5795 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6fd6cb42-cbf1-427b-a8aa-f8f26501081d · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 9 shows the UI page of our user study
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f80025b9-0df0-4d73-a614-011c00d5989c · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 10 to show more qualitative comparisons between fine- tuned LLaV A model to commercial SOTAs
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce8877c5-f5bf-41fc-b1a7-c7c842b22480 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 11, we shows more examples of diversity on object categories, camera-object relation, and background con- texts
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b09579b-8e9e-4749-9a60-76040f022433 · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad2ed1f9-ff4d-4f55-9408-d2fbb7629d6e · outbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.