Pith. sign in

Paper Citation Record · LEDGER

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation

As of 7 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2507.08513.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08513 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:24:55.227295Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22aedcb8-1c69-4983-b2da-fb13d0a3b8c5 · outbound

This paper cites Claude v3.0, 2024.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Claude v3.0, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.970840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:26.349606Z digest=sha256:51de094cac9cc697eac63744cb32911503fb33506e01120dc56140491a04785a

Observation 23ebc90b-877b-4010-9729-4f7bafea106f · outbound

This paper cites ShapeNet: An Information-Rich 3D Model Repository.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation ShapeNet: An Information-Rich 3D Model Repository

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:26.443277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:26.443277Z digest=sha256:1f92d3e45d25c26b23380deb7ce0660a9529e931cfdf7620f73097bcad9ca708

Observation ee36dacf-f13c-487a-9253-a849d2a1a051 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:26.519860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:26.519860Z digest=sha256:9cd650f8db819436cb1de65a1f03e690a5eddf1e5b91693da7615da85faaebe3

Observation c74c8ca2-29e0-459b-8ba7-1b55b8ab2e6c · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:26.584334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:26.584334Z digest=sha256:8e38b94cb162263a3cf47a6917c02de1ce0dd92b2ae8e6d912d22d58537ec7a3

Observation d46bf874-ec93-4731-aaae-f46aba4d136a · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.359030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.359030Z digest=sha256:e799eea251181bfed9952aa102208609066624a9bb576b2640ccee2bc0f0c68d

Observation a214032c-b48a-448e-b668-586e250128f4 · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.450996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.450996Z digest=sha256:6120a97b4ba51a1240466b012a87b43952a7c380ff80b855e0f090d428f5b42d

Observation e78d2225-ca41-4d40-9ac2-964e0bcdc054 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Gonzalez, Ion Stoica, and Eric P

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.501932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.501932Z digest=sha256:2fa5f4e882d862b6695a79197b32137bff81e17778a82daf58d2e7200b632b4a

Observation c0a5888e-9cac-4437-b9ef-a5f1d2f6639a · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.640240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.640240Z digest=sha256:2b73618ae626a0220a55edcf42abc467b5fed8dbf1a11d47016521757a2371bb

Observation d840d744-f67b-413a-a7ef-e6193407eb5d · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Objaverse: A universe of annotated 3d objects

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.904748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:50.721353Z digest=sha256:f83ce672c3744badb2f27ce71a75179d834a46484f8ae8ee283fae6deca53118

Observation 9ace1d4d-271b-4046-a81c-b2968a3a5186 · outbound

This paper cites Objaverse-xl: A universe of 10m+ 3d objects.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Objaverse-xl: A universe of 10m+ 3d objects

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.887292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:50.780968Z digest=sha256:32a06c58c984640aed64be5e09f65cbae974c85a3a8b57113c0393de3454dc07

Observation 67c8ed4e-ed69-437f-8081-08a895706f20 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.875281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.875281Z digest=sha256:db668d55d4fad09d545001189d3f7aec1d02857ac7e043ef9528d5e304a7de07

Observation e86d4e51-015f-42ad-bb9f-5d43029e79b4 · outbound

This paper cites What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.966995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.966995Z digest=sha256:3162145df89c3e1e4f2808a133a9f8efe36d4c99efa90f8b8ef22b601bcab008

Observation 336ad5c2-4776-4d06-ae92-5c7c304cd5a7 · outbound

This paper cites Training on Synthetic Data Beats Real Data in Multimodal Relation Extraction.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Training on Synthetic Data Beats Real Data in Multimodal Relation Extraction

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:24:55.779515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:51.055242Z digest=sha256:a47bbf41d841a65b29e315f9d5370f37dbc61cffc2b613a994f6e0a9c5e50542

Observation 4e757f10-5bb8-4a03-bdcd-822a6fc97c6c · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.096889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.096889Z digest=sha256:515358ad25b949fb2653f73b1235074689091eaddba546f42591de872a9d88be

Observation c730f22a-0a89-4e9e-8226-a4143ac99b18 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.868412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:51.180173Z digest=sha256:81f81b7abd352ec294176337669c7fa2728a62367ac4a99667fe1bf58c43964c

Observation 5dd674ed-5472-40dd-b3fa-5fe03143854e · outbound

This paper cites Kubric: A scalable dataset generator.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Kubric: A scalable dataset generator

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.852158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:51.252372Z digest=sha256:fda040fb098bfb3d6a03c9086ba5fdbffbe7f6ef4f3e64520bb7ce30a3ef1fc3

Observation 9356ccb3-194f-4f3a-b89f-5c7733380849 · outbound

This paper cites Regiongpt: Towards region understanding vision lan- guage model.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Regiongpt: Towards region understanding vision lan- guage model

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.832620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:51.328234Z digest=sha256:cfd9cbc2e36c197b04ae61b9eb501af35eb48fbc1a62b2bba99b217039f500cc

Observation 953075ef-eae8-48c0-b500-13bd6b9b84be · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Lvis: A dataset for large vocabulary instance segmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.814743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:51.398073Z digest=sha256:4d48992c1177884e0e4419a4fe106500598a2f7adda06d2c9c41c9aa27defb3f

Observation 4cacf1a3-b313-423c-a996-4dda7392e2a0 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Vizwiz grand challenge: Answering visual questions from blind people

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.445774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.445774Z digest=sha256:f37bd9b8b35bb3d9c5482584e80804492a92de22a4258fd9088bf26972cde8e0

Observation 9577be88-486f-4125-b157-813f5b0e6346 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.489571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.489571Z digest=sha256:b40f2358153d06ca1db4f8c53b165955f7fa0a9329c1c9b321e13d2e72ef13bb

Observation 2bf75061-9736-4928-9110-4fe24c39642c · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.541947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.541947Z digest=sha256:ebe0d6f9ddecff93a0854d967e3040733d6c17d67ee72ffd26b50675c955f1fd

Observation 75415d62-cf08-44fe-a973-1581acc6461e · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.592721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.592721Z digest=sha256:3839e96e119e1e088d09aae22aaa98d4e8a576613c53d3899faf4b7196628faa

Observation 04d06e66-6699-4701-9838-9557421ee249 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Denoising dif- fusion probabilistic models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.640713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.640713Z digest=sha256:8e034218cf25a771cb6120d8df9f60abe5207a58b7e6e863f5ad60ed9d022f74

Observation 31ab4b62-2a33-42f6-898d-1e9c7a8e186b · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.756961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:51.712219Z digest=sha256:78ad34131c88f8b814b2ddd372c70997d1ce598a43e8c06f47a5c26e80bf9f37

Observation 089ba357-1387-443c-8879-f522d36d1506 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.739583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:51.765806Z digest=sha256:2c1a5a554ad2fc28118e9e39bd4e4f49c49072cfb810730ca225437f1917e98a

Observation 095a08c1-e4e2-4643-8893-59ab52b7ae9d · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.825630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.825630Z digest=sha256:2b16fa88910964f1ebcd635ca67da94efe01e9c22a796499da85e5ef654826ab

Observation 8614b7fd-4116-423d-9fa5-881825ac47d2 · outbound

This paper cites Learning Action and Reasoning-Centric Image Editing from Videos and Simulations.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.875150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.875150Z digest=sha256:b5b3859e3cb9e6285844ca0a8fca5c6b554d785647782fd6cb3a01c7c6c2d69a

Observation b730cbda-5bef-43f6-8798-4771a4d6016f · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.915011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.915011Z digest=sha256:a9ee3f68713a4ce7c575ea2389d5782f7fad38e5b326a7116a0c2ea40458864f

Observation 0b07225c-fcd3-491d-a3f7-858971ef0997 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.958093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.958093Z digest=sha256:8fc96009a20ad090e8b642b3c8eb9cc70972c2e01941ae0525f647bc8749455a

Observation a10dd858-9346-4875-933a-7566f1fd5aac · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.066023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.066023Z digest=sha256:ac5a31618a8cb6f62fdd269e246547acc01ef69c9b42f6668b776770784ebc25

Observation e92f5ebe-1327-4337-bc68-0d2a152b7e54 · outbound

This paper cites StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.156204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.156204Z digest=sha256:526bb2c89012b47053ae1be523f0b81eff773ab92c56232b2cb2c43df4822acc

Observation b9b3ccb0-a351-44e7-9cea-820c995e8ae4 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Vila: On pre-training for vi- sual language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.687602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:52.337217Z digest=sha256:68661a093dbc27b0b06e7c596168ea501d0fd496435d88c9ca4c671e41ddb0ed

Observation f35e224e-dd1e-4a27-b02c-93a3aee0d20e · outbound

This paper cites Microsoft coco: Common objects in context.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Microsoft coco: Common objects in context

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.445832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.445832Z digest=sha256:c30eccd5c6e33983fda1eae0e3e0b65cd5c6c39132d84db0850f4cab13a943c7

Observation 814c3bba-122f-4c28-9e23-cae8c89f935f · outbound

This paper cites Visual spa- tial reasoning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Visual spa- tial reasoning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.659790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:52.521478Z digest=sha256:c80f4700826010097c8bf4bb0f176b0f02f850071dbb152a329e5e9cbdf9495c

Observation 42b1530c-475b-427e-a737-908dacbf949c · outbound

This paper cites Improved baselines with visual instruction tuning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Improved baselines with visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.642709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:52.581229Z digest=sha256:14056e369dc61233387f1b7e8299780cd1e7b1e5fef07a2afad11444e04bb78f

Observation 48815cf8-cfef-4993-a696-fbfce8c671f3 · outbound

This paper cites Visual instruction tuning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.651502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.651502Z digest=sha256:70119de9509edef1939e0fe639bb1b68c56e5c627ebf62f9b618489732ce43b8

Observation 4d370933-8331-4859-a900-570bb89c1077 · outbound

This paper cites Clevr-ref+: Diagnosing visual reasoning with referring ex- pressions.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Clevr-ref+: Diagnosing visual reasoning with referring ex- pressions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.617137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:52.714850Z digest=sha256:019f17a2b1ae44e2102ff3c117289ba075820d2b3bff7b0d69174735c81edc97

Observation 0607d074-6d96-4b08-9d9c-b9ba782c61a4 · outbound

This paper cites SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.771144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.771144Z digest=sha256:fde80486bcfe698c73838b120057b10c26fc3f46452b39abe1488d6c52e66d49

Observation ed303b27-b88d-462f-a43b-73855064240f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.830815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.830815Z digest=sha256:3476a15a1ae70f1326091bf983f1f3d3d0b416eb067146cc5c9c075e5b68c59f

Observation ab1081b8-88ad-4e2c-802a-313c49b41b59 · outbound

This paper cites Generating images with 3d annotations using diffusion models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Generating images with 3d annotations using diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.588673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:52.873041Z digest=sha256:5ff96e563fd4f8c664b3f9c0d9c981c971e462d93e9fd5454c72d554f38e9c10

Observation a6cd09cf-de71-4d48-9097-e578dad25472 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.570935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:52.914267Z digest=sha256:af379266f909177be6aecca3254c0c104f5e367c23ee5ca3eef342d069fdc374

Observation 4dd85980-213a-4cb5-871d-50e3964e5b8b · outbound

This paper cites Llama 3.2 vision, 2024.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Llama 3.2 vision, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.552750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:52.948441Z digest=sha256:1f9edce58c1fd4b173429077a2490e5eb7b1140acbe16506d717202e791fc62a

Observation c94a93ae-c02f-4e07-a0b7-2b79fc18a6e4 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Ocr-vqa: Visual question answering by reading text in images

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.982475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.982475Z digest=sha256:ae6db0d6fd2d8103fe596cb66aa5285067b0c911c47c351f0595130c718955c8

Observation 1ae0ea9c-8307-4059-b057-82dd1cdba90a · outbound

This paper cites GPT-4 Technical Report.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation GPT-4 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.028696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.028696Z digest=sha256:7dd9d17fb4dace402950fce9b21cccdfacfa2dc47e308631db11e7d5562c1ae5

Observation 1661cc17-f57a-44b8-b047-43babe783fc7 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.060999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.060999Z digest=sha256:3e4f018e22cadb27d47fbf518ad590dc36e6c08ee9b0ef5f64ee5955efd5d98e

Observation ca9c2755-f9d7-435a-8a1f-717b50699f4e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.119915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.119915Z digest=sha256:2cdcf72b1d252a901b025faace867016117209a6caa863aa31e51f2991306e8d

Observation 4052a144-3fa7-4a01-a69c-8fa3c4b4411a · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation High-resolution image synthesis with latent diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.524957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:53.198135Z digest=sha256:3fa64d148265b69c5062b05ecc6ae9fdd5feffdc1b98d6ee33aa2e3cb5eef44f

Observation dd505345-76be-404b-9b10-275b36819337 · outbound

This paper cites Imagenet large scale visual recognition challenge.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Imagenet large scale visual recognition challenge

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.303467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.303467Z digest=sha256:ae2a4e16536eb82bbf05f2bf6e2e80fd37678f171aa3b178cf28a41e9f2ab9ae

Observation 9875cc28-2f98-468c-a8d9-b47623ed61cb · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.390031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.390031Z digest=sha256:b45c95598ac8a36178f8af6c5e0f490ecc72204689ee9479ff193136ab7c36d9

Observation d002c5df-03fe-4763-a28f-23eee2621d6b · outbound

This paper cites Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:24:55.438877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:53.424779Z digest=sha256:52fd25453105b6bc3ddfd018c428881da41e73729f7c767d3f4ed50709b62fad

Observation d4797100-dcb6-4087-bc71-4461a0599388 · outbound

This paper cites Towards vqa models that can read.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Towards vqa models that can read

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.489465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.489465Z digest=sha256:69296cc974f18428d6ac46e983e606cec1eaca6878e97567c94fbe8ea9a6dfea

Observation b8aa1ea9-f8e2-4318-9320-7e0eb3c59cde · outbound

This paper cites Denoising Diffusion Implicit Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Denoising Diffusion Implicit Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.538670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.538670Z digest=sha256:52745391011ee3c734087b2391daf7b890583765300e8b485b759c4a0d05aed6

Observation c9d583e9-9e00-4bba-96a4-659e5838cd43 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.577193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.577193Z digest=sha256:5753a1975af973bac9869961ed3c63f3c23c874ef5bf6b3133b3b3926b9e8193

Observation 4b332fb1-d509-4478-b476-b450b80f5090 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.451740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:53.637886Z digest=sha256:c738672d53f91e8a5da42e87e9be8b2e323f0184e6b012206d8d991e7dcca2e4

Observation 57934c97-273e-42da-9534-3393f50de742 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaMA: Open and Efficient Foundation Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.692677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.692677Z digest=sha256:7d0603ea4de6537eeef8dc933096dc26903dc2da6b4a370b172e24240e21e24e

Observation 0e37eb4b-4166-45b0-b6df-e75df5fd21a9 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.783559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.783559Z digest=sha256:eb102a78e4f3bd3f1147d3eefb41ba83d1eff561aad0ac4f2ef09006e377ab15

Observation 8423d23e-82ee-47e6-9527-49274baf82e9 · outbound

This paper cites Imagen editor and editbench: Advancing and evaluating text-guided im- age inpainting.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Imagen editor and editbench: Advancing and evaluating text-guided im- age inpainting

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.419298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:53.842559Z digest=sha256:15877fa684de65ced19d404c108723e6aaac46dd125f21162e4116ce489a3c74

Observation fd5b5f6f-9e11-4231-831c-620c733cfe23 · outbound

This paper cites Mebow: Monocular estima- tion of body orientation in the wild.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Mebow: Monocular estima- tion of body orientation in the wild

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.390347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:53.966799Z digest=sha256:25488a27c9cccbf16828f7d42170758b93e6dc1b12db3bd6bed67c911361df86

Observation 03e92d3b-bf6b-4034-992f-3cc3dd659b52 · outbound

This paper cites Beyond pascal: A benchmark for 3d object detection in the wild.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Beyond pascal: A benchmark for 3d object detection in the wild

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.090317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:54.042393Z digest=sha256:180b62b21a9e578d55548a82204c260e9c7149be656fc904370b17de8bd4b1b5

Observation 117e2623-f045-42d3-91da-bc5c6bedd6b6 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:57.505282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:54.129706Z digest=sha256:5bf5364799032b59034007fa0979736c7c4e6d6f7456633a05c928dc275ec454

Observation 1a3ab2d6-e4ce-4703-be7f-c99d03379062 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.208155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.208155Z digest=sha256:8fd8dd89ae835159be72c6bf0bc4cb7be403ba4672449c9fea7f102b7fc3b837

Observation 6a79d9d7-e38a-465d-85c8-7c341d6176ed · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.285272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.285272Z digest=sha256:90a50e2bb07147801047294fa1e6f2125af5c1acd54495df3e15138cb7d00009

Observation a1ae6438-1357-4865-a2aa-6fea4b191f89 · outbound

This paper cites Modeling context in referring expres- sions.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Modeling context in referring expres- sions

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:57.281558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:54.353308Z digest=sha256:e470f80590db86f5ffc98fbef736f6d63a453f3bbf8846b4ecf6714d9274c1b2

Observation 720d82eb-b1da-4161-9cf6-1fba16e76b26 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.449939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.449939Z digest=sha256:4c37e2936d631d6c363501d87e8a0f0d451fb9e29824a970fe73fee9a0e2fa2d

Observation 63f700ad-48dc-4077-98d8-270a195668c4 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction- guided image editing.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Magicbrush: A manually annotated dataset for instruction- guided image editing

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:57.170158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:54.528019Z digest=sha256:4a9645bfcc683fa5b31f418fb9a178cdb61eb1387d16a53d926beaff7f4ead7a

Observation 8f26794e-2a13-4c39-9435-d62985f49828 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Adding conditional control to text-to-image diffusion models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:57.037280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:54.606014Z digest=sha256:6bb9cd78202573776a01a1a2436ac621a7bceff239f47fd705379659c845cf41

Observation 40b253b0-067d-4f80-b54f-90b17167e97a · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.680277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.680277Z digest=sha256:c3f5052676b06ff8a2afa3ae683f5eb9d200cc743a8d09df9886cb5b9321ed3f

Observation c59fd0ac-0659-49c7-88c2-f56a3efe1277 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.749240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.749240Z digest=sha256:abb119b4cbfb9cf29d09c8d4c75cfc8cae3fe889e7ca7e7ce27cb0d6e869caf3

Observation f97af725-1353-4477-a6e1-299e62af70b7 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Semantic under- standing of scenes through the ade20k dataset

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.904192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:54.790724Z digest=sha256:a7810a96eeaf105c16e08e1c9d3f71900a6f335b8823f453a4cc6b459eaf5745

Observation 7931a257-a464-4c0c-b3df-4fbbc23f9357 · outbound

This paper cites Corresponding section in main paper is Sec.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Corresponding section in main paper is Sec

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.816980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:54.830972Z digest=sha256:5cc90feb820ca382ad2cab89a921cbf831a453812d4ca8bdd0a47eebd4f5cf1f

Observation 4260af33-793b-4be7-b525-314b47663afb · outbound

This paper cites front" also means.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation front" also means

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.728795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:54.869213Z digest=sha256:38b5c170ab187ffbd66505da55e4b0d587c20ff14626c41ecadefaf63bf5666e

Observation 662fd4c4-96c5-4aa1-a11b-8950e0b57382 · outbound

This paper cites SYSTEM_PROMPT_FOR_GRADING_MLLM_RESPONSE =’You are a helpful and precise assistant for checking the quality of the answer.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SYSTEM_PROMPT_FOR_GRADING_MLLM_RESPONSE =’You are a helpful and precise assistant for checking the quality of the answer

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.651563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:54.929420Z digest=sha256:5a58e97c598c46c44861028fdf314af09a8b19fb02329e6ea2b87076a2c7f4e6

Observation 11274d10-18db-4291-b450-0af94b271366 · outbound

This paper cites 6, we show an example of using ImageReward [60] for the dataset curation by evaluating the alignment be- tween generated image and text prompts.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 6, we show an example of using ImageReward [60] for the dataset curation by evaluating the alignment be- tween generated image and text prompts

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.549201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:54.992499Z digest=sha256:fac084007dc2aed4c7c9265353bdb1543b6b9f649d650ac8ab9c96767dd87bbd

Observation a833ed89-9c6d-4823-a8ec-80624261cff8 · outbound

This paper cites 3, we perform quantitative comparisons on general image visual quality between different DM backbones: SD V1.5 and SDXL.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 3, we perform quantitative comparisons on general image visual quality between different DM backbones: SD V1.5 and SDXL

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.468624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:55.032116Z digest=sha256:336399573173ca073783c4c548671fe45d04edfa18e173ee637913acb54ff922

Observation d8e97d91-b1ca-4777-87d7-625d4fce5795 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:56.391242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:55.074253Z digest=sha256:04e8d892ba9ca2372c18cfde7d34069a4eca7c4f4bd2bfabc828dfd86f8b5e94

Observation 6fd6cb42-cbf1-427b-a8aa-f8f26501081d · outbound

This paper cites 9 shows the UI page of our user study.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 9 shows the UI page of our user study

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.320779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:55.116586Z digest=sha256:9a452b845a93f6e77f17f94a05b77281f8fd6f748c3626ba2e495420189396a2

Observation f80025b9-0df0-4d73-a614-011c00d5989c · outbound

This paper cites 10 to show more qualitative comparisons between fine- tuned LLaV A model to commercial SOTAs.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 10 to show more qualitative comparisons between fine- tuned LLaV A model to commercial SOTAs

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.259557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:55.153359Z digest=sha256:c8d5ee96eafae3d9527def6596ae2118b43b3f95edad1482a1bac84633b04555

Observation ce8877c5-f5bf-41fc-b1a7-c7c842b22480 · outbound

This paper cites 11, we shows more examples of diversity on object categories, camera-object relation, and background con- texts.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 11, we shows more examples of diversity on object categories, camera-object relation, and background con- texts

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.181755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:55.177843Z digest=sha256:6a01826eb3c9458a41a9c6ca7065851e7dc0e621191fb08cd5bb2e52d57593ee

Observation 7b09579b-8e9e-4749-9a60-76040f022433 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:56.077268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:55.205611Z digest=sha256:36437f06d5ed96d891a72383136aa25a3baa02795eeae7b67f3fa1c3abd9da63

Observation ad2ed1f9-ff4d-4f55-9408-d2fbb7629d6e · outbound

This paper cites an unresolved cited work.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:55.986032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:24:55.227295Z digest=sha256:ca54272a2eedfd73d796c528c56da7424fb435aee47c88e85019f161e42f7c5f

Pith citing papers

No inbound Pith citation observations are available.