Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:18:37.087313Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2412.16771.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:18:37.087313Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e38155e0-4d05-4cb7-bba3-222faacc11d6 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e34f2373-749a-43c2-bd2f-0c814adb5909 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4f962b5-abea-419c-b305-f453d600d2ea · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Flamingo: a visual language model for few-shot learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f2e9893-daa0-49c3-af16-b6a92ef2f91e · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Vqa: Visual question answering
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2a6a0260-a01c-4683-b7dd-6547a4b2878c · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 459c0c06-554b-4ee1-9749-d88fe2a3e16b · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0312c211-c292-45d2-8d1e-e3a8193549b3 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization wav2vec 2.0: A framework for self-supervised learning of speech representations
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fb064cbf-cd5a-49cc-87e5-603c74f6b0ab · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b40f0961-f67c-428f-a555-36d36ea718e3 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Introducing our multimodal models, 2023
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dbfaacb0-d65d-4cd1-b8b6-b015e61627af · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Language Models are Few-Shot Learners
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16c793de-ab37-45a7-a931-0b41ccdc0a2a · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b59c0da-a7c9-499f-b363-7351c3a535b4 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization A simple framework for con- trastive learning of visual representations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d94c25f4-a825-4bcb-ad53-fb20ee5ee581 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Improved Baselines with Momentum Contrastive Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c5ab8c-8f7b-494c-a089-4e946acde688 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Qwen2-Audio Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e1541a-ce3c-4549-a307-2e3c55995d51 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Simple and controllable music generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation aa87ad16-16f0-4646-883c-dcc999720988 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aefeb02f-3139-4ab8-b43e-941f9011e016 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization The Llama 3 Herd of Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc0100a1-bd19-4a57-ba10-50f819f943cc · outbound
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 09b7180b-b8da-4e02-9e10-587fb06203fa · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bde0622f-a211-413d-8959-07c100cfb1af · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b63b6cd6-1c1b-48e6-b72f-eeb59c7f6661 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Aligning ai with shared human values
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ade39451-50ce-4a8e-8cae-3dd90e9a4bf1 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Measuring massive multitask language under- standing
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5c309c09-ffe4-4cd5-bccb-a0a92548dc84 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Gaussian Error Linear Units (GELUs)
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 619ad864-1964-48fa-b65d-d0c55c5fded3 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Scaling up visual and vision- language representation learning with noisy text su- pervision
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 92a0d0e3-6832-4380-8d6f-873fcd73df42 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Chat-univi: Unified visual representation empowers large language models with image and video understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f9579513-82f7-4417-9f25-5b6b763b5cba · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Shamma, Michael S
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6d3628b0-1b2f-4458-acf8-b765d56e114e · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Lisa: Reasoning segmentation via large language model
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c374cde4-be6e-40cb-94e2-9aa4d54dd174 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb2d7cb3-d6b3-4062-9d5c-a4690339954e · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large lan- guage models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 42bd135b-0d8c-4205-bfa3-f35f6e8eb26a · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Competition-level code generation with alpha- code
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a3a2b95a-0413-4d0f-bc5a-23d8b7738aa9 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Microsoft coco: Com- mon objects in context
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation befac284-10db-426f-978c-4c4a32521ff7 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Improved baselines with visual instruction tun- ing
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 51897060-a786-4823-a143-68342533368b · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Visual instruction tuning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 08f9ab68-0c0e-413f-8bfc-30e8e2a42d1f · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Mmbench: Is your multi-modal model an all-around player? In Eu- ropean Conference on Computer Vision , pages 216–
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e80ef3e3-5524-4240-8ffb-5cb335fd151f · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Decoupled weight decay regularization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ecd81a06-f3e6-4d79-88e5-753a2c66ae88 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Learn to explain: Mul- timodal reasoning via thought chains for science ques- tion answering
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 62583d58-409c-40de-a2b0-4508efadae9d · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Cheap and quick: Effi- cient vision-language instruction tuning for large lan- guage models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2dc24cef-ef1f-4e3c-8b6e-f9ca1c774691 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 332fd706-58dd-4250-a898-23da7f04fab7 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization GAIA: a benchmark for General AI Assistants
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3e50154-f6f7-496a-85c2-100c9f8c7a52 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Foundation models for generalist medical artificial intelligence
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c680a341-7a5c-4e14-b967-c2edc00707e4 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0f8a3fc0-c640-43a8-a42a-8e65c678cb67 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Hello, gpt-4o, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0ce719a4-2f4a-41d2-8ea9-c490021cf772 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Learning transferable visual models from natural language supervision
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa93a7ea-1275-4055-aa81-06599b4edbb6 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Ro- bust speech recognition via large-scale weak supervi- sion
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 85f195c7-d613-470f-9948-0f78451472b1 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Language-based action concept spaces improve video self-supervised learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bc2f41eb-709c-42a5-8be3-8d5f40051593 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Learning to localize objects improves spatial reason- ing in visual-llms
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7e0adb7c-32b2-46be-9894-b53d69f2805a · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85b2b88e-6beb-47ad-b1ea-80d8555e1b7a · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Laion-5b: An open large- scale dataset for training next generation image-text models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation aff80d5a-e074-4bfa-b9ae-63fb309a0faf · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Hugginggpt: Solv- ing ai tasks with chatgpt and its friends in hugging face
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 443f7f8b-8b91-40a6-adac-1cab8eb7ccc3 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization SALMONN: Towards Generic Hearing Abilities for Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb7cbeb8-0834-4276-b634-1a2f9605f789 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Movieqa: Understanding stories in movies through question-answering
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 90bdbea7-9691-4b58-a967-112ba279d948 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization LLaMA: Open and Efficient Foundation Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 617d7226-bf4e-4f9d-b76f-9ee0a892f1ad · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f6bf28b-fb9d-4fc6-aca7-cb34fa842092 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Efficient utilization of large pre-trained models for low resource asr
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 30e37c54-f640-4ccc-a9bd-ad19b8ff641c · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c67679e-2587-46fd-bf77-1435bcda9beb · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03f6d93d-8bfd-40e1-918a-47977c128466 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Visionllm: Large lan- guage model is also an open-ended decoder for vision- centric tasks
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 18aff06f-53b5-412a-a9ff-8edc99353f64 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Finetuned Language Models Are Zero-Shot Learners
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 737edb95-c53c-4ab0-8ade-cea8be816922 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Chain-of-thought prompting elicits reasoning in large language models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 16bf530b-c262-4a9a-acc7-fddb5e19921e · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Hard prompts made easy: Gradient-based discrete opti- mization for prompt tuning and discovery
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4a30ec71-fd78-4fa6-9f6a-c1f04bc3bda5 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b0996d9-4110-46c0-bec6-68ca8e9c6c90 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization MM-BigBench: Evaluating Multimodal Models on Multimodal Content Comprehension Tasks
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7b41eac5-f975-4886-90f0-0b84e28863fb · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Tree of thoughts: Deliberate problem solving with large language models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4b999423-825e-491c-abad-e975a6d51b8a · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization From image descriptions to visual de- notations: New similarity metrics for semantic in- ference over event descriptions
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5fe61c81-6674-44ca-8e78-62ecac9f7c70 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0ea930b-99e6-4255-8723-f3812a007cab · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 606538b6-dfe2-4648-a2ce-557b2765661c · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 027c6b6b-388f-45ea-8a9f-9d62e9dc50dd · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0066eee-12ee-4eeb-a1e7-bffe7c37470e · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language mod- els
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8269c6d5-70a5-47be-9d32-af2f1d2ba8ac · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization P Xing, Hao Zhang, Joseph E
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5ce1ee04-c69c-478a-9513-2225377861f5 · outbound
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.