Pith. sign in

Paper Citation Record · LEDGER

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization

As of 24 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2412.16771.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16771 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:18:37.087313Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact1
  • verified fuzzy37
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e38155e0-4d05-4cb7-bba3-222faacc11d6 · outbound

This paper cites GPT-4 Technical Report.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:35.523111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:35.523111Z digest=sha256:b4a671b0064a0691a6b8e8e1fd8ffcef1fca9dcf80b7dc86c3c8d674b6860197

Observation e34f2373-749a-43c2-bd2f-0c814adb5909 · outbound

This paper cites The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:35.527376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:35.527376Z digest=sha256:d0785282c3a08587f7198fade8078bebfb8053d0f838d3eec6f2f4224840d921

Observation f4f962b5-abea-419c-b305-f453d600d2ea · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:35.531190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:35.531190Z digest=sha256:e0aade98748b38f64e86ccec6f1d8278af064a1f73cfee1b25b34eb5c9d8c848

Observation 1f2e9893-daa0-49c3-af16-b6a92ef2f91e · outbound

This paper cites Vqa: Visual question answering.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Vqa: Visual question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:39.314884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:35.557843Z digest=sha256:edcaaffc4a61b2a0a14e66424171402894b789ea3d9394b08b1f73a32244d8a3

Observation 2a6a0260-a01c-4683-b7dd-6547a4b2878c · outbound

This paper cites SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:35.659415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:35.659415Z digest=sha256:32faef4af904783dfebf85a3ad1572df8e1cfcad5ec80a8d3fbf680916af72e8

Observation 459c0c06-554b-4ee1-9749-d88fe2a3e16b · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:35.739442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:35.739442Z digest=sha256:81605f0a8183af440dd18f1703fe1d4b6e63f0a2c69f3935ae2846ec7539a237

Observation 0312c211-c292-45d2-8d1e-e3a8193549b3 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:39.224993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:35.816729Z digest=sha256:4c98ae034bf996de0d51243997001fa219e18250b5f4cd073cfda4ad2411dd6c

Observation fb064cbf-cd5a-49cc-87e5-603c74f6b0ab · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:35.820975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:35.820975Z digest=sha256:532803e7d9a00f00609f28e461b0678dcf45f20f9b5cb756771e04de85c404b4

Observation b40f0961-f67c-428f-a555-36d36ea718e3 · outbound

This paper cites Introducing our multimodal models, 2023.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Introducing our multimodal models, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:39.213648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:35.825021Z digest=sha256:941d7709c43abd6857c52bdc79af80f0c2a1f50ddaaa02824003b2cc0d3f7bee

Observation dbfaacb0-d65d-4cd1-b8b6-b015e61627af · outbound

This paper cites Language Models are Few-Shot Learners.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Language Models are Few-Shot Learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:35.829831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:35.829831Z digest=sha256:2424545f714b86aed853527f34161cdc187a77b883e5f6567e5c899ab05725fa

Observation 16c793de-ab37-45a7-a931-0b41ccdc0a2a · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:35.835469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:35.835469Z digest=sha256:68e06f5ba35d220aa5c5ba67e23a6fec431d031a21d94a7b56496a51a1d12ede

Observation 9b59c0da-a7c9-499f-b363-7351c3a535b4 · outbound

This paper cites A simple framework for con- trastive learning of visual representations.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization A simple framework for con- trastive learning of visual representations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:35.840253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:35.840253Z digest=sha256:7d06335d17135e800f819b7718f47965e10c4c1255865376b50b60123f6e65d8

Observation d94c25f4-a825-4bcb-ad53-fb20ee5ee581 · outbound

This paper cites Improved Baselines with Momentum Contrastive Learning.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Improved Baselines with Momentum Contrastive Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:35.844751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:35.844751Z digest=sha256:b6657201b5d16e2e356d522397b92f11763fce6dc62e30091247a4af21f7ba8e

Observation a8c5ab8c-8f7b-494c-a089-4e946acde688 · outbound

This paper cites Qwen2-Audio Technical Report.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Qwen2-Audio Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:35.849986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:35.849986Z digest=sha256:468d7a013283adf60a9925f71016e01d548a385f9d81882ac99baa0b08aa7a1c

Observation 05e1541a-ce3c-4549-a307-2e3c55995d51 · outbound

This paper cites Simple and controllable music generation.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Simple and controllable music generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:39.088067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:35.940954Z digest=sha256:b2fce736333799b7915dccd159660f221338ba36966f21f0f3cbc8b301efc7e7

Observation aa87ad16-16f0-4646-883c-dcc999720988 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.014991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.014991Z digest=sha256:d8df67803cc7669dac100bb6332eb97bfaea07e5e87f84084ea02479087f9b4c

Observation aefeb02f-3139-4ab8-b43e-941f9011e016 · outbound

This paper cites The Llama 3 Herd of Models.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization The Llama 3 Herd of Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.068257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.068257Z digest=sha256:f279b486cceb7c4660ed0cf58fb88697307423779329b05e59f27795a1abccee

Observation dc0100a1-bd19-4a57-ba10-50f819f943cc · outbound

This paper cites Dwyer, J.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Dwyer, J

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.930197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.171275Z digest=sha256:957ae577b503806c04c355ee932711fd96bc6cf1316911e02af4e7c2f5a22289

Observation 09b7180b-b8da-4e02-9e10-587fb06203fa · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.175024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.175024Z digest=sha256:cd542afa903818d3b98851aeb399c8389d6166814c1ad32779aebbac0276be95

Observation bde0622f-a211-413d-8959-07c100cfb1af · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.919117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.179323Z digest=sha256:f0cda8f5c7f1de2d073cce53f45c00a0eab1ad6e7be947e94ac4ba568e5b3622

Observation b63b6cd6-1c1b-48e6-b72f-eeb59c7f6661 · outbound

This paper cites Aligning ai with shared human values.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Aligning ai with shared human values

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.906366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.182415Z digest=sha256:e146ea6261badaa6e589f673aa5c7508d4a427b578989c4088c75f41d6105e3b

Observation ade39451-50ce-4a8e-8cae-3dd90e9a4bf1 · outbound

This paper cites Measuring massive multitask language under- standing.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Measuring massive multitask language under- standing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.835147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.186260Z digest=sha256:83f9fc089988b35bcb78ed697e8b306ecf73338d3cea6f63a4c7772392679344

Observation 5c309c09-ffe4-4cd5-bccb-a0a92548dc84 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Gaussian Error Linear Units (GELUs)

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.190036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.190036Z digest=sha256:71b43790597baf1068ebe697de157d7ce921bf36286d1f49b9fdea1cb3946741

Observation 619ad864-1964-48fa-b65d-d0c55c5fded3 · outbound

This paper cites Scaling up visual and vision- language representation learning with noisy text su- pervision.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Scaling up visual and vision- language representation learning with noisy text su- pervision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.774400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.193706Z digest=sha256:690b537cee4475cc12ac101699f4e03b4e810381ec8f7eebece6bd8106b61fae

Observation 92a0d0e3-6832-4380-8d6f-873fcd73df42 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.761351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.197131Z digest=sha256:fe54ee9d61ba8e45e001ff5da13ee31b158e940800ae12093c7e409a66ccc1de

Observation f9579513-82f7-4417-9f25-5b6b763b5cba · outbound

This paper cites Shamma, Michael S.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Shamma, Michael S

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.745834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.201470Z digest=sha256:93639b3b44b1894a7e06e7d830c7161cbf9e5e312868e8e8f04b2e4c51a15f71

Observation 6d3628b0-1b2f-4458-acf8-b765d56e114e · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Lisa: Reasoning segmentation via large language model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.714763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.204703Z digest=sha256:ac6d751e310cb4c6301f1e8aa57077fc39822241f4f1d3caa0abf0d82690db4e

Observation c374cde4-be6e-40cb-94e2-9aa4d54dd174 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.207966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.207966Z digest=sha256:8d4682523ef2815ac903e2b1603af09cdc57d33616c2c3bc6159a9e96702609c

Observation fb2d7cb3-d6b3-4062-9d5c-a4690339954e · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large lan- guage models.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large lan- guage models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.703410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.212507Z digest=sha256:e13bcdca830935cf0c8ef541d33eb2e9cd04148f643e2d910a5e7d921a9f22d5

Observation 42bd135b-0d8c-4205-bfa3-f35f6e8eb26a · outbound

This paper cites Competition-level code generation with alpha- code.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Competition-level code generation with alpha- code

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.692353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.260389Z digest=sha256:1cb45bc3bd13f2f0754f14b6539e45a26b6aa2d2727ba7af5719570f6599670e

Observation a3a2b95a-0413-4d0f-bc5a-23d8b7738aa9 · outbound

This paper cites Microsoft coco: Com- mon objects in context.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Microsoft coco: Com- mon objects in context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.644791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.327097Z digest=sha256:51d3645ffe051c8956dd34c21fa97dd79c11debbf671035daa264fdc13dc83cc

Observation befac284-10db-426f-978c-4c4a32521ff7 · outbound

This paper cites Improved baselines with visual instruction tun- ing.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Improved baselines with visual instruction tun- ing

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.634180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.331170Z digest=sha256:c9b0ad7e3e031a495e42199a5eb6297b05c0d60462f795b9bab4599cee495d1d

Observation 51897060-a786-4823-a143-68342533368b · outbound

This paper cites Visual instruction tuning.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.621462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.334875Z digest=sha256:caaeda27daf89857ac8248231ac969c61b2c5974c249b1a89e08292d5c430e9e

Observation 08f9ab68-0c0e-413f-8bfc-30e8e2a42d1f · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In Eu- ropean Conference on Computer Vision , pages 216–.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Mmbench: Is your multi-modal model an all-around player? In Eu- ropean Conference on Computer Vision , pages 216–

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.582846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.339480Z digest=sha256:5d87bfd0bd7e6c0f74ca375d116258cee9b4de7c57a5633de777043e811ad92a

Observation e80ef3e3-5524-4240-8ffb-5cb335fd151f · outbound

This paper cites Decoupled weight decay regularization.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Decoupled weight decay regularization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.537588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.343615Z digest=sha256:60ea841591689f57beff1cac60c61ceb229a5b425eec62284ffd03b44802eb00

Observation ecd81a06-f3e6-4d79-88e5-753a2c66ae88 · outbound

This paper cites Learn to explain: Mul- timodal reasoning via thought chains for science ques- tion answering.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Learn to explain: Mul- timodal reasoning via thought chains for science ques- tion answering

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.524984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.347388Z digest=sha256:e2618303a1896669bcbaecb4d629eab85e9221dc41322f0d59fd844f997826bd

Observation 62583d58-409c-40de-a2b0-4508efadae9d · outbound

This paper cites Cheap and quick: Effi- cient vision-language instruction tuning for large lan- guage models.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Cheap and quick: Effi- cient vision-language instruction tuning for large lan- guage models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.483032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.351603Z digest=sha256:4c7219a8f97bde10c99ff3c87da3accb3ca9646455dc2ac6b58797eb5305304c

Observation 2dc24cef-ef1f-4e3c-8b6e-f9ca1c774691 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.453926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.355061Z digest=sha256:f41b9276a6adb5467898ff48309b1d7f342a2114802c4e96d56e5a6cdd3cf852

Observation 332fd706-58dd-4250-a898-23da7f04fab7 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization GAIA: a benchmark for General AI Assistants

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.358685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.358685Z digest=sha256:0c4066781f6a3246ebbda38bf6f34a45089e64c38d098114a222b7eb8619a62e

Observation a3e50154-f6f7-496a-85c2-100c9f8c7a52 · outbound

This paper cites Foundation models for generalist medical artificial intelligence.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Foundation models for generalist medical artificial intelligence

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.362496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.362496Z digest=sha256:bc68a81a06023344a2c519d35375c9e62982779002d87319cc4b098db2111c9d

Observation c680a341-7a5c-4e14-b967-c2edc00707e4 · outbound

This paper cites an unresolved cited work.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:18:38.407829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.366623Z digest=sha256:d576a6b3aa611f000845db0d8354a97b15501ab2721e54fdd8e445ae33a60ffa

Observation 0f8a3fc0-c640-43a8-a42a-8e65c678cb67 · outbound

This paper cites Hello, gpt-4o, 2024.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Hello, gpt-4o, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.358839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.370776Z digest=sha256:404e16509ad6aa42304c3de3180a7b99b9ee3cccaef00424503baeed3b848ca7

Observation 0ce719a4-2f4a-41d2-8ea9-c490021cf772 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Learning transferable visual models from natural language supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.447836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.447836Z digest=sha256:42ab4facbaa49bdfe971078a7783660e49e2cc86db74d7f62ae4d045c90b4896

Observation fa93a7ea-1275-4055-aa81-06599b4edbb6 · outbound

This paper cites Ro- bust speech recognition via large-scale weak supervi- sion.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Ro- bust speech recognition via large-scale weak supervi- sion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.342125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.471471Z digest=sha256:f10387a3f9ce53fa432524728c3547b28522e4ed2817c1a1a3ec2b473d9b79d1

Observation 85f195c7-d613-470f-9948-0f78451472b1 · outbound

This paper cites Language-based action concept spaces improve video self-supervised learning.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Language-based action concept spaces improve video self-supervised learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.259066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.555199Z digest=sha256:7487752482a2e55b4675801e818ca408f0dd2e966588cefd3e6b9caa8642c21b

Observation bc2f41eb-709c-42a5-8be3-8d5f40051593 · outbound

This paper cites Learning to localize objects improves spatial reason- ing in visual-llms.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Learning to localize objects improves spatial reason- ing in visual-llms

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.130223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.618456Z digest=sha256:a3c075f1b9d84928c2443fdc50e6d94777bc1f842d991e588586efd95552e771

Observation 7e0adb7c-32b2-46be-9894-b53d69f2805a · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.711479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.711479Z digest=sha256:e496d464f9f738c294f67afb8c50ae86bca582b5aeef09326b3255347f411a68

Observation 85b2b88e-6beb-47ad-b1ea-80d8555e1b7a · outbound

This paper cites Laion-5b: An open large- scale dataset for training next generation image-text models.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Laion-5b: An open large- scale dataset for training next generation image-text models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.118720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.715389Z digest=sha256:623a8c789e407593edd16432becea289acfa5f9574b7e030e5ebaefaa632657f

Observation aff80d5a-e074-4bfa-b9ae-63fb309a0faf · outbound

This paper cites Hugginggpt: Solv- ing ai tasks with chatgpt and its friends in hugging face.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Hugginggpt: Solv- ing ai tasks with chatgpt and its friends in hugging face

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:38.106690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.719242Z digest=sha256:07f3327f2222e82775f3ef4413cee27c2b79a2ae81448287150933fade4d7eaf

Observation 443f7f8b-8b91-40a6-adac-1cab8eb7ccc3 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.723285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.723285Z digest=sha256:8e59339c96673a71624a51e99f7d400c4cb470722000fe9f232a824f615e85bb

Observation fb7cbeb8-0834-4276-b634-1a2f9605f789 · outbound

This paper cites Movieqa: Understanding stories in movies through question-answering.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Movieqa: Understanding stories in movies through question-answering

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:37.993812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.726704Z digest=sha256:b708c8ff8b74ca608516c5b2c715df90df5fbb1806747e4a7298227d75e4d283

Observation 90bdbea7-9691-4b58-a967-112ba279d948 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization LLaMA: Open and Efficient Foundation Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.730242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.730242Z digest=sha256:b6de7a8e96ee2c1c9eda2366691b29bf56f73ea01b43ff85b55d26bf099e826d

Observation 617d7226-bf4e-4f9d-b76f-9ee0a892f1ad · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.733806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.733806Z digest=sha256:617d106f515adfffc861c26e7238ab3ff45b746fede964e531f290fd069b80fa

Observation 3f6bf28b-fb9d-4fc6-aca7-cb34fa842092 · outbound

This paper cites Efficient utilization of large pre-trained models for low resource asr.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Efficient utilization of large pre-trained models for low resource asr

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:37.885573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.737901Z digest=sha256:ffc3b30f3faaa14c749f3662252ff60a910fbd0490b945a419ba97f2c12f4107

Observation 30e37c54-f640-4ccc-a9bd-ad19b8ff641c · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.741509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.741509Z digest=sha256:aa79dc669b190e88fce45a3f508b008b1f39804e7c27cad0da2e09a9e5ca9735

Observation 4c67679e-2587-46fd-bf77-1435bcda9beb · outbound

This paper cites BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.746475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.746475Z digest=sha256:cdfda51ff28d48ea0ca7b417ab0d430a38cf11425a0a949daeb913fbff148f5e

Observation 03f6d93d-8bfd-40e1-918a-47977c128466 · outbound

This paper cites Visionllm: Large lan- guage model is also an open-ended decoder for vision- centric tasks.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Visionllm: Large lan- guage model is also an open-ended decoder for vision- centric tasks

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:37.804424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.750409Z digest=sha256:8b30610e1be975e3c72fe2fc217a26f5346e2de158ee46003cc3467eba45dbde

Observation 18aff06f-53b5-412a-a9ff-8edc99353f64 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Finetuned Language Models Are Zero-Shot Learners

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.754966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.754966Z digest=sha256:c7abd4ba485a4643039498a9b0f2388a57be3dca1a49e330c09260f05772917d

Observation 737edb95-c53c-4ab0-8ade-cea8be816922 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Chain-of-thought prompting elicits reasoning in large language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:37.733320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.783781Z digest=sha256:56f26261f1687377f2493c9e8fa5748a5f67d277b17325390359230a43ca8fae

Observation 16bf530b-c262-4a9a-acc7-fddb5e19921e · outbound

This paper cites Hard prompts made easy: Gradient-based discrete opti- mization for prompt tuning and discovery.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Hard prompts made easy: Gradient-based discrete opti- mization for prompt tuning and discovery

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:37.719930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.840648Z digest=sha256:6329805d116234fe641b0a520abdcffb22026cb482286445454938a9ca8afdf3

Observation 4a30ec71-fd78-4fa6-9f6a-c1f04bc3bda5 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.962840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.962840Z digest=sha256:d8d93cce78e3a27125c1e958a5e05ac55f0519ed9d13a8ba07efeb164e85b537

Observation 6b0996d9-4110-46c0-bec6-68ca8e9c6c90 · outbound

This paper cites MM-BigBench: Evaluating Multimodal Models on Multimodal Content Comprehension Tasks.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization MM-BigBench: Evaluating Multimodal Models on Multimodal Content Comprehension Tasks

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-11T10:18:37.165466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.967176Z digest=sha256:46468223d777cfa0a121e7691a9673918f392f17d0978bb222add55a27d5c755

Observation 7b41eac5-f975-4886-90f0-0b84e28863fb · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Tree of thoughts: Deliberate problem solving with large language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:37.708765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.971941Z digest=sha256:d6da8371163d36178706d6c6c32ebeb665f658ede555a6d61332bd87e4226772

Observation 4b999423-825e-491c-abad-e975a6d51b8a · outbound

This paper cites From image descriptions to visual de- notations: New similarity metrics for semantic in- ference over event descriptions.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization From image descriptions to visual de- notations: New similarity metrics for semantic in- ference over event descriptions

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:37.597024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:36.976357Z digest=sha256:b192b955e4fd7d3b27e6eb315ee5be1de52f37a58077586da8b33666d6fea411

Observation 5fe61c81-6674-44ca-8e78-62ecac9f7c70 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.979839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.979839Z digest=sha256:51db8b24c9a9c6c31062cd5f8b0dd4652014d773d25012ce940ca7add7f10987

Observation a0ea930b-99e6-4255-8723-f3812a007cab · outbound

This paper cites Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.983631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.983631Z digest=sha256:f8c926d9822a7515ae947c233a96632ceb0894869e3a05ae7fb707fb17210b0f

Observation 606538b6-dfe2-4648-a2ce-557b2765661c · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.988004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.988004Z digest=sha256:e275f7867d4639996dcaac0880510a692a88dba25c7d456f0628ab553ca8f02e

Observation 027c6b6b-388f-45ea-8a9f-9d62e9dc50dd · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.993184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.993184Z digest=sha256:12e093f4b3bb01e05eb2c6563121a7a83d9d0c97867cafffbd639a48658ad73b

Observation f0066eee-12ee-4eeb-a1e7-bffe7c37470e · outbound

This paper cites Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language mod- els.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language mod- els

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:37.445003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:37.079563Z digest=sha256:6f958fe6cd7fa764e87c5e4f85936cb5fcd0a0ad73a5e7580c555a94eb8ea03b

Observation 8269c6d5-70a5-47be-9d32-af2f1d2ba8ac · outbound

This paper cites P Xing, Hao Zhang, Joseph E.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization P Xing, Hao Zhang, Joseph E

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:18:37.430388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T10:18:37.083251Z digest=sha256:a8eb5c4c033ac86ea71db0f967e4bb59b810163348a6d40e7ad0157dd94ba847

Observation 5ce1ee04-c69c-478a-9513-2225377861f5 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:37.087313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:37.087313Z digest=sha256:cb0438de93cd46ccb203514db32643033559d96d3f83e0824e7c9c4462a66c87

Pith citing papers

No inbound Pith citation observations are available.