Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:04:52.215193Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2412.14672.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:04:52.215193Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T22:43:59.364150Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T22:43:59.804431Z
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ff26ac3c-2ec2-4b9b-b637-fdf2a6d28b04 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b785016-af4d-4cdf-a077-8a09aa752174 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79ff357f-cedc-48ae-bb0f-8d2f1fd9a83b · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 086681d9-b567-4699-8bf1-9efd29a320b5 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Pixtral 12B
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21f79a20-16f0-4cbd-973a-e0dc8733b4c9 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25bd53eb-7e2f-493b-a3b1-d5be48c71d4a · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Counterfactual Samples Synthesizing for Robust Visual Question Answering
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d013adf8-4aa4-4ac0-938e-11bc648430d2 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4523d83b-2c4c-4166-a6e6-7b116df66198 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d77d9c3-70e1-4baf-9538-91718103eb7e · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0666ad8-9fba-4350-963b-0d05427ba75c · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00b0214-ba4a-412e-b0af-9aa0ad99af60 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability VizWiz Grand Challenge: Answering Visual Questions from Blind People
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e195d1-a1c9-46a9-af2a-e0598529f90f · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 588d9830-f8d3-4911-ab87-f66fbfd714b3 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability CARETS: A Consistency And Robustness Evaluative Test Suite for VQA
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 25e83590-4380-4b1b-bce2-b9805c464335 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e6a4ebcc-a84f-4388-9c9a-3e8118d2de7b · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Segment Anything
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 732c757c-2c04-4772-94f6-5a945e81bfbf · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff5aa1cb-fb42-426c-8a2d-82ae7e0838f7 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df637137-126f-43fb-bab6-7d1e2229d7ef · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3668c9ff-9dd0-4184-bcf4-e80c21f57c19 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Evaluating Object Hallucination in Large Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ef6ac3a-6597-4aa2-bebf-1d544b7dff6c · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Lawrence Zitnick
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb66e70-f996-40af-872c-001138625e7a · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability A Survey on Hallucination in Large Vision-Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39ee3285-833e-4add-a686-d912f619e119 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Improved Baselines with Visual Instruction Tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f323718-d450-4c55-8bed-172dbad347e1 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Visual Instruction Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 191cf6af-fb1a-4d4c-9ade-50e125814f88 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 801e4692-ca58-4b05-97a6-9fe1d533d203 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e9e88e-e424-478e-a5d8-9665eb2d0920 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability MMBench: Is Your Multi-modal Model an All-around Player?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5d1c835-e0d0-4809-bafb-a95430c61cba · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96b1b133-1c89-40ba-98d5-31d2859efc88 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1f3442f3-acd0-4c0f-98bc-c85dd2461612 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 477db756-fcc0-4676-b872-a93dafad9a44 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability GPT-4o System Card
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e92f48e5-7166-4b4b-a811-c3ef3bef3007 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c7a83cb5-39d8-4aa0-ba2c-fea099eb736a · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 371cc130-2cee-47f1-a5ba-c6a461305200 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d6be9a74-80d8-46b0-b060-465cb8ae1c48 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52af9d7a-05ae-4d5d-b21d-12bf9f74bc1f · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Towards VQA Models That Can Read
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfbe8a9f-4006-4824-b6ed-d0a7354d8e5e · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abbc4de3-1c5b-4c28-b4fe-8204ce3391c9 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 93a2aa39-1e4e-447d-a95c-7273cc3286af · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f05a9183-52c1-4f12-8a34-5c313c92f165 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 24989e07-aa24-4fb8-9597-2e3ff12b3a2f · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8629bd3f-cf61-44d9-b430-66ba0cd1febb · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 71b4fdb7-0531-4569-a1af-0737022e9ae9 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd4ecfb5-c249-4f82-85ca-74485ab90d3d · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 20ebc7e6-acc7-4e94-be64-df7a57e123bc · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability online" 'onlinestring :=
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ff82dea-cd58-4a20-8abc-a25ad7ab2784 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability write newline
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a4a4c0-0b50-4f99-812f-05df1ef81407 · outbound
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability write newline
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c959b2-d24f-4ed5-abe9-a793187fa366 · inbound
From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.