Pith. sign in

Paper Citation Record · LEDGER

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability

As of 20 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2412.14672.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14672 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:04:52.215193Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:43:59.364150Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T22:43:59.804431Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff26ac3c-2ec2-4b9b-b637-fdf2a6d28b04 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.026061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.026061Z digest=sha256:f01c90de9303c1712cafeade5a9b11fe49e5c0b4e59ab7743e5840155f6864e7

Observation 4b785016-af4d-4cdf-a077-8a09aa752174 · outbound

This paper cites VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.031554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.031554Z digest=sha256:8e0b0fec2f60c19e51b68f9ea6197de436ae941a9e790516bd86f362a8194503

Observation 79ff357f-cedc-48ae-bb0f-8d2f1fd9a83b · outbound

This paper cites Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.036033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.036033Z digest=sha256:eb44a50f2a628b830ded7cfcecc16830116df4099a7733d8cc3bf1316d261f7c

Observation 086681d9-b567-4699-8bf1-9efd29a320b5 · outbound

This paper cites Pixtral 12B.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Pixtral 12B

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.041003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.041003Z digest=sha256:2d14f661ac2f00d96759c9214b5f37cd42ea5b64639d0af4e8258ffb81040e6a

Observation 21f79a20-16f0-4cbd-973a-e0dc8733b4c9 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.045294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.045294Z digest=sha256:4172ac2ff5b95acca2a2f301e46490a598e0295e3362b55c2d93af99c95df474

Observation 25bd53eb-7e2f-493b-a3b1-d5be48c71d4a · outbound

This paper cites Counterfactual Samples Synthesizing for Robust Visual Question Answering.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Counterfactual Samples Synthesizing for Robust Visual Question Answering

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:04:52.535312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:04:52.049551Z digest=sha256:2b340c294e18c76308069433b4ac5b1388321d1f296d86dc4c519e5e551e670f

Observation d013adf8-4aa4-4ac0-938e-11bc648430d2 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.054245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.054245Z digest=sha256:cdb6aba98b929f05116777c4aa9f0db9f52096072a52292350dcdf05d1b7c3e6

Observation 4523d83b-2c4c-4166-a6e6-7b116df66198 · outbound

This paper cites an unresolved cited work.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.058176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.058176Z digest=sha256:36af707c3476ecdd0daa4691410f742e9d1ec3a90b5393bb187b937e13976376

Observation 8d77d9c3-70e1-4baf-9538-91718103eb7e · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.062108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.062108Z digest=sha256:8b857fa7ca6dd83ea8636b046b3967b83a5788d437b4f5cc84bb4a817d7aee9f

Observation c0666ad8-9fba-4350-963b-0d05427ba75c · outbound

This paper cites Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.066207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.066207Z digest=sha256:05b59aec35c88e5f57fe9db27a777c0413e82a985fc3f24dd48b199ba953cd12

Observation b00b0214-ba4a-412e-b0af-9aa0ad99af60 · outbound

This paper cites VizWiz Grand Challenge: Answering Visual Questions from Blind People.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability VizWiz Grand Challenge: Answering Visual Questions from Blind People

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.070422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.070422Z digest=sha256:33fa4be5b21ae32be1ff0e4aaf8589c1824011c3b0323e3ffdf17c07027db055

Observation 63e195d1-a1c9-46a9-af2a-e0598529f90f · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.075173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.075173Z digest=sha256:2d55d88a42cb6380b11258451aeaaf4dffdf37c42d91c28ce4e105a29131708b

Observation 588d9830-f8d3-4911-ab87-f66fbfd714b3 · outbound

This paper cites CARETS: A Consistency And Robustness Evaluative Test Suite for VQA.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability CARETS: A Consistency And Robustness Evaluative Test Suite for VQA

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T12:04:52.460368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:04:52.079237Z digest=sha256:59382f8b2140ba8e41fec74bcc26730ded452028a8164a731bfd3486e567dafa

Observation 25e83590-4380-4b1b-bce2-b9805c464335 · outbound

This paper cites an unresolved cited work.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:04:52.738740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:04:52.083326Z digest=sha256:c5033e95b14a06c610a29cbe48132f5eff37197a2bdf0036c0a3e31e779409f3

Observation e6a4ebcc-a84f-4388-9c9a-3e8118d2de7b · outbound

This paper cites Segment Anything.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Segment Anything

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.087048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.087048Z digest=sha256:00fbc98e920588d0bf5062980dd2b8ec3b641c8294803c635944aff7aaefbf6d

Observation 732c757c-2c04-4772-94f6-5a945e81bfbf · outbound

This paper cites an unresolved cited work.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.091059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.091059Z digest=sha256:59a3adfef49e90be303bcb91225279f5891ca5e93ae1897f7ca3cb85b090395a

Observation ff5aa1cb-fb42-426c-8a2d-82ae7e0838f7 · outbound

This paper cites NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.094678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.094678Z digest=sha256:3fb68ec7321bfadd2d372b5089798d727ee06a72d9e4fb1b8ea339363332ae0a

Observation df637137-126f-43fb-bab6-7d1e2229d7ef · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.098594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.098594Z digest=sha256:6026578d6edd5538bf37e7099e08b07bbb9edf7e611db519f00b8c93d1f457ec

Observation 3668c9ff-9dd0-4184-bcf4-e80c21f57c19 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Evaluating Object Hallucination in Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.102635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.102635Z digest=sha256:a3f99b61b52a3b394d5fcb620bd979608daa9e9091f360e8c27e308ff5eef699

Observation 3ef6ac3a-6597-4aa2-bebf-1d544b7dff6c · outbound

This paper cites Lawrence Zitnick.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Lawrence Zitnick

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.106768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.106768Z digest=sha256:e4d258eadf7d80dcb53001c7d8527263b56a965a08ba69572d0d4d0ecb5a3969

Observation bbb66e70-f996-40af-872c-001138625e7a · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability A Survey on Hallucination in Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.110559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.110559Z digest=sha256:4730d50a6550202d694a9290278ee98aafc2c52bdf167be9fb1192f709ad392d

Observation 39ee3285-833e-4add-a686-d912f619e119 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Improved Baselines with Visual Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.114495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.114495Z digest=sha256:db4c80f3adfc5fdc030830e547d361a43c1984d2b0bfcd0f85b918f6bd5947eb

Observation 6f323718-d450-4c55-8bed-172dbad347e1 · outbound

This paper cites Visual Instruction Tuning.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.118482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.118482Z digest=sha256:af5ac5b1bb6245c91fdfdd3dac7f0f81531da7bbda79e42ac9e885dba22e82d3

Observation 191cf6af-fb1a-4d4c-9ade-50e125814f88 · outbound

This paper cites an unresolved cited work.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.122548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.122548Z digest=sha256:883e71187f6bb3950a8a7503945dd27dcb9b8a90edcd70076bdd1aec6b214c71

Observation 801e4692-ca58-4b05-97a6-9fe1d533d203 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.126458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.126458Z digest=sha256:9742141d191f2dd08af20fe208263e883f675edecfdb9e2d785f4e64d93baa12

Observation 91e9e88e-e424-478e-a5d8-9665eb2d0920 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability MMBench: Is Your Multi-modal Model an All-around Player?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.130457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.130457Z digest=sha256:429845bad34b28c01be98c7e2390cc4671132f4a679791924ee38e227a855d33

Observation b5d1c835-e0d0-4809-bafb-a95430c61cba · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.134535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.134535Z digest=sha256:d7d8c7beb9e0f9f0a0c19a5cb2c1100635ffee1006e4e2fdca3487b8da05ebf1

Observation 96b1b133-1c89-40ba-98d5-31d2859efc88 · outbound

This paper cites an unresolved cited work.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:04:52.706640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:04:52.138886Z digest=sha256:c2846d9f517c154e42c97041e75bc9b6e8e43e153df2ab36ab9b10590578bec3

Observation 1f3442f3-acd0-4c0f-98bc-c85dd2461612 · outbound

This paper cites OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.142891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.142891Z digest=sha256:df8b8a1ca217922d85d9a2ae3c92c8d5f4828b60890b7259c17b048fa966e3ef

Observation 477db756-fcc0-4676-b872-a93dafad9a44 · outbound

This paper cites GPT-4o System Card.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability GPT-4o System Card

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.146972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.146972Z digest=sha256:4c6635331bd9834f5aaad1481f533bcca1fb4fc648bd53b9fd425c64d5111d1d

Observation e92f48e5-7166-4b4b-a811-c3ef3bef3007 · outbound

This paper cites an unresolved cited work.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:04:52.694991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:04:52.150777Z digest=sha256:d7e941bc776bedfc38f2b1dd01ff6bb90a80c291e9ba5163ae1324eab02f2519

Observation c7a83cb5-39d8-4aa0-ba2c-fea099eb736a · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.154295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.154295Z digest=sha256:f8ec9c05b3f8cbe19bf14792fcc2d47d62b038e44cb4baa4d18a75a08028ad4c

Observation 371cc130-2cee-47f1-a5ba-c6a461305200 · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:04:52.683828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:04:52.158052Z digest=sha256:0e29afe150d4a02e99c644da0d982825a6fd1dde64a61917b18e40b03090d51b

Observation d6be9a74-80d8-46b0-b060-465cb8ae1c48 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.161891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.161891Z digest=sha256:22c4d706103509463b343fff18e7ed9b5795ad5aed1cb9e0820e4b1883636024

Observation 52af9d7a-05ae-4d5d-b21d-12bf9f74bc1f · outbound

This paper cites Towards VQA Models That Can Read.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Towards VQA Models That Can Read

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.166016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.166016Z digest=sha256:26f13d4daf8a6b079ed4891b25726a519d99e8cdd178a339ffa9c29f0ac19aa2

Observation dfbe8a9f-4006-4824-b6ed-d0a7354d8e5e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.173953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.173953Z digest=sha256:7d63f7b1a6bbf2865aabe2844814c0a23a150c5d245237e12a9f51df45ed53be

Observation abbc4de3-1c5b-4c28-b4fe-8204ce3391c9 · outbound

This paper cites an unresolved cited work.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:04:52.671405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:04:52.178665Z digest=sha256:f6625a2ad73a85f20024d7c2dfe7dae2c67b589a5172dbbd5c965204e66e5670

Observation 93a2aa39-1e4e-447d-a95c-7273cc3286af · outbound

This paper cites an unresolved cited work.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:04:52.660201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:04:52.182477Z digest=sha256:9c7c6b505221953316f156dcf6959083cc26903612ff4427d9df939deabce982

Observation f05a9183-52c1-4f12-8a34-5c313c92f165 · outbound

This paper cites an unresolved cited work.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:04:52.648902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:04:52.186537Z digest=sha256:38207d345cfb04382092725c09b1ac4e616133664e2a0510047d9d0ed269054c

Observation 24989e07-aa24-4fb8-9597-2e3ff12b3a2f · outbound

This paper cites BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.190410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.190410Z digest=sha256:8496e1f7cbdf7629ffeefa711787f726efb1e79e14ae3e5e3e4cc3433049ea6c

Observation 8629bd3f-cf61-44d9-b430-66ba0cd1febb · outbound

This paper cites an unresolved cited work.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:04:52.636115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:04:52.194732Z digest=sha256:58f6d33dc91f2a53413fac3f0959c0d8cfcd73a5db275fed4dbf0c2c2154f877

Observation 71b4fdb7-0531-4569-a1af-0737022e9ae9 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.198451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.198451Z digest=sha256:aec341b373e87e81974b26656ae5ef561f2cd69b0a475ae5971dd69ddf7a88b3

Observation cd4ecfb5-c249-4f82-85ca-74485ab90d3d · outbound

This paper cites an unresolved cited work.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:04:52.623618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:04:52.202390Z digest=sha256:393f0e3a1674ad7ec1b08e3c2ff7bd6b814f2c0042bfc7a7a601db08c5f4ffcf

Observation 20ebc7e6-acc7-4e94-be64-df7a57e123bc · outbound

This paper cites online" 'onlinestring :=.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability online" 'onlinestring :=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.206310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.206310Z digest=sha256:894e2c936d1ede5a3a96c0cb3109fbb7c2ec649f2c8a40eef8caf5f0c46a7209

Observation 3ff82dea-cd58-4a20-8abc-a25ad7ab2784 · outbound

This paper cites write newline.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability write newline

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.211093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.211093Z digest=sha256:0f26fc286beeaadcff1905c909402cec0cd8237fce653891cab437f1ec1e211d

Observation 80a4a4c0-0b50-4f99-812f-05df1ef81407 · outbound

This paper cites write newline.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability write newline

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.215193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.215193Z digest=sha256:4d80c72e325596dff756001eb995a2181885a94ed7c95b9c20d8c37dc712f2b4

Pith citing papers

Observation b3c959b2-d24f-4ed5-abe9-a793187fa366 · inbound

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs cites this paper.

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T22:43:59.809376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T22:43:59.364150Z digest=sha256:9241923e8b2fb7e7ba6caef8adb8636a6d1f03d468d7334ca3c49ceda919e1cb