Pith. sign in

Paper Citation Record · LEDGER

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation

As of 12 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2412.09817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09817 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:45:44.870849Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab2c5cd8-8afd-46f6-aff5-a1120b759a2c · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.659646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.659646Z digest=sha256:0f0d714eda0f007859abdfff8c4da30a015bfe055167419b0f010696ddd6fd2b

Observation 26459d2f-3047-470f-9171-71096ab547f7 · outbound

This paper cites write newline.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.664145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.664145Z digest=sha256:12bed87ee3f0abaa4ad94696a1f696209ca4fa695f00feb2411dd471c78cbe09

Observation aaf457e3-f2df-4deb-8169-82c9c15c62a4 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.497001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.668737Z digest=sha256:feff27b8f450d1f2c55845fc5c973f1fdb1d3a6ce971df7734584c6e47feabf1

Observation de3f5a70-9bd8-4bca-bd00-eb1f083de74b · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.672692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.672692Z digest=sha256:53ac1fe3a9438890f6f94eab565bc13b19887e9d55a6c1d91ae5eea676d2be8b

Observation c065c839-cc18-4729-a2f2-bd9c412471da · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.677283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.677283Z digest=sha256:a3b90c58f7cdacb0556ef8665d6f91ebc4e1aa6fb45b31898cdff6f242d6fcde

Observation 8deb9786-40bc-4052-b46b-5b4eeeda68e0 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.486002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.681789Z digest=sha256:ee0500a0bb4d767b08ee97c679b9396529d1b7e5b951ae7498d0e1d6f616b42d

Observation dc2b24e6-c43d-4ecb-b453-762a7e6c9de8 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.689734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.689734Z digest=sha256:ecfc377335833c5b554e8e546d8353729856df9d40bde109808fa3f43852d350

Observation 6b751465-a209-4e61-9e02-4ab9df02e2e7 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.693657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.693657Z digest=sha256:872fa6e92c9dc6abaa588eae6f0b7d05b4f0e4d0fb4707e312f73490a36f2890

Observation 18c75d0a-b2f2-4773-a2f8-8fc7c8dd3284 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.697445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.697445Z digest=sha256:086c9894dd4d4ae9e1332667ec9429aa089f22f91f65553297dba7054483ba11

Observation 3c3a9258-cdbb-4d5b-92a2-039e3c1d37d9 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.701291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.701291Z digest=sha256:089d04d5e9dbe09f016f957a39af161263e3e33bf07ab3d7d88d4ce373c6cd4d

Observation 4adf3b1b-e2f3-4a84-8a27-eb4d1093642b · outbound

This paper cites K.-p.; Qian, T.; and Wang, S.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation K.-p.; Qian, T.; and Wang, S

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:45.462402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.704823Z digest=sha256:696ede652b2057bb68287d369934fa02d8aa73db96d070c1d042466ebfb07e75

Observation 57a9ad2d-a67b-4079-a171-e7daf591b99f · outbound

This paper cites VSE++: Improving Visual-Semantic Embeddings with Hard Negatives.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation VSE++: Improving Visual-Semantic Embeddings with Hard Negatives

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.708251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.708251Z digest=sha256:81c01230f165e4b67beab21d9fb085fa76e90b36fa624b46fda3b4129843b8ca

Observation b3883893-5b76-4d5c-a712-3e810113d082 · outbound

This paper cites S.; Shlens, J.; Bengio, S.; Dean, J.; Ranzato, M.; and Mikolov, T.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation S.; Shlens, J.; Bengio, S.; Dean, J.; Ranzato, M.; and Mikolov, T

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.711987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.711987Z digest=sha256:a815ebc87555b425cf6c5dd093e49a904df2186005cc831a2f62276847455131

Observation 7821a4e7-f3e3-4e19-926c-50d9ee9197ce · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.715548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.715548Z digest=sha256:a4f4376464df7a29984c75df3927a8290ce4d8c7ef15e52ae324d901961bbbf7

Observation 4a80bfb5-b0b2-4573-94b9-876f2eb89f3b · outbound

This paper cites OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.719020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.719020Z digest=sha256:51a9520bbb88c2d0bc4669393645318e2add2872de0697fba47d8d70a7424468

Observation bce94d8c-11e7-4d67-beda-510fc5ad6544 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.431894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.726564Z digest=sha256:7dc0ee2aea0cf5060edf306ae5d27ebc04dae729b99bd7c19a02c4c39bbe03c3

Observation 34dbc0fd-30fc-44bf-afd8-40fdd24a42f4 · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.730495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.730495Z digest=sha256:ed7d073e23ef813d6d1180a2dca7dfe3986b9c912a849cedcb2280023d45e657

Observation 5d6d876d-f52e-4c50-b6bc-c7d7e7457558 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.736267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.736267Z digest=sha256:ffdac04ea022cfa88e110776cd5c9c1aa34ce82a5c0ef3512f5679dbbab6efe4

Observation a09af1a3-2b9b-46bd-ac49-81e4258ca51d · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.740539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.740539Z digest=sha256:eebb1b09ca2712dd3cf60f2e63d97d19f65fe0bb46d361dd372ccf4c44ad5cef

Observation 0f706992-c754-4cc4-8ed3-cfa581e8a94e · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.743967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.743967Z digest=sha256:5f28309929434d38c420631cd2959d8bea48b11aea9886341f9de7ace938ad7b

Observation 3d81ac7a-ff0c-4cff-8b33-21f531c92a70 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.409110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.747739Z digest=sha256:86b283c4f4da78c7c47c24eaa43098dbde035c1e89474e14761d542bd1651eef

Observation 9bc62403-be46-45df-b7b9-ac5c02775dc1 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Improved Baselines with Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.751352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.751352Z digest=sha256:137a696cd2d9c37e7dde4f29e7c430a6605ed3fda3fc7f8032c8e9e5726a379c

Observation 620194f9-467b-4a8a-85a8-c0e622fdbf23 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.398877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.755283Z digest=sha256:7394f75ab6bae2cd3e557b50ba8d6e733798b3439b912cc3753b42f5ee8d100a

Observation e42b04e2-881a-4cf1-9295-94638494492f · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.758732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.758732Z digest=sha256:b1e0e269bab8d303fdda140adf3ceb28716e98710bcc2cb343baa37072dbca31

Observation 76bb4713-f1db-4249-a7d3-f824a59d7501 · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.762111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.762111Z digest=sha256:d13b6087c3dff2a7fbf83111077f97708594d16d485f8324532727c4ffdabda2

Observation d88a1c08-ceea-4557-8d90-ebd642c51052 · outbound

This paper cites J.; and Yan, Y.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation J.; and Yan, Y

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.765864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.765864Z digest=sha256:b9d7fddb2741110048af91a5c14ed9addf97d3b040451dbb0e248bcc5693b55f

Observation 7b70f671-3c40-4dfb-8340-20523a88b89a · outbound

This paper cites IMAGDressing-v1: Customizable Virtual Dressing.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation IMAGDressing-v1: Customizable Virtual Dressing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.769254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.769254Z digest=sha256:d29b2d4405ecd7156578d5a5432e29491bb00ecf069441424c90ac74f503735f

Observation b5ed5497-9a08-4ec5-b7b1-52e0524d14a6 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.773148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.773148Z digest=sha256:5af8990ac1de88956f12b82f1c91f2a54141056a7947c800b33da2d698dbb4d1

Observation 2a03c5ce-a27c-4224-ad62-437f84aff12a · outbound

This paper cites Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.776695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.776695Z digest=sha256:58fb497309b63fd7ac5beef153543a5c47351544265ce2d04c8227ef8fec73ec

Observation d74a7b21-5708-4231-a500-46c57347085b · outbound

This paper cites ???? Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation ???? Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:45.368028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.780400Z digest=sha256:4395bcf988440ea49aadec0b58a7b2538a447b4b6bc3944f71df8820d79d162e

Observation 3e395ca2-e21a-4d8c-a5a8-a44655a5da63 · outbound

This paper cites Parrot: Multilingual Visual Instruction Tuning.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Parrot: Multilingual Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.784099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.784099Z digest=sha256:7f9e16cf7a8c531117b68e264133f3f1479ff87f45adff7aff7793ae4f2d4bda

Observation 34a5d01c-c8e4-45f6-aac7-c14682479013 · outbound

This paper cites PILOT: A Pre-Trained Model-Based Continual Learning Toolbox.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation PILOT: A Pre-Trained Model-Based Continual Learning Toolbox

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.788078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.788078Z digest=sha256:072424eea9b5828a8f4848cf9b3972323cd0869bbd942aa2c6644ffe103ff2b8

Observation 396e2e59-ce9f-4521-8481-8b8377aa5703 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.356442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.791824Z digest=sha256:a7f00ea6def8d593bc4b0cae889b9e81583dca62e3f3504047917417b2d0c810

Observation 56b8f89a-496f-47a4-b0a1-88e60bc87b48 · outbound

This paper cites TextSquare: Scaling up Text-Centric Visual Instruction Tuning.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation TextSquare: Scaling up Text-Centric Visual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.795528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.795528Z digest=sha256:27e340cd32930ea68acc2238226122111429f4eba49a5c826796a265f01ebd53

Observation 642c7b47-686e-4b37-9202-3497554fc21a · outbound

This paper cites MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.799374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.799374Z digest=sha256:88eba92d512281ff80bdab714c18ba34ff06aef956b72cd6a8e1728266ccaf75

Observation 442acfa7-1bd9-4235-8824-f33b9a04493c · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.345348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.803209Z digest=sha256:7e45ccddfaf35a1e6731cd53f07bd3bf4443c5e10b7073301b8b4733aaf79d14

Observation c3ab9fa4-b95d-4483-8a88-3c089406603b · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.334384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.806264Z digest=sha256:e3156a8d9c6a7f7a4dbb5bb3ebc658c4a2e6f62f6165d4724993cac5a2cefa9e

Observation 56a4b260-275a-45f5-b687-9d6ec4439724 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.323492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.809370Z digest=sha256:595bd1b4c8f74d5c3875aece1bad3a9a6c3c8c61a5845898648a843906112f87

Observation 83f60d85-ae1a-43ce-bfa9-083fdd2e72ff · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.812686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.812686Z digest=sha256:9aa0a44501342231f8ecf25fdbd0d44369143f8f30c366c8245b6b7c485e963b

Observation e68e8702-e122-46c0-9500-451de0b6bcb1 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation CogVLM: Visual Expert for Pretrained Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.816068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.816068Z digest=sha256:2a634175d843f46dd49b3e2c3d798db089318919325ff80349a4eba4c64de8e5

Observation f8382428-2d73-451c-8dae-0309364dc75c · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.819685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.819685Z digest=sha256:2c05d9bc4eebc3a0fda36a99f872f9df37d94b5db6c48cfcf592f3eb5de10e8f

Observation f71c19a6-29b3-4554-89fa-06a4eb34e964 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.823310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.823310Z digest=sha256:b6931e59fafe803ddb9e278c92b0e3617a70ff7de260f3bf631569dbed2e8ad5

Observation fc69dfbd-bff1-487c-adec-93489945bde8 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.290479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.826541Z digest=sha256:692a92fb841b702651b1e752af59c5f7925f283682141e9bfdab69b08b071940

Observation 1b992b59-3a11-4214-8781-0888aec3cbb7 · outbound

This paper cites Instance-adaptive Zero-shot Chain-of-Thought Prompting.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Instance-adaptive Zero-shot Chain-of-Thought Prompting

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.830014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.830014Z digest=sha256:5c8c1fa3d390d2263db1ccd88003624f2d4e1b7c0f4368bbc63eb8c667f1cbc8

Observation c2a20de0-3f71-458b-b134-5d1d6262bc07 · outbound

This paper cites TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.833743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.833743Z digest=sha256:d5d8a0938cb388fe2aa04e788b8683f73cf168dcd55963d62c63bddbda89fce0

Observation 5afc7bfb-70f5-45c0-86c0-7b265c26c9ad · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.279234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.837610Z digest=sha256:3c250c41e3ed86a88956cfe0537d75bb91f917424792766956e6946dbaf46fb2

Observation 0695b377-4d15-41ff-9a21-6d63b07ac2d6 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.840954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.840954Z digest=sha256:7731f40fc7d4f2b04efcc91bb2427bbe64130e59659a5d4a1ed080ac9a37b5de

Observation a995798a-d4f3-4859-90c3-06a9585b31c1 · outbound

This paper cites Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.844437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.844437Z digest=sha256:05c6588acb6b5b08506d1255b3277ca400696636b46fb76685086130a77181e8

Observation 74983426-dfa4-4d02-85e0-4281e23a878c · outbound

This paper cites From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.848057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.848057Z digest=sha256:a4ac9812ee5f5f3045872c757d88f178760bacea4ac158b5ef4aa37f41ff5ed8

Observation 47904241-94ad-457a-85e9-7447f8ee7bd7 · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.852046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.852046Z digest=sha256:4503060e73977841ee7aad34d4e295b139547cd85a16b2c9ecaa94af97d46e46

Observation 46d43404-51d8-4228-9e7c-5db68ea67d8e · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.260897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.856179Z digest=sha256:73492ea10d6e88cfdc959e556d282dd85bdab131510bbbc18bebfb64d6e5360a

Observation 6ea5cc41-ee29-46f1-ba0f-10ee75153277 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.249757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.859985Z digest=sha256:61c912bcb86d80858fe2f955642b987e836274fde6c4e3bed5accb4e6a73af9e

Observation c4bba911-bde8-4b85-a860-85edb8eeb5a4 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.238750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.863687Z digest=sha256:fb2a4d7c9b3f9afb1eccde5299f59ea7044fcc9b5211bcef7db7f4f5db427dbd

Observation 3478655e-384c-403c-8842-4eb1ba3418d7 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.227225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.867315Z digest=sha256:bc4e91ff6e678b1fd280a22986918541dbf31382e2ee5c1f7055bd03a639034d

Observation 29ae8978-1afd-4aa4-9e5f-a58527ee7e6d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.870849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.870849Z digest=sha256:2ebd9a5d35090991e2cb41454515514b274456acd14e092c45febb7b401ec7e6

Pith citing papers

No inbound Pith citation observations are available.