Pith. sign in

Paper Citation Record · LEDGER

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens

As of 13 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2412.09919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09919 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:38:23.397829Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved19
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6e269592-11af-445e-aa6e-a0cf055d1045 · outbound

This paper cites Accessed: 2024-09-30.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Accessed: 2024-09-30

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.534530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.085961Z digest=sha256:780ec9942dd5adb224897109eb4872b18634677be32a58560dd55c67906ba697

Observation 5cce60a5-a4a9-4ed6-90ef-b513cafce434 · outbound

This paper cites Flamingo: A Visual Language Model For Few-shot Learning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Flamingo: A Visual Language Model For Few-shot Learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.520243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.090969Z digest=sha256:4f075ba6b6f4ee56c96ece4db530f7944410982c3b54af9c1bf658bc52971820

Observation 5b1a2745-d8f5-40b8-8072-e37ceeccd44d · outbound

This paper cites Deepspeed-Inference: Enabling Efficient Inference of Trans- former Models at Unprecedented Scale.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Deepspeed-Inference: Enabling Efficient Inference of Trans- former Models at Unprecedented Scale

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.505394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.095827Z digest=sha256:e35f41c73698bcc15308bbb9f82d4bb8dfeeaca2d2a197a078375fa442217930

Observation 0ed46894-0977-4b78-b1a0-fcc333a4411a · outbound

This paper cites Qwen Technical Report.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.100815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.100815Z digest=sha256:d23626af3c7e5c20c9c23c0dd15b1c5001694ec91ab991002ccdb7bc47def38d

Observation 24e6e467-6e9d-48d9-84df-282eb0896330 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.489843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.106163Z digest=sha256:ab1cae9d49252ac700237904e1a7e66e136c8fe781415ef3ef93bce4d297cf26

Observation 66863b66-26b1-4de9-999c-5059196ed2d0 · outbound

This paper cites Token Merging for Fast Stable Diffusion.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Token Merging for Fast Stable Diffusion

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.472817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.111065Z digest=sha256:0cf92e8c8b7aaaf907f56ee99682577bf3a50629fe8094be35545038d27308e6

Observation 165db5d1-38ec-480f-bb5e-65fb6fdd29d3 · outbound

This paper cites Token Merging: Your ViT But Faster.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Token Merging: Your ViT But Faster

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.456825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.116301Z digest=sha256:ec19642424eab0fd9183b6851b3de719a4abf8f4c886a66530d0508c38ecb082

Observation fb4e4d0b-64f3-44c5-bbd6-066c0d878c26 · outbound

This paper cites Dif- fusiondet: Diffusion model for object detection.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Dif- fusiondet: Diffusion model for object detection

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.440804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.120851Z digest=sha256:9709f74148b9f12124eab85b394571f70dc6857daaa7f8c32a5f5bb76f0cbb5d

Observation 97cfd238-8e33-4990-bfe2-38f7dded1efd · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.125642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.125642Z digest=sha256:dc4afc006dab50af68685392b84f59afa7db04a24d097044a7473bfa8f0fa99f

Observation 4a20ac0e-7c12-4615-828d-320b0218b5bb · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.130731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.130731Z digest=sha256:3d7cccd5b948fe79d75ebbd7c915aa8fc9c474b93cab41b37c976404726edbe5

Observation 7c5aca54-ace4-42a8-8118-b1de65e80b0e · outbound

This paper cites The Llama 3 Herd of Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.135791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.135791Z digest=sha256:f66b2efe0726da92b96227e426e8cca2a3606ea968415234fc81a17704188831

Observation 2330e1df-0570-4baf-9e50-13fe6758d348 · outbound

This paper cites ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.422363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.141476Z digest=sha256:9a701bd4d97e46fa95b2f3da8ed475f5607d3653656d43c13b2821cc95bf8a7f

Observation 6c8bf51d-4b97-4f92-8a27-f03d3374c844 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.146380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.146380Z digest=sha256:4190722b17aec66de51569239aa2ffe553140409517239a19f01ec023d3c147a

Observation 8781f6cb-a499-4491-b7e3-42c24537ec2a · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, 2024.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.407378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.151391Z digest=sha256:3294b879c90898d401b473dd078f42838f2379df9dc2f9b2cf1dffc69eae1797

Observation 3f706293-107d-4993-bbd4-e83e824ee93d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.155758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.155758Z digest=sha256:c8338df6ad7c453a4f3792f6ac1e4b049353f1182047f441a89def4dfb1dc502

Observation e8f1a4dc-1609-4fb8-9d2b-03606732a319 · outbound

This paper cites Making the V in VQA Matter: Ele- vating the Eole of Image Understanding in Visual Question Answering.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Making the V in VQA Matter: Ele- vating the Eole of Image Understanding in Visual Question Answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.392962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.160786Z digest=sha256:d4c06e731eeedef2f3e47d88179effad912d0f5d6809392ead093cf4698d63f7

Observation f1228d8d-5a4c-4164-9e72-a3f607853442 · outbound

This paper cites Vizwiz Grand Challenge: Answering Visual Questions from Blind People.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Vizwiz Grand Challenge: Answering Visual Questions from Blind People

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.378034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.165606Z digest=sha256:ec3028c4bec1ae97640a8d671ac17605c0ed7309fafff8f01b67cb3242e979a6

Observation 6a00835a-f980-4756-a9df-70f5f278827b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.170289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.170289Z digest=sha256:2e4f1f0bed899af50ba1e70e702dca6b1f46c1e3608d3c22229fcb94ae6555f2

Observation 858c5d64-c8a9-495f-be26-980af22e3251 · outbound

This paper cites Vision-based freezing of gait detection with anatomic patch based representation.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Vision-based freezing of gait detection with anatomic patch based representation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.363992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.175362Z digest=sha256:99723690d0355539485fc2acbbf51d16ae6ef44db388ccff2c00b982946f031a

Observation 5b56b643-4a71-4cfc-9c8a-74f0e0f86360 · outbound

This paper cites GQA: A New Dataset for Real-World Visual Eeasoning and Compositional Question Answering.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens GQA: A New Dataset for Real-World Visual Eeasoning and Compositional Question Answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.349286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.179941Z digest=sha256:5545139ab51b3b1698c88d7c870da951a36cb729f3c6f7f9e3d0e0e05f21c8df

Observation 7794ffa5-db67-469b-b98c-b57fb734b362 · outbound

This paper cites LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.184668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.184668Z digest=sha256:2f7eac4517bcd5abf7a8cd892328a9195077eab8f60ca443ee770f65e122facc

Observation d3431e70-d808-426f-8882-589bf490cab8 · outbound

This paper cites Chat-univi: Unified Visual Representation Em- powers Large Language Models with Image and Video Un- derstanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Chat-univi: Unified Visual Representation Em- powers Large Language Models with Image and Video Un- derstanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.334087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.189447Z digest=sha256:232341ed92fa0deb18522e3c63749e52910448a366daca2c8c004b120b45c3f7

Observation e95e7ffa-4f76-4850-ac58-5a023780b521 · outbound

This paper cites ReferItGame: Referring to objects in pho- tographs of natural scenes.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens ReferItGame: Referring to objects in pho- tographs of natural scenes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.319453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.193513Z digest=sha256:650923a54470c3c5a3dfb82c768f5b152852e1153593715158366e8e9fb247a5

Observation 050cfcae-719a-4d2b-aa14-f9261a73546f · outbound

This paper cites Shamma, Michael S.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Shamma, Michael S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.303994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.197464Z digest=sha256:f7c13ef32192a2e29de252274f169367801742cad19bd2f6ec6f589df73c6b4b

Observation 68eec932-30cf-47ef-bfd9-ffb94fc0e74a · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.201279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.201279Z digest=sha256:4a788af2edb440f3da73f9ddff1f16f3fb429486d43985f233d5a6ba59d9c3d5

Observation 988303e2-7d51-4828-bbcf-77e5927944d9 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.289534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.205522Z digest=sha256:ff76b977d2a00468465831721254bece9c2eb9fd6615d5e5977321083d74144b

Observation 06f1b6c1-cfe4-4fcd-b12b-9167853b30d2 · outbound

This paper cites MVBench: A Comprehensive Multi-Modal Video Under- standing Benchmark.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MVBench: A Comprehensive Multi-Modal Video Under- standing Benchmark

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.274403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.209832Z digest=sha256:e525ee62c33923541a77aabc7d61071b534e7b688ffee02cd15f119475e3f899

Observation 518762f7-23b3-4235-92d4-4c1344c7aecb · outbound

This paper cites VidToMe: Video Token Merging for Zero-Shot Video Edit- ing.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VidToMe: Video Token Merging for Zero-Shot Video Edit- ing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.258141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.213617Z digest=sha256:338037ffe32d0e1dc9fd1e37aad680122b05b0603457b52067c7dfa6f286c2ef

Observation 51dd0415-e64e-4783-9a9e-fda1710687c5 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Evaluating Object Hallucination in Large Vision-Language Models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.242917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.217336Z digest=sha256:f6d4d78bc276cf2eb64b1e6db0c33423361925cc9f8a549ee7a77440f540c0fb

Observation 92a16398-b545-44e4-8979-57ae413ca2be · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.227794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.222025Z digest=sha256:d30659cc886e73b07feb1023236e3211a826e30c235a897c20102ba539dfcb72

Observation 2e3f976d-c981-44d5-a858-17164d855aee · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.226456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.226456Z digest=sha256:26abeb13524fafa091911825d2a2edb37d8c959c45dfcefcb68e9eb5e66cb36d

Observation 3f5c0778-e792-4e59-8a65-efc0971f955e · outbound

This paper cites VILA: On Pre-training for Vi- sual Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VILA: On Pre-training for Vi- sual Language Models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.212782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.232050Z digest=sha256:df0c09f9a754f95541297c8bf37ca38d72a97f9da0e3aba846d0fa9b63e1608a

Observation 32ebfc29-6c55-4996-aae1-a76ca8583fbc · outbound

This paper cites Visual Instruction Tuning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Visual Instruction Tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.199799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.237633Z digest=sha256:8faf66fcf63f6de0e01e2fd73176f60234fd545c902118a4dbf5764cc507e2b3

Observation 07d0d442-44ae-4980-a08c-383de4b60c55 · outbound

This paper cites MMBench: Is Your Multi-Modal Model an All-Around Player? In ECCV, pages 216–233.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MMBench: Is Your Multi-Modal Model an All-Around Player? In ECCV, pages 216–233

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.186301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.242415Z digest=sha256:6950552062f51f0216643e560608da0a05948e7be7421573e0ab546f8e33d53a

Observation 7a9a20a6-08bb-4b17-9cb3-637fd6a280f8 · outbound

This paper cites Decoupled Weight Decay Regularization.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Decoupled Weight Decay Regularization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.247628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.247628Z digest=sha256:f22bc1b1d9adafb204a2d39b552f16334ae6818013fbbb3e2f4be6942028154a

Observation 7e65f279-9a63-403a-824c-3ac0e044ba28 · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.171875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.252816Z digest=sha256:251aeee4b412a652c2e1973d4ab08055a952d1025bc91832dca781c6883cdcd0

Observation 548e78ee-8e54-4b6a-8702-cc6a77bc6639 · outbound

This paper cites Autoregressive omni-aware outpainting for open- vocabulary 360-degree image generation.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Autoregressive omni-aware outpainting for open- vocabulary 360-degree image generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.156331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.257615Z digest=sha256:ec11077e1670d164a7dc5ba58ee084d5b11e4b7a29bfccee49c8e663ea8a9b12

Observation e4ff90b6-9572-4e9e-bb04-3ebbee2a1dcc · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.262723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.262723Z digest=sha256:6aaa989c04b843512d6290cfa7e4db2be2511a6826b0c537155a981b4975a76f

Observation f691b681-a29b-4662-b668-5adb143c233a · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.141744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.267826Z digest=sha256:3aaf104f05e22de7f13715595e6507f9f85b33ef85652bf38e5022745cf5b663

Observation 1380e099-43ab-480b-8e23-5aa7f9838e04 · outbound

This paper cites Egoschema: A Diagnostic Benchmark for Bery Long-Form Video Language Understanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Egoschema: A Diagnostic Benchmark for Bery Long-Form Video Language Understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.126565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.273657Z digest=sha256:06f2215f37e543cd8716aeabc21975d66ce82ce92f80d5df6b46ede9e7a98207

Observation ddd3961b-6767-4666-8164-aac1912eaec5 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions, 2016.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Generation and comprehension of unambiguous object descriptions, 2016

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.111537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.278963Z digest=sha256:d09a354e32d58baa1ca558d3df85a432a816725c4434cb30347a39371725fee9

Observation 185ae4d6-ef84-462f-a676-5b1346e15f11 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Ocr-vqa: Visual question answering by reading text in images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.096536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.284075Z digest=sha256:a05958d2255f5e6530e999c0c04bf03465a33aa4327321dc60c3b907c58b4141

Observation 0ae5b099-655e-46a9-b95f-99c036a5e7cf · outbound

This paper cites Perception Test: A Diagnostic Benchmark for Multimodal Video Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.082201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.289095Z digest=sha256:8912ec0ccf1530012f33cf610cd6ea6088b20c2aed9d8ac24fbf767c97fe264b

Observation f9c5f32d-f17c-4c41-a43f-f69bcfcbb382 · outbound

This paper cites Learning Transferable Visual Models from Natural Language Supervi- sion.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Learning Transferable Visual Models from Natural Language Supervi- sion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.065917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.294097Z digest=sha256:38df6749700e3f842a799f894dfbc6e17a3f49967ce19d3436c037a1ebe761e5

Observation f605098e-75ae-43b5-9fdf-74ae1e73dc5b · outbound

This paper cites A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge, 2022.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge, 2022

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.050404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.298810Z digest=sha256:9c5b940347fe540a409d7e1ac54ab0ea82ea935d427b67af57506bd77552186d

Observation f567e276-c483-4e2b-ac6b-118259650460 · outbound

This paper cites Llava-prumerge: Adaptive Token Reduction for Efficient Large Multimodal Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Llava-prumerge: Adaptive Token Reduction for Efficient Large Multimodal Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.303588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.303588Z digest=sha256:4dfb52e74cae291738ed8a803b780f8cd9617865fdd53d701e360b98968ab28e

Observation b534e526-2aa0-4464-8f86-c197afa5c564 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.034898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.308314Z digest=sha256:d17cd6afc87185e4de927cdb394906f11c51d163ba5229d704c4222408906bd6

Observation 0203478c-b5fc-41e7-a016-a25cda20fc82 · outbound

This paper cites TextCaps: a Dataset for Image Captioning with Reading Comprehension, 2020.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens TextCaps: a Dataset for Image Captioning with Reading Comprehension, 2020

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.020566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.312965Z digest=sha256:6c542ef8378658693dee66b31a2d8518b720680a2cb9eb4f0288e8fd39614c1b

Observation be74fcc7-6dc7-4e10-a015-78bf9100fdce · outbound

This paper cites Towards VQA Models That Can Read.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Towards VQA Models That Can Read

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.006513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.317610Z digest=sha256:102fa38474de375ab959ade8fc89fa59b2100c9fe29cbca5bec76fc6c2daec80

Observation 4c438198-f11a-4cd1-a214-35405fb7ca7a · outbound

This paper cites Moviechat: From Dense To- ken to Sparse Memory for Long Video Understanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Moviechat: From Dense To- ken to Sparse Memory for Long Video Understanding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.992233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.322443Z digest=sha256:1adca572f170e343acb76ec8244ef320131aff1dc03f441529a1537cd3beb65b

Observation 2d20aec9-ae63-4483-a3d9-132da3fc07ce · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.326515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.326515Z digest=sha256:895590c6bd25974f6409dc80d4148c5120cc117d84085a3d2f723f5e2d8b57b9

Observation 927d1f5a-eb39-4d25-a0d6-5dd9d7e7c996 · outbound

This paper cites [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.330686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.330686Z digest=sha256:7d3c1cb7f05d20fd5efa138725eca5a4c2df09ac1e515ce77c2fa360c00e69f9

Observation a58606c6-b27b-4f42-bc8c-f950dcfea81b · outbound

This paper cites VisionLLM: Large Language Model is Also An Open-ended Decoder for Vision-Centric Tasks.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VisionLLM: Large Language Model is Also An Open-ended Decoder for Vision-Centric Tasks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.975446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.334738Z digest=sha256:5b40d0c923a7e3853e454c8691f220d7bcd390f02dee25c6e5e3d3fb50a911fb

Observation 40d9be43-8b8c-4ab3-b5f0-119fd11e7d43 · outbound

This paper cites LongVLM: Efficient Long Video Under- standing Via Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LongVLM: Efficient Long Video Under- standing Via Large Language Models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.959253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.338794Z digest=sha256:3616571d587fb15ff3e4753dbd994fcb4b41332ce8df499d4091025ec7af0cd2

Observation e7dd3e38-dc97-4301-ab75-fa784b31980b · outbound

This paper cites Video Question Answer- ing via Gradually Refined Attention over Appearance and Motion.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video Question Answer- ing via Gradually Refined Attention over Appearance and Motion

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.928832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.346774Z digest=sha256:2b510e43d3d05d3798d3acad2e3e6458fab12f953ac859814332e84606ff1db4

Observation ad4deac9-1809-4d49-8119-ad02323728de · outbound

This paper cites Qwen2 Technical Report.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Qwen2 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.351205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.351205Z digest=sha256:7ad8c4582644e2f8ff587a537556fd156f0c699d8013fe87780dd707e10c817c

Observation 1440bf94-3c60-4b0a-a6ba-06e2acd78b34 · outbound

This paper cites SurgicalPart-SAM: Part-to-Whole Collaborative Prompting for Surgical Instrument Segmentation.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens SurgicalPart-SAM: Part-to-Whole Collaborative Prompting for Surgical Instrument Segmentation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.355997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.355997Z digest=sha256:a170bcd216d3fdb4aa166fc25689722998f4ecb8e64f1f208e8307c5beb05bcd

Observation 8bf437b6-282a-4b7d-9fd5-9bfd9b010e8e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.360911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.360911Z digest=sha256:67b00dbc22041f8d5e6439f0915fa5e09a07084bffd526f01c808fa03f528e76

Observation de2b71cb-7949-42ce-9489-ce89004ba5e7 · outbound

This paper cites LLaV A-NeXT: A Strong Zero-shot Video Understanding Model, 2024.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLaV A-NeXT: A Strong Zero-shot Video Understanding Model, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.914436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.365795Z digest=sha256:3cde00f897a540cb1075a6f9375e2d1da328763b7f25f3d82c41a3f1f3070102

Observation 2e7fbe65-0514-494c-b5a9-b5f9956e60c7 · outbound

This paper cites Needle In A Video Haystack: A Scalable Syn- thetic Framework for Benchmarking Video MLLMs.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Needle In A Video Haystack: A Scalable Syn- thetic Framework for Benchmarking Video MLLMs

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.897803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.370070Z digest=sha256:41cc04d23297b6cada1863885a3a6d12549caf3b058fa64bafb232b53febe238

Observation 1d57da18-8074-49fb-bc71-4fbfdcd2012d · outbound

This paper cites Clip in medical imaging: A survey.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Clip in medical imaging: A survey

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.880456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.374598Z digest=sha256:1e98a6cfa68d1adf2d8e84404260ae15239a2143ba52ca1adc446c5a08459976

Observation 4e9f2669-99f0-4716-809f-fdfd699558a7 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.379009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.379009Z digest=sha256:f0ed30bc545739201680e98a72ad30aab6277454cdbc35ae538129438b628542

Observation 4f8d223f-0972-4550-9f92-406505997ab6 · outbound

This paper cites A closer look at the cls token for cross-domain few-shot learning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens A closer look at the cls token for cross-domain few-shot learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.864704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.383842Z digest=sha256:c69c3151f67dd552fdde70b649305d792e5e1f520ae7e6aafeb68bebbbaea88a

Observation 1a170a45-5858-417e-ad21-d9e763a9e8ce · outbound

This paper cites Training Details We adopt a two-stage training strategy [9, 30, 33], dividing training into pretraining for modality alignment and fine- tuning for instruction tuning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Training Details We adopt a two-stage training strategy [9, 30, 33], dividing training into pretraining for modality alignment and fine- tuning for instruction tuning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.849232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.388255Z digest=sha256:95e8fd6e135705a34d13bb2a88edfceaf3e2866f248348f9a6914d1d3ffcd434

Observation be9a685a-eca2-48a4-a5dc-bca19ff2bb72 · outbound

This paper cites Additional Discussion on Different Frame Se- lection Features.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Additional Discussion on Different Frame Se- lection Features

Reference 66

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T16:38:23.834550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.392641Z digest=sha256:d8086ee5b5722d878fdbbe01b29e13d74836d0e127b07bcc70385fa41c41011b

Observation f95bd0b0-8aba-4b16-b736-e932672b0924 · outbound

This paper cites Game Science.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Game Science

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.817966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.397829Z digest=sha256:6a77241869021bc50123b68c8e5d405c5b583425e2f66b5e708b2368838c617b

Observation 3039831d-a05b-48b2-935a-8e8423126b96 · outbound

This paper cites an unresolved cited work.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Unresolved cited work

Reference 470

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:38:23.943843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T16:38:23.343023Z digest=sha256:6513143e92b59126824f2dbf5b39d524262368060351ca69e461f38d3a99bd43

Pith citing papers

No inbound Pith citation observations are available.