Pith. sign in

Paper Citation Record · LEDGER

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency

As of 13 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2510.01009.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.01009 v3

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:22:35.965901Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1fa159f-0ae2-4cd6-827d-62e3c9089184 · outbound

This paper cites GPT-4 Technical Report.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:32.864573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:32.864573Z digest=sha256:73ab9477ef0f1dd40ef657aeaea9eaa183e07ac73030a55e8a7710b47f2fcce2

Observation d28358f7-6e0a-471c-bb90-8db79045d677 · outbound

This paper cites Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:32.906404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:32.906404Z digest=sha256:7e5da4449d75d4873424b3c081b94c35067abb3e9858bf79baed9e280582b01f

Observation 8eb72ff8-da18-4084-97c8-d18d3098e792 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:32.959339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:32.959339Z digest=sha256:b06fd61cd6dfcf32fae758a38ea176a49f6569da4349709d36dc970e4e6c492e

Observation f771fd04-7b2d-4444-8768-5e7858c4cd8f · outbound

This paper cites Vivit: A video vision transformer.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Vivit: A video vision transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.042395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.042395Z digest=sha256:b24d7132e0c4534e648a1188610ab00db6045934ae8fb5dfc58ecb0588476446

Observation ab416c63-d1fa-4e14-9e0a-437d18f683c7 · outbound

This paper cites Goldfish: Vision- language understanding of arbitrarily long videos.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Goldfish: Vision- language understanding of arbitrarily long videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.104568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.104568Z digest=sha256:16a46f8b063596aa0b56fb689b0c6d1d47db91bcfafb4fa5f90c8e50b2522c8e

Observation 13fa6ef2-939e-4d02-822d-ce5f3cc4761f · outbound

This paper cites Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Machine Learning, pages 813–824.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Is space-time attention all you need for video understanding? In Proceedings of the 38th International Conference on Machine Learning, pages 813–824

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.188025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.188025Z digest=sha256:19b2c1931b1bc44fb9ad5d410ec87e7e228c0fa256b610be817efd99e2723f2c

Observation dc12f8f4-c088-4f66-94bc-2f9d30b69dd3 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.250105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.250105Z digest=sha256:fdd3a989b855784fbc1a45034725fe223e36b2048fa118ff5cab4dcbee2d54b9

Observation dcccc084-5db0-4ccc-bb78-657ac82e8683 · outbound

This paper cites Vindlu: A recipe for ef- fective video-and-language pretraining.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Vindlu: A recipe for ef- fective video-and-language pretraining

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.305690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.305690Z digest=sha256:c048c6ee453d2ac64ab87d1c306c6a83ad521995e6e15bd7617560a0f48c21c6

Observation f8df1c74-c858-4cbb-bfa8-f46a40823712 · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.363413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.363413Z digest=sha256:0fd4aa52a25437123a1fbf62e18906b70cba21433c6de3a88462e0d73622072f

Observation 0be4c9d6-cf02-4784-a35a-fa43352f0d1f · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Qlora: Efficient finetuning of quantized llms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.420144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.420144Z digest=sha256:066dbc3d92ee15f6afa4380735ef2d71e8d25c38fd248ac8f4d0adb0a296e14d

Observation bfaa2350-16c1-4ef5-8d5a-ba4f979d49a2 · outbound

This paper cites Knowit vqa: Answering knowledge-based questions about videos.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Knowit vqa: Answering knowledge-based questions about videos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.471412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.471412Z digest=sha256:89ac66ff59df39a81cfa4645d21046c067ea3b59afbc42456a00ac93c584e41f

Observation e13ff2d9-b68e-4c4d-9bce-c3afb6d3f4b4 · outbound

This paper cites AGQA 2.0: An Updated Benchmark for Compositional Spatio-Temporal Reasoning.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency AGQA 2.0: An Updated Benchmark for Compositional Spatio-Temporal Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.529728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.529728Z digest=sha256:0420727a8edfdedc90267e5a998d78cabfa293e82defd23c5733978e76096468

Observation 1a62ce53-ac7a-4488-ad1a-10bb4b1ceaf6 · outbound

This paper cites Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge mem- ory.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge mem- ory

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.588211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.588211Z digest=sha256:081f28fc2a5d2114418eb62a284fdedd797710a2d5a920fcc8146cfe339852d6

Observation a829eb62-cfc0-4e14-8b83-8a4677c108ee · outbound

This paper cites Ku, Qian Liu, and Wenhu Chen.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Ku, Qian Liu, and Wenhu Chen

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.629803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.629803Z digest=sha256:103df3d8f64d7baf61d132d690cd718049ae1031b0279c0887e67d29c336afe0

Observation 33fb768c-b02c-49e9-9b1a-c66d0889da50 · outbound

This paper cites Visual question answer- ing: Datasets, algorithms, and future challenges.Computer Vision and Image Understanding, 163:3–20, 2017.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Visual question answer- ing: Datasets, algorithms, and future challenges.Computer Vision and Image Understanding, 163:3–20, 2017

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.689185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.689185Z digest=sha256:49d2dd3c2da8ef68964da1e295ca6135e732913029b0cb7b45bc80bd45b7f3b4

Observation 1704e858-929a-4b94-aeda-2aff16fb65db · outbound

This paper cites An image grid can be worth a video: Zero-shot video question answering using a vlm.IEEE Access, 2024.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency An image grid can be worth a video: Zero-shot video question answering using a vlm.IEEE Access, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.773614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.773614Z digest=sha256:656f4013a956d22ba53ab5cbac340f7b19601678150b561f012f4a65cc3aba00

Observation d0619942-7f50-4460-b84f-b62a7f3495d8 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency TVQA: Localized, Compositional Video Question Answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.816672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.816672Z digest=sha256:7522b34a0a74be17fff3f9acc3cf43ff921d8fa68e281edc6231658ab3265ffa

Observation a56dd022-f6b6-4b89-b253-71bc8c61c996 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.874949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.874949Z digest=sha256:af1d2463ba534bd140cf0ec29a05b22d90d076c5c16b5c52139d20ff65a00d4b

Observation 96318be8-ddff-4ddb-88fc-bc9c5f7b8dd6 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.924245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.924245Z digest=sha256:461fc8794c1d754bba2fd28bc8ad3f6b9f40448803ee222df2536816f20079fb

Observation 2ebd77bb-2af6-41d2-9d3b-68ddd1a43fd1 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:33.982304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:33.982304Z digest=sha256:846d604d69c2a6d439b4f3b7842ef668a005cd8672f134e114d732c0e51ecce7

Observation 4b3eb857-445b-43f8-b62d-fb3afcfac7d1 · outbound

This paper cites HERO: Hierarchical encoder for Video+Language omni-representation pre-training.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency HERO: Hierarchical encoder for Video+Language omni-representation pre-training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.060873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.060873Z digest=sha256:edecdc6c8ea83ad6bdaed5ac8ff7fe822269f7858f4bb186f6c27c0bc2bbc13e

Observation 230df3f3-e26b-4782-94b9-2932af801295 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.EMNLP, 2023.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Video-llava: Learning united visual representation by alignment before projection.EMNLP, 2023

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.096438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.096438Z digest=sha256:ed4f3568b6fb09c7c7572a9f2bb453175420e55c848dbe98ce192ba42bf53577

Observation 776dcebe-c3dc-4c7a-8907-93b65501d099 · outbound

This paper cites Common- sense video question answering through video-grounded en- tailment tree reasoning.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Common- sense video question answering through video-grounded en- tailment tree reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.150045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.150045Z digest=sha256:c3237b8a5ddbd4bd0c838fb6344fee9e211ab6872d4df4c2946b9661106cb332

Observation 91238d44-23f5-45e0-9537-c046fabea754 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.208012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.208012Z digest=sha256:e09732c6e03f2388500c70a90e98da6db1d015c5e68dbf76b500b7799b7a4a1d

Observation cc2e1ca2-8b63-4f5e-add4-4a821f048960 · outbound

This paper cites Question-instructed visual descriptions for zero-shot video answering.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Question-instructed visual descriptions for zero-shot video answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.252847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.252847Z digest=sha256:b7ab59ff8a7f3749ec883eae4776293504f9b02f2c1dfcabd1d77a7e724c3855

Observation 986407c1-5c37-447b-a358-0ae5bb9f5ce9 · outbound

This paper cites Streaming long video un- derstanding with large language models.Advances in Neural Information Processing Systems, 37:119336–119360, 2024.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Streaming long video un- derstanding with large language models.Advances in Neural Information Processing Systems, 37:119336–119360, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.335998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.335998Z digest=sha256:9c6e62f3d84abd46c5f1979e08bba288a4ebf69b5c4bcbe3f5e5f81b3e967337

Observation 35b22632-d826-47b1-b4fd-0e6c1e51a8f9 · outbound

This paper cites Direct prefer- ence optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Direct prefer- ence optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.387610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.387610Z digest=sha256:2caa4274fb29b74c7c5886fb327ad73c8a317056274de481656594433a6dbe46

Observation b4624298-6b57-4ac2-a74f-d0f676cc4057 · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.441656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.441656Z digest=sha256:b21ffb926df13ca8c97e4351f9f5436818cdc43582de71887d11e31d1fbfb3a3

Observation d22ebfc7-ef6e-418b-af03-7c5fdf45f922 · outbound

This paper cites Action Recognition using Visual Attention.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Action Recognition using Visual Attention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.505843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.505843Z digest=sha256:81be0534170488421aa927d1a4f26ba7e26d29a14cf4b8106d8a6cea89a42953

Observation e1c2e0aa-9720-440b-b197-e7c5170677dc · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Moviechat: From dense token to sparse memory for long video understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.560993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.560993Z digest=sha256:642a76588ea6b6e24da9735aa452f257231cbe02c272c4d20657059acf5a3196

Observation dbb01734-afaf-43e6-9f72-eb607114c636 · outbound

This paper cites Modularized self-reflected video reasoner for multimodal llm with application to video ques- tion answering.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Modularized self-reflected video reasoner for multimodal llm with application to video ques- tion answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.572884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.572884Z digest=sha256:1a7f0718c55a95e29ca16c93eb1c3c918bf4fda40992dcd0ea4dd4bab7c9bddf

Observation a26c2812-3d26-407c-b46a-4376dada978a · outbound

This paper cites Movieqa: Understanding stories in movies through question-answering.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Movieqa: Understanding stories in movies through question-answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.602802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.602802Z digest=sha256:8ceed0e9c54a839a2f8160811bac9a6b5675a7e003ea3f40f8b5946943045c51

Observation 02b5ffba-2c35-4879-a1f6-5f492720c4b7 · outbound

This paper cites MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.683325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.683325Z digest=sha256:ac147d77ffbd74d3602c8110a9dcf06fb3e032134de656933eaa43ca646fd46f

Observation 65826b61-f3c3-4dc0-a3cc-ab8cf5ebf585 · outbound

This paper cites Video- mae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural infor- mation processing systems, 35:10078–10093, 2022.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Video- mae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural infor- mation processing systems, 35:10078–10093, 2022

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.801271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.801271Z digest=sha256:9f1ca1c5ba64888f5daccddf21d4e2c25ec5d891af30a78cfad13c04ed007d6b

Observation 9a34a74f-3e2b-405c-9619-8b161a62fbdb · outbound

This paper cites Fastvlm: Efficient vision encoding for vision language models.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Fastvlm: Efficient vision encoding for vision language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:34.935141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:34.935141Z digest=sha256:1f158dc0b5297680129127014b6aff1829f66d5becfa08a1f3b5ab99d8efaebd

Observation b6b1023c-47f2-4c5e-8eae-f3c77288b4ea · outbound

This paper cites Temporal Segment Networks: Towards Good Practices for Deep Action Recognition.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Temporal Segment Networks: Towards Good Practices for Deep Action Recognition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:35.057234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:35.057234Z digest=sha256:da541790f414edc7b70a37e7b35d0589d08195d51d8efa13d70f6ddeb21fd619

Observation 49598a56-397e-4a7c-90b2-9d5a7e7f9b95 · outbound

This paper cites Vila: Efficient video- language alignment for video question answering.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Vila: Efficient video- language alignment for video question answering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:35.180439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:35.180439Z digest=sha256:3f2464528fec273ff14efa8bed47e227f1b61380c18e51bd5eac841b0d0a2b56

Observation 4cab872b-e3ab-488b-b404-c942bc03d25f · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:35.256568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:35.256568Z digest=sha256:96f8930c29c5af359e3ef60c3bc333b4c05ae1b7e27177db0997dc75d83ef8bf

Observation dc9d193d-3232-42ab-9b75-3bd67ae6e484 · outbound

This paper cites Longvlm: Efficient long video understand- ing via large language models.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Longvlm: Efficient long video understand- ing via large language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:35.344552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:35.344552Z digest=sha256:ec4877747ae6eb803f02d632308a9700d08f95c840fee66f1fe1f9b7c10a219c

Observation bccc7492-cf70-4dea-a402-134cf0aea9e6 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:35.455844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:35.455844Z digest=sha256:8cd90dc4b299c1ce1563e4dcbd89cd02d566045fa726071631d19b5d4841950a

Observation 29b384d6-847f-4ca8-9baa-7cd7870bde0d · outbound

This paper cites Adaframe: Adaptive frame selection for fast video recognition.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Adaframe: Adaptive frame selection for fast video recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:35.550061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:35.550061Z digest=sha256:6410e0d42fd782a1e87d5459fd8b782bd6a517f7a99828bf6fbc55b53b6e78cc

Observation 286b7aeb-d09e-4542-85c9-27b73dd22711 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining tem- poral actions.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Next-qa: Next phase of question-answering to explaining tem- poral actions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:35.619397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:35.619397Z digest=sha256:90633ef4ed51722cec80718798555104762d59edf927309ab52592e4f9fc7466

Observation 1228da27-da17-4897-b1a4-2f6bdd2c8883 · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models.Advances in Neural Information Processing Systems, 35:124–141, 2022.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Zero-shot video question answering via frozen bidirectional language models.Advances in Neural Information Processing Systems, 35:124–141, 2022

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:35.694555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:35.694555Z digest=sha256:957f0244d8d7d700e029d79e5b3f845e28a6e39d3add43c7f2f8d11540ecfae0

Observation 6d6a8ca9-fca7-4d50-a6d2-bf099951be2a · outbound

This paper cites Qwen2.5-1M Technical Report.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Qwen2.5-1M Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:35.773784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:35.773784Z digest=sha256:e54157bd07eb7672ab2d04ef889a55256311bcfc5e5931baf8ba71616af3cfb0

Observation 73be145b-6fc5-4477-a6ee-694af2d996e8 · outbound

This paper cites Enhancing Long Video Question Answering with Scene-Localized Frame Grouping.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Enhancing Long Video Question Answering with Scene-Localized Frame Grouping

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:35.874760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:35.874760Z digest=sha256:154651cf849a0c875fd028def8db5d9abe32f55e0e157db2ffaa65d94d6eec33

Observation 71ef35a2-c38d-4799-9ae0-68ffb6073932 · outbound

This paper cites Self-chained image-language model for video localization and question answering.Advances in Neural Information Processing Systems, 36:76749–76771, 2023.

POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency Self-chained image-language model for video localization and question answering.Advances in Neural Information Processing Systems, 36:76749–76771, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T13:22:35.965901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:22:35.965901Z digest=sha256:c890c2eb9eba7b5edbd0ab398964b72b380bdd44df1940f28d7dd94e8e9d6179

Pith citing papers

No inbound Pith citation observations are available.