Pith. sign in

Paper Citation Record · LEDGER

Building a Precise Video Language with Human-AI Oversight

As of 4 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 1 inbound Pith citation observation for arXiv:2604.21718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.21718 v2

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T00:37:31.858728Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T02:10:27.595446Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-12T02:11:15.529838Z

Reference resolution

100 of 109 outbound references displayed

  • verified exact53
  • verified fuzzy37
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9eedcf7-ea0e-4e25-a36b-53136ad5baff · outbound

This paper cites Critique-out-Loud Reward Models.

Building a Precise Video Language with Human-AI Oversight Critique-out-Loud Reward Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.566261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:99cc28403e2c58f8874b5b0a759edc3ba60fc97e99c60a42d81085b79af18c9d

Observation a15b36a9-931e-4aae-af9e-b21bc85e7a8c · outbound

This paper cites Cycle consistency as reward: Learning image-text alignment without human preferences.

Building a Precise Video Language with Human-AI Oversight Cycle consistency as reward: Learning image-text alignment without human preferences

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.594783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:45ab3fbee04732e0458c967d406b7033af121426852462afb7c10f338bf55715

Observation 06f4607a-d2a0-41d7-9c47-5a6e580fd043 · outbound

This paper cites Qwen2.5-VL Technical Report.

Building a Precise Video Language with Human-AI Oversight Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:46:04.578883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:e8bbc068654bd3c7731bc66aba3bde1002ae56d9261cac1b3f3859e9435783a6

Observation ec221864-2935-4a42-aa07-75563e7a935b · outbound

This paper cites Improving image generation with better captions.https://cdn.openai.

Building a Precise Video Language with Human-AI Oversight Improving image generation with better captions.https://cdn.openai

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.762533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:3371e9fb5026caa3bcc33b8dd61be4951db671e91920b6433dc2fd9634f1e190

Observation 50419c15-4d9e-43c2-949d-5d1335d125af · outbound

This paper cites Measuring Progress on Scalable Oversight for Large Language Models.

Building a Precise Video Language with Human-AI Oversight Measuring Progress on Scalable Oversight for Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:01:41.422297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:0520cd63b869de857165418d79873cd11f137baa7d84fb208634e39c237d6466

Observation c697a69a-b83f-45aa-8222-9158dc93e9e7 · outbound

This paper cites Activation Reward Models for Few-Shot Model Alignment.

Building a Precise Video Language with Human-AI Oversight Activation Reward Models for Few-Shot Model Alignment

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.573192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:921751f0a51bf807549a026934a93368a8e4bf1ca365f9e2a5d7ae079ddef14e

Observation 9fdf966d-545e-41f9-a686-acde1a8d4700 · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

Building a Precise Video Language with Human-AI Oversight AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.576239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:ac32b62bddaa62fd4264daa0c976af46de164324ca8df56fa6744ae5aa29d1b6

Observation 31541708-eb04-487d-9adf-4426ebe3b289 · outbound

This paper cites Agneet Chatterjee, Rahim Entezari, Maksym Zhuravinskyi, Maksim Lapin, Reshinth Adithyan, Amit Raj, Chitta Baral, Yezhou Yang, and Varun Jampani.

Building a Precise Video Language with Human-AI Oversight Agneet Chatterjee, Rahim Entezari, Maksym Zhuravinskyi, Maksim Lapin, Reshinth Adithyan, Amit Raj, Chitta Baral, Yezhou Yang, and Varun Jampani

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:46:04.581628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:bff163e8040add2b3db12d91c7b2fea6975401a82521e7181707406d3a697fca

Observation c00a7c3a-e33c-4d74-beb9-b97b8187c1d4 · outbound

This paper cites How people use chatgpt.

Building a Precise Video Language with Human-AI Oversight How people use chatgpt

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.622035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:5136ba0312d52360be5d4ccdbe2ac169f6b78ab43b08ab02c0b482ad39b4a6bc

Observation d81cc469-82bb-4da0-ba32-304f0a53c70f · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Building a Precise Video Language with Human-AI Oversight MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:2560a4640b41f63e406d716c09d4ca528f2b4698b660f10aefee7c6b1826bb37

Observation 1b03c787-31d2-4f1f-ab7c-09ba1e959249 · outbound

This paper cites SkyReels-V2: Infinite-length Film Generative Model.

Building a Precise Video Language with Human-AI Oversight SkyReels-V2: Infinite-length Film Generative Model

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:23:04.486750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:c0ae0322cc59037ef3eb7beea83f22d2695d8ad13d7db03d960e3bff5a447fce

Observation 8d6fb856-8573-444f-8a44-1678397164f0 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472–19495.

Building a Precise Video Language with Human-AI Oversight Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472–19495

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.606384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:8abc1abf0004d22478be16b756a88337029af30b3eea85b9096db0c6750cc1aa

Observation c5d467a7-3f64-48a8-8920-ed7d4b1e35a8 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross- modality teachers.

Building a Precise Video Language with Human-AI Oversight Panda-70m: Captioning 70m videos with multiple cross- modality teachers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.742056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:e315c9f1654a01e0ac45a7af55c3523767b43f1188979373519003394b4388cd

Observation 0e9efa01-f19c-43b2-af33-481834f54bb4 · outbound

This paper cites Dress: Instructing large vision-language models to align and interact with humans via natural lan- guage feedback.

Building a Precise Video Language with Human-AI Oversight Dress: Instructing large vision-language models to align and interact with humans via natural lan- guage feedback

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.746356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:48ee2cee62976e7abd9107c751d881ff0d5444fd93149862c22960148e78e163

Observation 90b75ce0-2758-4e5f-b709-91725c335095 · outbound

This paper cites PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding.

Building a Precise Video Language with Human-AI Oversight PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.390564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:e92cc922cb66c43ebe5a8d27834a71ec281bd941f37625d02799b4451a21da58

Observation 401c45f7-e49b-406c-bab4-dd2801b4cc3f · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Building a Precise Video Language with Human-AI Oversight Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:46:04.491750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:a364316235b3986f595343fdcb280fe38f99bfb692971c47bff684d51f044f69

Observation 1d78293f-a1b4-4659-b132-aec9bab34279 · outbound

This paper cites Synpo: Synergizing descriptiveness and preference optimization for video detailed captioning.

Building a Precise Video Language with Human-AI Oversight Synpo: Synergizing descriptiveness and preference optimization for video detailed captioning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.488544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:d3eeff1fad86538a7c35e5cd688b3694cd1ff1bc5566d532f725268c22125446

Observation a6a72b4d-aad8-4f9e-a7fb-06cdb0c0d91c · outbound

This paper cites Ties mat- ter: Meta-evaluating modern metrics with pairwise accuracy and tie calibration.

Building a Precise Video Language with Human-AI Oversight Ties mat- ter: Meta-evaluating modern metrics with pairwise accuracy and tie calibration

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.750340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:52cd1665a4ea6cc1e8eb64f13465c33b705207b1a5011dbc2241aeb8deb4d1a1

Observation ed2839be-4249-4b15-8cd0-bcfac256da9c · outbound

This paper cites Mo- tionsight: Boosting fine-grained motion understanding in multimodal llms.

Building a Precise Video Language with Human-AI Oversight Mo- tionsight: Boosting fine-grained motion understanding in multimodal llms

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.474762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:5c559f492506ef978926ba3df8d2cb1e5d1ca61001334438f1f49282ea87a6d0

Observation 1a3dc349-b33e-4776-859b-94bd6403d8a1 · outbound

This paper cites Improving clip training with language rewrites.

Building a Precise Video Language with Human-AI Oversight Improving clip training with language rewrites

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.766452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:c8fb045adab96d4c5706e8d7bea141989ef24e19afcb11a4c25e422d5fb5af19

Observation 2e6f6cf1-5234-48f3-8b61-d07c1bd0f19d · outbound

This paper cites Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline.

Building a Precise Video Language with Human-AI Oversight Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.497600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:f632a06fa44c8f2e85733ba9e4c6cf80bef544c5f78f72a5397eeb8f10cd830c

Observation 1e91b538-f976-4b34-a3de-f53d3bdf5476 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Building a Precise Video Language with Human-AI Oversight DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:46:04.444596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:50a6b3a0e282c01fb2a9d273eeca8b153a7b31e9245918740835abacf50187e8

Observation 4eab8fb4-88dc-4e4d-a727-82e952cacafe · outbound

This paper cites Captioning images taken by people who are blind.

Building a Precise Video Language with Human-AI Oversight Captioning images taken by people who are blind

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.610019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:c7fc919d2752b52846f2edab1dcb3ac016ead1cee0d7a25e007989a914edf431

Observation d8a87605-6d4d-45aa-8f76-c29c9c799c21 · outbound

This paper cites Miradata: A large-scale video dataset with long durations and structured captions.Advances in Neural Information Processing Systems, 37:48955–48970.

Building a Precise Video Language with Human-AI Oversight Miradata: A large-scale video dataset with long durations and structured captions.Advances in Neural Information Processing Systems, 37:48955–48970

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.629694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:ff4853a1caee9d4f2e996e0c46a356bc3a13e8d47606cf7ce464665d3cf8431d

Observation 4a43f467-6b72-4538-8332-b8e9ab49a570 · outbound

This paper cites TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos.

Building a Precise Video Language with Human-AI Oversight TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.355382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:cc16e4bdb0db7f91d28eda1d9b4dd96272219d001770c822845d97e36a0b95d9

Observation d76ca6e1-d434-4986-8296-1874eb0f27f6 · outbound

This paper cites Dense-captioning events in videos.

Building a Precise Video Language with Human-AI Oversight Dense-captioning events in videos

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.625831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:fe2a4224f0475a2475c38adbe90f522a751061415bb8d55b87e48c125adecc66

Observation ae0e9b7c-b088-4cc8-b11e-3f19f7350c8d · outbound

This paper cites Videopasta: 7k preference pairs that matter for video-llm alignment.

Building a Precise Video Language with Human-AI Oversight Videopasta: 7k preference pairs that matter for video-llm alignment

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.362744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:740aca72d1de85fc165121d3984bfb5cb9ee88bca48e04e51d3415a15be72642

Observation 116576c4-bd88-43ba-8753-b4f5db295734 · outbound

This paper cites Mmr1: Enhancing multimodal reasoning with variance-aware sampling and open resources.

Building a Precise Video Language with Human-AI Oversight Mmr1: Enhancing multimodal reasoning with variance-aware sampling and open resources

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.449001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:7a10fe6c98533aa0d884347e2bae29045bc908d8622b215736773208ed70a995

Observation ac5efaae-6185-4e1b-870f-79d132f06187 · outbound

This paper cites Evaluating and improving compositional text-to-visual generation.

Building a Precise Video Language with Human-AI Oversight Evaluating and improving compositional text-to-visual generation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.738018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:91152a6047b74430b35cd81c815ecfec26b25d1e9de4452dbc7af318d557f073

Observation bc7a5f59-65ab-48d1-8a02-19a18dca9090 · outbound

This paper cites Naturalbench: Evalu- ating vision-language models on natural adversarial samples.

Building a Precise Video Language with Human-AI Oversight Naturalbench: Evalu- ating vision-language models on natural adversarial samples

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.618555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:ae1248ef266bc12e775e572d3f81b2a794680f35a5ca0b02b991f5fcd50bbbac

Observation d56732cd-baac-4ee8-89e9-ba0e76b85cfc · outbound

This paper cites Fire: A dataset for feedback integration and refinement evalu- ation of multimodal models.Advances in Neural Information Processing Systems, 37:101618–101640.

Building a Precise Video Language with Human-AI Oversight Fire: A dataset for feedback integration and refinement evalu- ation of multimodal models.Advances in Neural Information Processing Systems, 37:101618–101640

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.659959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:0c9b0a309ea019b0293ccb3cd748f66bf03322d4737351b9b4138aaf9c78d0b7

Observation 1da19f52-9a29-4949-acb7-04330ab36821 · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

Building a Precise Video Language with Human-AI Oversight Describe Anything: Detailed Localized Image and Video Captioning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.395869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:c3ed4d14e7cd9877ca819be92a203ce1d99f15296015601180dedf5d4d280a55

Observation b5f31781-152d-4b68-9fe1-3bff1c292f0c · outbound

This paper cites Multimodality helps unimodality: Cross- modal few-shot learning with multimodal models.

Building a Precise Video Language with Human-AI Oversight Multimodality helps unimodality: Cross- modal few-shot learning with multimodal models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.614146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:a22547ed00462c73b673507f96b0959db2db2748e534536d734c2ca6813e604c

Observation ed0e9454-7441-4148-89dd-c43c9464f95d · outbound

This paper cites Revisiting the Role of Language Priors in Vision-Language Models.

Building a Precise Video Language with Human-AI Oversight Revisiting the Role of Language Priors in Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.507451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:729524e89c656dd64c29c1af103ccdca22da550bf7e58367bcf5fa20b3b4c8e4

Observation 610a4531-40be-463b-9e7c-4064aefe9996 · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

Building a Precise Video Language with Human-AI Oversight Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.455477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:e8500bf2387bcc6047a3719e25ca81252fac45ff555114d07c17e06c81097394

Observation d78503c1-5dde-4e15-bf05-6da70b76daa8 · outbound

This paper cites Towards un- derstanding camera motions in any video.

Building a Precise Video Language with Human-AI Oversight Towards un- derstanding camera motions in any video

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.754022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:a11b6e479f633b606455f7ddacd40e073a5be221545e48bef8a131ac5ccb34c2

Observation 02e83265-309d-4208-9ee6-014d57a575ba · outbound

This paper cites Language Models as Black-Box Optimizers for Vision-Language Models.

Building a Precise Video Language with Human-AI Oversight Language Models as Black-Box Optimizers for Vision-Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.380042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:555d4f58cad38ff51867a2df723af09d909977c36e8e916767446cecaef42659

Observation 920dda8b-60d4-4540-ae58-5e2bcc7c36d2 · outbound

This paper cites Inference-time scaling for generalist reward modeling.

Building a Precise Video Language with Human-AI Oversight Inference-time scaling for generalist reward modeling

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.366143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:f3806dff5f811d549e179f655f799539f74de66b4c17930d7b95eb1f5a855716

Observation 0ced2900-fdf0-48fe-bf91-46c0cb81480e · outbound

This paper cites Liu, C.-W.

Building a Precise Video Language with Human-AI Oversight Liu, C.-W

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:46:04.458720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:0c94e1446da71ad907eb1e4a8bd16fa057d102c005f3edfa53db899089556416

Observation aaae4784-7ddb-47e9-b285-17460b58f9c0 · outbound

This paper cites Omni-captioner: Data pipeline, mod- els, and benchmark for omni detailed perception.

Building a Precise Video Language with Human-AI Oversight Omni-captioner: Data pipeline, mod- els, and benchmark for omni detailed perception

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.423071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:15e3bcfd1707b9e7d24a1e7d68971e747683d885776893f0ed0edfe253fddb68

Observation 555fbe6a-3994-4b86-8d91-31d91fd74cb3 · outbound

This paper cites Self-refine: It- erative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594.

Building a Precise Video Language with Human-AI Oversight Self-refine: It- erative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.640879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:34cc19d8a0fa4612f8163cf3603fe6f18ed02252d9e706047288971fd074c6be

Observation b357e0b2-48e4-4405-bdae-47f7827693cb · outbound

This paper cites Native language pro- motes access to visual consciousness.Psychological Science, 29(11):1757–1772.

Building a Precise Video Language with Human-AI Oversight Native language pro- motes access to visual consciousness.Psychological Science, 29(11):1757–1772

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.584261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:ff48d61753b9f837557c13eaa7902fb3a137cad0eab02ba276face611f9a8976

Observation a5806370-9914-4158-9a1a-0b5fc8847c67 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

Building a Precise Video Language with Human-AI Oversight LLM Critics Help Catch LLM Bugs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.557005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:6312866ba1e27fdbee0d51917e0ba4a25e90dc043f82185d05658e597cfd28cb

Observation 36e73833-b17f-49e4-9a20-48e792f14e45 · outbound

This paper cites VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking.

Building a Precise Video Language with Human-AI Oversight VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.484250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:999358e442e77a1182458d39f16182f6daa72032e632558e1c32cf7c18c7202b

Observation 642668d0-5545-4f39-b802-9518ff0dcce8 · outbound

This paper cites Enhancing few- shot vision-language classification with large multimodal model features.

Building a Precise Video Language with Human-AI Oversight Enhancing few- shot vision-language classification with large multimodal model features

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.703165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:aef5cad163ac2265b5d422e4a9c0d8ee9b0bac0f658aa4c7e657b548da63cd43

Observation d0ecbb6b-09db-4cb7-b11d-87f97e9d8521 · outbound

This paper cites Language can shape the perception of oriented objects.

Building a Precise Video Language with Human-AI Oversight Language can shape the perception of oriented objects

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.652518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:a0acdcffe7d8d953cf162b394ebd9fbe4a7e649c98176d57abcea1499dc35b43

Observation c91a6621-de1f-48c0-8eb9-48ac08e1dac3 · outbound

This paper cites Docci: De- scriptions of connected and contrasting images.

Building a Precise Video Language with Human-AI Oversight Docci: De- scriptions of connected and contrasting images

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.730074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:17eb8dd4b25ff829a0141d4fe956f26671517211d6f8a897f5ab8286a40ddb53

Observation cce94aa3-61de-4abd-bcf8-78fd22ff39a1 · outbound

This paper cites GPT-4 Technical Report.

Building a Precise Video Language with Human-AI Oversight GPT-4 Technical Report

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:46:04.494657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:ffcc08384911e20e20fc7b8cf798a3b555fa971eeb1648db70b2e47ad086f95f

Observation fab64950-2ef3-416f-8252-261540142da0 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744.

Building a Precise Video Language with Human-AI Oversight Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.722707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:128df91df8c89ae6653b6b73622d582965c26f4a4ae3ef58b337f207735f59c1

Observation 09d4df88-1d67-471a-bf52-5a46939c7e47 · outbound

This paper cites The Neglected Tails in Vision-Language Models.

Building a Precise Video Language with Human-AI Oversight The Neglected Tails in Vision-Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.553346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:e1db9bb3d3c139ac97d7cc0aab2a0bf3045dad5179f68d8fa220d499a3424749

Observation 6a627dc3-bee6-48ab-98de-cd8fc0408b04 · outbound

This paper cites Direct prefer- ence optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741.

Building a Precise Video Language with Human-AI Oversight Direct prefer- ence optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.637352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:1319291b20de93656b0cf7a620096aa1b32cdcec071bfa3d97bc4565422d5f78

Observation 43039de7-bff4-4315-9231-b565572fde46 · outbound

This paper cites Moodio: Making anyone a professional video studio.

Building a Precise Video Language with Human-AI Oversight Moodio: Making anyone a professional video studio

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.633332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:cce7189281fba81fdd8349ed1357d8cb5e5b1abec04c4e6b4dfb2735cb30a615

Observation abe315a5-7067-4893-ad38-89e05322cd72 · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Building a Precise Video Language with Human-AI Oversight Self-critiquing models for assisting human evaluators

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:25:41.981797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:1a3273c70761925b70fe7c9783ad66714d97cc1630d674b1b14e4416e6b91dce

Observation f404c9b1-1fe6-4cf1-abb8-991087306050 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Building a Precise Video Language with Human-AI Oversight Proximal Policy Optimization Algorithms

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:46:04.374998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:cad2bca2a0298eaac2ea82d6bdf80b9c392f7d7efd43ee9a943ae0bce04ad329

Observation 3638098c-5729-459b-afa4-0a1fd9a4fc43 · outbound

This paper cites Transnet v2: An effective deep network architecture for fast shot transition detection.

Building a Precise Video Language with Human-AI Oversight Transnet v2: An effective deep network architecture for fast shot transition detection

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.784433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:a2eb6daf56d9ad84b1cf0b8ad9857445833829c1d90fa424638b2a9cab61ab2d

Observation 3f1654f2-972b-4eba-9820-27121f0a7708 · outbound

This paper cites Univ of California Press.

Building a Precise Video Language with Human-AI Oversight Univ of California Press

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.719004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:8ec176af3baacea3937397106468dbcfb2c3f15d952203e1689bda45fed561b6

Observation 1fdf3072-c114-448c-b7b4-a0eaf1dd41b5 · outbound

This paper cites Going beyond one-size- fits-all image descriptions to satisfy the information wants of people who are blind or have low vision.

Building a Precise Video Language with Human-AI Oversight Going beyond one-size- fits-all image descriptions to satisfy the information wants of people who are blind or have low vision

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.663455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:6ee313b900b43c7c9cf179bf464bc36ae9c5eb065a3e363c4e8f44c51bc8792a

Observation a1876222-65cd-496b-a2f7-0b2893d96e6d · outbound

This paper cites video-SALMONN 2: Caption-enhanced audio-visual large language models.

Building a Precise Video Language with Human-AI Oversight video-SALMONN 2: Caption-enhanced audio-visual large language models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.532001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:2c9f213ae37e528d9bb91939fa3a08cd9b13385589017c6c912d2df309977708

Observation 7c73f769-4a9d-49d2-aae9-80d07f013e11 · outbound

This paper cites Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence.

Building a Precise Video Language with Human-AI Oversight Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.406418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:a0d3c793c09762ee46ed99b219bdc8c6f02bd07155524d8763153be103032360

Observation e991aa20-8fa8-414d-8276-9af26d05fb64 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Building a Precise Video Language with Human-AI Oversight Wan: Open and Advanced Large-Scale Video Generative Models

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:46:04.437267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:b0d1c0b430c73cfd696d0c48e40c05e473c5c94464a2849a398d1ea5838133fc

Observation 3b876ced-1ea4-4dad-ac02-ba66176a193c · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

Building a Precise Video Language with Human-AI Oversight Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.401243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:d3a7f2ac07c17e70b6359b7df943bb677a30877a18c81554a933eb4d9b29a7cb

Observation 7c85acd8-b1ef-479b-8583-72beb5f4d15a · outbound

This paper cites Spatialvid: A large-scale video dataset with spatial annotations.arXiv preprint arXiv:2509.09676, 2025a.

Building a Precise Video Language with Human-AI Oversight Spatialvid: A large-scale video dataset with spatial annotations.arXiv preprint arXiv:2509.09676, 2025a

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.527564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:a9c937435ccd97e3eb40cf0f8c976b4c3a09c28f773f687e38afa84ef494d75b

Observation 2ac2a1e3-f852-4bfd-a894-90c20f47c1a1 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Building a Precise Video Language with Human-AI Oversight InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:46:04.471178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:afc0f64937fed3abc466ef7ce6cd240f67597a577dea9fd271756f2a751d0be5

Observation 17f1de07-8b54-4c5f-8b14-063ac5ec519a · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Building a Precise Video Language with Human-AI Oversight Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:46:04.500300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:e8a49f6f40566ce11e4776205f5852cc14d7634471b24a915ffa8867c3915cec

Observation 06990ab1-b734-4eab-9a9e-73a7c6db852f · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

Building a Precise Video Language with Human-AI Oversight SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.410651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:18d9a3ec8043a67cbf2fac6306db9bae2384d8fec9dcfd4a113c40e4b148dbc7

Observation d92826e6-801a-40ec-8eb5-c169f641af7b · outbound

This paper cites Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate.

Building a Precise Video Language with Human-AI Oversight Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.441298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:c75108f384f7f7f67733cf5d6c741497c668deb4bc60a661299178bbb30381f6

Observation 67b79438-5040-4fcf-bf0e-f45dcbcbc217 · outbound

This paper cites Perception in Reflection.

Building a Precise Video Language with Human-AI Oversight Perception in Reflection

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.452323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:2fa5b670aca64ecd3a65bcccd41cc579876e7d659a3ae21930825189f74b617c

Observation 5b274cdf-a283-410d-8051-6404ee0f3811 · outbound

This paper cites Russian blues reveal effects of language on color discrimination.Proceedings of the national academy of sciences, 104(19):7780–7785.

Building a Precise Video Language with Human-AI Oversight Russian blues reveal effects of language on color discrimination.Proceedings of the national academy of sciences, 104(19):7780–7785

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.649141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:5030258e57b9ba69e02ac360c76b88676e886cbcf582acf7e9eeef59cfa779f4

Observation 6f3619df-454b-4890-9bdd-b361fc6f3df6 · outbound

This paper cites Tractatus logico-philosophicus.

Building a Precise Video Language with Human-AI Oversight Tractatus logico-philosophicus

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.726613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:b9d4d2b09e83329126346725a1412a8f447422ba0eadaea1478855218951b8d3

Observation f4b0ca41-deb5-4562-8a85-b082ccdc5040 · outbound

This paper cites UGC-VideoCaptioner : An omni ugc video detail caption model and new benchmarks.

Building a Precise Video Language with Human-AI Oversight UGC-VideoCaptioner : An omni ugc video detail caption model and new benchmarks

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.428524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:0ab910cc2ce82af1527c5f3d54128a9f914e559ddd7f675b888523e1ebe20f40

Observation 3cc8b25a-e96a-4bff-9df5-1e7b0918f531 · outbound

This paper cites Visco: Benchmarking 12 fine-grained critique and correction towards self-improvement in visual reasoning.

Building a Precise Video Language with Human-AI Oversight Visco: Benchmarking 12 fine-grained critique and correction towards self-improvement in visual reasoning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.715516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:136161e69dc3c59f2c8171a98a7dcbcda3940a55465c5e2e306505a430de9a49

Observation 986a1ed0-7cae-4c83-9a50-51d8501ef1e2 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Building a Precise Video Language with Human-AI Oversight Msr-vtt: A large video description dataset for bridging video and language

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.644783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:efaa904f6d293deeec26bf4b296c67c295f8573666ff904cdd68f05360a8504c

Observation aae24a09-cf0d-4e5a-b627-1f12c7a66bfe · outbound

This paper cites UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions.

Building a Precise Video Language with Human-AI Oversight UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.545079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:b7154a6f4e870123ebeaca3733452b49c392f4f6aac76c95a5031a5c8e873631

Observation b6976d6d-24d1-4bf3-8451-2e2fd8e87daf · outbound

This paper cites Videochat-r1.

Building a Precise Video Language with Human-AI Oversight Videochat-r1

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.515301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:9e6290b5e9c37db400d1a202dd337f1ed7ffc3e556e59320455a37271818a253

Observation cdc0f7b5-e307-49ef-99c0-3247e225d0e2 · outbound

This paper cites Qwen3 Technical Report.

Building a Precise Video Language with Human-AI Oversight Qwen3 Technical Report

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:46:04.385125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:23aaaa3ed026db5cff64c6b16d2bfc4e4cb895eab5c2d157661bfabc5e03eac1

Observation 41e5262b-1c59-46c4-a7c8-f586ef0ccd47 · outbound

This paper cites Kwai Keye-VL 1.5 Technical Report.

Building a Precise Video Language with Human-AI Oversight Kwai Keye-VL 1.5 Technical Report

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.503586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:ddaba9b0523391181ba2600a9044eed13f6fa016b7fa34818c2b2872b3acf63d

Observation c70b1334-2a5e-4107-8010-c0a8cfe02d17 · outbound

This paper cites Vript: A video is worth thousands of words.Advances in Neural Information Processing Systems, 37:57240–57261.

Building a Precise Video Language with Human-AI Oversight Vript: A video is worth thousands of words.Advances in Neural Information Processing Systems, 37:57240–57261

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.675429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:059a1908318583db65161b8f03d1cf5cc759dedba66a6f817db3e284b1b72fee

Observation 0301289b-53ba-42bb-a78a-9253bb63f7ef · outbound

This paper cites Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models.

Building a Precise Video Language with Human-AI Oversight Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.512494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:f0ff5fbe5d397d2caeff1e7632b4a01464f337e5eba7aa2a566520b11126ec4f

Observation 446f8eea-975c-4e69-9a51-32bf8ecbe9a6 · outbound

This paper cites Omnivinci: Enhancing architecture and data for omni-modal understanding llm.

Building a Precise Video Language with Human-AI Oversight Omnivinci: Enhancing architecture and data for omni-modal understanding llm

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.541934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:a377b8d219a01975d9035d0f915fa540d79ec69c392a0e964f5e16199446644f

Observation a0d3fcfe-5e05-4f60-a8bf-6992b3793198 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Building a Precise Video Language with Human-AI Oversight DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:46:04.550286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:080f06a53ab835a16709f4a5999841fe49b4e84b593ab26adad63df5d994fee8

Observation b9131628-ea5e-4c0b-a9be-d8164d89402e · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

Building a Precise Video Language with Human-AI Oversight Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.671283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:007d51730027592216dd12111da79c9149e0b85f62bcfc85ef9e4a344e9d4c86

Observation 470b0ebe-bbbc-4c34-b282-3575218a909c · outbound

This paper cites Diverse ai feedback for large language model alignment.Transactions of the Association for Computational Linguistics, 13:392–407.

Building a Precise Video Language with Human-AI Oversight Diverse ai feedback for large language model alignment.Transactions of the Association for Computational Linguistics, 13:392–407

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.707689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:aea3d92b00ee556c8525e11d62ad29304063a5d09764ac5aafb0a78545c1fbc9

Observation 58c58590-3bd8-48e6-b25b-a0b1448d2c79 · outbound

This paper cites Self-Generated Critiques Boost Reward Modeling for Language Models.

Building a Precise Video Language with Human-AI Oversight Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.461763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:bb9f318d51f4ed77a7a9ad8daff12b15fc9b06dad2a9d1bf6152ac207597f1ea

Observation 64567a72-0e1e-4f4b-9842-6764df4e8f78 · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

Building a Precise Video Language with Human-AI Oversight Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:01.041344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:7d7f7360c541d6cb823daa181e1fdebb3c546114c08304d21c802cd25991f681

Observation 1240c1d5-88f3-4bf3-b071-43b958528b03 · outbound

This paper cites Critic-v: Vlm critics help catch vlm errors in multimodal reasoning.

Building a Precise Video Language with Human-AI Oversight Critic-v: Vlm critics help catch vlm errors in multimodal reasoning

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.679253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:7e83efdc4016c70a65d52674a1343d4a33c14f22de944a5817f9b5a2f93a8d70

Observation c4bf2f72-1de4-411c-b375-54d1efc47545 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Building a Precise Video Language with Human-AI Oversight Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.359535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:3a827ed94b75b0918f5f3d1959f78943a59c018798904f06646ee259147e9e7f

Observation 567f4ded-f209-43b9-981a-2df0fcec0af6 · outbound

This paper cites VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation.

Building a Precise Video Language with Human-AI Oversight VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.547812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:dbe9d5a241d784136d77a6502a8408152f39144b44575c153c5d58f24b6bd327

Observation dd93ae88-6690-4b52-b09e-642ad95b154e · outbound

This paper cites MM-RLHF: The Next Step Forward in Multimodal LLM Alignment.

Building a Precise Video Language with Human-AI Oversight MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.467982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:04ce54a90cc43b513442b3fa6d3d0347be068d03e71efef5f0b8cca05d6d524c

Observation 59e1817f-572d-48ce-b27c-881f952ef19c · outbound

This paper cites OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference.

Building a Precise Video Language with Human-AI Oversight OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.414865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:c171db753ec70ccfbc8d6989eac20bdf6a75db43fa36c3137783d65c7dc41fae

Observation 642faeba-3ec0-4052-a2a4-953c14d80d2d · outbound

This paper cites OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward.

Building a Precise Video Language with Human-AI Oversight OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.477814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:d34505c78827411d8aa40fc754e2159c2b620d4adb686378fc403661e07db89e

Observation ba0644b0-a956-4225-9c19-6a4fc48eabad · outbound

This paper cites bird’s-eye view.

Building a Precise Video Language with Human-AI Oversight bird’s-eye view

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.682997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:da360cfd370e7710dbcce96daf572389543ccb3bac9eff011712503ed9c8f969

Observation fe87bb66-a9f6-41aa-aae5-bf6fb200a33f · outbound

This paper cites Roughly one sentence may need addition, modification, or deletion.

Building a Precise Video Language with Human-AI Oversight Roughly one sentence may need addition, modification, or deletion

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T10:57:53.691451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:b90cf572a4d3743a7b034c2ae4854e09eb1362a9d68946f755b5ac38440badef

Observation 77b4effc-78f4-44a0-8048-dcce58e74724 · outbound

This paper cites an unresolved cited work.

Building a Precise Video Language with Human-AI Oversight Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-05-23T10:57:53.777110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:48bce1f1f9353b95429485423c910cb25c8e167f37e6d919e86160782138a146

Observation f1bef335-57cd-4d22-b89d-6447275cb736 · outbound

This paper cites an unresolved cited work.

Building a Precise Video Language with Human-AI Oversight Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-05-23T10:57:53.780646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:fa00727ae5ee9413f58fea11c6c0308fdd3572f094d2c7ff81a232505fc0a0f7

Observation 94537efb-c622-4f5c-82cc-d70971f8815c · outbound

This paper cites conveying changes in actions, behaviors, environments, states and attributes of objects, and camera movements between adjacent frames.

Building a Precise Video Language with Human-AI Oversight conveying changes in actions, behaviors, environments, states and attributes of objects, and camera movements between adjacent frames

Reference 95

Resolution
malformed identifier
raw_fallback, observed 2026-05-23T10:57:53.687474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:6f5ba944ab67af6c56162b8601505503318eed7cf2b6afa24c71e2deebc86454

Observation 48330530-902c-4040-b965-af0d8c596305 · outbound

This paper cites an unresolved cited work.

Building a Precise Video Language with Human-AI Oversight Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-05-23T10:57:53.602156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:3bab55ec59a31a12166c5802979ff976bbe61f60809fdf1b3a699978bb00aca8

Observation 8f109f96-40cd-453e-be17-b4d4c78d26c5 · outbound

This paper cites an unresolved cited work.

Building a Precise Video Language with Human-AI Oversight Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-05-23T10:57:53.711652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:c588a0960c25d41a3e348999c23cc933b080acb8e15e852b23934eb2300236db

Observation 57f7b7a2-bcb2-414d-8bb9-bdb056cc7933 · outbound

This paper cites an unresolved cited work.

Building a Precise Video Language with Human-AI Oversight Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-05-23T10:57:53.656327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:73c2aa9a81f9577984892fa18a3beaf84c811fe1ba25e3cdc8b3de502b9d74d6

Observation 5611d61e-26f7-413a-9b26-6e87c13383c8 · outbound

This paper cites an unresolved cited work.

Building a Precise Video Language with Human-AI Oversight Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-05-23T10:57:53.733996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:97b1bf91d3fd167edbe1f5e184fe3ce84084f821e5588c6b03c15683dd49e7a3

Observation 0b17e300-459f-46b3-8ac2-73fb3146fc98 · outbound

This paper cites an unresolved cited work.

Building a Precise Video Language with Human-AI Oversight Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-05-23T10:57:53.787950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:deb18f78654e226e1f67fea105fa8eff5ec20f40046883b615e0cbfe83dec600

Pith citing papers

Observation 0deed888-ebdf-4f65-b6a6-ab5d3db5e0b2 · inbound

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs cites this paper.

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs Building a Precise Video Language with Human-AI Oversight

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-12T02:11:15.531611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:10:27.595446Z digest=sha256:9a8fac587b7a27ab087f34c7f5080b9886b3c07b72bbe521c2f0fb7691def868