Pith. sign in

Paper Citation Record · LEDGER

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision

As of 22 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 2 inbound Pith citation observations for arXiv:2506.09445.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09445 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:54:26.160193Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T19:27:29.843866Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T20:50:17.341265Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 295dd0d9-aa54-4642-ade9-a0719bb25b76 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Lawrence Zitnick, and Devi Parikh

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:34.418059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:19.680017Z digest=sha256:a5041d81528855a3f6f21235476b629cfe9d1dd5f7ee7e808a18c8df974cfd58

Observation dbfd2f08-e292-4686-afd5-270197474f72 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Collecting highly parallel data for paraphrase evaluation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:34.333400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:19.778807Z digest=sha256:8ec2d8fb048940f574132cfca9ada973bc5ca60126c3c5781fb39451f97e1619

Observation 69255cb3-6be0-487e-900b-675897cbf0dc · outbound

This paper cites ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:19.896286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:19.896286Z digest=sha256:5c4cdba5573f7ad4e36c79b4b84378ad40764d2940c6bf515e6481776cac64a8

Observation 90400441-fc8e-40fe-9c8b-3998db443ccb · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers, 2024.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Panda-70m: Captioning 70m videos with multiple cross-modality teachers, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:34.237655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:20.016894Z digest=sha256:6f8c63e60724dad8d1b88d5bfcb0bb4c3e1cf5ebe5915f3c6e8eca156ae323c9

Observation 6185ed6d-c959-4901-950d-7ca6a72578a8 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:20.167152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:20.167152Z digest=sha256:9595b95afe8b1e30fb02a2971df90b971e6473e73cffd740b2e8f859b00607eb

Observation d8b9c55e-899b-4743-a83d-9e8ca37bae95 · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Activitynet: A large-scale video bench- mark for human activity understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:34.076786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:20.270167Z digest=sha256:7a49a22287f5a005abe3dfea35e659f9cc31a051b44c317e9546a1b2e80b71ec

Observation 398bbdde-d18c-4de5-b979-78f40a9cd347 · outbound

This paper cites Video2Commonsense: Generating Commonsense Descriptions to Enrich Video Captioning.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Video2Commonsense: Generating Commonsense Descriptions to Enrich Video Captioning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:54:26.479918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:20.381417Z digest=sha256:dedce16584cd6462dbfb11961afa3690d3596606f1547d0b90a96979ed9d2ca7

Observation f88788a8-7de8-4a93-a3c9-6c1371b03331 · outbound

This paper cites Slowfast networks for video recognition.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Slowfast networks for video recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:33.931610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:20.482759Z digest=sha256:aa1b6358dc50b641c87a926a7edff58ed88644fbb7fee7297bdfd8f8a034a536

Observation 1df26929-12ed-4478-9ee8-ccdb9c334890 · outbound

This paper cites Knowit vqa: Answering knowledge-based ques- tions about videos.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Knowit vqa: Answering knowledge-based ques- tions about videos

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:33.818123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:20.615316Z digest=sha256:d9a0c95feb9ec067274e357464339747e9fbec217c1a044bf126260c6293935b

Observation 8df026c7-8c35-489e-b126-318d61e725f8 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.Int.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Making the v in vqa matter: Elevating the role of image understanding in visual question answering.Int

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:33.693793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:20.738505Z digest=sha256:330f5517baa5907859d55ce97bb066ee4b9bbba858e80f5cda780229b5f8e173

Observation 842d6db9-0eb5-4b53-80ff-7aa67f8fdd71 · outbound

This paper cites Question generation via overgenerating transformations and ranking.DTIC Doc- ument, 2009.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Question generation via overgenerating transformations and ranking.DTIC Doc- ument, 2009

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:33.567100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:20.848575Z digest=sha256:822a34b57d173447ed6b3acc619f27465e13b67b396f3348c57ff77456e61698

Observation 13654fdb-269d-40ae-ba16-66444ae84746 · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Lita: Language instructed temporal-localization assistant

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:33.428533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:20.961326Z digest=sha256:7cf3bc71af842699939599eec3c989daa93db5b30f9e64e06b2c8109a2e043e8

Observation abf0ed84-e72e-40fa-9906-adf2804bde05 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:33.312863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:21.132050Z digest=sha256:79181abbf3dbb2c1837533b407d9327c3585d5a712134a0a424838acb344c7a7

Observation 6a85ecbe-79f8-4846-9e38-94921541bd3d · outbound

This paper cites Video question answering with spatio-temporal reasoning.IJCV, 2019.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Video question answering with spatio-temporal reasoning.IJCV, 2019

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:33.164247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:21.220455Z digest=sha256:c3138e868731ac1d04c272ac390650fc92a717470b41c01fd75f2f261ebe93fd

Observation 17d00146-0b54-4bb2-a0c2-c15f28ee0375 · outbound

This paper cites Video question answering with spatio-temporal reasoning.Int.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Video question answering with spatio-temporal reasoning.Int

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:33.052950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:21.342540Z digest=sha256:78f5e66e0ec400733bd1b5f5ef171b8826734a35b1d916b6ef72260192167199

Observation 396bfc1d-bd7f-46a5-9bdc-f94a35de442d · outbound

This paper cites an unresolved cited work.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:54:32.896787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:21.453425Z digest=sha256:1db0159cb030702849710943852909f8d0cdb0dbc90c11b553709f17704b8aca

Observation 989df242-7993-46d1-956d-025e065fe126 · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:21.541697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:21.541697Z digest=sha256:77e79493a53bc92d80f31cdc285cef3078cf0ecaae5113687a92405a7b17ad4c

Observation a947470e-043f-4e9f-802f-60f82cec8818 · outbound

This paper cites Slow-fast auditory streams for audio recogni- tion.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Slow-fast auditory streams for audio recogni- tion

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:32.692985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:21.641461Z digest=sha256:620e354b2bb8af820688353256472beb97e26319da0197e079396a88ae1b3178

Observation 317d255f-cf6f-4107-98aa-938d7e1a3fd4 · outbound

This paper cites Open-vocabulary video question answering: A new benchmark for evaluating the generaliz- ability of video question answering models.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Open-vocabulary video question answering: A new benchmark for evaluating the generaliz- ability of video question answering models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:32.496952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:21.735840Z digest=sha256:2efdb325724e3093277d55fa07ebb453ce3ddc22ec12157705c130599d30f934

Observation 27941dcf-fb06-4b19-9303-447e49a28a46 · outbound

This paper cites Dense-captioning events in videos,.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Dense-captioning events in videos,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:21.887761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:21.887761Z digest=sha256:6f390c9d498cc56c9e4d06ecbdb6638b9b5bccde5006a954e437a9a33c064535

Observation 6c7564c8-9381-41f9-a311-96fe165a6d9b · outbound

This paper cites Meteor: an automatic met- ric for mt evaluation with high levels of correlation with hu- man judgments.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Meteor: an automatic met- ric for mt evaluation with high levels of correlation with hu- man judgments

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:32.293998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:22.009516Z digest=sha256:78023e8c2771f1cb3f580e5dcfb6277a32d14e1b02a9e11f0037d406cb834e7f

Observation 0a497587-db97-46b0-9033-2e8a78d6e4ea · outbound

This paper cites an unresolved cited work.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:54:32.079746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:22.113233Z digest=sha256:9a65dbcc83c598ff3bffc81173ac439fdb4234b17af8a120b53887406ebd1355

Observation 7a3d4347-22c3-45f6-b5e3-5d413ca918bc · outbound

This paper cites Berg, and Mohit Bansal.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Berg, and Mohit Bansal

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:31.878707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:22.180864Z digest=sha256:47d6234b60daf6b851c46c611e62aad477e589e62142c548be33a67d729805bd

Observation dafab7fb-eddf-4b09-a8f5-610c60bdfd5c · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision VideoChat: Chat-Centric Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:22.296324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:22.296324Z digest=sha256:a307425fd4384d8882b4952d54d26386313c14b1da43de13f91cb0c2d03de981

Observation 4e198a4e-1809-45a8-b7bf-6613d7c96de8 · outbound

This paper cites Equivariant and invariant grounding for video question an- swering.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Equivariant and invariant grounding for video question an- swering

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:31.721168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:22.418442Z digest=sha256:9796e639c15f2c6b8e40bcd69f8c6eea35ed9286ffddafd69fd6973368003935

Observation 75daa1ab-0f59-4c5e-956e-4b331188f5b0 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:22.532998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:22.532998Z digest=sha256:2b28bfca080a09d1e8d3f6dd67b4b78f3f1a56f4ae1bf73505aa8d3f8891d6aa

Observation 9ebd6a51-b375-459c-9fd8-2d3e7399286d · outbound

This paper cites Visual instruction tuning, 2023.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Visual instruction tuning, 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:31.584089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:22.641056Z digest=sha256:4b92acead7d7093bea6bc767ec671b8b6d4aecfc65097d69a9d6fa7bc20a06f7

Observation d6b62272-aa05-49f2-ab5c-7b48a399a281 · outbound

This paper cites Decoupled Weight Decay Regularization.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Decoupled Weight Decay Regularization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:22.723878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:22.723878Z digest=sha256:6e4e3c05e7e1dc184fb74bc7b8ceade8745ceea43ee1a23e39406036902a5ffd

Observation 19ff832e-c4df-4b85-97a6-f8774ebcc7b1 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:31.384271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:22.833731Z digest=sha256:fec92b4da45957d081885ce3ebb4af5b09885a861ac626a05e5f8b41b50e9553

Observation 4f627028-808f-422f-810c-eca3d299bd70 · outbound

This paper cites The surprising effectiveness of multimodal large language models for video moment retrieval, 2024.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision The surprising effectiveness of multimodal large language models for video moment retrieval, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:31.189409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:22.955323Z digest=sha256:6d52a01948acba6e67c1a2eaa090b1887e75d26b55adb87188228303be3bafe8

Observation 64856e89-4e94-41d9-8ee4-9d815b6b4fc7 · outbound

This paper cites Exploration of masked and causal language mod- elling for text generation, 2024.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Exploration of masked and causal language mod- elling for text generation, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:31.013217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:23.062143Z digest=sha256:75b54847cce201893acee9a62c1bbe16355ad534d9f55ed142caa1e76806b030

Observation 309208f7-276d-4872-aad0-e50276115ceb · outbound

This paper cites an unresolved cited work.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:54:30.811022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:23.139081Z digest=sha256:e048eed5187ed9b34d06359780f374cf01834b17682ba0de7d5bc7c28623c887

Observation 2b130726-f460-4ba1-a023-f63a332b68aa · outbound

This paper cites Gpt-4 technical report, 2024.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Gpt-4 technical report, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:30.641711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:23.238338Z digest=sha256:f6f0d489a33ae3fb34d3f75f5e543320a12cd21fc23cb30a6f687899b7ada1ab

Observation dcc712b2-7a5c-4c81-b8c9-28be9b8543f1 · outbound

This paper cites Streaming long video understanding with large language models.Advances in Neu- ral Information Processing Systems, 37:119336–119360,.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Streaming long video understanding with large language models.Advances in Neu- ral Information Processing Systems, 37:119336–119360,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:23.322717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:23.322717Z digest=sha256:eab3ed80bffc46fa535da1c94bb4d1215f541b1f170d3a3679e39297d895221c

Observation e2332bd5-d03b-4bf1-8ac9-518fed5ec492 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Learning transferable visual models from natural language supervision, 2021

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:30.447637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:23.450152Z digest=sha256:3fa93d43d91f80924332e80ad03d1f448512773cfd3dbdf8476e357e0a5c8315

Observation 269d11b5-0262-487a-a750-ea1ff097651f · outbound

This paper cites Designing network design spaces.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Designing network design spaces

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:30.264765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:23.530831Z digest=sha256:1b38ed428203d039f102f60946e7fd355d9aed717537512f81051c45aab467ca

Observation bd97e098-8a9c-4d4d-b5ed-441402fe2cb8 · outbound

This paper cites Concept- net 5.5: an open multilingual graph of general knowledge.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Concept- net 5.5: an open multilingual graph of general knowledge

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:30.068310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:23.630808Z digest=sha256:315e353fe3450ca6032c6421f1772ba6e8b2ed37b9ce42903ba4369fe5219738

Observation 2283dc94-5ee1-4a52-abc1-d9a008a36dbf · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Vipergpt: Visual inference via python execution for reasoning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:29.892922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:23.765043Z digest=sha256:01dfe4ccf5925d6605666ed971c0f832a4b3a4fede95eebc0d38bd6b9df07d3c

Observation 3e38dfc3-968c-4ec3-a0e9-8650d27c77b9 · outbound

This paper cites Movieqa: Understanding stories in movies through question- answering.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Movieqa: Understanding stories in movies through question- answering

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:29.731567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:23.869719Z digest=sha256:63e040795c9b79457027c7c8668406054e9e97071d8f0568ef3fd3ce4ad7c08e

Observation 63986425-5191-4e5d-9ebc-9770e7241a2a · outbound

This paper cites Gemini: A family of highly capable multi- modal models, 2024.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Gemini: A family of highly capable multi- modal models, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:29.536874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:24.039039Z digest=sha256:84875481c0e7b050d52727365733a9e77776706ce30b7ce866002b8a79456b7d

Observation 65feca4f-fa28-4416-864c-9f761a81bdd5 · outbound

This paper cites Llama: Open and efficient foundation lan- guage models, 2023.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Llama: Open and efficient foundation lan- guage models, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:29.394008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:24.198578Z digest=sha256:58bd9fef93d444512ae9a4b95a501c3201524e46d06ab717c5679290bc1aa5d9

Observation dedf4efa-c102-4797-b018-9d9b7f70afe7 · outbound

This paper cites Grounded-videollm: Sharpening fine-grained tem- poral grounding in video large language models, 2024.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Grounded-videollm: Sharpening fine-grained tem- poral grounding in video large language models, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:29.231800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:24.345618Z digest=sha256:36decffb48c288afc1b857c90cf5eaaece197dd4b3d63061f221016249524904

Observation 6441a79e-3fa6-448a-ac90-99b4a81b0d0b · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:24.446133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:24.446133Z digest=sha256:9963c0b014a1684b94b2a77ef464ecbe94c86e8e60579490ab05e30a9eab55b9

Observation 392bf377-9646-4ff1-87ec-15f8401b57b1 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:24.495885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:24.495885Z digest=sha256:73a0f161e80ff83b227be95993395af74d3b41c630c5e444b697d0958826dc4d

Observation b0beb357-0bc5-4803-9b93-2e37a27c3837 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:24.595020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:24.595020Z digest=sha256:cf03c612af2a7baacb90e95efaacececed1c126e2a42637e416242ec8199053d

Observation 1e07e1e8-ef33-45b4-90a0-73e24488a849 · outbound

This paper cites Number it: Temporal grounding videos like flipping manga,.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Number it: Temporal grounding videos like flipping manga,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:29.062288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:24.666423Z digest=sha256:eab68ef02ed8c2979103146ace60d01e1c0eb085acbc6d25dcc18f5bbb84c51b

Observation 1e498db1-61bb-4662-9211-c4c727c6e716 · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Audiovisual SlowFast Networks for Video Recognition

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:24.733808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:24.733808Z digest=sha256:c40a2a9b823b221eadc1e4c9b529ae09b6a00003fede3fc210bb3dda8823cb1f

Observation 94bc6eb1-5e3c-4cc2-9075-a3ab29bb0e25 · outbound

This paper cites Next-qa:next phase of question-answering to explaining tem- poral actions, 2021.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Next-qa:next phase of question-answering to explaining tem- poral actions, 2021

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:28.877640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:24.816412Z digest=sha256:82a2d231b7b1319e32602ccd8a603d31f817df28d800b441b66d2c6f7599402f

Observation 07a5d2d2-f1e6-42d7-9258-5a7156cc420c · outbound

This paper cites Videoqa in the era of llms: An empirical study, 2024.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Videoqa in the era of llms: An empirical study, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:28.708921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:24.908233Z digest=sha256:b748deaa54caca44e014b732510c2e28f55fee793f38326cc47a659e251c9525

Observation 7da56493-0950-461d-b803-f649a5524bdd · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Can i trust your answer? visually grounded video question answering

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:28.527173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:24.981743Z digest=sha256:544f830555d702e94911146492e72afc0fec7be7681fe020660c6ba2f018b9dc

Observation ac06e2f6-e250-43bf-89dd-d33edce741c7 · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:28.342463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:25.054818Z digest=sha256:394c8e96186c837777b9b59a04a3e78a0183d9a0e8fd77476b721a5c3c72c7bd

Observation 1cc3b771-f2bf-4bc6-8df1-6f1759cfc0c1 · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:28.181719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:25.106092Z digest=sha256:e60db2e59ba494136677d642a5f85006d7b9758bd08554f7b8164b05a1e3a8db

Observation 3d3537fb-8c5e-40a6-aa05-71d250a2ec1c · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Zero-shot video question answering via frozen bidirectional language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:28.007775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:25.206524Z digest=sha256:73210447e927cdb1ee46f3b03eb9f688b383fa352b21ec2fb56649f2a28b3acc

Observation 67ab510d-630e-44c0-b607-603febc1636d · outbound

This paper cites Tubedetr: Spatio-temporal video ground- ing with transformers.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Tubedetr: Spatio-temporal video ground- ing with transformers

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:27.746855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:25.311819Z digest=sha256:26d7b945816c9d35365ed166a2c0460eeb1750bf5c2ca35c025a97eb9ad484be

Observation 87a72b27-beb8-4b75-98a7-8558e7049f72 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Self-chained image-language model for video localization and question answering

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:27.568804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:25.399743Z digest=sha256:9490070124451451470e461fa481306504e177300374ccbc2cc4686f3cf57fcf

Observation 0e2d1948-2c99-4c33-8452-57918793584d · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:25.466900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:25.466900Z digest=sha256:b4f4d3b4306fac91de4f17fc84c72505d1d93f264af8814a515b5778a634110e

Observation 04154e34-90dd-4a1d-91ae-9313f0881b31 · outbound

This paper cites A sim- ple llm framework for long-range video question-answering,.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision A sim- ple llm framework for long-range video question-answering,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:25.560743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:25.560743Z digest=sha256:f7ed4f8b9c9618239de6025e6475fb0a3b34ffee9d685363fed7a32cc8708521

Observation 4328a720-fb21-4fd0-8d8e-c73098de5137 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:25.661857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:25.661857Z digest=sha256:307fba43cfbd0ccfb91f1eb181b94fb4bd32cc8778688ac2908b7c64f0fd2257

Observation 3cd26946-7f05-47c6-aecf-50432a7d58eb · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:25.743338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:25.743338Z digest=sha256:4efe985ee898cdf15850b6347c57a546ff8b61ecddd84b246ee58c6a48ce7fd7

Observation 45455f66-1722-427d-af2a-559d5503c96e · outbound

This paper cites Video question answering: Datasets, algorithms and challenges, 2022.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Video question answering: Datasets, algorithms and challenges, 2022

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:27.383451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:25.846782Z digest=sha256:89c9578e6ef6dd878c0ccafd2a3777226424e6ddf0b825a12aa717f5a08c90cd

Observation 82184906-7451-42ee-8420-da441870c981 · outbound

This paper cites Training Details The additional settings we use, including hyperparameters and implementation details, are shown in Tab.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Training Details The additional settings we use, including hyperparameters and implementation details, are shown in Tab

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:27.194726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:25.914216Z digest=sha256:7750fbd1be894dde0c2072f549555f5c948a70a525873fe726fdb4d346969097

Observation d0128375-3087-4256-98ce-522db323b035 · outbound

This paper cites Multi-scale vision-language connector In this section, we conduct ablations on the MS-VLC and analyze the performance under various conditions.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Multi-scale vision-language connector In this section, we conduct ablations on the MS-VLC and analyze the performance under various conditions

Reference 62

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T04:54:27.002436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:25.979743Z digest=sha256:f5a1e323a11ed0439535e9f2d3b2d2eb5c2c4c3c38bca82c7fd8de8d768e4dbf

Observation 9deab42f-f34e-431b-bc13-70d52a8a980a · outbound

This paper cites We achieve a METEOR score of 0.498 on the ActivityNet-QA dataset.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision We achieve a METEOR score of 0.498 on the ActivityNet-QA dataset

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:26.838113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:26.059057Z digest=sha256:207a660488f936fe12b6bc55244c113af8b227a92152413c24022573109e5921

Observation 7b8ccd30-91db-4a43-8fd3-0ba1f9bf4de2 · outbound

This paper cites 1 and the ActivityNet dataset in Fig.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision 1 and the ActivityNet dataset in Fig

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:54:26.675471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:54:26.160193Z digest=sha256:b71147d0eb0af2c23901717759527a6f2ae27cb9fe305b2c39e728a866e81230

Pith citing papers

Observation 6d894c32-3ecb-45dd-a86a-c6762ad69e19 · inbound

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking cites this paper.

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:50:17.342680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T20:48:44.933542Z digest=sha256:79204f79d62f0040a1bbb56dbe872223496025729b5a9e2a3cfca5ac426ab78e

Observation 23dcea82-4846-49e1-aa59-a48588907cf1 · inbound

Towards Temporal Compositional Reasoning in Long-Form Sports Videos cites this paper.

Towards Temporal Compositional Reasoning in Long-Form Sports Videos TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T19:27:29.843866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:27:29.843866Z digest=sha256:e003e20c6382358b813fff6dde0e00f52e1301105c0c6851efe6bb244438333d