Pith. sign in

Paper Citation Record · LEDGER

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions

As of 21 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2504.15918.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15918 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:20:36.814781Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:55:40.858859Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T19:55:41.494044Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy56
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 374a4168-67ea-4a1e-869f-e87f082f25bd · outbound

This paper cites Let the llms talk: Simulating human-to-human conversational qa via zero-shot llm-to-llm interactions.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Let the llms talk: Simulating human-to-human conversational qa via zero-shot llm-to-llm interactions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.748460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.527975Z digest=sha256:49eb21eff38199acc2d34eafd6b289e7205b38f93fb227424a2ab4d4d6fe5697

Observation 22db37f0-4f1e-4017-8f82-570e4c6272eb · outbound

This paper cites Gpt-3-driven pedagogical agents to train chil- dren’s curious question-asking skills.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Gpt-3-driven pedagogical agents to train chil- dren’s curious question-asking skills

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.733529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.532976Z digest=sha256:0bbbb9ca6f43e6031c7c99f72fc7dc0a211ee2ba0c9131d05ed3b1fa98039ed7

Observation 7f313568-e6ce-4351-9666-bc3fd9f7be53 · outbound

This paper cites Synthetic dialogue dataset generation using LLM agents.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Synthetic dialogue dataset generation using LLM agents

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.718277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.542625Z digest=sha256:8b8d14c3beff43ae68fd45defb15cf0d80268effee11def5e695671bee255203

Observation c53b39f3-0777-4c6c-ae17-307129651210 · outbound

This paper cites Self-RAG: Learning to retrieve, gener- ate, and critique through self-reflection.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Self-RAG: Learning to retrieve, gener- ate, and critique through self-reflection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.704823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.546830Z digest=sha256:7fb95cfa8090389a502c0ac89a9d0d1e5c081ec22f75eb87c20b983f698e3343

Observation b246b9ed-ed5c-4c4f-8c0b-7542e5664981 · outbound

This paper cites Vindlu: A recipe for ef- fective video-and-language pretraining.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Vindlu: A recipe for ef- fective video-and-language pretraining

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.691689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.551451Z digest=sha256:25333ba94b3164a6ecd25881cc04773e5c25227c4c2afcea94247620ceb21b74

Observation f78488ca-a08c-4610-860a-fa13bf3b7da9 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.677740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.555912Z digest=sha256:aee47baae79c84ebe76da82c0b73a53e2af01bc1d13bdeead18c2457a7ff711e

Observation d4e2d27b-0db0-4ee8-8de9-4fe33dd54639 · outbound

This paper cites Learning musical representations for music performance question an- swering.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Learning musical representations for music performance question an- swering

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.664594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.560944Z digest=sha256:84bb982620c2a75b2f2bdbd6a08d49e212a0699b078ba93803ed402331be1e39

Observation 63b297e1-fe3f-4019-a850-75c978a95f2b · outbound

This paper cites From local to global: A graph rag approach to query-focused sum- marization, 2025.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions From local to global: A graph rag approach to query-focused sum- marization, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:20:36.565314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:20:36.565314Z digest=sha256:85602693dddeacecfc885eeee8fa05eeafafcc6e7c672eecccd650901286db80

Observation 4c94dc9e-240e-4906-b76f-5fbf51003fa9 · outbound

This paper cites Hi- erarchical modeling for task recognition and action seg- mentation in weakly-labeled instructional videos.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Hi- erarchical modeling for task recognition and action seg- mentation in weakly-labeled instructional videos

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.642522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.569190Z digest=sha256:a76d837ad0df3a7387152f95278e5ee0cb360545cdfb71ed11b8743ca59a06cb

Observation 748b7648-ecff-480e-84c0-7c4e87d9f227 · outbound

This paper cites The llama 3 herd of models, 2024.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions The llama 3 herd of models, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.629741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.573141Z digest=sha256:89442e7376c714dc1413f71496804ee1c32d4d52de677fe1591a70de1845058b

Observation 5347dc67-045d-4d09-ba8c-e7ddf45deb7e · outbound

This paper cites A dataset for medical instructional video classification and question answering.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions A dataset for medical instructional video classification and question answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.616947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.577318Z digest=sha256:5e5c456552950f9c8e6248f7b22afcbabb31a28f7a7af3bcd364ff73d32ddfb9

Observation 91416d4b-a8d6-475e-8187-e9189351de46 · outbound

This paper cites Does prompt formatting have any impact on llm performance?, 2024.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Does prompt formatting have any impact on llm performance?, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.603416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.581178Z digest=sha256:b30f6cd4d1806bf92ef0a6b3b3575403b07f6e90a7c2402e3d5e9d031727d646

Observation 00828faa-55ba-4925-bc58-a08ecc24d507 · outbound

This paper cites Active retrieval augmented generation.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Active retrieval augmented generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.590466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.584984Z digest=sha256:38f5b8ea4e0b25a2a5241c369a10ccaf0675b418cf8533330ba9ac31dbb854fa

Observation c14ac5e1-4d67-4e30-b8dc-c7d996676847 · outbound

This paper cites Instruction-tuned language models are better knowledge learners.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Instruction-tuned language models are better knowledge learners

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.577242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.589303Z digest=sha256:205d285480b28708ebbb31191fb22f99be1cea99d71ee18390d96fe60d686a77

Observation f9916b4b-48d8-43bd-9574-c4a4533f7681 · outbound

This paper cites Prospector: Improving llm agents with self-asking and trajectory ranking.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Prospector: Improving llm agents with self-asking and trajectory ranking

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.564547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.593788Z digest=sha256:a65a03c6d5f14c3d48732068529a1ed56741b57b30c5194759f45638813856b4

Observation 43f47a9b-20e8-41f2-993d-20afe710b60f · outbound

This paper cites Qube: Question-based belief enhancement for agen- tic llm reasoning.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Qube: Question-based belief enhancement for agen- tic llm reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.551487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.597956Z digest=sha256:55cb65697bdf1f9b03dd741b1591a120a6494d177536c1a35a27fab1c9b06c8d

Observation 3804e5d9-186a-4d2c-9946-050f6173c8a0 · outbound

This paper cites Dossier at medvidqa 2022: Text-based approaches to medical video answer localization problem.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Dossier at medvidqa 2022: Text-based approaches to medical video answer localization problem

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.538923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.602048Z digest=sha256:8471170bb39fdc697e527b1c7b4a2f719e6163b78f3028c2a1efd30b8d454256

Observation 50443342-cb3f-4e24-8f78-9a414171e901 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.525180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.606024Z digest=sha256:795c246c928004601289befee7903eaddff47e23380505451fa70bdeccb93a63

Observation f6a2fa42-a3d2-4c4b-9428-594b9f8ea07b · outbound

This paper cites Overview of the nlpcc 2023 shared task: Chinese medical instructional video question answering.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Overview of the nlpcc 2023 shared task: Chinese medical instructional video question answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.512438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.610261Z digest=sha256:2fa848ee796d3a20d339ecde7467e6cbdb97aa400dcb438d9115569eb932a083

Observation ccc36e41-004d-4ccf-8243-8bafc8025ea6 · outbound

This paper cites Learn- ing to locate visual answer in video corpus using question.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Learn- ing to locate visual answer in video corpus using question

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.500002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.614436Z digest=sha256:c566df1b960d286da3ab8af9f49d2ff4644077549395ac8db9e2ca0130ae524a

Observation 8db75a8c-6d12-4255-a964-3871bf5c37f8 · outbound

This paper cites Overview of the nlpcc 2024 shared task 7: Multi-lingual medical instructional video question answering.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Overview of the nlpcc 2024 shared task 7: Multi-lingual medical instructional video question answering

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.488240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.618454Z digest=sha256:2997af87d51564dffbda936dc69995ccc265563a11b50534e40bd39e09c88a19

Observation e8e4e8cb-90cd-4e4d-b011-2a760b0224ea · outbound

This paper cites LLaV A-neXT- interleave: Tackling multi-image, video, and 3d in large mul- timodal models.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions LLaV A-neXT- interleave: Tackling multi-image, video, and 3d in large mul- timodal models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.475490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.622545Z digest=sha256:84f3ac0b245dce72c37a3e99d1857a904362a93704929031c5faebd786cbdd9f

Observation a61b2af5-83d9-4e47-989b-7e3688b0c3c7 · outbound

This paper cites Hello again! llm-powered per- sonalized agent for long-term dialogue.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Hello again! llm-powered per- sonalized agent for long-term dialogue

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.462916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.626573Z digest=sha256:e112956c35674d1be5bc04e5bc67a44b9247907f959f855ef1c6f011d64cdd3d

Observation 62560dd2-b484-42d0-9cef-abe1e29eabfc · outbound

This paper cites Miller, Sumit Chopra, Marc’Aurelio Ranzato, and Jason Weston.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Miller, Sumit Chopra, Marc’Aurelio Ranzato, and Jason Weston

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.449870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.630524Z digest=sha256:f403a1471505787349b52bf0ab774f2f563d6b3c15f3b4d56b84e4da21951bf0

Observation f0512055-eb1c-454a-bbe1-d8ec78f8413f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.437412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.634348Z digest=sha256:599d41a13f7947445262e957a80b7d7a8492af36dd39cbac0866d48586829f8e

Observation ad59e12c-3519-4dd5-bf81-900b4449e4e0 · outbound

This paper cites Mediq: Question-asking llms and a benchmark for reli- able interactive clinical reasoning.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Mediq: Question-asking llms and a benchmark for reli- able interactive clinical reasoning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.425026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.638189Z digest=sha256:8ead4a2f48ee35ad04b5dd7c60d5f6c5ccddb2cbc7581e9057e37a412daa34e8

Observation 23dd8833-7f3c-4db8-8d4d-290c4fdf9bde · outbound

This paper cites Towards visual-prompt temporal answer grounding in instructional video.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Towards visual-prompt temporal answer grounding in instructional video

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.412467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.642097Z digest=sha256:ab495577eacc9f8c1847452a4ca4a20867a8d6bc06fa05a3a18d95429468c0a9

Observation 0684a7af-088e-493d-8ccd-f8bfc04efeb3 · outbound

This paper cites Fuzzy multimodal graph reasoning for human-centric instructional video grounding.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Fuzzy multimodal graph reasoning for human-centric instructional video grounding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.400182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.646068Z digest=sha256:94879afe475893d2c9b91707c099e0631e9c5c08ba72a2f20992944674c15c0c

Observation 3d07cffb-231a-4d5c-8da1-90f06ecfad9a · outbound

This paper cites Video-LLaV A: Learning united visual rep- resentation by alignment before projection.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Video-LLaV A: Learning united visual rep- resentation by alignment before projection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.388062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.650188Z digest=sha256:9f693e24a2b3a7867cd3bb016574e7b7a824cdb045dd5a80395e863e61cdaa4a

Observation e55680a5-bdd6-4ea7-864f-f7bc762f4cc9 · outbound

This paper cites MM-VID: Advancing Video Understanding with GPT-4V(ision).

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:20:36.654208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:20:36.654208Z digest=sha256:c79cb47cfe50d7b86aa6e50e133cba67b993eb947716c42fcb8a4c8dbd2009c8

Observation d94ee299-635c-4b12-8f37-eedecc69b9f3 · outbound

This paper cites Integrating video re- trieval and moment detection in a unified corpus for video question answering.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Integrating video re- trieval and moment detection in a unified corpus for video question answering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.375645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.658730Z digest=sha256:db691b1dd43a3f13432b679eb1ef9887f463d34908434de5045ab9cf4d6a2c9d

Observation e2c386cd-d852-46f2-ba3c-93d63162317e · outbound

This paper cites Query rewriting in retrieval-augmented large language models.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Query rewriting in retrieval-augmented large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.363395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.662875Z digest=sha256:d18775e06b79e8bc1fb7f449a37728db2b1be3352a15d5a66fa7ee5cd3f57ba1

Observation f864ce1e-b94d-4a23-a846-8d0c520141c1 · outbound

This paper cites Query rewriting in retrieval-augmented large language models.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Query rewriting in retrieval-augmented large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.350743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.666715Z digest=sha256:ef40a7c861bda1feb375ee86c610ded1817e2b5061ea788af98c0dc1272021b6

Observation d21d769e-e83f-422b-9eed-8170612d5f60 · outbound

This paper cites Query rewriting in retrieval-augmented large language models.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Query rewriting in retrieval-augmented large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.338387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.670940Z digest=sha256:d16554187e40b2de8a6e9fff24011800f613c96abb5dfde4a0a323b7df9cf53f

Observation 7895c9db-c17c-458a-a711-13867f0a5104 · outbound

This paper cites Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Video-ChatGPT: Towards detailed video un- derstanding via large vision and language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.325734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.675181Z digest=sha256:ac006e22fa207a9bf9f6facc1b56bfb681dbe376bebe231825c423d6e518655f

Observation e7be9eef-d2c8-4355-ae84-f2b2a706849b · outbound

This paper cites Learn- ing to retrieve videos by asking questions.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Learn- ing to retrieve videos by asking questions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.312778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.679098Z digest=sha256:9084a538b76758bdc968b5f49e70e2f8a115b0127dcd631e0d05b2f66fef06de

Observation ff58acc1-ae80-4ec4-b399-47bb3a5bd7b7 · outbound

This paper cites Evaluating very long-term conversational memory of LLM agents.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Evaluating very long-term conversational memory of LLM agents

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.299656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.682949Z digest=sha256:ec03a8c7edc3272ee2f005cc18dbfdc3c275345a4c8dd477b5c945c33e994df6

Observation 81d1e1cd-7d60-4336-8f28-5f3fd932fc83 · outbound

This paper cites Chatvtg: Video temporal grounding via chat with video dialogue large language models.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Chatvtg: Video temporal grounding via chat with video dialogue large language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.285328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.686998Z digest=sha256:338dd750b954aa1b1ceb711f32ef803b8bfbbcfbec25dbfa6a1da5b2e6e4d338

Observation cdc2c8e6-59ca-4b12-a0f8-2902c406c899 · outbound

This paper cites Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.271985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.690881Z digest=sha256:87569f5ac4ee76e55d9e6e92db8e13c12ced9b34d4d83b69ae0f8d576de2c855

Observation e6724848-406a-4423-93f6-0613a5a69aad · outbound

This paper cites Hybridrag: Integrat- ing knowledge graphs and vector retrieval augmented gen- eration for efficient information extraction.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Hybridrag: Integrat- ing knowledge graphs and vector retrieval augmented gen- eration for efficient information extraction

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.259409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.694767Z digest=sha256:8e80975f764adb07cfb70e2ff0ba16e684e07c85318520bb5c8f1e14179d36cb

Observation ad4036a8-04b3-4c01-adfc-a0350c22b212 · outbound

This paper cites The art of creative inquiry—from question asking to prompt engi- neering.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions The art of creative inquiry—from question asking to prompt engi- neering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.245849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.699028Z digest=sha256:e26c014f269b1495172a1c34fc6a6d8b79e347cf322081ee9c9e499824d8d58c

Observation 87bb036c-d5d1-4f8b-ad84-4d30a12d02dc · outbound

This paper cites Rewritelm: An instruction-tuned large language model for text rewriting.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Rewritelm: An instruction-tuned large language model for text rewriting

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.232739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.702967Z digest=sha256:048b25d198c76dc386e0fff16c78283f78048049848dd69ab018af19ba4e2307

Observation d51c9aa3-e9ab-49ba-b99d-b6a2f21bdeed · outbound

This paper cites Toward expert-level med- ical question answering with large language models.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Toward expert-level med- ical question answering with large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:20:36.707141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:20:36.707141Z digest=sha256:1b4c8eba71e5e281abc349f2cedc43cdf3c8671a979b587ab5ad921f837fca5e

Observation 7b33975c-4740-4086-82f7-74fd6544341c · outbound

This paper cites R-Bot: An LLM-based Query Rewrite System.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions R-Bot: An LLM-based Query Rewrite System

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:20:36.711018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:20:36.711018Z digest=sha256:fece5fdcb070b22cd1ed08623457800fde254688239420dd0aec2ade0167c344

Observation 2db706f0-c2c7-41f5-86b4-7cf04f2d15e1 · outbound

This paper cites Rag-adapter: A plug-and-play rag- enhanced framework for long video understanding, 2025.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Rag-adapter: A plug-and-play rag- enhanced framework for long video understanding, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.210580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.715172Z digest=sha256:cacd499efb623bfb64c38dd95dadb7c56b5f53438553fa069a309b22e0b06d53

Observation b05d4636-12f1-4763-b6a2-d002bf7ea955 · outbound

This paper cites Grounded-videollm: Sharpening fine-grained tem- poral grounding in video large language models, 2024.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Grounded-videollm: Sharpening fine-grained tem- poral grounding in video large language models, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.196420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.719038Z digest=sha256:aac13e50c6c95f0fe30cf38841f5b2ec6bd728379f670cadc92b0b1dab22dba4

Observation 6b139709-f645-429f-b2cc-a1f145559caf · outbound

This paper cites Llama3-8b-chinese-chat (revision 6622a23), 2024.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Llama3-8b-chinese-chat (revision 6622a23), 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.183283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.722799Z digest=sha256:106ae66bdb76e5967f774400396cb25d7e22803aa364a95613dbecefa27f0a2a

Observation 0d46acfc-a995-4a4b-8370-79052ca73c57 · outbound

This paper cites Smarter, better, faster, longer: A modern bidirectional en- coder for fast, memory efficient, and long context finetuning and inference, 2024.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Smarter, better, faster, longer: A modern bidirectional en- coder for fast, memory efficient, and long context finetuning and inference, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.170641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.726563Z digest=sha256:a6acab796359c67e7d75cd4bd7e96d2ee66acb8e6bafec2e2400b5cf3df74b86

Observation 5ea07396-0bdb-444e-a5d7-b488446568f3 · outbound

This paper cites Visual answer localization with cross-modal mutual knowledge transfer.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Visual answer localization with cross-modal mutual knowledge transfer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.158153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.730623Z digest=sha256:7b861b8baed5d596ea3a50cdebe6c8ba55965eaebd5ed74465c3a63208de62fc

Observation 427beb43-e6d2-4524-b9b2-dca724669d09 · outbound

This paper cites To find where you talk: Temporal sentence localization in video with at- tention based location regression.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions To find where you talk: Temporal sentence localization in video with at- tention based location regression

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.145319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.734575Z digest=sha256:d4a66f5f33b1c9f0ce4810409f04b655ff30a35ad46923a4fbaad1e6f0edf9ed

Observation ccd2417a-355c-4030-ab57-848da9991451 · outbound

This paper cites Videollama 3: Frontier multi- modal foundation models for image and video understand- ing, 2025.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Videollama 3: Frontier multi- modal foundation models for image and video understand- ing, 2025

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.132803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.738666Z digest=sha256:09569444c6e20d8d998e7960e469068b6860e48c081b5f048e8ff26afb79b238

Observation e5a247c4-ceb5-4692-b423-ddeb8ffcdbd2 · outbound

This paper cites Multi-scale video super-resolution trans- former with polynomial approximation.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Multi-scale video super-resolution trans- former with polynomial approximation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.119764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.742531Z digest=sha256:41b653e353f151ae293ac219800fb52d30f285c7d17f6486dcdc665a59396d27

Observation f5f76ab3-71f7-4fd0-a3ce-9c84f4562286 · outbound

This paper cites Span-based localizing network for natural language video lo- calization.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Span-based localizing network for natural language video lo- calization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.106066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.746225Z digest=sha256:b648496cbf72baa0c3291838946df818a10e82e56970ae563d2f64dd17d2c4bb

Observation b918aff9-c91f-4d66-9545-ab83d94212f2 · outbound

This paper cites Natural language video localization: A revisit in span-based question answering framework.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Natural language video localization: A revisit in span-based question answering framework

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.092617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.750149Z digest=sha256:bf7e238fe46142e4ec2e30406e5374a2b29d30a939809b6b0885e6559fb9c5af

Observation 55cd55bb-2f3c-4061-a05e-6ef191752d7d · outbound

This paper cites Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:20:36.753928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:20:36.753928Z digest=sha256:a414d68ffb87e93d1127657b9265599fdb80d32bff8cfc9d09af4747f8752344

Observation 3777ac06-d384-45f1-b073-021c50ce107e · outbound

This paper cites haotian liu, yong jae lee, liangke gui, di fu, jiashi feng, ziwei liu, and chunyuan li.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions haotian liu, yong jae lee, liangke gui, di fu, jiashi feng, ziwei liu, and chunyuan li

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.078965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.758258Z digest=sha256:fe7c3ad98b80b0c51939f1c66ecb0314445eb249ddb932c9766a03cb45b36d04

Observation 63938cf7-aa15-40ac-b172-d429bad4748b · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Llava- next: A strong zero-shot video understanding model, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.065596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.762242Z digest=sha256:6eee80f62214ac597bb6f6e930a81b54f0cbba82936ec733da759e6e9eddb3fb

Observation 1a6632da-4ab1-4f46-9dff-82ef0755be52 · outbound

This paper cites Medrag: Enhancing retrieval-augmented generation with knowledge graph-elicited reasoning for healthcare copilot,.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Medrag: Enhancing retrieval-augmented generation with knowledge graph-elicited reasoning for healthcare copilot,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.052644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.766354Z digest=sha256:3928920606294ce3dfa3e8976818eec16b0d67c5de38b7be5e34f1e819200029

Observation 9f655ec8-399c-4b55-8bd6-ea936e0015c9 · outbound

This paper cites Training-free video temporal grounding using large-scale pre-trained models.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Training-free video temporal grounding using large-scale pre-trained models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.039473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.770827Z digest=sha256:3226b5d2adea9810c41e4a74c98f27137ead96bd43ad04ad0b174b322139835a

Observation f3b91129-c2ac-4d45-a988-57b19e93abfd · outbound

This paper cites Cross- task weakly supervised learning from instructional videos.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Cross- task weakly supervised learning from instructional videos

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:37.026495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.774782Z digest=sha256:3b145b60345e9ad2027c26a8406cc3180360929ab04b8567980074f3da5314f4

Observation e686d75a-7c3f-4a5e-b5d8-7ab0fc9b60f1 · outbound

This paper cites an unresolved cited work.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:20:37.012280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.779065Z digest=sha256:fa765ab7809b614de267d42f17f06298e7e7fc6e3512302f3d717d68af8b3457

Observation 66befbfc-e568-4043-a2b3-bbe1b9cfd1ac · outbound

This paper cites an unresolved cited work.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:20:36.998459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.783102Z digest=sha256:114e730679ef4aaf932154d23185b9592a21a35e81836d60ada1464ca18bc469

Observation 6ee3bf54-6834-4d2a-94be-644cce6e19bc · outbound

This paper cites an unresolved cited work.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:20:36.984519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.787546Z digest=sha256:b1dc0bcc9d5474eab5a73840ff13fd4570373292eb9649162e8f7aed0d6f71be

Observation c96aefb2-c9bc-43b3-ad18-8aa7beaad72c · outbound

This paper cites an unresolved cited work.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:20:36.971680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.791458Z digest=sha256:fc53006fec3dcf39ced13bc8382ae3997b53cc19c1f73cbca4b638114dd63ee6

Observation 59128a0d-0256-4213-b472-4e348969f07e · outbound

This paper cites an unresolved cited work.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:20:36.957560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.795291Z digest=sha256:e4ae2c6b45e1faf25a74ca555b05420622fc1f352aeb1af6a9ca7c215ce56334

Observation 7635f7af-5976-4c5f-83ac-4a7fd0d8f209 · outbound

This paper cites an unresolved cited work.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:20:36.943259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.799230Z digest=sha256:4480227a3e0bdc93ab4a3a6774e95f2d7f282f228729114c255e6b76ff5453f1

Observation 28dae07d-0667-401a-bc5c-87631e57f60a · outbound

This paper cites an unresolved cited work.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:20:36.930034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.803307Z digest=sha256:fe9db8e8a3f085cdb0dd965873cd0b2fd4d13cbc96a7c13c9ea49cf319568bf7

Observation 21ec0141-1324-4470-8476-af9dcadf9ade · outbound

This paper cites 是”或“否”回答的问题,比如你可以用“你的 意思是...?.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions 是”或“否”回答的问题,比如你可以用“你的 意思是...?

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:20:36.916434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.807197Z digest=sha256:1a61113264511e3ce74a72e96af2b07cd308a9e5c4c6c00a693112127cc88ffb

Observation 05d4d6d5-66cd-4fa1-88d3-90e9dbe248bb · outbound

This paper cites an unresolved cited work.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:20:36.902952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.811003Z digest=sha256:7cda2e6186315ec58e94ef33d0dffe38a466d90abe58ce5b260344057ccb4f4c

Observation 6fd4f728-8858-4620-a0b9-82d1a72ecd34 · outbound

This paper cites an unresolved cited work.

Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:20:36.889453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:20:36.814781Z digest=sha256:33e9ca837cd045ac4bcfd15bbabc9807e40a6982b6630105473bae25eb37df25

Pith citing papers

Observation 88150503-529d-4efc-a184-218b1afd31e3 · inbound

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding cites this paper.

M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:55:41.566893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:55:40.858859Z digest=sha256:72fc4e02bffdf4b05c4db2db8b3a25df2bcd91542fe8b5207319d5ac09a9a2fd