Pith. sign in

Paper Citation Record · LEDGER

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models

As of 14 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2608.09666.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09666 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:10:34.717993Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact4
  • verified fuzzy45
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0ce8ec41-0033-4d84-ae98-2b7e0883dc5b · outbound

This paper cites Qwen2.5 Technical Report.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Qwen2.5 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.273827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.273827Z digest=sha256:83881cd0901e157899111650eaaad2cfd6c2dda001f4093a7e87b42455cec220

Observation 67d9b3e1-a1aa-46af-969c-78eef55aa7fd · outbound

This paper cites Evaluation Agent: Efficient and promptable evaluation framework for visual generative models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Evaluation Agent: Efficient and promptable evaluation framework for visual generative models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.546122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.281730Z digest=sha256:0aadc86a7f8986c0d38f7d923c83fb2c2a1e210a8519956d9ec8bf0b406d3816

Observation d51bb55d-ef9d-4530-b004-e4ac10c922c1 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.288711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.288711Z digest=sha256:b734cefd5d77cf2d017f4098eb92a49d335403a412e859846ad52bb7254fd483

Observation a948c38c-dd87-4d02-a0ee-1618637cd98d · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Gans trained by a two time-scale update rule converge to a local nash equilibrium,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.295523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.295523Z digest=sha256:9120f0d6766ab0cda851fccad2f72f0fe767c2c74d93d02753bb51ed6c587d20

Observation 8683498f-8e6a-449c-af39-1306ee3b6cb7 · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.514802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.301963Z digest=sha256:4ed886be0f35115741a35b49f086f53c72d5fccb82fa8bb7ccb66d281ca60fb7

Observation 56ace4f0-c158-4dda-9b57-846df8fd8610 · outbound

This paper cites VBench: Comprehensive benchmark suite for video generative models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models VBench: Comprehensive benchmark suite for video generative models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.309036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.309036Z digest=sha256:08f414349b07f8ab15fa2505478a2ff1093a709263539ef87ca219dd800814ea

Observation 1926fd13-2cc2-455f-8a17-e65e72cd35a9 · outbound

This paper cites Denoising diffusion probabilistic models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Denoising diffusion probabilistic models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.486838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.316725Z digest=sha256:10343dc4a9d1467f2169dd41d1021e6975b4d6356046627785e0503a65dab401

Observation cb7107c7-370c-4dba-8432-49fb0ed245a0 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.323455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.323455Z digest=sha256:20871a695d37d3e6ad6612107b223f2a9a6911803a1e6b4da5d367f1bdf16e59

Observation aa6da551-973d-4fe0-bc41-6a9f6bc117a2 · outbound

This paper cites Panda- 70m: Captioning 70m videos with multiple cross-modality teachers,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Panda- 70m: Captioning 70m videos with multiple cross-modality teachers,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.456921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.329660Z digest=sha256:ca69044606e247e8bfb56c206f28c6e899b1dd3135a8938bbcd168187fc9d2fe

Observation bfe0e96f-3486-4fc1-9e1e-b1765795bc47 · outbound

This paper cites Video enhancement with task-oriented flow,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Video enhancement with task-oriented flow,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.437694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.335500Z digest=sha256:8b76bccdb1785dcf03c433fdde229e28de6903d9a66cbdd5d7fba201bc82e174

Observation e3b501df-378b-46d2-9218-9ca7e8f17a1b · outbound

This paper cites Vbench++: Comprehensive and versatile benchmark suite for video generative models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Vbench++: Comprehensive and versatile benchmark suite for video generative models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.420165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.341243Z digest=sha256:e21147cc929f0155251100f289c256afddfce2e58ac12bd49731ef822416e2f0

Observation 466618b1-645e-4339-a2d3-f6335af20e77 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Evalcrafter: Benchmarking and evaluating large video generation models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.347482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.347482Z digest=sha256:876afc8685ab9da8ded3e89de4c449b1afadc5feafc277f5955dc0eec69e8769

Observation 0b97338e-38c6-4012-abba-3d1f69e9101a · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Llamafactory: Unified efficient fine-tuning of 100+ language models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.390417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.354222Z digest=sha256:bdadbe87b26b75c42ba54b4dd958c696f6c43631f0771c92d00d0a9f4fa4792d

Observation e14bbe5b-0663-4032-a70d-f4411876dc4e · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models High- resolution image synthesis with latent diffusion models,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.359831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.359831Z digest=sha256:6efbe8330f49f1e46995920c3bc245dbcf3317fae2e02dd7b71f7e1321b9257c

Observation 4692b390-534f-43ac-90d2-916e6ee2979d · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Sdxl: Improving latent diffusion models for high-resolution image synthesis,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.365122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.365122Z digest=sha256:de1f9ef8eefe11a36a626dee4b8a4ac399040b6986e4c892c6c0d34b2303909c

Observation 2b9bf8ef-553d-4166-b1b1-586c0d30964c · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Scaling rectified flow transformers for high-resolution image synthesis,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.348653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.371146Z digest=sha256:3495e4ce47d913fd7543fc27997643d1c08f526c1ee05db19b191fcc355d99c7

Observation 891eab10-9724-4dbc-9f45-efea4e269154 · outbound

This paper cites Latte: Latent diffusion transformer for video generation,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Latte: Latent diffusion transformer for video generation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.330334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.377213Z digest=sha256:0c28242142c7c262b77d8917a5e1070f7bf476d036ff9924d21827a7905d8f0c

Observation 32131614-b96d-4de5-9171-82b66ac2b60f · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models ModelScope Text-to-Video Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.382533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.382533Z digest=sha256:a53a19aa76bdb0d8464b99a29c29a24a6315dc5b8a72674abf17a45b2af71780

Observation a7703b7b-ae9a-4d25-bdf8-919011621982 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.311625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.388679Z digest=sha256:bad1df66073963234c64a9181b83c8ce8ebe7c99c2b1f4fbd40fcee2f27558ec

Observation cb222d48-952f-41e6-af15-0d0ed195691b · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.394196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.394196Z digest=sha256:be490f30e98b9ae7f724273bf1a57a34332de452208bb26c747bb5363fb6485a

Observation 3c65ca84-5049-493e-b634-5488567a0402 · outbound

This paper cites Lavie: High-quality video generation with cascaded latent diffusion models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Lavie: High-quality video generation with cascaded latent diffusion models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.399728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.399728Z digest=sha256:7187b8f3b38c34ca72fdde337997d57f01bfec3a3529b7b4ec00d45ff9741d05

Observation 9ef1105c-3b0d-4988-9341-7a894ac0c23b · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.405066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.405066Z digest=sha256:c17a42e194e3f7ef642eeaf075ee5afa9595ba1d56944abd8682da897e45a311

Observation e2a1f3a3-abfd-4173-8acd-7a4248f1b9c0 · outbound

This paper cites Holistic evaluation of text- to-image models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Holistic evaluation of text- to-image models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.277628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.410559Z digest=sha256:f057991c3d3504a9cbd9d39b1e308eecae10805f471c36c492db17df3047bbd8

Observation 1cb80ab0-f00c-4da4-be17-8c706ec867a1 · outbound

This paper cites T2v- compbench: A comprehensive benchmark for compositional text-to-video generation,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models T2v- compbench: A comprehensive benchmark for compositional text-to-video generation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.257829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.416354Z digest=sha256:307e6d7080e76d8b6dd19918334d7a2b12f4ecf28eb2de988b6ebfade5567d46

Observation 1980a47c-d758-4e6c-a947-3e4be4d134ee · outbound

This paper cites RealUnify: Do unified models truly benefit from unification? A comprehensive benchmark,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models RealUnify: Do unified models truly benefit from unification? A comprehensive benchmark,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.239214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.422723Z digest=sha256:57d69d3261f1c97a56db188a1aeec306ca4b137c5d2ba29a517d795ed4af3a01

Observation 1a877a63-62b2-4459-ab31-4171aa424c08 · outbound

This paper cites Uni-MMMU: A massive multi-discipline multimodal unified benchmark,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Uni-MMMU: A massive multi-discipline multimodal unified benchmark,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.220201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.428074Z digest=sha256:3a100af929563d0da3719c3bf9197342fe420e3e0cbc0a14cd79054b45610909

Observation a76b7fe1-b4a2-4447-b7a3-44b6ba3eae7e · outbound

This paper cites Unison: Benchmarking unified multimodal models via synergistic understanding and generation,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Unison: Benchmarking unified multimodal models via synergistic understanding and generation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.202093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.433565Z digest=sha256:f9a9c685a7d095b90a0152258bc01187ced0e3605b01d54ffd8ea4997c2819e2

Observation bb6af445-5f93-4967-a2db-20498bdaca42 · outbound

This paper cites Multi-dimensional evaluation of text summarization with in-context learning,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Multi-dimensional evaluation of text summarization with in-context learning,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.439026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.439026Z digest=sha256:95d763bf18e548171690131cf06d34a5e4ea2c61d0c1a51f8bf74e0f8824e211

Observation f360cad3-a1a7-4b5c-820d-9cabe5ff4e23 · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Can Large Language Models Be an Alternative to Human Evaluations?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.444699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.444699Z digest=sha256:3a14d15303030a40832850eb556a13ae9ffa0f84bae13a2f24bbcb29b96a278e

Observation 5138a724-6483-4c4a-8de8-b6b274f43bbf · outbound

This paper cites Gptscore: Evaluate as you desire,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Gptscore: Evaluate as you desire,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.183048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.451791Z digest=sha256:ea2ca26117575faa279004631355602573fe5708fa03c9df168ca893ee8bf2f5

Observation 5f82c26d-891f-4dca-9760-8b0e6f1bffef · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.457472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.457472Z digest=sha256:5c0755b2c4808c018b28a77117453b9f3a053f5d34159fba4f1e63d035134d2a

Observation 27b06c18-8f3c-4548-affb-c52c9f8ec0e5 · outbound

This paper cites Exploring the reliability of large language models as customized evaluators for diverse nlp tasks,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Exploring the reliability of large language models as customized evaluators for diverse nlp tasks,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.146744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.463686Z digest=sha256:344a4694aba18d74300d531e7ffd42a5871653b49d238e84542c1bab0381e595

Observation 3cfd6982-883c-45c2-a497-0c5ae3031012 · outbound

This paper cites Au- tonomous evaluation and refinement of digital agents,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Au- tonomous evaluation and refinement of digital agents,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.127889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.469999Z digest=sha256:3f6ac4b9f61be76d231a8fe90e621cae4e95a807631b9528a8e508f4eb0c19c4

Observation d9a95d08-7651-4dd2-8b5f-8e93d92bbc5f · outbound

This paper cites A survey on large language model based autonomous agents,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models A survey on large language model based autonomous agents,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.475664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.475664Z digest=sha256:1b387f8e12b76b986ffbfcf69884f1cff88c1fe117c07fcffb1712060868851a

Observation 317db75e-6f79-4bc0-9b4a-c7b907d44b51 · outbound

This paper cites From generation to judgment: Opportunities and challenges of LLM-as-a-judge,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models From generation to judgment: Opportunities and challenges of LLM-as-a-judge,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.096131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.481986Z digest=sha256:d45419882580deafe384b83c3f2f96e002a68e6d603ebf576a4a41cdecc341f4

Observation 23bfe681-a927-4fd4-b297-1f8a6d9c75c4 · outbound

This paper cites Agent-as-a-Judge,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Agent-as-a-Judge,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.487928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.487928Z digest=sha256:2bf5ca137291cd5884bfef4c2a70c340e0c37668a91530e1d69b49e106dbcd2f

Observation 97456063-a9d5-4723-93d8-aa98e61173d0 · outbound

This paper cites Agent-as-a-Judge: Evaluate agents with agents,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Agent-as-a-Judge: Evaluate agents with agents,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.077458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.493042Z digest=sha256:051d73f02c464055cee7a0d1cddb3c11eb6539dcadf0fde90a7a6b892a804df2

Observation 6cfbbefe-85e5-40ea-b8f4-66c4e846f40b · outbound

This paper cites Adaptively profiling models with task elicitation,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Adaptively profiling models with task elicitation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.059211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.498563Z digest=sha256:8eb80f243bb73c5e5cea3b100dbc1b31127a37594437554fb5f2153117061ab9

Observation 069b477f-27f9-436d-ae62-8c70bda07030 · outbound

This paper cites Mind2Web 2: Evaluating agentic search with Agent-as-a-Judge,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Mind2Web 2: Evaluating agentic search with Agent-as-a-Judge,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.040552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.504275Z digest=sha256:3aefa527fc674e9b539fd841aacf0ea197851e0e000259ee7c63a4fd908db289

Observation 59f83f50-386a-4d89-b8ad-59be708d796a · outbound

This paper cites One-Eval: An agentic system for automated and traceable LLM evaluation,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models One-Eval: An agentic system for automated and traceable LLM evaluation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.509937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.509937Z digest=sha256:671b7508bb339d8c4c5cdbe4b1d70985032544d8bdb16dd0e591ae104379f13a

Observation cb060ea8-c4f9-45ba-a6ee-15d2b48e2974 · outbound

This paper cites AgenticEval: Toward agentic and self-evolving safety evaluation of large language models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models AgenticEval: Toward agentic and self-evolving safety evaluation of large language models,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.021413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.515912Z digest=sha256:9ec0bd3ab254979fefa9bd57ca303385957f187dbaf910beab1be5496b9feeb8

Observation 1c1d5067-e92d-446a-b227-13c7d55f2d7b · outbound

This paper cites Evaluating Hallucination in Text-to-Image Diffusion Models with Scene-Graph based Question-Answering Agent.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Evaluating Hallucination in Text-to-Image Diffusion Models with Scene-Graph based Question-Answering Agent

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:10:35.145783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.521533Z digest=sha256:f4e6a9f63f77a94466fa044d09c70b979d17e18834c189eb90710f6a68d40e9f

Observation ef08f582-8260-4a79-b72f-d8ce8c51fe6b · outbound

This paper cites A unified agentic framework for evaluating conditional image generation,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models A unified agentic framework for evaluating conditional image generation,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:36.001964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.527852Z digest=sha256:bea52d3e0dc8681521cdbd1c32973b4a97295ad367912253fd2d413294227370

Observation f8f1d3d3-c7d6-4506-a6e1-fcb7cc07e9f9 · outbound

This paper cites EdiVal-Agent: An object-centric framework for automated, fine- grained evaluation of multi-turn editing,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models EdiVal-Agent: An object-centric framework for automated, fine- grained evaluation of multi-turn editing,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.534020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.534020Z digest=sha256:8bd2f32383d6e8d16b9b3703a9d91986d20ec08a7873c0c6d8b9892e97af696b

Observation 2765068d-88ee-48b8-9d9d-6331df4fb633 · outbound

This paper cites RewardHarness: Self-Evolving Agentic Post-Training.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models RewardHarness: Self-Evolving Agentic Post-Training

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:10:35.013589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.539760Z digest=sha256:1402ca6381adbc94e15180e6bb8e06f3b41988ec2d2c9c7c952bdadc4d4032f5

Observation 3a1ae944-f003-49cd-bb5a-9c262d9b31a4 · outbound

This paper cites VideoGen-Eval: Agent-based System for Video Generation Evaluation.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models VideoGen-Eval: Agent-based System for Video Generation Evaluation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.545499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.545499Z digest=sha256:e6d28c9393f6a2f52e7f4217035f2b6272d7c1ca50095513d54bd0f23b0bc63e

Observation a568ed33-ab06-4704-816f-5183ef1ddc14 · outbound

This paper cites VQQA: An agentic approach for video evaluation and quality improvement,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models VQQA: An agentic approach for video evaluation and quality improvement,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.551263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.551263Z digest=sha256:0f61e0f74a93b8264e1857f2796d8521c9eb9cc951b7b1d98e7e4743bb5665b8

Observation 0a055ea0-48d5-4d28-8c40-4dbf0e375dd3 · outbound

This paper cites VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:10:34.850890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.556555Z digest=sha256:e36d7b28be682b42e46df19a3efcfb336f452da17710de17493eda37ac3250ed

Observation a4915c07-cd28-4da8-9f87-231455740a72 · outbound

This paper cites AgentRewardBench: Evaluating automatic evaluations of web agent trajectories,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models AgentRewardBench: Evaluating automatic evaluations of web agent trajectories,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.981707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.562246Z digest=sha256:67a31650dc398f125769c8824882a918b636926ec7d539660541101d28ee31ea

Observation b805a161-23ef-42b2-899e-54b29dcf7edd · outbound

This paper cites AJ-Bench: Benchmarking Agent-as-a-Judge for environment-aware evaluation,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models AJ-Bench: Benchmarking Agent-as-a-Judge for environment-aware evaluation,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.961740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.567520Z digest=sha256:2cdb612e9d69efffcf31fb7d335c7e4afca405072309737be7f8911bf32fa926

Observation 55fbe9c5-8aad-44a1-9f4b-1665da83e741 · outbound

This paper cites Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:10:34.811346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.573818Z digest=sha256:a6c2ff4eaa590c52462b7be2f641f0d12a1e1f7a55dd64e94a107dc3c07b3ad9

Observation 4892a728-5e49-4714-986d-8ce6d38c1493 · outbound

This paper cites Establishing best practices in building rigorous agentic benchmarks,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Establishing best practices in building rigorous agentic benchmarks,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.940994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.581031Z digest=sha256:6aecabb9af63cf5f3f97bdb510742e509559ba92eabd2a1a16946707ead142ba

Observation f5b72079-8318-4c8f-9af9-c07a9d4f51c0 · outbound

This paper cites Webarena: A realistic web environment for building autonomous agents,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Webarena: A realistic web environment for building autonomous agents,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.920062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.588041Z digest=sha256:5e1520a0d038b0afea78ec4f642d12e0ca3ba27c0ad9214d6adfb71cf5b68de8

Observation 884f717c-c04c-4882-8921-73d07daa2807 · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.901109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.595042Z digest=sha256:1fe45d16eb5f52ac12c9c9a8ca02aa32e8068b6296ec748e9202d0a8bd8a37b6

Observation 5ec24db2-4335-42c0-8707-af4d624159d8 · outbound

This paper cites Mobile-agent: Autonomous multi-modal mobile device agent with visual perception,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Mobile-agent: Autonomous multi-modal mobile device agent with visual perception,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.881632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.601006Z digest=sha256:920db0fe0212ade64273bf18ddceb7f3fcbf9d690b715057898cec61a82b0fca

Observation 0779f335-e4bc-40aa-a9e8-7f7f13f1c724 · outbound

This paper cites Omniact: A dataset and benchmark for enabling multimodal generalist autonomous agents for desktop and web,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Omniact: A dataset and benchmark for enabling multimodal generalist autonomous agents for desktop and web,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.863054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.608570Z digest=sha256:25d4c9f69de49cb1f69e46722e173ef59369c00bf551211e704b1634f72dd61c

Observation d20b6cec-7b9a-4352-983f-8470dfa8f8c4 · outbound

This paper cites MMInA: Benchmarking multihop multimodal internet agents,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models MMInA: Benchmarking multihop multimodal internet agents,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.844739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.615477Z digest=sha256:9b2aef1f22dce974bb2aedd9dde07431943dd15b444d2774bee3ab8df81eb539

Observation 15e09f98-67a5-4fc6-abea-43c5421f0cea · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Chain-of-thought prompting elicits reasoning in large language models,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.621032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.621032Z digest=sha256:91ec1ac71e27ea0cd0dc8bab50f79fef3dbcd09811d7c24ce1a70e56af8dcb5d

Observation 394a836c-5f31-404a-938a-26428e3f7bd7 · outbound

This paper cites Large language models are zero-shot reasoners,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Large language models are zero-shot reasoners,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.813842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.627573Z digest=sha256:cdc96c99bed8c1822be00c1bbfc0adc0003ad5fdfbc850fae1d617df7ececa84

Observation 62dfbfd9-5c76-4a20-b42d-f4073384bb7c · outbound

This paper cites React: Synergizing reasoning and acting in language models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models React: Synergizing reasoning and acting in language models,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.795668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.633742Z digest=sha256:78a126e6469e6b3c57e57562cc1c9acbddb3e6845d00ff0faf836bcf00adb6da

Observation ec597b3a-aba7-4ed7-9f8a-611813990e26 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Tree of thoughts: Deliberate problem solving with large language models,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.639977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.639977Z digest=sha256:a795dc801e58fc8078d8de0780a454c06386d18396fe49e9ffc87b9fc152b430

Observation e55cfbbf-f400-478b-991e-9584cd3f4c0e · outbound

This paper cites Algorithm of thoughts: Enhancing exploration of ideas in large language models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Algorithm of thoughts: Enhancing exploration of ideas in large language models,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.761248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.645742Z digest=sha256:33009c8c6f2e5edda7ec8722bb42d752be0e7d56c7a69c8316f94b297958ce58

Observation a892988b-bd30-44b1-9cfe-1025c985f153 · outbound

This paper cites Graph of thoughts: Solving elaborate problems with large language models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Graph of thoughts: Solving elaborate problems with large language models,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.651497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.651497Z digest=sha256:2f7af4a038f9c1f43407fc889187c09075d98c272b2024859956ac10c289f35c

Observation 76b191dc-095c-4924-bd9b-b451039044fc · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Self-consistency improves chain of thought reasoning in language models,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.727152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.656757Z digest=sha256:6e723d7e12283b946fd8c7ab5bede74c96f302e4f28573577a48b64d9036ed19

Observation bb21bdbe-597c-4960-935f-d2c26ea82ddb · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Reflexion: Language agents with verbal reinforcement learning,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.707568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.662278Z digest=sha256:89a5a90ac515c855ccedb0de487b8e640251cfae87e7b9c71df3487f10615f4c

Observation 57c7b42e-ed26-4ad6-927a-87324c483f98 · outbound

This paper cites Api-bank: A comprehensive benchmark for tool-augmented llms,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Api-bank: A comprehensive benchmark for tool-augmented llms,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.689207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.669165Z digest=sha256:0ea96f2cca91ebe3277e7e608f1f8bd63e16a459821ce869f159c4b6cebda3bf

Observation c67a8f2d-6640-446c-acfe-2e2b536b9b3f · outbound

This paper cites Toolllm: Facilitating large language models to master 16000+ real-world apis,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Toolllm: Facilitating large language models to master 16000+ real-world apis,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.670378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.674523Z digest=sha256:0816f201492d942e510de72fb860d18e384af211e6e3fe800273733ee3dc6f02

Observation b15b39b5-b8e5-4919-81bd-aa13ab787c5d · outbound

This paper cites Toolformer: Language models can teach themselves to use tools,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Toolformer: Language models can teach themselves to use tools,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.680055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.680055Z digest=sha256:5073313170de544ea10afc37cb7432ecad98ac7e635dc5ce2f00602b738f6ce0

Observation 8154be5d-4f83-4782-a13c-cd724398a4cc · outbound

This paper cites Gorilla: Large language model connected with massive apis,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Gorilla: Large language model connected with massive apis,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.636703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.685517Z digest=sha256:111644f04d68244383c74d65932903f1c73e05aa25b30fc102497ef805ad461a

Observation c8a17faa-7dd2-4f7f-a83c-43090825a05b · outbound

This paper cites T2i-r1: Reinforcing image generation with collaborative semantic-level and token-level cot,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models T2i-r1: Reinforcing image generation with collaborative semantic-level and token-level cot,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.614941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.691190Z digest=sha256:3ebe54d379839e1284cc547c02a929e226e906f02403b2d96c8934ddfd39d1e5

Observation b61177a4-7c74-4244-b54e-69e585e58440 · outbound

This paper cites VChain: Chain-of-visual-thought for reasoning in video generation,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models VChain: Chain-of-visual-thought for reasoning in video generation,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.595771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.696235Z digest=sha256:e981065d7ac167555e0eed3020cef56e697cb6a701b7ec24981687425ea0fc10

Observation 8882e30d-a1af-4e9b-9776-c976ddb3d579 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.701245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.701245Z digest=sha256:785bc302187d5fa382e1f999019adfe7b57dfb7001da0ce895ca1cf758a683da

Observation e7707f2b-9fd1-4c3b-acad-42cc919d19f1 · outbound

This paper cites An illusion of progress? assessing the current state of web agents,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models An illusion of progress? assessing the current state of web agents,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.577136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.706817Z digest=sha256:914cabb8e73d8c9c91358e8e2f61dd9627f22e73220996064e75c01b42ed150e

Observation d77e4872-7f73-4950-94e3-8a347616e4d0 · outbound

This paper cites PaperBench: Evaluating AI’s ability to replicate AI research,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models PaperBench: Evaluating AI’s ability to replicate AI research,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.558779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.712208Z digest=sha256:0539a3c4758f9e488a138315f31f8b3c7ae8d576e036cc4cd3d9a5c0c26a1787

Observation c4903075-9b2a-4fab-a38f-4b73a59b0173 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning,.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Clipscore: A reference-free evaluation metric for image captioning,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:10:35.539681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:10:34.717993Z digest=sha256:49210ca9ce41423e0a8e788a68050a8015b5f12943372b7bf6a4be6645122f3a

Pith citing papers

No inbound Pith citation observations are available.