Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:21:45.118956Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2608.04670.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:21:45.118956Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fa7468e2-a0f4-4914-bcae-ec7edfd4b7dd · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Lewkowycz, A
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d1fe282e-2c47-455a-9284-f52d84fc41a1 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Chang, X
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e731ee4e-e7ee-40ad-a049-d58207ca6228 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3d39048b-ef78-4893-b00a-e0898b2c173c · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2726dda-cc6f-41a1-acf9-d8277c2e3b49 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32961ddf-e62a-4474-9931-7139f5ae7dc2 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1664dfb-7a8b-43ae-9d8f-25dcb1d79db5 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e4a11fcf-645a-4aaa-98c9-6ec210916ee1 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 409df814-db12-4d26-87d8-298b6d9b4a98 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb55a51e-25bd-48a6-a8a3-2e1a23e3fbd3 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 44e5fd0c-1344-4e75-a719-9bb3ce192717 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf15a65-4f2c-4d03-a251-96731052994a · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark On the Measure of Intelligence
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5841a2f2-84ac-4dcb-b08f-127a485e314a · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ca2b51a-5480-4db7-9ac3-05a694cf0493 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Attanasio, P
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 50242c54-3c37-4d0a-8311-5e6ccd0e23b5 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Evalita-LLM: Benchmarking Large Language Models on Italian
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3becf2ce-aee1-4e6f-97dd-7a5fdfa9ea04 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0b1b81e7-71ef-4cc2-8dc1-049d87926a49 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Tedeschi, F
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3fa49c86-3971-4e5c-bc3c-ff461ea6fd2c · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Khoshtab, D
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b2676a61-15f0-4206-a9e0-216573980de9 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aae267c-af0e-4576-99bd-5c6f3ded4d1f · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b6c82470-0b85-460f-bb34-6dd550e20627 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7dfba636-070c-4208-a330-61e7b51bf8c4 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Donthi, M
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d7760f35-722f-4c16-9172-439e74990dff · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6daa9457-2d54-433c-bd73-1eb68edea5b5 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Proverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 09ce6256-3f37-4da6-9d04-c4342578370a · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Caramagna, I 200 proverbi italiani più belli e famosi (con significato), 2025
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 38eded67-ba3d-48fd-939b-b3ca4da3b4a5 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark URL: https://openrouter .ai/, accessed: 2025-06-15
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e487e10a-0c4b-457c-b815-695b8298c026 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark GPT-4o System Card
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e269fac7-1e18-4536-b981-0c539fa1a946 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark DeepSeek-V3 Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bfe8cad-5a95-4ff0-b287-f69a764c903f · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecb7b7ea-3c87-443d-beff-9246f08d6f9d · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Qwen3 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a462dfd-05cb-47d9-9ff0-71c05c9a5435 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Gemma 3 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af5a729b-cfb3-4a72-876e-fd28f7d1accb · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Orlando, L
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 60171fa0-2dd6-4f60-9313-81f3b1345f32 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark URL: https://huggingface .co/ mistralai/Mistral-Small-3.1-24B-Instruct-2503
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d3a5a1f8-39b5-4461-92ce-32202342d426 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 005ae45e-b53a-4848-bbfa-7a72dad54bc4 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Etxaniz, G
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69c0ce65-5727-4e35-b690-39617bbbd0ea · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Ranaldi, G
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f768a60d-86bc-491b-a4d5-762fe23e970d · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ebb7780-776f-4606-a588-55645408e891 · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark Reasoning Models Don't Always Say What They Think
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d10fc8b-a605-42a7-8afb-100a10f1f11c · outbound
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.