Pith. sign in

Paper Citation Record · LEDGER

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains

As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2607.18438.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.18438 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T15:29:29.577032Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 199fe0a8-c4c4-4def-b0fc-d153adc8b48b · outbound

This paper cites ARC Prize Leaderboard.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains ARC Prize Leaderboard

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.070728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.070728Z digest=sha256:6a84e4c11e21e7b9f41ac06dc2bab050a47727242c6b7ed3b2b9e253b1f15d6d

Observation 98b12e11-c705-42d5-8d7e-4dd29783267c · outbound

This paper cites τ 2-Bench Telecom Benchmark Leaderboard.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains τ 2-Bench Telecom Benchmark Leaderboard

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.142006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.142006Z digest=sha256:c6ac00bf1125286d8c279178938b163df2e054507b47262ceb2dc02a23791e3f

Observation 65cf5065-a338-470a-ab92-4fda93111e5c · outbound

This paper cites Humanity’s Last Exam Benchmark Leader- board.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Humanity’s Last Exam Benchmark Leader- board

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.285997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.285997Z digest=sha256:498d46ba5324ac3573dea1c2f1c8eca6b8b5f796a44f03b342ace8c5683e841b

Observation c5941784-cf22-451b-b93c-d91dc20f6a0d · outbound

This paper cites Introducing Claude Opus 4.7.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Introducing Claude Opus 4.7

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.431791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.431791Z digest=sha256:b471e49c8c7d35d59010d65bda221ef5884591eb20e68feb65a8ab942634b69c

Observation a92a0fe3-69d0-4dbc-b7df-5392c0d50925 · outbound

This paper cites Claude Mythos Preview System Card.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Claude Mythos Preview System Card

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.571603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.571603Z digest=sha256:1872a0139be2281dc4d8a92278ba35d174b921c657aec78ac22c8464b63ebe95

Observation 46c37784-cc22-4005-b93c-d405ec4e3649 · outbound

This paper cites an unresolved cited work.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.682318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.682318Z digest=sha256:1f1a8fc6ea26423292abbab5071c48840d808ed54d99a495b9710fd6603af130

Observation 25c5f1eb-6ec9-490e-8c4c-ecdc93a38316 · outbound

This paper cites Introducing GPT-5.5.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Introducing GPT-5.5

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.797304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.797304Z digest=sha256:3d0375c06dc058d8425035dad2eb6181599ffbc901def6a5f9efa5368c409817

Observation 86e1502d-9276-42a8-bfd5-614bd3e637c9 · outbound

This paper cites Grok 4.3 Beta.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Grok 4.3 Beta

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:23.914038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:23.914038Z digest=sha256:b0411d65a125b5d1ed76634b577d9286308de32574bfe414ae4eee8ba14a01c3

Observation eb367314-54ce-45e8-a229-cb20f3c374f6 · outbound

This paper cites Qwen3.7: The Agent Frontier.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Qwen3.7: The Agent Frontier

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.079864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.079864Z digest=sha256:273bbaf442cde1399593ce8c291befd4bf3b61c8a0933ae208a2f6dfb6828efc

Observation aa2f134a-7b0c-48a9-8009-f1286765e82f · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.224184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.224184Z digest=sha256:ea3d9226e2d974a853bca2fd6fad124866a32f78c85632103ef5ab0d0a8b2c42

Observation 79d822aa-4ddb-4df3-bd1e-ec4c29589558 · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.341223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.341223Z digest=sha256:b4c2b2d6425417eed9ca436a8550429a75c6d753369725aeece4f4dee396b1ff

Observation cf52d656-99a6-493f-ad55-56598de2dd88 · outbound

This paper cites AgentFrontier: Expanding the Capabil- ity Frontier of LLM Agents with ZPD-Guided Data Synthesis.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains AgentFrontier: Expanding the Capabil- ity Frontier of LLM Agents with ZPD-Guided Data Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.460230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.460230Z digest=sha256:05e14bac13d294899683d018343a5c2d82ae57bcc50e53bc2338e2bef03d2af5

Observation be277fbd-77a6-48c9-8b0c-12d907e8234a · outbound

This paper cites Measuring Agents in Production.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Measuring Agents in Production

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.584013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.584013Z digest=sha256:cd74a1d6f943795c4781ad4efdd5d6c3c46aaaa90ed30629b3b99a0c86c4f420

Observation 34aec3da-e1c8-4667-92ca-5fb377f91ae9 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.722570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.722570Z digest=sha256:8ddf65f8fbc298910812acd658047d41225d8a6dcf7c0b1308c4b8f93e3aab50

Observation ca1b3420-6791-41f2-bfa6-4db8fbad9a91 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.846792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.846792Z digest=sha256:3376afb4afeca55c7729bc664e23a26a72e8ff3e08547b25c5c321d1761defb5

Observation 164f11a7-8477-4b31-afce-87a851f1deff · outbound

This paper cites SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.950535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.950535Z digest=sha256:c542142d0e861a0e26580cb94e2fa35caff9d1e7ff50781ef446bd0416d4e3c8

Observation 1afa2c31-87b0-4c15-858f-f118faad26df · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.040725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.040725Z digest=sha256:3b58c6ee26a0731ac3ba4b7273e5dacc24ee7c14b846f30fd72d46ecc92bad7a

Observation e6019c4f-e0c5-4d44-8dc9-dfb79436910d · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.119434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.119434Z digest=sha256:1a1bfaaefd65192eb8d13465de5286160baf7d8fc8d5a6a41a6f190448343bb8

Observation ba79076e-70f6-42aa-b0c2-6211e74cc6e6 · outbound

This paper cites GPQA Diamond Benchmark Leaderboard.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains GPQA Diamond Benchmark Leaderboard

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.207610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.207610Z digest=sha256:96c9b2834167720e907c701c46ab269e0d16e091fa16847106bfe5631242403b

Observation a682bf96-755a-461d-9ae9-0677552ab193 · outbound

This paper cites MATH-500 Benchmark Leaderboard.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains MATH-500 Benchmark Leaderboard

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.254961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.254961Z digest=sha256:f283766b0a98feac5b7e3c31ac66e006e567883f6e4b98da3181ad319673858f

Observation 10e94e7b-1cd9-4635-942e-8306e0a16a0c · outbound

This paper cites https: //artificialanalysis.ai/evaluations/aime- 2025.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains https: //artificialanalysis.ai/evaluations/aime- 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.310495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.310495Z digest=sha256:b6c9bb59cb88ef6af86db06ba2c870e811c178e562e97962039f18eec0f3948d

Observation 61df446c-4935-4572-926a-f55994545daa · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.488347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.488347Z digest=sha256:1deafd84ee243387863e5cc160b8a577b3d5896767794536fd0dd6f2c8728ec1

Observation 0450d210-6410-4c0d-b99f-b0f43082cc93 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.602195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.602195Z digest=sha256:2cd76c33f02ca78e585d902df2cc44aa23f1096fc3cfb70887e57e453b267fed

Observation 9a02249a-32d0-481c-aaa8-e66eae75cdca · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.746583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.746583Z digest=sha256:916e43a0ba85c66fb2dce8d5a72398be0f8daa981017b24cd4553c908472d9b9

Observation 3e48715e-bf87-4070-8042-f448aa32867b · outbound

This paper cites API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.830028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.830028Z digest=sha256:58c0292afce078cc827e585896029e223aa28b1ee609c45bf8adb8b814c3f865

Observation 75bdcedc-0a18-4811-958c-30d6e7c4cf80 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.978104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.978104Z digest=sha256:90911e5b95a23b18a5eef624f60964778addad8499c73baaa11405895e53fa68

Observation e0e9443d-03ac-42f0-b36f-7d9d8a8b9786 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.141497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.141497Z digest=sha256:ac9e0ee86808f0b607444edc28399d7695f4c65257bca52029e111b9514ef217

Observation dbc74aea-868e-4afc-9ae7-3903b33d138c · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains GAIA: a benchmark for General AI Assistants

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.306719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.306719Z digest=sha256:52e90d260b2a847460e854afa11fb3f7dbe6929e148696dcfa83367677210554

Observation 6721ea98-8a24-4d57-944e-251f92ceba4c · outbound

This paper cites APEX-Agents.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains APEX-Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.450130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.450130Z digest=sha256:9392629c64cefe5fa3d59d34e92ff11900c4428b4fe9107223d435a1e204b326

Observation df722cae-32da-4264-9865-75bf17af0167 · outbound

This paper cites Are Your LLMs Capable of Stable Reasoning?.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Are Your LLMs Capable of Stable Reasoning?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.561138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.561138Z digest=sha256:cc5bd1fc2c602e8163e533e3e5c3b9d559ba9550fb72c628ba250d1a746a88e6

Observation e43b92a4-133a-4a15-8c65-8b8bb364b280 · outbound

This paper cites Project Glasswing.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Project Glasswing

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.693813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.693813Z digest=sha256:777c28f1181b431ecadb64051775738c6cfd7c330869e285437ded048fdb1d75

Observation a0ca0af4-d1a1-45b7-a902-76095d03c743 · outbound

This paper cites Claude Mythos Preview: Anthropic’s Frontier Model Explained.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Claude Mythos Preview: Anthropic’s Frontier Model Explained

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.838864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.838864Z digest=sha256:f9cc94687490c278f332da4918524ffac5b20beb872ebaae2ea1a6b231079860

Observation 54ef6eb1-c340-4cc9-8b34-a0f704039f99 · outbound

This paper cites Holistic Agent Leader- board: GAIA.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Holistic Agent Leader- board: GAIA

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:26.968437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:26.968437Z digest=sha256:8a8a5dd84f496461f90177585de1ce5a683268c1ed5bee0a13b523546c20751a

Observation c8223312-68e0-4dc1-b75a-1ce7a7689831 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.108781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.108781Z digest=sha256:1a9529de9a68f9de26bc2f4b5fe2d534e820404f7da3bf189d94577b634d800c

Observation dd0ef2fa-538b-4ceb-ab22-42df4ed8a441 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.236946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.236946Z digest=sha256:be2a417e000ab5d1b664686640c739816f816c5f5dc2ed1c9f749724800c3e03

Observation 832cdfff-61ee-400e-9c61-dcf0b491a644 · outbound

This paper cites Investigating Data Contamination in Modern Benchmarks for Large Language Models.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Investigating Data Contamination in Modern Benchmarks for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.374252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.374252Z digest=sha256:3a2a155b17f2bda9f7936f3909d139ace78b939adaf301a80ca56aa1d17f2b19

Observation f2cbc7d4-c6f3-4dc7-9736-64b0026e5fb3 · outbound

This paper cites Anthropic API documentation.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Anthropic API documentation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.505166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.505166Z digest=sha256:92e9a52d1cdf48c8b77d570585af2d6bb54deb41e2ed83b2b13b11a1e5d8d1d9

Observation bb4d9cfe-0542-4d59-81fc-e38bc39c98dd · outbound

This paper cites Gemini API documentation.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Gemini API documentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.579197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.579197Z digest=sha256:18d7caca851b6379478d3970177d735f088f540ac7444e4bbf67291a0a988e94

Observation 2f09ac12-5d15-46c7-9323-d81a42746256 · outbound

This paper cites OpenAI API reference.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains OpenAI API reference

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.635227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.635227Z digest=sha256:30c2b40f0b6b32a4a0fee6ce5f41f3d437f6c583d9e6415ea4418a912d54d74a

Observation 9015a179-9740-4a4e-9b9e-e292861079af · outbound

This paper cites Adaptive thinking.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Adaptive thinking

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.700320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.700320Z digest=sha256:0767d33ec57d33a6920410a588e7445505297e43fcde694d33693a79d59aa215

Observation ff719fe2-494c-4888-88d9-75f23ea9945f · outbound

This paper cites Gemini thinking.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Gemini thinking

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.703737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.703737Z digest=sha256:6e129801b726ba52638aa541fa0f0909768352874235c4139df13406a7d706b6

Observation b9c86265-215c-42b4-98af-cade56917493 · outbound

This paper cites Reasoning models.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Reasoning models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.729612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.729612Z digest=sha256:f3cf97f729ae8cb03d29225d962ebfa8d6d609bf46b1a989066827cd8286ebc5

Observation da1afee8-2254-41c8-a250-dc17ace86eb9 · outbound

This paper cites AA-Omniscience: Evaluating Cross- Domain Knowledge Reliability in Large Language Models.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains AA-Omniscience: Evaluating Cross- Domain Knowledge Reliability in Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:27.872812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:27.872812Z digest=sha256:a277d400cfd8058703b292066ddae5698cfc9cfbdbb08ab11e65c9004105c430

Observation 686673b7-628e-455a-817a-91a3608e4013 · outbound

This paper cites Context Length Alone Hurts LLM Performance Despite Perfect Retrieval.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Context Length Alone Hurts LLM Performance Despite Perfect Retrieval

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:28.016995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:28.016995Z digest=sha256:26317057465fbdd72dab23ecf293d7bc90181a7ff9bc1b0539c74d3a25ac35a5

Observation 924eb265-a4e8-4033-8bd0-665f4c8775e2 · outbound

This paper cites AA-Omniscience: Knowledge and Halluci- nation Benchmark.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains AA-Omniscience: Knowledge and Halluci- nation Benchmark

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:28.123820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:28.123820Z digest=sha256:835c0b2f7ea13a6d6bcd4c5a107b7b5f3ea26776b2c0a515c1cd9af81e2e2003

Observation d31f90aa-ea5f-4814-b6eb-2506aa063eb5 · outbound

This paper cites Why Language Models Hallucinate.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Why Language Models Hallucinate

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:28.235129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:28.235129Z digest=sha256:486bc5a71e0f459c8013b49400c348806f9ec9ae8bd46860dd5ac0495add4181

Observation fcd7a8fc-2e9e-4786-b250-2b065eb18798 · outbound

This paper cites Claude Fable 5 and Claude Mythos 5.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Claude Fable 5 and Claude Mythos 5

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:28.335808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:28.335808Z digest=sha256:fe625ff18b1abfb5d6491773fc60bba3b0c02b85fa79db6eed382b1f1008a1db

Observation 80a38cb3-1ead-4f0f-9e68-19d31dca1107 · outbound

This paper cites Artificial Analysis Intelligence Index.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Artificial Analysis Intelligence Index

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:28.505901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:28.505901Z digest=sha256:c7bfeb7db6abb46de165ebd4076bf5f0f5abdb4a44ce0d51504d2dafebc87e5f

Observation bf7420a9-0ec0-4c21-92ab-0a5a112220d2 · outbound

This paper cites an unresolved cited work.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:29.037989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:29.037989Z digest=sha256:9b48a88fa271ce3cbbffa4e34a498c2e93ebd7e8c89ed071f81d6b49a1011dc7

Observation 744271cb-af13-4ff9-bc2c-d1eb992b9f24 · outbound

This paper cites Artificial Analysis Intelligence Index: Methodology.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Artificial Analysis Intelligence Index: Methodology

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:29.254414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:29.254414Z digest=sha256:b38a5f0df6fb46ca6c94432979a8d0398653167a2eb45babb94f068201ae6a65

Observation b76a908b-d87a-4bce-a245-d025cda6e789 · outbound

This paper cites AssetOpsBench: Stirrup Agent.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains AssetOpsBench: Stirrup Agent

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:29.460428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:29.460428Z digest=sha256:867f49cdb1560da4b4473cf1b4017223246271236ac41b6d8ab03a64b388bd93

Observation a169a86c-7b9a-4bc1-8c97-bfb852aaa9e5 · outbound

This paper cites Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:29.577032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:29.577032Z digest=sha256:8c7781016e66c5c92390c3c49ec40f064faa686476fad12dbfbe4a9fe6eaa4fb

Observation 1d788fbd-d216-4eab-9658-8c918d5a007b · outbound

This paper cites Humanity's Last Exam.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Humanity's Last Exam

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:28.838974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:28.838974Z digest=sha256:e540025bc3b2014365302fa3d3dd2e65c4966b3e90d3dc2b8c2b3b2969bfd3a5

Pith citing papers

No inbound Pith citation observations are available.