Pith. sign in

Paper Citation Record · LEDGER

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks

As of 1 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 2 inbound Pith citation observations for arXiv:2604.11778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.11778 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:09:07.114453Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T01:54:10.297931Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-01T06:16:24.888925Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved12
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5f231b57-86a5-48fe-b968-864e781eff5d · outbound

This paper cites an unresolved cited work.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Unresolved cited work

Reference 1

Resolution
parse uncertain
raw_fallback, observed 2026-05-17T16:50:00.054632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:ad2b9a04515ee5008a82e2e315378fb5f38332d7c92b0150bff7b75701141119

Observation 60a60b64-f0d0-476d-9f80-f6ef5588bc30 · outbound

This paper cites an unresolved cited work.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:50:00.079054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:f79f2abc91b5ae1103c0d350dded928e9be0630fdcceee77a5e71e66e7dee4ac

Observation 99eecef3-820c-43a4-8a56-7b8031b53d8f · outbound

This paper cites an unresolved cited work.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:50:00.051508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:bba9a2288263dce6d9310e13078e24230ea1d60f7802b30be2aa4fe140bee3e2

Observation 30bbc333-6f03-4288-b29c-bfbf7e2d172a · outbound

This paper cites an unresolved cited work.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:50:00.045333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:d9dc70e71568cd0f4717f518b75e9b2003d6433c7a268ab9504dc4db9e31a76a

Observation 3cc22ba3-1315-4e74-ab48-1d3db397cdba · outbound

This paper cites an unresolved cited work.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:50:00.038911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:43138645af29513b8661b180a19269f7eb5be604a8061d479e089e7f6fdfd853

Observation 71abe3a5-efc4-4a9b-9136-6e61dd4c5832 · outbound

This paper cites an unresolved cited work.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:50:00.058078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:0cb608788e514c0a7e74d24824d7288f2e1e42c901dc9a17262ef4b89cb1c557

Observation 26fa5e3a-8b08-436e-8a95-9af0da2adf1a · outbound

This paper cites an unresolved cited work.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:50:00.068726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:aa5f7fd7038e034b9136a9c88253d8f60b801dd176ca8519b886bdda8c042c7e

Observation b2e3c2d4-a7a8-40bf-b70d-176dc2684400 · outbound

This paper cites Among them, there was one person who ordered espresso.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Among them, there was one person who ordered espresso

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:50:00.061823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:efe8d614ae8cd16e5dba1b8440796ed800b7640e6071ede0a1638ea4f4838b16

Observation ff24f900-9379-4be2-a4a8-f91c7b9f9e25 · outbound

This paper cites an unresolved cited work.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:50:00.065573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:9fd49daa168f28ebf8652c764ca1c1548ce464bcc6afde55abd22fed7a73aa81

Observation ae159013-4ef5-432e-82f3-ef0a49b2c5f2 · outbound

This paper cites an unresolved cited work.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:50:00.042242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:fcb68bb3c835b137842b92a5b681c7e25a2fac23bfee23bf67e6110da81ec840

Observation d69b5c4c-9671-433e-8259-7ddcccbd900e · outbound

This paper cites an unresolved cited work.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:50:00.082248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:f4c488d1469e7489d9a4a675faedd97716846fb8cbaff25c327db951f6dab75e

Observation 86a291f3-6929-4c6b-ae69-4adfeb1cbc22 · outbound

This paper cites an unresolved cited work.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:50:00.048413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:155eb9e805a6abd3c69d6fc1bb8cd46cb81765a7ec15bc957e360b75a2e86576

Observation 6fcd68d2-b7da-484a-a3f8-7a21107b98c8 · outbound

This paper cites an unresolved cited work.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:50:00.071733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:07a3becf6f3c902a2ce83e31d84ae38ed3f12da1a5c6ebedc237f4694be3c0f4

Observation cee4f2b8-0995-456d-b057-891ffbd2a569 · outbound

This paper cites an unresolved cited work.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:50:00.085456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:313e44ff3d283809cbe27b16e6997c53145dd957e6a8d1324186feafe3b07884

Observation 03374e5a-4991-441e-99a2-4c4afd0f327f · outbound

This paper cites We’ve lost a friend. The murderer is likely hiding in the crowd. Please be careful.

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks We’ve lost a friend. The murderer is likely hiding in the crowd. Please be careful

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:50:00.075935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T16:09:07.114453Z digest=sha256:a9681a53a54dc4d0a5dbdbb2b998b4ed198382aab52a9b5ac1e91577eedc7b6e

Pith citing papers

Observation 3cd5a9ed-3e9b-44bb-8875-079919e4b148 · inbound

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space cites this paper.

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-01T01:56:12.212710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-08-01T01:54:10.297931Z digest=sha256:1d374c57fac5f6932196b7627e7e9fb62934b0dc5d5e495b3a229e885b9dfa77

Observation 60424a0e-76ca-48fd-aae6-0c05591112d9 · inbound

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents cites this paper.

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T19:49:39.646529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T19:49:39.646529Z digest=sha256:3163b57c173a8d4686de64a3479cfcfba4e5663d7edee034da00d4e3e51f170f