Pith. sign in

Paper Citation Record · LEDGER

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

As of 18 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 16 inbound Pith citation observations for arXiv:2501.18099.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18099 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:44:45.284564Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.749685Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:56:47.550212Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc164337-acc7-4956-877a-17e8d889379b · outbound

This paper cites In this case, the function should be named ‘separate_paren_groups’, take a single parameter ‘paren_string’ of type ‘str’, and return a list of strings (‘List[str]’).

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge In this case, the function should be named ‘separate_paren_groups’, take a single parameter ‘paren_string’ of type ‘str’, and return a list of strings (‘List[str]’)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:44:45.642788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.196446Z digest=sha256:e031953b380236ed2cbcd72572a44e3956d376e62033ca9a98af60af8c123fcf

Observation fc3a7f91-7d8e-4a5c-9fbf-50b7ce7359f5 · outbound

This paper cites Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T00:44:45.185254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:44:45.185254Z digest=sha256:e84d93c88b0d17ce63effd390e9a961cbd35e8c941e3e85a3e3da3133f368bd1

Observation aac8be9a-6223-4361-9a41-8afb072704ca · outbound

This paper cites RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T00:44:45.190432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:44:45.190432Z digest=sha256:ad9002cb017c9773e34bcae9b0b469a08f8d3c7bb520f3915b28a222ec88f7a7

Observation 6c727024-d7a7-4fb0-a190-0a7021a27b6e · outbound

This paper cites an unresolved cited work.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:44:45.596828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.210742Z digest=sha256:a268096db2720c928e68de10ba5abc33866f6164817a2fc7452818fb416b62ca

Observation b3765267-df56-4930-8297-8710464bb7ed · outbound

This paper cites an unresolved cited work.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:44:45.627559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.201183Z digest=sha256:5726da71fc897bad5292098cf313b8ac7196837db9f900487c90f776897b1d95

Observation 3c784ed5-bf1a-4cc1-8bea-3c432a48f152 · outbound

This paper cites an unresolved cited work.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:44:45.611864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.205885Z digest=sha256:cdaac5c03999252b06c77f728a1d3826ea4875496841bb7ad71eec2e10ef37d1

Observation 89bcefaf-c433-427b-a933-45091b76336a · outbound

This paper cites Check for proper use of comments, variable naming, and function structure.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Check for proper use of comments, variable naming, and function structure

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:44:45.582498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.215938Z digest=sha256:bc16c5cbf5905de340eb1f2ee510c823602e91413390628f3ccf50039be85741

Observation b759105a-60cf-4967-a623-5aaaba004efb · outbound

This paper cites Figure 6 Example of a plan generated by EvalPlanner for a coding problem.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Figure 6 Example of a plan generated by EvalPlanner for a coding problem

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:44:45.566267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.221383Z digest=sha256:b915144a7c3363b8e6e31be52fbf77fa9115256e5695ac8e81f324a071f33c93

Observation 3816d3f3-aae8-474b-8d1a-c784f297a9a6 · outbound

This paper cites an unresolved cited work.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:44:45.549255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.226813Z digest=sha256:778ef31d990d6360b83c2732cefbd8208ffa4f1935459ac00ba2fb9afa9c9873

Observation 83096b4d-49d9-4e99-9ac5-26dd6c64ba1e · outbound

This paper cites an unresolved cited work.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:44:45.532363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.233586Z digest=sha256:3606d0a2daa2307534d4a243a8545b70a5dc44554e601175046aa357aeebe144

Observation 7002d404-313b-4e1c-93b6-39d5fff4302e · outbound

This paper cites Knowing ̸ A = 14◦ and ̸ C = 90◦, we can find̸ B.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Knowing ̸ A = 14◦ and ̸ C = 90◦, we can find̸ B

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:44:45.515774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.238466Z digest=sha256:3005cd53f1b448e37cd2fe0d84d12badaab4bc69a33094e93efc51358062340a

Observation 0f0d993d-f168-4ffa-95e1-fbf6f4142c36 · outbound

This paper cites an unresolved cited work.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:44:45.501054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.243695Z digest=sha256:9a40d21e2b5be617bd503f570c92e80681f06ebe31aad7419d8cea4bed0e1053

Observation 14acd250-4b11-4805-a9c4-fa62e802e47a · outbound

This paper cites an unresolved cited work.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:44:45.486770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.248724Z digest=sha256:4396e4af4f65ee2104fbf42bbfb5e1435f3ce5d92c907dec618fe784a34b02c8

Observation c09d396a-ea1c-4fd0-b038-110f1c87bc03 · outbound

This paper cites an unresolved cited work.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:44:45.472885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.254474Z digest=sha256:d9f8517da9cc849ca1f563bb46aa3abdf3a133a7cc53aa7703bdfdc6057dac34

Observation 2caba5da-62dc-44d2-9480-fb7f52be9935 · outbound

This paper cites an unresolved cited work.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:44:45.458337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.259298Z digest=sha256:66884ff136f242e398d68af172d9ed9d3fbcc2964fcab004407766155c5120d0

Observation d5fdd007-237b-4342-ab8a-202890c98abf · outbound

This paper cites an unresolved cited work.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:44:45.440784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.264359Z digest=sha256:e958461c346887a7ce9cf618697cac7cef1cb616c0ed08091ff56ea9f048876d

Observation 850dda19-7eb2-4a4f-83fa-72047b29d339 · outbound

This paper cites * Avoids perpetuating harmful stereotypes and discrimination.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge * Avoids perpetuating harmful stereotypes and discrimination

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:44:45.425482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.270117Z digest=sha256:586cc44ac48a735cea49e0f902294c45005ad724532bb5d652a26caaa1b1e687

Observation 7a7ffd55-29bb-4999-bb09-13417d47b104 · outbound

This paper cites an unresolved cited work.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:44:45.409898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.275053Z digest=sha256:6b0d94c64f5b8c2bfaf2979e4a60031d1af404078ec0499155084ffc519ef14d

Observation 61129db8-2202-4ccb-bbb0-37fd128bf0f7 · outbound

This paper cites * 2: The response is somewhat discriminatory, but attempts to be respectful.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge * 2: The response is somewhat discriminatory, but attempts to be respectful

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:44:45.393993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.279607Z digest=sha256:975b93e860f77a72b3422f73a23a005ca84da0b96980b0195ec88f5275b1f703

Observation c034370c-da07-49c1-95ca-68d96c87c9a5 · outbound

This paper cites an unresolved cited work.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:44:45.377163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T00:44:45.284564Z digest=sha256:341d096f904a6a79ccb380d98ba7d8c82aa121fd06b6d172da92299380dac7db

Observation 9d13f608-d47d-456d-a61d-ec3b2068612c · outbound

This paper cites The Llama 3 Herd of Models.

Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T00:44:45.179119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:44:45.179119Z digest=sha256:3ffefec383ea6e3c4736552df4a9695d3cf73ed59c5a8ecedfbc56a1626e3c7c

Pith citing papers

Observation 54c8fb5b-035c-4630-bfd9-ced69c32ee5a · inbound

A Survey of Scaling in Large Language Model Reasoning cites this paper.

A Survey of Scaling in Large Language Model Reasoning Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 165

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:22:09.202737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T21:20:07.238992Z digest=sha256:e175bb8f22506fabb24fd0c7e5c9f1a98987be08d4e24127f74b99114b7a4496

Observation 09153926-605c-47f7-8fc8-4015d7e61189 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 291

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:43.749685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:43.749685Z digest=sha256:aaa005926abd84cb2e62d828e0f43b3ca975003b60234c0424e8b9bbda6d5f83

Observation ce64f43f-f56f-4ceb-a41b-946b213c5fc1 · inbound

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models cites this paper.

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T05:40:09.399353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:40:09.399353Z digest=sha256:2223f66d9da9b8af3cada4a7abbcb5b1023a940a806f2a0f7dd23274a3257b12

Observation 2dc70ec6-264d-45fc-83fe-70a0567d2b40 · inbound

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge cites this paper.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.636826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.636826Z digest=sha256:5831cf10fc14a920fb33b030e2b3ba77df804b81c1040f523b5626a659ae5873

Observation 701ad2f1-fb4f-4a3e-a3d0-0bbfb0027a9c · inbound

AI vs. Human Judgment of Content Moderation: LLM-as-a-Judge and Ethics-Based Response Refusals cites this paper.

AI vs. Human Judgment of Content Moderation: LLM-as-a-Judge and Ethics-Based Response Refusals Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:32.458219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:32.458219Z digest=sha256:ec3b2fd20ca14b1169217120a25ba1d97e864c21a51d15efefee736a8daee747

Observation 39e72395-1558-4edc-b94e-87487b779188 · inbound

LLMs for Customized Marketing Content Generation and Evaluation at Scale cites this paper.

LLMs for Customized Marketing Content Generation and Evaluation at Scale Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:04:00.287318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:04:00.287318Z digest=sha256:5e2cb53b47dd041d37b91893df4e4dcb049b14c05c08abb50b1cefbdc1cf6b0f

Observation 626726f4-65da-4bf8-9c85-17976b960876 · inbound

Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillation cites this paper.

Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillation Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:40.965435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:40.965435Z digest=sha256:d4d1655f8b1d47d2b17102074887509d004fab21dc065c19bb23f60f3fac75f0

Observation 2055a67a-652b-4683-a375-b6c463c0399d · inbound

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.093199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.093199Z digest=sha256:778d77b9007b706257509847eb58129110eb636fadb3123dd17c1a6819ea61e4

Observation 8b37e988-2b23-49d4-8efc-859f6ac5108d · inbound

LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers cites this paper.

LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:03:08.127743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T17:01:47.603706Z digest=sha256:b8d32d20b74ecc075f9fdb5cbcb8f7c071a1b1c39ad25701a69e6a18f1a9c8b0

Observation 636125d6-9a04-4850-9dd3-4878cc946113 · inbound

LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection cites this paper.

LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:45:48.088784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T19:37:19.360378Z digest=sha256:30d6f7a92429ddfd353d34c51b4a445044a3131a8c84b441a8f682f972042426

Observation 5f890e66-3a40-4204-a50f-f731830dcea3 · inbound

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training cites this paper.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:20:57.479472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:1ec886bfe61bc49b9e3998ee6d5339f269a1d901b349241349808cb9ce6daa4a

Observation 747f9058-ca33-4cf5-afe8-9836c47004b1 · inbound

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints cites this paper.

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:15:29.423793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T14:12:45.438246Z digest=sha256:778615d132a57d37708ffb4d6381337c93507a759d9a54b664ccf5fc51d5dd5d

Observation bc22ae39-7286-4b5c-8a06-5298f4e1d6f4 · inbound

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge cites this paper.

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:23.805931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:19:00.238602Z digest=sha256:ed2955fdefc97164b01bc74d2969cee9673cf06b894b0b822e9720eb97759449

Observation a335bbd2-9190-45a0-ae7b-07cd06343c15 · inbound

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling cites this paper.

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:27:07.043395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T02:25:58.830629Z digest=sha256:cee8fe28b235f8ba2e3f00a664da1246071818d5c848357b2d6f4a5ed5146223

Observation c6b5f40b-3b95-4ffb-bcbb-8a39bc641019 · inbound

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling cites this paper.

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:47:58.147912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-14T20:46:38.558668Z digest=sha256:6d4ab77ad0d25f598dd732ce7541be2359f760ee3dd0b08e0e9a1e97fd8c0c6d

Observation f90c6286-572f-408a-ad7c-bbe402d251db · inbound

Plan First, Judge Later, Run Better: A DMAIC-Inspired Agentic System for Industrial Anomaly Detection cites this paper.

Plan First, Judge Later, Run Better: A DMAIC-Inspired Agentic System for Industrial Anomaly Detection Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:56:47.551595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T06:31:44.815249Z digest=sha256:8ccceb2898914f3546816dcdd7e0929917e0a0330aa0a1942100a34e9dc2f1ad