Pith. sign in

Paper Citation Record · LEDGER

Coupled Variational Reinforcement Learning for Language Model General Reasoning

As of 10 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2512.12576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.12576 v3

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:43:58.144921Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3497387c-12ee-49bb-801a-3631be7a3bb7 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Process Reinforcement through Implicit Rewards

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.836145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.836145Z digest=sha256:34c4fa7d55834da080262da6bdf78171e236a0c76304f3fe0bb29ddb9cdfd60b

Observation 1a5dfe83-8427-40a7-b41e-8dadf5934bd3 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Coupled Variational Reinforcement Learning for Language Model General Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.952300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.952300Z digest=sha256:29638a01dd16de3b75843b3685a7616c84c897cb682a85de645465565e94de98

Observation 40e1fcc6-1eaf-44c7-925f-cc05714f5343 · outbound

This paper cites Auto-Encoding Variational Bayes.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Auto-Encoding Variational Bayes

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:55.183222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:55.183222Z digest=sha256:37fb6de0b96379a3d070b753626586ece8786412cb90b4caf8c35b3fdd2cea27

Observation 6c5b8c08-1601-4292-88d2-341ce52d9524 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Solving Quantitative Reasoning Problems with Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:55.364185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:55.364185Z digest=sha256:0d708efe4e08f02f277f361175cfdb9729a942732a43968e39bd9fb2f704b301

Observation 2290000d-250f-4ada-bb46-e38ec905da98 · outbound

This paper cites Ling, W., Yogatama, D., Dyer, C., and Blunsom, P.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Ling, W., Yogatama, D., Dyer, C., and Blunsom, P

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:55.743062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:55.743062Z digest=sha256:1dae3b4bcf09ab1e777d174186328e189ac8077f92a47d11adf694d6b7c72149

Observation f3cbd91c-f7c8-41f4-983f-73cbf847eeea · outbound

This paper cites General-Reasoner: Advancing LLM Reasoning Across All Domains.

Coupled Variational Reinforcement Learning for Language Model General Reasoning General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.063243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.063243Z digest=sha256:a41137fbaaf03ef976394f04f11874308a6dd58f851affeccf0ee648c6436dce

Observation 37f9f366-a0f3-4470-bc4d-9f6148da8e13 · outbound

This paper cites Con- tains 32 math questions from May 2023 SAT.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Con- tains 32 math questions from May 2023 SAT

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.213416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.213416Z digest=sha256:de01655dc08088345261ba6a0328144e47eeb0241b8d6c899d84146ee06787f6

Observation ca7b58c3-e35f-410a-ae60-dae9cfdb3602 · outbound

This paper cites Training language models to follow instructions with human feedback.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Training language models to follow instructions with human feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.345427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.345427Z digest=sha256:287c575d32080b52e0451a8f68c822e62c1fe1d37673ea5e56c807fcdfb2e937

Observation 70c9bf1e-8376-4080-876d-81ff21b63c2a · outbound

This paper cites Qwen2.5 Technical Report.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Qwen2.5 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.535548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.535548Z digest=sha256:bfcd895d229dfd79e23a91fddebaa0b5f72c449a19f0352a4cf57641b0a73f8f

Observation 03e68c57-8a5d-4486-ac96-719212089c52 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.609302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.609302Z digest=sha256:529359eaa6e9959bda85dfcfb067553e7c16662bb7bbf01dcc5ca2096b53f359

Observation cbc0d829-1989-4f9e-a83d-b0726aab6594 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Coupled Variational Reinforcement Learning for Language Model General Reasoning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.708534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.708534Z digest=sha256:a4ad3d1f8fe92dfb546bb35dbf8d0d2d55af79d30635a6a9cda6ae35648fea65

Observation e54c9422-7261-488f-8b32-076fb86ce597 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.772053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.772053Z digest=sha256:ddf3078fa57d39dc4b4a8ef8a6f2888b784b8aece14b09479b3062a623967eab

Observation 9d8ee088-8d68-4994-a5ee-8f3e5d9904f4 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Coupled Variational Reinforcement Learning for Language Model General Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.842532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.842532Z digest=sha256:0f2c8aaf96ec0ed8cc4b98946b726f4b539f66134548e3b31f8fc500ebc724ab

Observation 571222e9-820a-477a-ac3c-11adea58878b · outbound

This paper cites Defining and Characterizing Reward Hacking.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Defining and Characterizing Reward Hacking

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.945344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.945344Z digest=sha256:c6a46e0e9747aa57e429524426f16778fc371a49e3848ce2979da81b53bdd55a

Observation 9793575a-5edd-46e4-a648-7609fd8e5b4f · outbound

This paper cites Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.074776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.074776Z digest=sha256:a6166fb208bd0a0c5ba210fc83b17a245010c8e6156343117ed08c118e95addf

Observation 4097a892-2d58-475f-8c93-65fb44437048 · outbound

This paper cites Accessed: 2025-01-23.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Accessed: 2025-01-23

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.273034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.273034Z digest=sha256:f97b8b43ab78ec9b60cb947756f58a0f8c8aa0a899aec2046c8460ed60c30a49

Observation 1bb6e91e-f99c-434f-94c0-21869bc5e718 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Coupled Variational Reinforcement Learning for Language Model General Reasoning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.348223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.348223Z digest=sha256:89a4e68d564986c1e78c051ca0bd6233d9a9cee68f8bdebbceaf1fd4e97e7750

Observation 905da558-301d-4278-97b4-c12f208f8ab1 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.528598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.528598Z digest=sha256:bc63bcf9f710762be40771e69d194140d42956f766a4a04157d06c83cb42b12b

Observation 02882fce-fdb2-4124-8bec-9b46003db01d · outbound

This paper cites Qwen3 Technical Report.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Qwen3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.687482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.687482Z digest=sha256:f1813f5643161652a9ab3492821416162cd9cb000891f0a0f55328696845a6a5

Observation 6b66f49f-04bf-4706-aedb-d3cab0e19629 · outbound

This paper cites Self-Rewarding Language Models.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Self-Rewarding Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.799392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.799392Z digest=sha256:21be65be742a145ea18d2df0454c008f5a062b2b8af3ac485aa527d3a6c0cfb9

Observation aae23b8d-5928-4e2a-9202-60cc7bc8c165 · outbound

This paper cites MAmmoTH2: Scaling Instructions from the Web.

Coupled Variational Reinforcement Learning for Language Model General Reasoning MAmmoTH2: Scaling Instructions from the Web

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.892407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.892407Z digest=sha256:76205643228fb326da3a147f41754a3ae35cb568e9d415da78938052eb0eadd9

Observation 42998865-32a2-4ce8-880b-7e501606d255 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Coupled Variational Reinforcement Learning for Language Model General Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.954210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.954210Z digest=sha256:b6eaba30fbf7a8efbc3926821d22e8d782c8da3aab3f9bbb167a762916f64a49

Observation 4ad521f6-daf6-494d-ba12-97f25ef4b993 · outbound

This paper cites Evaluating and Improving Tool-Augmented Computation-Intensive Math Reasoning.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Evaluating and Improving Tool-Augmented Computation-Intensive Math Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:58.027815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:58.027815Z digest=sha256:fa10f8dd8e0b3e03d4066c1119e10e1f76775f6d905d171f6036b10479df1074

Observation 60c549bd-ce14-43bf-b44b-f2a72e3a1bf3 · outbound

This paper cites Zhou, X., Liu, Z., Sims, A., Wang, H., Pang, T., Li, C., Wang, L., Lin, M., and Du, C.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Zhou, X., Liu, Z., Sims, A., Wang, H., Pang, T., Li, C., Wang, L., Lin, M., and Du, C

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:58.082657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:58.082657Z digest=sha256:6382f7d1edb362568b47dc4e863f7aacbc73e4cb2d01538966e4dbdb14592212

Observation 44544868-b68d-4992-bbe6-c746fd65ae14 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Coupled Variational Reinforcement Learning for Language Model General Reasoning TTRL: Test-Time Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:58.144921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:58.144921Z digest=sha256:aa4a477cff3d4ee8b8ad643bb2f5924aec7016b1660a0b53e0ef1890fc992fd9

Observation 7358163e-dae2-4d17-86cb-5f3ed4e362ae · outbound

This paper cites Importance Weighted Autoencoders.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Importance Weighted Autoencoders

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.075263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.075263Z digest=sha256:e5ab2497a4fb93b3ccca2968db77241d58d12ffeba7d508e5181e52b4bae0f71

Observation f7cb0a4a-63d4-41e5-bbd3-3ede3214bafc · outbound

This paper cites Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:55.921050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:55.921050Z digest=sha256:6e930967a5b1e7d303d9bf3c95c1fe9526c0a4d10a85b501f3b0f816fb5dd41d

Observation c261e508-4925-434a-9d17-c4254acb339f · outbound

This paper cites A Stable Variational Autoencoder for Text Modelling.

Coupled Variational Reinforcement Learning for Language Model General Reasoning A Stable Variational Autoencoder for Text Modelling

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:55.552728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:55.552728Z digest=sha256:598e1fbc45af04a701669a3dfb66489e84085a565fdd68b3907fb2caba8ef32d

Observation 2ce08a28-16e3-462b-afe8-c0500fc395d6 · outbound

This paper cites Denoising Diffusion Probabilistic Models.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Denoising Diffusion Probabilistic Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:55.056382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:55.056382Z digest=sha256:49935e95b0dd1a024f21f765c0f8ef51a467847f0b146123498fffae8baced27

Observation 7cb0e51f-c2a5-4a38-9c2a-988bb9120003 · outbound

This paper cites Greedification Operators for Policy Optimization: Investigating Forward and Reverse KL Divergences.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Greedification Operators for Policy Optimization: Investigating Forward and Reverse KL Divergences

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.214636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.214636Z digest=sha256:4926c1364717c11c43dc7d3a7dfe9a9e27d664e3cc87ebf8c5d9b92e4d070812

Observation c0545efa-1a4e-4f6a-8a2b-a39b66788a1a · outbound

This paper cites TheoremQA: A Theorem-driven Question Answering dataset.

Coupled Variational Reinforcement Learning for Language Model General Reasoning TheoremQA: A Theorem-driven Question Answering dataset

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.710371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.710371Z digest=sha256:dc26c43b51001916535a0283ebee5a62ae4ed49b18d48a2af400477c565e3bb9

Observation e026c71f-ed98-4660-954a-bb8ca40016ef · outbound

This paper cites Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.497117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.497117Z digest=sha256:81ddbdef79047854f3106be129e34e649bd44a765756b23bee3eccba886b47ac

Observation 5b33c3c0-dcd0-458f-92a1-1ca91077afc9 · outbound

This paper cites Bootstrapping Language Models with DPO Implicit Rewards.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Bootstrapping Language Models with DPO Implicit Rewards

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.381657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.381657Z digest=sha256:44dc938fcbb91ba2ac6d5038f05aa5b12d91adf14494c1ba4ae1fedbecc96787

Pith citing papers

No inbound Pith citation observations are available.