Pith. sign in

Paper Citation Record · LEDGER

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents

As of 11 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2605.08721.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.08721 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:59:13.198980Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cad5851-df3c-427d-b3b2-78398015ea4e · outbound

This paper cites TextArena.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents TextArena

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:46:43.663160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:3974f6253fe7dddbe4cdc7a05c764cd454fd6c8deeff1ba2f2ce975ba3f0cc8e

Observation 82090acd-4b38-49fd-97da-c4d435347163 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:46:43.680046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:e932fcf4ed9da04fe3bfa2ccc80d51679558d8420f7494edf760babe2ca3b769

Observation 431c2fb2-bbba-4a20-8eec-e496e562c7f3 · outbound

This paper cites Qwen3 Technical Report.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents Qwen3 Technical Report

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T06:46:43.668302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:713a7334cde716ddf4a6aff7ac9d5981d9bd9e3f52a1b3fc3fb7f21a216205d3

Observation 8a7a4dd3-ee0f-49d9-8b87-4e708599fdad · outbound

This paper cites We select the hyperbolic tangent kernel: σ(t) =P(S |δ (t)) = 1−tanh(δ (t)).(19).

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents We select the hyperbolic tangent kernel: σ(t) =P(S |δ (t)) = 1−tanh(δ (t)).(19)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:16:44.263645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:f35ffcce9557fd217fedf90ef8f6596708201bf814f134f34d7b28f37d43a1ab

Observation 9c580a0b-0795-4208-a5c3-f1019166d0ce · outbound

This paper cites (20) Substituting these specific kernels yields the instan- tiation used in DEPT:λ (t) =σ (t) ·γ (t).

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents (20) Substituting these specific kernels yields the instan- tiation used in DEPT:λ (t) =σ (t) ·γ (t)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:21:42.514143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:5edbd063e3339dcad81fd9befb930ad77078a98983dc5eed7bbaec1a65be0bfd

Observation 16cb49d1-f37f-4326-b177-64804ece2077 · outbound

This paper cites Thus, the advantage for the M dominant samples ap- proaches zero.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents Thus, the advantage for the M dominant samples ap- proaches zero

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:16:44.270619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:028a5d35ee5ec103b611d3b7f4fe888d9903aaa01e10c7e8ece92f7ac0da39c6

Observation b39a3ee6-2275-412c-bb41-1156f134116a · outbound

This paper cites Since Vmax ≥R p(τ) for τ∈ D dom, the term (Rp(τ)−V max) is strictly non-positive.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents Since Vmax ≥R p(τ) for τ∈ D dom, the term (Rp(τ)−V max) is strictly non-positive

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:21:42.496179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:612a44594dd26262c83751dae2d769dee601cd0c60871a3ef2664178d3f20c06

Observation 48f3a4ff-26f5-4f7d-a4af-4ef869cfb9bb · outbound

This paper cites The term (Rp(τ ′)−V min) is maximized, as- signing a high positive weight to these sparse signals.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents The term (Rp(τ ′)−V min) is maximized, as- signing a high positive weight to these sparse signals

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:21:42.509147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:3e7730daa8f55a33ca75d7edd057fedcc4dd49fa5a4e1456fbcff7abc7d6f3cb

Observation 2f249552-9e3e-4e32-a3a0-9d0774f9e6c9 · outbound

This paper cites These benchmark cover a wide range of topics including algebra, geometry, and competitive mathematic.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents These benchmark cover a wide range of topics including algebra, geometry, and competitive mathematic

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:16:44.267528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:801104fb818ecc4aaf0b929b9c9f1d0b1dbadac4ac68ac69dc6960c60029926f

Observation bb0b6bde-3668-497d-bccf-2ae639e91b3d · outbound

This paper cites I think this is fair because ... [ Propose ] $X . XX \.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents I think this is fair because ... [ Propose ] $X . XX \

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:16:44.273806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:3b41d74709a52ca014e7cdf82bfd646f3ba02e7a1a45173b9e77a37f81ba95b3

Observation 302a32a7-d7b9-46a8-aa0b-d416678ab315 · outbound

This paper cites an unresolved cited work.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-12T17:21:42.503866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:3c583cebf608ef0bfaca702164d1b0a06057e1fed0dc0a95ccbcbcd49f3e5a93

Observation 49df0166-9fba-4738-aebe-e3279c1b1253 · outbound

This paper cites an unresolved cited work.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-12T17:21:42.543660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:8a6bf32865ca78c29253a042d2d0edc3cf24a1cb3a2d4acad81d20ce1fd4068d

Observation 448ad6b4-0ef4-4c6d-bf58-66017103a6bb · outbound

This paper cites comb " in a context that must naturally arise during the conversation . For example , if you are discussing hair care or grooming , the word.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents comb " in a context that must naturally arise during the conversation . For example , if you are discussing hair care or grooming , the word

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:21:42.557399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:f39be036b429cec577700ff18702f0ce8a152a5f99997d609cd30cbbaba5ad40

Observation 6978db0c-967f-475a-8092-6eda30fdc18d · outbound

This paper cites an unresolved cited work.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-12T17:16:44.277303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:8be82f705989c08226d7cbb2bddcfaf440429cdfa4a221a0ccbf90359bed26f2

Observation a087d721-eac2-4f96-a582-897383dd0a15 · outbound

This paper cites an unresolved cited work.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-12T17:21:42.560930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:eb4dafd5f618872a5b9026d6c2b1cc98f65b2d3ab60e864224de9b6015f00b5b

Observation ae9d90af-ed9d-494e-9b67-fe50b83ec2ae · outbound

This paper cites I think this is fair because ... [ Propose ] $X . XX.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents I think this is fair because ... [ Propose ] $X . XX

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:21:42.531411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:684557b44ab9b9b58ac052c4aea1ef93acc1ce36281ed183b791d3171e86bb8a

Observation 4960f535-30d1-4357-9f15-116136126590 · outbound

This paper cites - This means Player 1 would receive $0 .01 of the total $2 .00.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents - This means Player 1 would receive $0 .01 of the total $2 .00

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:21:42.535189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:94206d0c0d75d1a0f49edb557e4fc5169eaba8a2df60c235021b5cacbf3c1407

Observation 844158cb-a80c-40eb-8845-ccd689c1a0ce · outbound

This paper cites They would get only 0.5% of the total $2 .00 , which is $0 .01.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents They would get only 0.5% of the total $2 .00 , which is $0 .01

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:21:42.525509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:ad546d454f0bc768447a2905f380288c3393e8226150abbe8778bcb7223915ed

Observation 2bd470fe-ba43-4f03-9016-a32af0221924 · outbound

This paper cites - Accepting the current proposal would result in Player 1 receiving $0 .01 , which is far below their required $1 .60.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents - Accepting the current proposal would result in Player 1 receiving $0 .01 , which is far below their required $1 .60

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:21:42.538882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:93dde3eb5720f8b78199c3687330b0b607fedd993a8da36f77c67bbd2e51576a

Observation e963184d-5e71-4446-8eb8-8eebe0173228 · outbound

This paper cites - By rejecting , Player 1 maintains the option to propose a better deal in the next round or wait for Player 0 to make a more fair offer.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents - By rejecting , Player 1 maintains the option to propose a better deal in the next round or wait for Player 0 to make a more fair offer

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:21:42.553081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:fae67f21b9d3289e0c419e828aaba6ff2b1b4605795f73415f3829f0883d84e0

Observation f7968f5a-765f-4791-b4b3-d8594bb9754c · outbound

This paper cites Player 1 would be worse off than refusing to cooperate at all ( which would result in $0 .00 for both players ).

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents Player 1 would be worse off than refusing to cooperate at all ( which would result in $0 .00 for both players )

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:21:42.519542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:9cc3a1c02a410a9d8b6249f813284f3f7131ebab4328f97549928fe0bd73e619

Observation 5d6f9fcd-a8ba-481e-9017-3c333facb65d · outbound

This paper cites For example , a proposal like $1 .60 for Player 1 and $0 .40 for Player 0 would satisfy Player 1's instructions.

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents For example , a proposal like $1 .60 for Player 1 and $0 .40 for Player 0 would satisfy Player 1's instructions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T17:21:42.548185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:59:13.198980Z digest=sha256:a95fde276f5ff8ceb7ffce8f0523957562ebbcd35358d195c24250c075714ff0

Pith citing papers

No inbound Pith citation observations are available.