Pith. sign in

Paper Citation Record · LEDGER

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery

As of 19 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 2 inbound Pith citation observations for arXiv:2605.15412.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.15412 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T14:45:07.168330Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:14:12.505221Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T20:52:58.379963Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact33
  • verified fuzzy49
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ad706de9-3d00-405e-affa-71e7dad3d3ce · outbound

This paper cites AutoAlpha: an Efficient Hierarchical Evolutionary Algorithm for Mining Alpha Factors in Quantitative Investment.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery AutoAlpha: an Efficient Hierarchical Evolutionary Algorithm for Mining Alpha Factors in Quantitative Investment

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.272401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:c45e686d28eec517db2f4307fa02ba3dc4cbf0b933e6ce8690b49c9ac76153f5

Observation e781bf42-1358-4299-8970-2b8d308615ba · outbound

This paper cites Alpha Mining and Enhancing via Warm Start Genetic Programming for Quantitative Investment.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Alpha Mining and Enhancing via Warm Start Genetic Programming for Quantitative Investment

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.266455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:2d8ecd480456ccebaab02a4d3a6d4eab9780fff9aac460efd14339292089f5f9

Observation e21024ea-f425-4410-b0cf-d675a6fba8c8 · outbound

This paper cites 101 formulaic alphas.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery 101 formulaic alphas

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.648682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:ae2297d57e9007181a1122afec5c9130af5ddaeda1b891cea590a599db9edc31

Observation 1175895f-f6de-4c91-92cd-a73ec512ccc3 · outbound

This paper cites Multiple regression genetic programming.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Multiple regression genetic programming

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.693887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:13da9dd13da4a185210da81b6b4d0c00d57614d76220293d1ac16febb9042fc3

Observation 54b6fef3-1146-4bbe-ae80-2d70a2e05c0e · outbound

This paper cites Alpha discovery via grammar-guided learning and search.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Alpha discovery via grammar-guided learning and search

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.336651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:a9266e546c209d82db1758eda7003baac7bf97de48172f230c25954bf332d6e7

Observation 31cc164d-9f20-4563-b887-be948259d8cb · outbound

This paper cites Riskminer: Discovering formulaic alphas via risk seeking monte carlo tree search.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Riskminer: Discovering formulaic alphas via risk seeking monte carlo tree search

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.621250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:f86028402647ff890a57d3834476d938cd965f9c29fc577d3c2517a06459b7ae

Observation ffee040b-b83d-4847-b571-9887465c511a · outbound

This paper cites Generating synergistic formulaic alpha collections via reinforcement learning.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Generating synergistic formulaic alpha collections via reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.607853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:d0c725ab48ad9e8fd505e5e3d05751df86f70de5f2b0af0d6d5c84a049a93ac6

Observation e4baaaae-e0f3-47c6-826a-6d6693f28192 · outbound

This paper cites $\text{Alpha}^2$: Discovering Logical Formulaic Alphas using Deep Reinforcement Learning.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery $\text{Alpha}^2$: Discovering Logical Formulaic Alphas using Deep Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.217241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:c0840bf1cb5a15d5d983c36b5d6d37b058a417d511d88a67f87f189a26a7f160

Observation bd243977-0bc9-41db-a80a-f14fece1479a · outbound

This paper cites Alphaqcm: Alpha discovery in finance with distribu- tional reinforcement learning.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Alphaqcm: Alpha discovery in finance with distribu- tional reinforcement learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.625169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:07258cdbccc77d9ff2eaef1da6eb0a892c5274002c601eb6c5b9c243984a5822

Observation ba3258be-e696-4aa8-b0ba-1edb0572ddb3 · outbound

This paper cites Alphaforge: A framework to mine and dynamically combine formulaic alpha factors.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Alphaforge: A framework to mine and dynamically combine formulaic alpha factors

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.644755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:926b5780f9dbb87a53e94bb6af33fdf48b0c2968824a7965087dd9a356d59134

Observation bcbcdf6a-56ff-4321-9d10-842b64754d03 · outbound

This paper cites AlphaSAGE: Structure-Aware Alpha Mining via GFlowNets for Robust Exploration.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery AlphaSAGE: Structure-Aware Alpha Mining via GFlowNets for Robust Exploration

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:04:45.338223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:5bd3345d122c37e974aa7be55e96626fad3395fc80bbadc8b4f047dcd2675c59

Observation dec02696-8e36-4a3d-b868-8edfc9f5239e · outbound

This paper cites A survey of aiops in the era of large language models.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery A survey of aiops in the era of large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.684134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:0ad0dda9b49af4ce433ad686f5cf6297659cf8136b1613224ccc819651a2e0f6

Observation 0eec8e55-d3d9-4b5f-9eea-6eaba84072c1 · outbound

This paper cites E-log: Fine-grained elastic log-based anomaly detection and diagnosis for databases.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery E-log: Fine-grained elastic log-based anomaly detection and diagnosis for databases

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.611877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:fa9e90f9954571d6dbf206efc2af8d74fe74ccc4878faf4ea3d10542e1c7587b

Observation dbde6636-c8d6-4f6f-a217-f0b67d1d6491 · outbound

This paper cites Towards close-to-zero runtime collection overhead: Raft-based anomaly diagnosis on system faults for distributed storage system.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Towards close-to-zero runtime collection overhead: Raft-based anomaly diagnosis on system faults for distributed storage system

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.617376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:28ab75bc583531615b4c51b16564e6eede123790b20a0cc61ecae460f934e118

Observation b225c227-f418-46b9-bbd1-8c8efe59ac6d · outbound

This paper cites Multivariate log- based anomaly detection for distributed database.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Multivariate log- based anomaly detection for distributed database

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.652270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:6afc3210895669ba8b789b8b027ff4c21161eab1667d4ac371452f07484491f5

Observation e8cce409-02a4-4da9-bae1-736d4059cd47 · outbound

This paper cites Reducing events to augment log-based anomaly detection models: An empirical study.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Reducing events to augment log-based anomaly detection models: An empirical study

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.603732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:a418c8a44ff1550e3a6eb877f89b4fd76c8eed5ba497a1749c5529cfa29f0fc3

Observation 7262480a-dbba-419b-bcc0-14443108ba16 · outbound

This paper cites Scalalog: Scalable log-based failure diagnosis using llm.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Scalalog: Scalable log-based failure diagnosis using llm

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.605748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:f9270fbb29ee17dfd6c368a2fceed62c13f0cf39dd9a8665d5f5923b0769e97b

Observation bffb9279-de5b-4e4d-bfe3-d3b07fd8d513 · outbound

This paper cites AgentFM: Role-Aware Failure Management for Distributed Databases with LLM-Driven Multi-Agents.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery AgentFM: Role-Aware Failure Management for Distributed Databases with LLM-Driven Multi-Agents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.253998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:3e9252a138a6f6fffdd388a3bbfe60085fe0612d044e81772c6158cdf6deb636

Observation a8ea2feb-2f05-41c8-936e-0d9ca71d2598 · outbound

This paper cites arXiv preprint arXiv:2504.18776 , year=.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery arXiv preprint arXiv:2504.18776 , year=

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T14:47:36.263109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:53c767490cbf16f41f9859b6169d70326cb99c3f0cded357f891c4a07d3e89e2

Observation cbed4980-c3c7-4275-b7b0-935ac356ba52 · outbound

This paper cites Agentic Memory Enhanced Recursive Reasoning for Root Cause Localization in Microservices.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Agentic Memory Enhanced Recursive Reasoning for Root Cause Localization in Microservices

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.230561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:01122537e3218159ddf0052e539f4974e0dcf1d07126ed452959a54322949f77

Observation 7510fb1b-d0b8-4f98-8ff8-d09fff1c83b0 · outbound

This paper cites LogDB: Multivariate Log-based Failure Diagnosis for Distributed Databases (Extended from MultiLog).

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery LogDB: Multivariate Log-based Failure Diagnosis for Distributed Databases (Extended from MultiLog)

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.315659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:7c2a387936cbb46fc5871a2b85cf7892e1ba96318278262af144fe2a5229cbf8

Observation 51a10d05-bad2-4bed-ae36-158f68ecd2f1 · outbound

This paper cites Xraglog: A resource- efficient and context-aware log-based anomaly detection method using retrieval-augmented generation.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Xraglog: A resource- efficient and context-aware log-based anomaly detection method using retrieval-augmented generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.682246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:55ef4310fd01b6f948caae942fb7f3fe3c9464c3b90490345becabd2fac74b7a

Observation 146b82ad-a823-4d51-9708-1dedeeb56e66 · outbound

This paper cites A survey on parallel text generation: From parallel decoding to diffusion language models.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery A survey on parallel text generation: From parallel decoding to diffusion language models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.320979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:b64db4eada69ac85d5fad6d99032c255096335d0b900bc60210e23c3349f2164

Observation d2aeabda-f564-4e52-a0b1-1b231ea25bb7 · outbound

This paper cites Time-tired compaction: An elastic compaction scheme for lsm-tree based time-series database.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Time-tired compaction: An elastic compaction scheme for lsm-tree based time-series database

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.704664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:68c0d33fb5fb03a27879ac0a19ecab48878ab98f4b7c73c7c04d7145fee8b55d

Observation cecef55d-a8ab-4323-980d-659d71de94ea · outbound

This paper cites Separation or not: On handing out-of-order time-series data in leveled lsm-tree.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Separation or not: On handing out-of-order time-series data in leveled lsm-tree

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.710661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:b17ee23e0ee2800a14910e6a241db829064ca90a0a22101ca7225b6a447b5c47

Observation b5a30f81-95e5-4415-be16-90a05fe64a64 · outbound

This paper cites Adaptive Root Cause Localization for Microservice Systems with Multi-Agent Recursion-of-Thought.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Adaptive Root Cause Localization for Microservice Systems with Multi-Agent Recursion-of-Thought

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.295516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:d4c63f7db3726f6972e0275fe57ed8c10927fefaf5c40db97869d56a2ae8fab1

Observation 293a8778-4fb8-403b-9aab-84eee2fa95ed · outbound

This paper cites Ora: Job runtime prediction for high-performance computing platforms using the online retrieval-augmented language model.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Ora: Job runtime prediction for high-performance computing platforms using the online retrieval-augmented language model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.712629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:220230c2200bb143b67ea6384bce95345c1feb74113f0698b11a826e1c48d775

Observation be40c601-3745-4031-a8cf-d40a1e6c22ae · outbound

This paper cites Microremed: Benchmarking llms in microservices remediation.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Microremed: Benchmarking llms in microservices remediation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.331849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:20e0ad05c3b6cba81607b891b01cebeb0ab3d74c1c840157d16ea84f1f16aa64

Observation 02e9ae15-a085-4bf8-b708-c5388c5d870e · outbound

This paper cites arXiv preprint arXiv:2508.07173 , year=.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery arXiv preprint arXiv:2508.07173 , year=

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T14:47:36.281297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:7a23c01c3677b7352a41a9aade04b96fe106ce2120042f78e8f532e259a9393a

Observation 36013edc-89f9-439d-81b8-2ba86d44f319 · outbound

This paper cites Walk the talk: Is your log-based software reliability maintenance system really reliable?.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Walk the talk: Is your log-based software reliability maintenance system really reliable?

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.257289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:ed846ef5c2d7bed3294734711573a7b395a5da4ee9eaceb8858b2b697c739d48

Observation 71e4a089-2448-432b-a1d4-6acee72be0f2 · outbound

This paper cites d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T14:47:36.286767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:b8b207a3888d8f88efa3d18a9774d1654e31b71a1982ea34cef1a6db73812b38

Observation b81409c5-5a0d-4953-a896-57420c71ff32 · outbound

This paper cites Cslparser: A collaborative framework using small and large language models for log parsing.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Cslparser: A collaborative framework using small and large language models for log parsing

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.676998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:8de79c627cc88a4febf07f0171ccf4da6b61cc0b94b4a1a55003005aeb5a8b3f

Observation 3d6c43c9-90a4-41bb-a5f2-25d05d183aa3 · outbound

This paper cites United we stand: Towards end-to-end log- based fault diagnosis via interactive multi-task learning.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery United we stand: Towards end-to-end log- based fault diagnosis via interactive multi-task learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.209258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:2610d322c613f41418bee8e8a9e37742cf3f767c8a4533ff260431f3348efa20

Observation 4ddcc394-6fb3-450c-b87c-b2dfb1885b5c · outbound

This paper cites Hypothesize-then-verify: Speculative root cause analysis for microser- vices with pathwise parallelism.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Hypothesize-then-verify: Speculative root cause analysis for microser- vices with pathwise parallelism

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.243734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:a8f75d3ff4378f5ecafb0c354373eaa0a0f888f5576df4b89cdd52308349e8a5

Observation 1c3cb501-4de4-4bc4-bd7a-1b923e0f2b38 · outbound

This paper cites Uda-rcl: Unsupervised domain adaptation for microservice root cause localization utilizing multimodal data.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Uda-rcl: Unsupervised domain adaptation for microservice root cause localization utilizing multimodal data

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.629412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:0c62f0a1d347befcc7093f72eea9f79059b8023b034e9bd21df245dd96d93fdb

Observation 0b112839-e8ac-4dae-a22a-177d80b8e05c · outbound

This paper cites Aaad: Asynchronous inter-variable relationship-aware anomaly detection for multivariate time series.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Aaad: Asynchronous inter-variable relationship-aware anomaly detection for multivariate time series

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.619340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:ad9f3df7161421938ce78133def221a380a20928735b685562f07c8a996364ef

Observation 803c6591-bdba-4362-a8dc-f4cb7c2120db · outbound

This paper cites Logaction: Consistent cross-system anomaly detection through logs via active domain adaptation.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Logaction: Consistent cross-system anomaly detection through logs via active domain adaptation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.623196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:eba4a2f12d688123c53d65bd8a14faffd99710573bba881708ed8f4c98fd1997

Observation a914dd5a-e2e6-4d71-bc87-72b58fb5fa4e · outbound

This paper cites Runtimeslicer: Towards generalizable unified runtime state representation for failure management.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Runtimeslicer: Towards generalizable unified runtime state representation for failure management

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.342208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:8158fec6ce0cd8e64c1acc64ba045d2a74a53943b5b40e5434e52081a0ea4866

Observation c5a91968-3e0f-4bb7-a8de-f55508b5c774 · outbound

This paper cites Efficient failure management for multi-agent systems with reasoning trace representation.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Efficient failure management for multi-agent systems with reasoning trace representation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.212457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:ce7c6f2576d747f8ef78102c9fe81f18a426b1290c50e11b80468987c4d3388d

Observation 79545063-c56f-426f-95f6-3602a32376a1 · outbound

This paper cites E2E-REME: Towards End-to-End Microservices Auto-Remediation via Experience-Simulation Reinforcement Fine-Tuning.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery E2E-REME: Towards End-to-End Microservices Auto-Remediation via Experience-Simulation Reinforcement Fine-Tuning

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-19T14:47:36.289553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:079b0ea128d360a1b98683e449736688221dc4a1b5b3b143f2d0ceb9dda74976

Observation d3f61822-85be-498e-8c91-77b0af6e515b · outbound

This paper cites Coorlog: Efficient-generalizable log anomaly detection via adaptive coordinator in software evolution.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Coorlog: Efficient-generalizable log anomaly detection via adaptive coordinator in software evolution

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.639158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:cb4287c519dbb293aa4b3226eae44f5123ce7a7fd2e4cf917b895f68c7f5f117

Observation 62c3f35f-3748-4689-b30e-92f227fe6101 · outbound

This paper cites Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-19T14:47:36.250658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:a17a0ed97ea400e6aba10fb2db4bffab38788429d607aeb086c40ea1049a599c

Observation ef387fb2-e226-413d-a45a-58775ef38089 · outbound

This paper cites Alpha- gpt: Human-ai interactive alpha mining for quantitative investment.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Alpha- gpt: Human-ai interactive alpha mining for quantitative investment

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.714573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:5f14d38b9a2e7a9b0ae5bacafb61c24aafe6cc2350500e9ebad1e5270268be6b

Observation c7dec49a-8c68-41f3-92ac-9aa887f8d2a4 · outbound

This paper cites Can large language models mine interpretable financial factors more effectively? a neural-symbolic factor mining agent model.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Can large language models mine interpretable financial factors more effectively? a neural-symbolic factor mining agent model

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.633163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:ef73c26b85c590ad6af8f865c27087532df9592d7f55c6842fbf7aebd94e8b11

Observation ec5d7af7-bc6c-4748-a0e7-bbf54841cdcd · outbound

This paper cites QuantAgent: Seeking Holy Grail in Trading by Self-Improving Large Language Model.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery QuantAgent: Seeking Holy Grail in Trading by Self-Improving Large Language Model

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T14:47:36.325982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:1e17fbfe9f27b1246c888a2702dfee49884bd76f8ec5d4a365dd8faf930e04f0

Observation 00885b09-5999-43ce-a615-6645e3e3453b · outbound

This paper cites Al- phabench: Benchmarking large language models in formulaic alpha factor mining.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Al- phabench: Benchmarking large language models in formulaic alpha factor mining

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.700082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:73efa693ec8f37e00cbaebc851739593e69ffceabf933a2cd48ad42012400fd4

Observation a55db157-cea9-457c-8142-9c3553a7daf3 · outbound

This paper cites Alphaagent: Llm-driven alpha mining with regularized exploration to counteract alpha decay.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Alphaagent: Llm-driven alpha mining with regularized exploration to counteract alpha decay

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.631321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:7e6f7219928a9280dd4750ba3d255b19511f10e11fdb275ba61c8597cee04788

Observation a290ce47-72dd-4ce2-aa93-c54a4b1a2043 · outbound

This paper cites Navigating the alpha jungle: An llm-powered mcts framework for formulaic alpha factor mining.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Navigating the alpha jungle: An llm-powered mcts framework for formulaic alpha factor mining

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.701932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:28ead56775aeae4dfb02b587b45ab86779d36deea5a9de83565ab8f2d218f935

Observation 651fba28-30d7-47e5-9510-150ee0383d0b · outbound

This paper cites R&d-agent- quant: a multi-agent framework for data-centric factors and model joint optimization.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery R&d-agent- quant: a multi-agent framework for data-centric factors and model joint optimization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.635051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:2b367bc04db96cf853e44675ff20facf7fd0ebf16772e46754725973a477798f

Observation 455ff51a-0288-415d-a9fd-16ee47d07f08 · outbound

This paper cites QuantaAlpha: An Evolutionary Framework for LLM-Driven Alpha Mining.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery QuantaAlpha: An Evolutionary Framework for LLM-Driven Alpha Mining

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-19T14:47:36.310589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:cea7ef2e05043c947778efb8c8b9e492c90baf5f907b2ff8436ca71e6af1a3a1

Observation 4e2037a5-a597-42fa-961c-1b8900f6032a · outbound

This paper cites Deep reinforcement learning from human preferences.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Deep reinforcement learning from human preferences

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.637242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:2c9de74ddb09b6580457173a2691cf7420650bf301e944f9630a4a9121b1cb5c

Observation 007c2136-f782-4ffd-a8c6-58988cf53fdd · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Fine-Tuning Language Models from Human Preferences

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-19T14:47:36.239121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:b37029cc127e99fb5024d2ee9b2f491179d88a2b574ff384fa162650a45a7027

Observation 5bdc8b63-09bb-4124-9919-2cbb136e32f1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Proximal Policy Optimization Algorithms

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-19T14:47:36.247436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:730a41a26ff600c27b255e78332d54bf9bcd9ce1c21c4bbfc73b420a09fc5412

Observation 05b2011a-ab4d-4ba5-9c25-28fee6cb0c5d · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Direct preference optimization: Your language model is secretly a reward model

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.686482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:9e3c803bf03a0cfca58dd0f24b2ad5677f0cbb31dbd2884bf9edd6f1d3e7df1d

Observation 7f0d32ce-03ee-406e-aad1-4ae5d7fb3cce · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-19T14:47:36.275263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:c9128a3d8cfb412510610e36d47128978af900003125048c9d5e97b403ff4c02

Observation 9d0fe3e8-af28-4be7-b8fd-4fec8dc024a9 · outbound

This paper cites Large language model agents in finance: A survey bridging research, practice, and real-world deployment.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Large language model agents in finance: A survey bridging research, practice, and real-world deployment

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.609754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:6686f90578250a678b14a9552db30c6a8342ebc8fff8cc8293cfbcd7746bdfc6

Observation f4702f7e-7bd3-4936-9ef8-6d649dd24be9 · outbound

This paper cites Ectsum: A new benchmark dataset for bullet point summarization of long earnings call transcripts.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Ectsum: A new benchmark dataset for bullet point summarization of long earnings call transcripts

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.642915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:5b07394259eb6dc184c5645d02e0679d84880b3ef73516c8a963d4a655422e08

Observation 6e388a45-fe37-4315-b7e0-1a7f0e3a0e78 · outbound

This paper cites Ab- stractive financial news summarization via transformer-bilstm encoder and graph attention-based decoder.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Ab- stractive financial news summarization via transformer-bilstm encoder and graph attention-based decoder

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.706543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:7f755ffb5c212fa8f57475a2d1f375da0f4190bce07ac3384620932ffec87071

Observation 30cb479c-64d6-4552-b948-feb20c8db348 · outbound

This paper cites Finred: A dataset for relation extraction in financial domain.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Finred: A dataset for relation extraction in financial domain

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.718588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:d024689d2f1f9b5b4e9689eadce255d6187e51d7181590c3cb4a09d783bd9b37

Observation 3ed20da5-7ba5-4a7e-8c3a-4b1c0fa66307 · outbound

This paper cites Finbert: A pre-trained financial language representation model for financial text mining.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Finbert: A pre-trained financial language representation model for financial text mining

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.688371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:bb6e6eabdc69c422643daec4a9c7d1a1ad977ee29048ecfd63c51ea31e55f5c6

Observation 8d15e896-acde-4ef3-8f90-aea13c4d4da7 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery BloombergGPT: A Large Language Model for Finance

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-19T14:47:36.260094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:aa2bf30125b8775fdc32f6f285c8ef852190f06875fcf0cb1d9a65866c31e44c

Observation d529f88d-6074-4a06-8910-8ebedf170944 · outbound

This paper cites Fingp t: Open-source financial large language models.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Fingp t: Open-source financial large language models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.224261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:9e668710880e878493dbd0fa43556ab3dfa301c010048bf857a7f07da49ca497

Observation 1a5d46e0-8d14-4462-91a2-321b132110d2 · outbound

This paper cites Pixiu: a large language model, instruction data and evalua- tion benchmark for finance.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Pixiu: a large language model, instruction data and evalua- tion benchmark for finance

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.690252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:a1d40c29e6c9dbd3ce0cb4cc8467e8e6ca0c97b0503a861d4de5b3508a01c1a4

Observation 128cde12-d392-43dc-bc3a-4e71e3c7effa · outbound

This paper cites InvestLM: A Large Language Model for Investment using Financial Domain Instruction Tuning.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery InvestLM: A Large Language Model for Investment using Financial Domain Instruction Tuning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.269386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:c03dee215c192ce503291103018c62454b4c07d7246a077f5b2e4b203c792d2f

Observation 4dc23fde-bc29-4619-8321-39a964dc5eaa · outbound

This paper cites Fintral: A family of gpt-4 level multimodal financial large language models.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Fintral: A family of gpt-4 level multimodal financial large language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.680248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:5bf969ba1ffa49ca80adbba91f7371d15c91b07b1a93333f3d6c0ebbffbf06a9

Observation 4396d6d8-cf5a-4891-aa9c-c2e22139932d · outbound

This paper cites No Language is an Island: Unifying Chinese and English in Financial Large Language Models, Instruction Data, and Benchmarks.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery No Language is an Island: Unifying Chinese and English in Financial Large Language Models, Instruction Data, and Benchmarks

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.301523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:2c94f35ef0a587e2c4b84981aad61bf834549caafcf615b32fc79b084a7f5f50

Observation ba328254-a753-4da0-8037-53ec4e882af3 · outbound

This paper cites Fednlp: an interpretable nlp system to decode federal reserve communications.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Fednlp: an interpretable nlp system to decode federal reserve communications

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.692092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:64de05e25ad177b4c82db85b1d61ad9a73894bc9f83370daf971b7eaf2a47a0b

Observation b863898c-4775-4238-be4c-ced62623959a · outbound

This paper cites Trillion dollar words: A new financial dataset, task & market analysis.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Trillion dollar words: A new financial dataset, task & market analysis

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.698218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:aa83555387fdc37189ead5604c4ba9cc74275e1a7eaf40737d0a05759dd59148

Observation 7585ef5b-6ebb-449f-b62b-65e776718e4c · outbound

This paper cites Impact of news on the commodity market: Dataset and results.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Impact of news on the commodity market: Dataset and results

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.695810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:30a3fc0829124531230a31d4afdd8d1f3328afccddfc38ef319d469382cd1498

Observation 15ccc756-9e46-4bb0-abc3-b74266131a94 · outbound

This paper cites Harnessing llms for temporal data-a study on explainable financial time series forecasting.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Harnessing llms for temporal data-a study on explainable financial time series forecasting

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.708879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:1c1b25fff0bc1f1397fed4647e9aaf51feae9c5591c5e0773d19e0ddcd7f8c38

Observation ea6fae76-7cd6-435e-b7a6-7c66b4c731b4 · outbound

This paper cites FinTSB: A Comprehensive and Practical Benchmark for Financial Time Series Forecasting.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery FinTSB: A Comprehensive and Practical Benchmark for Financial Time Series Forecasting

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-19T14:47:36.236103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:91c9642af5129fcbb1dffaf9f898642c62acdeeb35e3dd0814a2dada5511cf26

Observation ec45ee6e-b130-4635-87f0-6dc9ceebacd6 · outbound

This paper cites GPT-InvestAR: Enhancing Stock Investment Strategies through Annual Report Analysis with Large Language Models.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery GPT-InvestAR: Enhancing Stock Investment Strategies through Annual Report Analysis with Large Language Models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.306912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:c71d6b95f2eb8ac6ae30b48f0c3f9b5169d2639a20ebc5038850be0b3c17def7

Observation 8bf974c8-aaeb-42a8-a378-31e23e86eac9 · outbound

This paper cites Finben: A holistic financial benchmark for large language models.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Finben: A holistic financial benchmark for large language models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.654130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:6062309dc4321846346a578cd36d5608a275a74f46ff5dcc4eba04d03f3d528b

Observation 40a7e525-b480-4997-935a-105273cbf5b9 · outbound

This paper cites Investorbench: A benchmark for financial decision-making tasks with llm-based agent.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Investorbench: A benchmark for financial decision-making tasks with llm-based agent

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.614736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:caf8779bed1b28d6883a95482b76091fd9b6a3b7d255cdc0d5857e8e9fdb8bb1

Observation f1d75017-7511-4705-8df4-9700e1c982c2 · outbound

This paper cites Strux: An llm for decision- making with structured explanations.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Strux: An llm for decision- making with structured explanations

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.658442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:502847e44cc337132aba87ac99a2d6bef99b4a5e4b7844948bfbfc7ab122265e

Observation 39f4cd51-20e1-4ae7-a915-383e9fb0d46b · outbound

This paper cites Finmem: A performance-enhanced llm trading agent with layered memory and character design.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Finmem: A performance-enhanced llm trading agent with layered memory and character design

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.716791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:d94139aede8e524853f5f35060eefbbf6017930067e3905c8efd8327ada35b30

Observation e25a5035-417d-4eb6-8e41-dc0f89be25e1 · outbound

This paper cites CFGPT: Chinese Financial Assistant with Large Language Model.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery CFGPT: Chinese Financial Assistant with Large Language Model

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.283958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:314cf71395cd383f794f41d8907212b200963aef0e0849ed9e2b20d65fb39fd9

Observation 917f3932-ecdb-48c5-a090-434528d50501 · outbound

This paper cites When AI Meets Finance (StockAgent): Large Language Model-based Stock Trading in Simulated Real-world Environments.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery When AI Meets Finance (StockAgent): Large Language Model-based Stock Trading in Simulated Real-world Environments

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-06-24T02:14:19.804532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:b4840e82cead1daf6f454cea8c085769c41839dad7d765081a947010bd0c786d

Observation abbc0d37-cebd-456a-8597-73cefdbd7a38 · outbound

This paper cites Tradingagents: Multi-agents llm financial trading framework.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Tradingagents: Multi-agents llm financial trading framework

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.650516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:4c525ede0db683252c2feb03790ee67bd16405db9c45371be39bde419abb94c8

Observation e9673466-28c8-40fa-b6be-dce388eed046 · outbound

This paper cites Convfinqa: Exploring the chain of numerical reasoning in conversational finance question answering.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Convfinqa: Exploring the chain of numerical reasoning in conversational finance question answering

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.656177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:7da93fa4e2ca924e098064fca33ac91c1d6d14c0671900985ebd940d4a369723

Observation 5974c9a5-044a-4210-b56d-295450fafa58 · outbound

This paper cites Finder: Financial dataset for question answering and evaluating retrieval-augmented generation.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Finder: Financial dataset for question answering and evaluating retrieval-augmented generation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.641052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:ed5cec9fa0e32bdbcaf02d03097b0036083e42a3b9eb4ce27970413b7a2ce807

Observation 8d8394bc-33bf-42cf-b0ed-e3379b17d7e1 · outbound

This paper cites Al- phafin: Benchmarking financial analysis with retrieval-augmented stock- chain framework.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Al- phafin: Benchmarking financial analysis with retrieval-augmented stock- chain framework

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.646803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:d181c80fa01110c8d3c9cc5668f999bd4cd2ab818d311d69a66f32773ed98cdf

Observation 11e9a090-f8ff-4b8e-9d4f-449edc6a637f · outbound

This paper cites Substituting human decision- making with machine learning: Implications for organizational learning.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Substituting human decision- making with machine learning: Implications for organizational learning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T14:47:36.627568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:c0e0bb826f1237487962efded232c3d2e54943c27127a24563b0dbc639ca735f

Observation 70ef0511-6a29-414b-9bb8-d5e933044967 · outbound

This paper cites FinPT: Financial Risk Prediction with Profile Tuning on Pretrained Foundation Models.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery FinPT: Financial Risk Prediction with Profile Tuning on Pretrained Foundation Models

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.292296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:561715057865c430293e4750fbdfdf4e49f02826aaeed5b95cc5ae3ccb2cebef

Observation b5a6723b-53a0-42af-84ef-c46f364078bd · outbound

This paper cites Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models.

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.221245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T14:45:07.168330Z digest=sha256:6917d09ec4956166fb9eb2a45f52a65047695491caa90faf31d2dec47809a72a

Pith citing papers

Observation a75f6ec0-3312-45a6-8397-ec5e75bab7a9 · inbound

Towards Autonomous Formulaic Alpha Discovery: An Evolutionary Computation Perspective cites this paper.

Towards Autonomous Formulaic Alpha Discovery: An Evolutionary Computation Perspective From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-04T20:52:58.385612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:52:57.876245Z digest=sha256:fb36a2ccd8a70516791a2af005c42f8a1be481121225ebb4983f820f939f9515

Observation 015d67b1-8258-4a82-9c8d-b4c900f82cc0 · inbound

AQuA: Recursively Self-Improving Quantitative Trading Research Agents cites this paper.

AQuA: Recursively Self-Improving Quantitative Trading Research Agents From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:14:12.505221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:14:12.505221Z digest=sha256:6038ef90b0809742012d894fd2a5d579113f2ca8bbb7ca8a5569f943b21e83fa