Pith. sign in

Paper Citation Record · LEDGER

Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2503.16974.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.16974 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:39.064313Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T01:56:27.114999Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation afce4c83-1dbc-448d-82ba-b27682fe53ec · inbound

EasyMath: A 0-shot Math Benchmark for SLMs cites this paper.

EasyMath: A 0-shot Math Benchmark for SLMs Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:39.064313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:33:39.064313Z digest=sha256:78a468fc7930318d9595a8c7a36a6db6a98c6f2c7b14bca88f0d4cee7581e611

Observation eda8e380-146e-4921-ac4f-6308acb1848a · inbound

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning cites this paper.

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:58:33.525577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T15:58:33.219451Z digest=sha256:75a67796478096ae522a566aa06a463f88f07cd53bbd6e40fd0924c9c843baff

Observation d42194f6-6da4-4a9b-bc1d-a72b0c8f9db0 · inbound

Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models cites this paper.

Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:47.834129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:39:47.834129Z digest=sha256:9a2cdbb91c818c236b2378d1a3db592a4905f3be5382d54e3d0215fd0670a083

Observation 13d0e133-3bb0-4eb9-8de0-dcf7bb612298 · inbound

Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis cites this paper.

Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:49.119437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:49.119437Z digest=sha256:fc50fbaa7ecf62dde340d260f31a2c7775b2d377a3243534081fef8afe65c46b

Observation 7525b13b-a51e-4a14-9485-0428f05dc043 · inbound

Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis cites this paper.

Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T06:19:09.680553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:19:09.680553Z digest=sha256:8f6b75bd104d345eb66425d58d742928866d956b7c5a82549b690dac862626a1

Observation 58b21271-5c75-4a4a-8ea2-59011ee69658 · inbound

Towards a Science of AI Agent Reliability cites this paper.

Towards a Science of AI Agent Reliability Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:02.023467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:02.023467Z digest=sha256:9e35f72f7bb1dcc0288dd14566738269f0cb7c2a88be9d45d4ea4c01ae5a0261

Observation 5bf674ad-7849-4264-a9ca-469bcab8a4e1 · inbound

Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models cites this paper.

Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.402158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T13:31:46.940449Z digest=sha256:720dff58816dd09a3dc6f122582971e2b54204a537c599fed584869665f6ba02

Observation 6b05796b-0d86-42d7-8102-e322530cb9f5 · inbound

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis cites this paper.

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T20:50:53.567415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:50:53.567415Z digest=sha256:e2b82b67f7ee64f2e509198b1d98bccd2913cd96fcd554a4a37da827f7c75f1e

Observation 6d4a2621-690f-4808-ae15-0689c3f5d5c0 · inbound

From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems cites this paper.

From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:45.811279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T22:27:30.692356Z digest=sha256:a7863e25770e53ab623de614409a7a9d18254b1d7f3f29944c61d970a7d33abf

Observation 6a8b0efe-dbe4-48e5-94be-133175b7d75a · inbound

Shapley in Context: Explaining Financial Language with Domain Expertise cites this paper.

Shapley in Context: Explaining Financial Language with Domain Expertise Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 284

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:56:27.118015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-02T01:46:33.186116Z digest=sha256:6454c69fe0a585ada80104a003e528e2059201e875e34de953960aed72e65f14

Observation d781a69a-bd23-4d97-b24f-630a92ca4bd9 · inbound

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making cites this paper.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:21.831933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:21.831933Z digest=sha256:d61f8515315d4b60eb099246e52541ee973f1f2bf8e1ac1995e2e2b2824c50d9

Observation f0e011b8-26aa-419b-9208-9ca2b9a07db7 · inbound

Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement cites this paper.

Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T12:15:36.327966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:15:36.327966Z digest=sha256:634502e5f7f31d30510c07a3fabc2ddfb1fb4cad7f5f460ed7f831bd4dd13917