Pith. sign in

Paper Citation Record · LEDGER

Phi-4-reasoning Technical Report

As of 4 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 48 inbound Pith citation observations for arXiv:2504.21318.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.21318 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T03:40:25.706499Z

measured 112 of 112 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:48:28.952743Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact31
  • verified fuzzy28
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 884eb5c4-2c6f-4fcb-8561-c79b42985f6b · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Phi-4-reasoning Technical Report Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.821543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:686682ee3b169d3e3018938341652271eb5b40b3ddc98adcfce690fa402b4ea5

Observation 70f4f4a0-2245-4405-8cbe-5b2ae0de7fa4 · outbound

This paper cites Phi-4 Technical Report.

Phi-4-reasoning Technical Report Phi-4 Technical Report

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T03:40:25.743418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:0c9dc43e3aa053256b7859af06249c0e00253c5c02147870dba06c1608c3bd13

Observation 0c522f1b-f36d-4fed-8c6f-7a155b72c24a · outbound

This paper cites KITAB: evaluating llms on constraint satisfaction for information retrieval.

Phi-4-reasoning Technical Report KITAB: evaluating llms on constraint satisfaction for information retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.894984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:eb39d87c4f589ef96feb8d35223cb19bc36dae8a3468a743e45a146de4bc0bba

Observation 405b4788-a173-499f-9cfe-f84552267446 · outbound

This paper cites Aime 83-24.

Phi-4-reasoning Technical Report Aime 83-24

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.898840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:65b4dfe1a0fe63305da441e10bc8372cdda3910396f4bc2ccf5e735984488992

Observation b3c289d0-e7da-4350-b714-d80962887c12 · outbound

This paper cites Aime 2025.

Phi-4-reasoning Technical Report Aime 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.902043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:334db3d643d0644c5da01d081d06c68e679c43d957825cc70cf34299d6f00f78

Observation 1902c78c-a273-469b-a505-0c7b317d526d · outbound

This paper cites Concrete Problems in AI Safety.

Phi-4-reasoning Technical Report Concrete Problems in AI Safety

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.866001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:8a812b9b54f88ebb40bd542c2cd4ac3ff4a30b042a741eb032bba2014c624304

Observation f20ae402-9939-4ef4-8cab-0cfeeb0f4147 · outbound

This paper cites Claude 3.7 sonnet.https://www.anthropic.com/news/claude-3-7-sonnet.

Phi-4-reasoning Technical Report Claude 3.7 sonnet.https://www.anthropic.com/news/claude-3-7-sonnet

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.909069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:6bf9b0f8855489c77fe0978c48724dccb31e61dddda592ca6d97826294a2550b

Observation 9194b23c-8816-44fa-a0e8-38533ac977b9 · outbound

This paper cites Chain-of-Thought Reasoning In The Wild Is Not Always Faithful.

Phi-4-reasoning Technical Report Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:44847b4b863052650d5f61d8901e7e6ec6eb089ebc49a3bf9824da440343e9a9

Observation 14ac81e6-d67c-4614-990e-1fb3a65762d9 · outbound

This paper cites Eureka: Evaluating and Understanding Large Foundation Models.

Phi-4-reasoning Technical Report Eureka: Evaluating and Understanding Large Foundation Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.874331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:d638673ceef15f87726b94ca6214ac614cc1e296d37bd89499630f416b5c00d7

Observation 59c63ade-c704-4bfb-a7bb-4bcd342082f9 · outbound

This paper cites Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead.

Phi-4-reasoning Technical Report Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.879473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:1b6754750dab58f797f7d0fcf4671bef83d2360af0fd2d4070086941881e2f1d

Observation 5e85f7d8-de6a-40db-82a5-23f435a81586 · outbound

This paper cites Matharena: Evaluating llms on uncontaminated math competitions, February 2025.

Phi-4-reasoning Technical Report Matharena: Evaluating llms on uncontaminated math competitions, February 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.921108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:23cc73f0fe7882f7f0002938cbc2999a7eb1c0eac7a6b2c46376d33efd7de52d

Observation cbe872af-1240-4f64-9c6e-3f24462e89c1 · outbound

This paper cites Designing disaggregated evaluations of ai systems: Choices, considera- tions, and tradeoffs.

Phi-4-reasoning Technical Report Designing disaggregated evaluations of ai systems: Choices, considera- tions, and tradeoffs

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.923933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:2bdf27254995bbfc6836c542c69a148ff8e39c156a3efc7daf28d5ab8442f19b

Observation 4b9ba938-2cae-4a2f-92a9-1091500c1dd9 · outbound

This paper cites Benchagents: Automated benchmark creation with agent interaction.

Phi-4-reasoning Technical Report Benchagents: Automated benchmark creation with agent interaction

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.885399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:42e9b9e4f4c9043c523cf2cbb3f736319f02bed805132cfa206d3d3abda511fb

Observation 975589be-8544-4620-88b4-f511fbc5888c · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

Phi-4-reasoning Technical Report Extending Context Window of Large Language Models via Positional Interpolation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.890787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:0afe0a90ae294b5a1bec700305318b4eefbb4f14752e5913e3885128723feb33

Observation aa6524b6-d3ce-4db7-8f1a-ed4512309cd7 · outbound

This paper cites Reinforcement learning for reasoning in small llms: What works and what doesn’t.

Phi-4-reasoning Technical Report Reinforcement learning for reasoning in small llms: What works and what doesn’t

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.931552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:646bb4df4020bd55ad1aa3f0d9dc97c35a8a279018a87078bcbf2fbe4e49d793

Observation 26d8073d-d36d-46a8-a2b1-0b2749b72f0c · outbound

This paper cites Reinforcement learning for reasoning in small llms: What works and what doesn’t.

Phi-4-reasoning Technical Report Reinforcement learning for reasoning in small llms: What works and what doesn’t

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.747904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:0dfba4d33895df04a74aa83bc0232be6862d43d0e278c357fd5ba2380aeb54c5

Observation a0d419ea-9ae3-4750-a7f9-0201382ee9fd · outbound

This paper cites Omni-math: A universal olympiad level mathematic benchmark for large language models.ICLR.

Phi-4-reasoning Technical Report Omni-math: A universal olympiad level mathematic benchmark for large language models.ICLR

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.936482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:0aef5628dfcca95af402c9e69d78edb2b8b57ae9b70606a040a2e72d2cf57f00

Observation b9aef3d6-f435-4f7a-9fc4-0d4910ffe383 · outbound

This paper cites Scaling laws for reward model overoptimization.

Phi-4-reasoning Technical Report Scaling laws for reward model overoptimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.938967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:af3d1110344c94997f1161fd8fd44c4cf93203023e6ad58bc9c56e843844bb5b

Observation 708c5db8-b9a9-4c31-8b02-4d88848d1fe1 · outbound

This paper cites Gemini flash thinking.

Phi-4-reasoning Technical Report Gemini flash thinking

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.940982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:f8ec6b54ff9e424b2c625db1c197a8db320aeb535ca736491fd3476a444b3203

Observation 14c552b0-c3f1-4f16-aa6b-e44d33d2395b · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

Phi-4-reasoning Technical Report rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:43:48.343981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:7d91ca87d7ebeddd7ea1b300fb60c120f0785b7b25ed2cc17cc0a716d5a303d7

Observation d2038791-e7e6-475e-aa8e-2a79b644e4f3 · outbound

This paper cites Textbooks Are All You Need.

Phi-4-reasoning Technical Report Textbooks Are All You Need

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.760971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:ba1d63a59b47af7144dfa2e638bfeb2c8dadd2917382330960c42b8ee04cc406

Observation caaa20f3-36fb-4418-b8d3-d12c69e68d78 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Phi-4-reasoning Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.767220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:bf88185cc320c18734a9c24a6c65301425f8e8b86938a87cacf59108eeeca092

Observation e1d81e83-35bb-4a4b-a9df-4b5d3e4ed44b · outbound

This paper cites Computers and intractability: a guide to the theory of np-completeness (michael r.

Phi-4-reasoning Technical Report Computers and intractability: a guide to the theory of np-completeness (michael r

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.952938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:56c0cc6fc45f0ed4b78d38a8a519fb9b1b08c44737fbbc990e16d1d2ab8a1f49

Observation 4af62f25-e51d-447f-b38d-9b1c5f2f522a · outbound

This paper cites ToxiGen: A large- scale machine-generated dataset for adversarial and implicit hate speech detection.

Phi-4-reasoning Technical Report ToxiGen: A large- scale machine-generated dataset for adversarial and implicit hate speech detection

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.955477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:671762a98c540d960cc356e48d9189d14499290ea97b3b567e2782c226c899bd

Observation 0c7d2f1c-30bd-4bd3-aa73-70f7fecdc5c2 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Phi-4-reasoning Technical Report Measuring Massive Multitask Language Understanding

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.773304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:b5fd247f1c2c9d95c1c159346015fd2ecf2c9fb3b25c767c48ba35c84f09dd1d

Observation b3cdae9f-de2f-470d-a3ca-7c7a8245630f · outbound

This paper cites A sober look at progress in language model reasoning: Pitfalls and paths to repro- ducibility.arXiv preprint arXiv:2504.07086.

Phi-4-reasoning Technical Report A sober look at progress in language model reasoning: Pitfalls and paths to repro- ducibility.arXiv preprint arXiv:2504.07086

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.778105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:68e43ea33f293339940cfadc2149457744e9d525ec8ff4ae46826a0b088bfa1b

Observation 97fc3336-691c-4ecb-9bfe-058c13bf4aeb · outbound

This paper cites GPT-4o System Card.

Phi-4-reasoning Technical Report GPT-4o System Card

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.782438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:150d5b16854646d541bcb96a715244928f3b9351bca5d5ed4dae1b49cacc4e10

Observation 879aa7e9-f46d-4dd6-891b-6f41031ca8e6 · outbound

This paper cites OpenAI o1 System Card.

Phi-4-reasoning Technical Report OpenAI o1 System Card

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.786192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:da1f9d6bc1a9917d0b654db1a03766c9c5840f5acee59e45cd6046b69ced9cd6

Observation 77fb31cf-ce3f-4c8f-a2c5-06955ff09624 · outbound

This paper cites Phi-2: The surprising power of small language models.Microsoft Research Blog.

Phi-4-reasoning Technical Report Phi-2: The surprising power of small language models.Microsoft Research Blog

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.969847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:4a0a2bfb880f1cee5839c53f3efc351c9c41897f68fb3c1d76722f827c925467

Observation 23a33d8e-2625-4bff-9419-3d0866e4a02e · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

Phi-4-reasoning Technical Report SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.790569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:00b53ed8cae73e7c8993e62678f653e0659e2fbf65996c2d132924d452dd5072

Observation 650992ac-952a-4c66-bf10-937a05982ffe · outbound

This paper cites Same task, more tokens: the impact of input length on the reasoning performance of large language models.

Phi-4-reasoning Technical Report Same task, more tokens: the impact of input length on the reasoning performance of large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.905646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:dd768a914951fb5e1012d3b860b34bd068a1d7d2db3d6d250e5fc5dfb221473b

Observation 1ae656e7-6a29-4e3b-9f02-e37a610bfacb · outbound

This paper cites Exaone deep: Reasoning enhanced language models.

Phi-4-reasoning Technical Report Exaone deep: Reasoning enhanced language models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.794268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:d0d039199a5200a8cd6e91b5768f7a5b6dd6b4e90d26d9d15b932a20f6f13ca6

Observation 06acca0a-4a6b-4ef4-adf2-8e2a0bccc721 · outbound

This paper cites Functional Interpolation for Relative Positions Improves Long Context Transformers.

Phi-4-reasoning Technical Report Functional Interpolation for Relative Positions Improves Long Context Transformers

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:40:25.798098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:492f4f7a4478cc80ce8de1dd84a5db128ed59b27d3ebac5e535a1de6053d571c

Observation 09202a5b-c0cc-48c2-9a0a-a4cd1f7c70c5 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Phi-4-reasoning Technical Report From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.802248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:5df0859a8187b83c2e2ed1887a24594b61c330b26dc8017990860cf65dcbb368

Observation 22d88ba6-9fb2-4916-9253-2b9225e41f06 · outbound

This paper cites Limr: Less is more for rl scaling.

Phi-4-reasoning Technical Report Limr: Less is more for rl scaling

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.926961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:56d5e3dd7ec869e6d9bd5ab67723f7ba156da1ffa690e22ddaeaad641b9ed2a0

Observation 067e08a1-e275-43f4-b043-771bcfa4cf1f · outbound

This paper cites Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation.

Phi-4-reasoning Technical Report Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.929212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:0ce9e6ef170de332baaf79c15cf78107cb99519ef38282d838edf175c511d354

Observation caa1ec0a-9bce-413f-a700-0b243edd05ba · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Phi-4-reasoning Technical Report Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.934122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:c1b15e3b5cae34b64e5315c5d639011b62ae9d52d928c3f611847545968ddf58

Observation 8b4f0a72-814c-4c09-bcf3-d5361d44a8f9 · outbound

This paper cites A framework for automated measurement of responsible ai harms in generative ai applications.

Phi-4-reasoning Technical Report A framework for automated measurement of responsible ai harms in generative ai applications

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.943637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:07c28b00a301b2ca01f9297adb84b898ee32b5c778c74e5cbdc9348c5356c32a

Observation 05cb8000-2abd-43a7-a6f0-9ed37764ac97 · outbound

This paper cites A Framework for Automated Measurement of Responsible AI Harms in Generative AI Applications.

Phi-4-reasoning Technical Report A Framework for Automated Measurement of Responsible AI Harms in Generative AI Applications

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.805809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:148cb8eda150773931a8de37935fcd8e0cb6a2eb534786c43a1978abff071d02

Observation bcab39d6-4e6a-41b1-830e-d5d23cf06dd6 · outbound

This paper cites Orca 2: Teaching Small Language Models How to Reason.

Phi-4-reasoning Technical Report Orca 2: Teaching Small Language Models How to Reason

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:40:25.809710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:e7d5ccec76d5dd09b8f337b00059378d5b1917e1181a3645e1512239437e0b92

Observation 48bba3f0-f6db-4db6-a44c-0abcb5d01cab · outbound

This paper cites AgentInstruct: Toward Generative Teaching with Agentic Flows.

Phi-4-reasoning Technical Report AgentInstruct: Toward Generative Teaching with Agentic Flows

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:40:25.813945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:f3d9930cf5d69331fbdd328c1df6fc7705e0d138c4dda066d1afd50f9ccaf7d2

Observation 62b830a3-b5e2-40bc-92a7-693df01b60cb · outbound

This paper cites Unearthing skill-level insights for understanding trade-offs of foundation models.

Phi-4-reasoning Technical Report Unearthing skill-level insights for understanding trade-offs of foundation models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.960938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:70c8ee0e1259757a7fe10e7558c59d32f5f97ef906ba8047e2860afb17457b75

Observation c3ce1455-2c71-4e32-9e01-7a65a732e463 · outbound

This paper cites Orca: Progressive Learning from Complex Explanation Traces of GPT-4.

Phi-4-reasoning Technical Report Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.817493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:cc291a0662b615cc9cebabfcdb8f12f6c8e1c6117f868ecd731f9e01f9f3d7fb

Observation b9e2937a-24b1-4aaf-870d-206decf3c76a · outbound

This paper cites Towards accountable ai: Hybrid human-machine analyses for character- izing system failure.

Phi-4-reasoning Technical Report Towards accountable ai: Hybrid human-machine analyses for character- izing system failure

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.967169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:d8ab8018aaf54a2f4b42d7dba4a4000ec266df7ef1686e64c966b38d654079b2

Observation 3bb89c3a-e54e-48bf-ac37-2afb70abcbc3 · outbound

This paper cites Openai o3-mini system card.

Phi-4-reasoning Technical Report Openai o3-mini system card

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.972146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:53142c9885022e0a5a43eb79c19b0edb86595c8c0ba5e7da44f59915df5f2235

Observation eda276f5-855e-4d1e-abfc-7d7e3999e968 · outbound

This paper cites Computational complexity.

Phi-4-reasoning Technical Report Computational complexity

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.912239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:6b3548408236f7a4f56cc4b461aa1f7d7d7ddad8339a2f926aa4f2024c14325f

Observation 2935f4e6-a434-4fec-a593-134d636562c4 · outbound

This paper cites Overreliance on ai literature review.Microsoft Research, 339:340.

Phi-4-reasoning Technical Report Overreliance on ai literature review.Microsoft Research, 339:340

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.915243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:6abf176ec1bd4b45a8bcb9e0fd9f8e0c8dc959fe88503b529fc2005921c1b038

Observation 2ef8e4b6-d6f6-4472-a70a-31803eb8d840 · outbound

This paper cites Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad.

Phi-4-reasoning Technical Report Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:40:25.739163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:39d19e0a29790004d7683a4ec1ebb7a907f67fb061f3af8ab3e29b98a96d5395

Observation c6d2135e-2de7-4c80-98ca-708e03964762 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Phi-4-reasoning Technical Report Gpqa: A graduate-level google-proof q&a benchmark

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.947365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:65e05c04c3302969c2618ff51f310acfb5847a2b2243c806e7e00a6c9211c597

Observation 07f4864d-994d-4a8c-bde5-534c44c89efe · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Phi-4-reasoning Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.825422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:31faff477d55799d3449ebe8c1f75f0d3e877c92d074effcba8af1e88c0ac79b

Observation dae5687a-7da1-497f-b6c0-301c4720db1e · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Phi-4-reasoning Technical Report HybridFlow: A Flexible and Efficient RLHF Framework

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.829028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:b3ffda75dafec22f179c40879b394a3a49eeaa06e0099d58b514a632bdd7e5ea

Observation d765b8ae-949d-4062-a282-f98a624ff0c2 · outbound

This paper cites Language Models are Multilingual Chain-of-Thought Reasoners.

Phi-4-reasoning Technical Report Language Models are Multilingual Chain-of-Thought Reasoners

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.832469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:ed7f820a4f98d9de904f0372db9c4dc39859f7040af8bec7b8de8af02c914a52

Observation 702cef1a-b9e6-492f-8b63-0a6cea967aed · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Phi-4-reasoning Technical Report RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.835676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:fec2cb2fae5f6f5782a98ff2efe5ed9644c1c3edce6eec7ebc7df34b1370a011

Observation 7d51f35e-89ac-4b36-ade0-9ed494a3c128 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Phi-4-reasoning Technical Report Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.840520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:5ebe0189ae5e20d825b25aeb940501fb96ef0d31e788e7a38abe77745e6d1f45

Observation af2cc9d9-95e0-46d5-98f6-78193ca15954 · outbound

This paper cites Open Thoughts.

Phi-4-reasoning Technical Report Open Thoughts

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.957980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:b8475178c3ef6b5c6ccca97deda524f942c580323ff9411bfa3968cbe5d366fa

Observation 2b4ff737-13ae-4ac6-a000-624913b1a9c2 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Phi-4-reasoning Technical Report Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.964635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:a7e7204f296c0dd61e37c340e10f446cd1b6d08d2647e269c63b43f3487fd9df

Observation 41c0b3e2-e67e-4c59-800d-c2526b0c2d1d · outbound

This paper cites Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.

Phi-4-reasoning Technical Report Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.843969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:928f355fca465c45f332b70fa2eb20d2b31f87da423a714c03acbb2e5d9e296f

Observation ed7a1c04-96b3-4eb0-a788-2bce9ef46613 · outbound

This paper cites Is a picture worth a thousand words? delving into spatial reasoning for vision language models.Advances in Neural Information Processing Systems, 37:75392–75421.

Phi-4-reasoning Technical Report Is a picture worth a thousand words? delving into spatial reasoning for vision language models.Advances in Neural Information Processing Systems, 37:75392–75421

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.950192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:6e3277066c22e6e0f37086c5d049fc15c99638bc06e87ee50179caaf5317de2a

Observation 21a94ca9-336c-4565-b90e-f089ea37fe87 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Phi-4-reasoning Technical Report MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.847444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:28fcaaa8997654decb7fe99831bfa3ee00d1f554d82d67a2239845158ae58849

Observation 8307ea11-0e5e-435d-a947-f20e74869e07 · outbound

This paper cites Qwen2.5 Technical Report.

Phi-4-reasoning Technical Report Qwen2.5 Technical Report

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.850758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:921bc5d163abe03e95aa8cde9ca1b0b85e592cb9b2edea8f2acdd1b05ae8a779

Observation 9c08c4c8-1eff-47cd-8605-016f003dce04 · outbound

This paper cites On the Emergence of Thinking in LLMs I: Searching for the Right Intuition.

Phi-4-reasoning Technical Report On the Emergence of Thinking in LLMs I: Searching for the Right Intuition

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.854460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:5f0dae2f39a9873f7a7861a18ac6a1dd8eb89ec5110fea9d7247b5fc12eab428

Observation e28cb773-4c87-4c67-a3e1-afcfa5363cf0 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Phi-4-reasoning Technical Report Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:29:59.463518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:7671e0bd470a95de982e8bb443f2522665b2fb7f12880e81ed6a288f3a8d0187

Observation e6fd1974-f1f3-4512-9ee8-bb2b60627811 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale.

Phi-4-reasoning Technical Report Dapo: An open-source llm reinforcement learning system at scale

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:40:25.918153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:9411391e4f8c5355220f8d59203ab4a79c162960252f3254d240af8810f8eb76

Observation 06a25481-bea5-48ac-a61f-9144bad29121 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Phi-4-reasoning Technical Report Instruction-Following Evaluation for Large Language Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.862410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:b8cd40be720e986e8051da69ab160e689216ae28779fac1ddb8ab139ae6413ee

Pith citing papers

Observation df745f2f-0866-4f33-b7f1-cbe7fbe5e9fa · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Phi-4-reasoning Technical Report

Reference 170

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:32:00.977531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:3bc6582d59f49cade911883392c4d48291bef760575545592c5025e1c4c14d09

Observation 3f78d8d7-b2a3-4632-9083-f15a89340d50 · inbound

MathArena: Evaluating LLMs on Uncontaminated Math Competitions cites this paper.

MathArena: Evaluating LLMs on Uncontaminated Math Competitions Phi-4-reasoning Technical Report

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T00:10:14.812539Z digest=sha256:863a0aba1ae480d680de362af52d8d3e412df87371878a1829ac3f1f20316be2

Observation 8a857573-5dcc-4e4a-8842-ec8372391689 · inbound

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models cites this paper.

Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models Phi-4-reasoning Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:27:07.269810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T06:25:12.799097Z digest=sha256:1c792de9e5bbaa9ad1cfa1e24992b8ebb7df901aae6d6a3c9649441d5dc819cf

Observation ca875656-cd14-4e80-81b1-8a0755f4523c · inbound

ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing cites this paper.

ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing Phi-4-reasoning Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:17:00.830862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T03:14:05.109011Z digest=sha256:cbe7cefcf6fb91d3999e95d28acd4300cff9e0de6d273aad30325e738ad27e72

Observation a4710964-eb33-481d-ab96-754f1a9cb7f5 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Phi-4-reasoning Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:24.461462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:c7d5f41a01380122191ee9374d67c4241df9ceb150b679e84a08544e330629e2

Observation 4c1e6eb3-6977-47c4-9c41-5057cd7f5a15 · inbound

Visual Reasoning Agent: Robust Vision Systems in Remote Sensing via Inference-Time Scaling cites this paper.

Visual Reasoning Agent: Robust Vision Systems in Remote Sensing via Inference-Time Scaling Phi-4-reasoning Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:21:32.839473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T15:19:59.478032Z digest=sha256:76899961439745802bbfc364fffecb3cc09d522ffccb87a250674ec22f9d790e

Observation 8a051f55-631c-4fe2-ba5a-ca5b49fa7303 · inbound

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation cites this paper.

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation Phi-4-reasoning Technical Report

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:06:13.858270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T10:04:39.223895Z digest=sha256:9e799895fdf5561bb086810bc1680c960bece6c590922571669b5ed71ecd8c8b

Observation 5d21b7cb-1dac-4395-ba9a-d8826d8c2876 · inbound

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models cites this paper.

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models Phi-4-reasoning Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T09:37:05.721179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:37:05.721179Z digest=sha256:3a44ef48e7a9a59a48f2cbf31a8ec09f40804e99e326b9a0d751c69f5c9e81d4

Observation 43239800-29bb-4948-ab42-335f3500ecd2 · inbound

Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment cites this paper.

Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment Phi-4-reasoning Technical Report

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:23:56.318846Z digest=sha256:ea3831c583733be622d348ab4acffa4467fa39d4981bd102e631b4647de52b50

Observation 99465e2b-03fe-4e2a-8ef5-ac45a23b13bb · inbound

Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment cites this paper.

Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment Phi-4-reasoning Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T09:22:42.385847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:22:42.385847Z digest=sha256:4c1d5671419f1482923a512000eae834a6626592cbc85ee9cd355802dedf35a0

Observation 62bd258a-3b18-4a7d-9287-0db06f442bc8 · inbound

The Geometric Reasoner: Manifold-Informed Latent Foresight Search for Long-Context Reasoning cites this paper.

The Geometric Reasoner: Manifold-Informed Latent Foresight Search for Long-Context Reasoning Phi-4-reasoning Technical Report

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:42:33.579501Z digest=sha256:75e154ea5ac12998f02fa85ef0d2268937598cf830945e13c5e8dcaeaeeed54e

Observation 1e350365-2b6d-4b50-8e89-17be021f38e5 · inbound

Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling cites this paper.

Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling Phi-4-reasoning Technical Report

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T09:53:14.045466Z digest=sha256:25e625ab62d7327af4c78e993a9a03b528639c8c4c8fe5ce2f7d776ee068aec7

Observation aad44ffe-f341-491e-aca1-c56bf64094cf · inbound

Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective cites this paper.

Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective Phi-4-reasoning Technical Report

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T02:34:43.248634Z digest=sha256:1849092630f44739b60fa1b1c74a9306e29c61f70c686e45f5f1da515fba429a

Observation 313f32af-5eab-45e7-b000-c5137af0bc82 · inbound

Learning to Evict from Key-Value Cache cites this paper.

Learning to Evict from Key-Value Cache Phi-4-reasoning Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T01:18:01.894936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:18:01.894936Z digest=sha256:92df4d21b108129527a1ec703f31a192b36867212038c0f0bfd273884ef68095

Observation ba6c202a-3788-4db8-9695-0408dfbc8dca · inbound

Author-in-the-Loop Response Generation and Evaluation: Integrating Author Expertise and Intent in Responses to Peer Review cites this paper.

Author-in-the-Loop Response Generation and Evaluation: Integrating Author Expertise and Intent in Responses to Peer Review Phi-4-reasoning Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T09:42:58.204848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:42:58.204848Z digest=sha256:f3b59335478e3092cc05c26302259033f31764d2277ec740a129e73a2f89d8a5

Observation a55d5afc-cbfb-46a5-be08-5fe0e1c730f4 · inbound

Ranking Reasoning LLMs under Test-Time Scaling cites this paper.

Ranking Reasoning LLMs under Test-Time Scaling Phi-4-reasoning Technical Report

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T13:33:18.767999Z digest=sha256:f0cc0e034c933d4b6c33dc7c3c3702527740c00b98e3312e5b52064bb9eca901

Observation c2f5d83b-9b0d-4ab0-a196-c9f7d925ff05 · inbound

Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations cites this paper.

Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations Phi-4-reasoning Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:24:08.468108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T11:23:26.852673Z digest=sha256:0cde14083a5d33788180a5566ac2783c7dc5960809a13407a8a01d14717db847

Observation 17b61778-41a9-4fb5-bb1b-cfbda779c868 · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning Phi-4-reasoning Technical Report

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:fe72307c920deda0464fb0778cb2089f297ed3eafc17396da4009e8398bf1eb9

Observation 7c3d0251-d398-4903-ad22-b42f4609d85d · inbound

Unified Deployment-Aware Evaluation of Open Reasoning Language Models cites this paper.

Unified Deployment-Aware Evaluation of Open Reasoning Language Models Phi-4-reasoning Technical Report

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:49:52.594972Z digest=sha256:c4f798aee6e98f961d70da1a8e1b1922ee70fb6b3356668f819bada8d2233d7e

Observation 6c6f73d2-f14d-4bed-b812-4956cf574142 · inbound

Unified Deployment-Aware Evaluation of Open Reasoning Language Models cites this paper.

Unified Deployment-Aware Evaluation of Open Reasoning Language Models Phi-4-reasoning Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T09:44:05.752285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T09:42:33.144129Z digest=sha256:ffc0edae8dac79ed1ad052f3eba3392656f9d0e45da898e3000d4306d059d9fe

Observation 31225b90-44ea-47b4-bbf4-f7091b96764d · inbound

ZeroCoder: Can LLMs Improve Code Generation Without Ground-Truth Supervision? cites this paper.

ZeroCoder: Can LLMs Improve Code Generation Without Ground-Truth Supervision? Phi-4-reasoning Technical Report

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:36:23.127141Z digest=sha256:68d519c73d81581d4f39beb37be2125774d7502fd99146f1f066f067b6709419

Observation 373706ed-23ef-412c-a036-4250ad0f5a2b · inbound

SeLaR: Selective Latent Reasoning in Large Language Models cites this paper.

SeLaR: Selective Latent Reasoning in Large Language Models Phi-4-reasoning Technical Report

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T18:27:36.132030Z digest=sha256:5f48d8dd3a11a8964e64364827978a607df8393664c5e94d33f740eb32e2e6ce

Observation b5991c9b-da4e-4930-94de-065ab2571ab5 · inbound

When AI Models Become Dependencies: Studying the Evolution of Pre-Trained Model Reuse in Downstream Software Systems cites this paper.

When AI Models Become Dependencies: Studying the Evolution of Pre-Trained Model Reuse in Downstream Software Systems Phi-4-reasoning Technical Report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:26:31.894139Z digest=sha256:9af63b6fa14f3083d2deb3a681f81e0735afcc3f3e4deb830c4c65bee1d1a904

Observation 9b53877c-5c07-4ecd-95b9-a1dc1add3399 · inbound

VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation cites this paper.

VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation Phi-4-reasoning Technical Report

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-07T16:46:01.749101Z digest=sha256:e5519393b5f1bcb65be1961bc8c2c861f88e6e2d4e797651caf73a440e6419a1

Observation 53a9fc33-c0a5-4561-aaa9-d6953f4530d4 · inbound

XekRung Technical Report cites this paper.

XekRung Technical Report Phi-4-reasoning Technical Report

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T20:55:10.400291Z digest=sha256:dc65ecfaae498ec5948eaa42f0c272f3c4737cebac93c0e25f4cbbc5852fb1da

Observation 2e4d2d91-0987-4f84-b766-cc5cca0c8ffb · inbound

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs cites this paper.

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs Phi-4-reasoning Technical Report

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T19:18:42.039311Z digest=sha256:1df781809cbc1ffdac3fbed538e0e6d420181680b970e104c2902c97f8bafad5

Observation ef6863f1-78cd-43e5-ade2-5797199b0003 · inbound

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs cites this paper.

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs Phi-4-reasoning Technical Report

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-19T18:22:43.544420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T18:19:01.759230Z digest=sha256:35aef48d54761b6a0ab85dc784b9aac42e9cc4d96f0024a2d8c0b586da5cb07c

Observation 9c0f06a0-e1fe-49c9-9cc7-5f7779f47031 · inbound

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding cites this paper.

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding Phi-4-reasoning Technical Report

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T16:29:05.186607Z digest=sha256:8d3390a8f1c6025f208dc1715490fdd21ab01245a97f8fa9fc9339abdd2b3fc4

Observation 79a9a8d6-b714-494b-bb18-01b84c5dda6c · inbound

Chain-of-Thought Reasoning Enhances In-Context Learning for LLM-Based Mobile Traffic Prediction cites this paper.

Chain-of-Thought Reasoning Enhances In-Context Learning for LLM-Based Mobile Traffic Prediction Phi-4-reasoning Technical Report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:52:07.910558Z digest=sha256:2ca7c9994f8500f89608e1f1d97415101e0a6a66e8670efff670210c3f7762c9

Observation 6acf9e8c-0c6a-46a8-87b0-0b2e76b51ec4 · inbound

TwiSTAR:Think Fast, Think Slow, Then Act,Generative Recommendation with Adaptive Reasoning cites this paper.

TwiSTAR:Think Fast, Think Slow, Then Act,Generative Recommendation with Adaptive Reasoning Phi-4-reasoning Technical Report

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.972991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T01:40:04.856714Z digest=sha256:2d491ce2aeca4e6fc7e3293e5c812bed877e93ae98f5e0b36c867143c8035034

Observation b38ec3d3-e27d-4ceb-9e5d-0d87f900f57f · inbound

Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making cites this paper.

Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making Phi-4-reasoning Technical Report

Reference 107

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:53:23.430395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T14:48:55.203993Z digest=sha256:24b4c97999348e083ed7ee663b13d794096cd19d6cbb1c6d26b894ed72ad89b7

Observation 26af29a6-5d40-4cc0-8069-809003dc0b7c · inbound

TRACE: Trajectory Correction from Cross-layer Evidence for Hallucination Reduction cites this paper.

TRACE: Trajectory Correction from Cross-layer Evidence for Hallucination Reduction Phi-4-reasoning Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:28:11.826210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T10:28:03.944764Z digest=sha256:8ef61a275a1def673490a22bb82f76c667b21bf74ace90077d6be9407e3828b3

Observation 1f7a6f04-65d1-4fd2-b627-b0c4b1db699a · inbound

Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution cites this paper.

Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution Phi-4-reasoning Technical Report

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:48:05.949143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T06:43:23.870127Z digest=sha256:64309975072591c0d0fb8c55bc53bad7864f96d70830f43d2dd5404aa66268e3

Observation fd29940f-c883-473c-98cf-fbdb501a3432 · inbound

CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning cites this paper.

CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning Phi-4-reasoning Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:28:04.988837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T05:25:51.655208Z digest=sha256:315576011b1d2a7d4f627d7a645e597e30caff8a6968a5d4215038735aa6fa0e

Observation 9f0cc0ec-4644-405c-ae32-c1e8124165df · inbound

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning cites this paper.

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning Phi-4-reasoning Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T08:11:17.051924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T08:10:55.720464Z digest=sha256:41adb85d332b70cb8c22059a9a0a290d855711155ee9511436b1ced5eba34fba

Observation 3c9db0b1-dd96-4e22-bbf4-4e04e3685f9f · inbound

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning cites this paper.

OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning Phi-4-reasoning Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:50:24.328443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T05:47:33.925413Z digest=sha256:8ecb715d15baf115eadd5d83a0695450d73c038fea30a806d6e63e1d4d944d78

Observation 7dbd29ac-5ccb-466c-8c22-9e5bdf90b7a5 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Phi-4-reasoning Technical Report

Reference 235

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.454774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:534da5bcfa7eed77d6f83e4e1e049126e155997ce47dab6a832cd0fb6d3bf34c

Observation 0dab8cde-a0a2-4e65-ba22-f98d44567788 · inbound

Rethinking Molecular Text Representations for LLMs: An Empirical Study cites this paper.

Rethinking Molecular Text Representations for LLMs: An Empirical Study Phi-4-reasoning Technical Report

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:06:26.456584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T11:18:47.561764Z digest=sha256:6277037bff74ba8bc6014e9c49bc2d4131725ad631514abd7bb65c513be5c1c8

Observation 731654d1-56e4-40de-9631-570d2e7fbbf6 · inbound

When to Think Deeply: Inhibitory Deliberation for LLM Reasoning cites this paper.

When to Think Deeply: Inhibitory Deliberation for LLM Reasoning Phi-4-reasoning Technical Report

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T13:36:59.359240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T01:07:47.737574Z digest=sha256:35e5b2d4ac776b6df7505ee2a4f62b869ed03c87047826dd6393506f9cd0093f

Observation aa0a2c45-257b-469e-9752-70b6c9fd9a24 · inbound

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning cites this paper.

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning Phi-4-reasoning Technical Report

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T02:17:34.947472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T16:02:04.749003Z digest=sha256:dc4967d430e959e61465cb6f160bc9d140ab5aea789842bc42cf8d01309aba03

Observation fc87d099-da85-4bec-8d63-b1f8c266294b · inbound

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It cites this paper.

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It Phi-4-reasoning Technical Report

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T13:10:55.926358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:08:57.218711Z digest=sha256:ff091007dfa7beace4b260aedb239244e479ee9deb9b18d52f6a7fb702cb57a6

Observation 7d1d8372-8a79-4d8e-9e06-6e1433bf4fa3 · inbound

Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs cites this paper.

Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs Phi-4-reasoning Technical Report

Reference 109

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T22:52:45.529339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T22:49:03.124658Z digest=sha256:063f8ca5cd9264dc969c1b23d6fd1a6b3a43db38916d4d4f8d9263fcae611a8c

Observation 23584aa2-013b-4936-9d59-1e74d3f2c407 · inbound

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning cites this paper.

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning Phi-4-reasoning Technical Report

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T01:00:19.860001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T00:59:50.038405Z digest=sha256:ec4694853672810c5a12670a26aa53e7f04512152bc65b65e7084fbb1649c123

Observation ab31bb19-93ae-4bcc-a5bc-98c156ae0781 · inbound

ThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought Graphs cites this paper.

ThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought Graphs Phi-4-reasoning Technical Report

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T09:24:32.205313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-30T09:23:26.386399Z digest=sha256:e63050d03b06e6e62757a929aac143f6c7e5eb821c7e1a789c77b851e59a161f

Observation 71834db7-0c7d-4d5c-b9db-c7e3ed0bd1ab · inbound

Benchmarking Large Language Models on Floating-Point Error Classification cites this paper.

Benchmarking Large Language Models on Floating-Point Error Classification Phi-4-reasoning Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:15:44.918943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T05:39:01.180338Z digest=sha256:aba333e554e2daf8392b4d41815115625773bd12992a20a8d3885b4a0ec5a198

Observation a40a5c98-377b-44d0-ad5a-56ecf542cd22 · inbound

On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens cites this paper.

On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens Phi-4-reasoning Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T10:26:08.486717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:26:08.486717Z digest=sha256:f9d0379f75bf863301a0f0ba56b6bf1ea8da8e7b78fcbb36313a83823200e707

Observation 5a1b21e5-e5f4-4521-aeb6-f6de3eb6b26f · inbound

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer cites this paper.

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Phi-4-reasoning Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T16:21:12.065335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:21:12.065335Z digest=sha256:b9430271b7c22e18603fa2da1e4f845aac3983a3cbd700c5b805303ee27f704e

Observation 35da46c7-9b11-403c-b1cb-b2cdd374ecdf · inbound

Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates cites this paper.

Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates Phi-4-reasoning Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:48:28.952743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:48:28.952743Z digest=sha256:144fc720c0e5ac063c96e0cc9b46fc29211b4e8e8de8e746bbbfdb8958552e4f