Pith. sign in

Paper Citation Record · LEDGER

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language

As of 19 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.09789.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09789 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T15:42:46.891891Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2ec28d7-6270-4dd0-ad33-bc8e518c2b07 · outbound

This paper cites Recent improvements of the particle and heavy ion transport code system—PHITS version 3.33.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Recent improvements of the particle and heavy ion transport code system—PHITS version 3.33

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:e9eae83933c5f9cbbe1edcc853d0187655270b9601827de388fd506d8d1423de

Observation 02ce4a4b-15f0-4585-b5da-7486e31b0fec · outbound

This paper cites MCNP version 6.2 release notes.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language MCNP version 6.2 release notes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:ba868917223ab94eb83ca01b021e2ee0ea832fa600f89bc7ceb2f2fa542182b1

Observation 293cf11b-73e8-4b70-844b-e12dde132c5c · outbound

This paper cites GEANT4—a simulation toolkit.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language GEANT4—a simulation toolkit

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:172d3a989aea23644718b6bdd146cd66a607f7c4eeaf3e810659d0fc6a1e6788

Observation 9985ea9a-af1b-4075-89c8-2bea58df31a7 · outbound

This paper cites FLUKA: A multi-particle transport code (program version 2005).

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language FLUKA: A multi-particle transport code (program version 2005)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:60670045c058b5a1a8f957c5be7797b2e3a42d3f3c59d7791514a590c178fc53

Observation e73e0443-d4dd-457b-ba88-fb90ae8f41e3 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Evaluating Large Language Models Trained on Code

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:0262a740f1f18673d6efda4710672b1cd9d3da2a914834153bdf4d15004e6a24

Observation 70ce1b01-f659-4acd-a1f2-7c2ae610a75a · outbound

This paper cites SWE-bench: Can language models resolve real-world GitHub issues? In: The Twelfth International Conference on Learning Representations (ICLR); 2024.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language SWE-bench: Can language models resolve real-world GitHub issues? In: The Twelfth International Conference on Learning Representations (ICLR); 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:eb2446eca49cb025dd3709ec1d5690fdb3b197468be22cb898fe6d530a191371

Observation a223f2aa-f381-4e21-a30c-4fe3f3136a0c · outbound

This paper cites DS-1000: A natural and reliable benchmark for data science code generation.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language DS-1000: A natural and reliable benchmark for data science code generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:ad3c39bc5d0c5e226f0dd1768a9f9043c900b7fd7a22087aa2ebec707fc2bddb

Observation 5369fd8b-2fd5-402d-bcdb-0abf8cf82d51 · outbound

This paper cites Invited paper: VerilogEval: Evaluating large language models for Verilog code generation.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Invited paper: VerilogEval: Evaluating large language models for Verilog code generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:9413230078a816d3f4427cbf493e830526a6c7343ab7aa7a559edb0d1a996b9c

Observation 3883d4ff-e314-4fe0-b238-3c68cd20ddf9 · outbound

This paper cites BigCodeBench: Benchmarking code generation with diversefunctioncallsandcomplexinstructions.In:TheThirteenthInternationalConference on Learning Representations (ICLR); 2025.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language BigCodeBench: Benchmarking code generation with diversefunctioncallsandcomplexinstructions.In:TheThirteenthInternationalConference on Learning Representations (ICLR); 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:b33deb36e5a5a44efc93afe61fc4f1fd295c52e57e0f82f7ab688fd603e13531

Observation aa598f56-4dc0-4f3f-9aff-e2ffed2d9ebc · outbound

This paper cites SciCode: A research coding benchmark curated by scientists.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language SciCode: A research coding benchmark curated by scientists

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:07b69735fca2fcb2b009d50a6e797f10fe0a80a90e4a68995250bd34394782d4

Observation 3f6db513-5eb9-4a35-bcd9-d5a82776a6fe · outbound

This paper cites Automating Monte Carlo simulations in nuclear engineering with domain knowledge-embedded large language model agents.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Automating Monte Carlo simulations in nuclear engineering with domain knowledge-embedded large language model agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:706e202bae3d64b6f6fa283413d649ebc253520e4a2ef827ce943c89f979d8f7

Observation 0eb98e7c-860f-4f57-8b52-c03e7e3daef8 · outbound

This paper cites A self-correcting multi-agent LLM framework for language-based physics simulation and explanation.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language A self-correcting multi-agent LLM framework for language-based physics simulation and explanation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:1201ea8ba9ec92dd3cb25af55c3faf8500058c4aa79e29aa37898c5f0d0414bb

Observation 9c8c1575-97e9-4691-aea3-c6f60217f7e0 · outbound

This paper cites Evaluating the performance of large language 19 models for geometry and simulation file generation in physics-based simulations.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Evaluating the performance of large language 19 models for geometry and simulation file generation in physics-based simulations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:f64a2bc0437906addf6e5596ea060ee8f098a9069826cda1d774f25e0556f802

Observation ac238909-decc-4a99-a2c3-24d9646e5dd7 · outbound

This paper cites OpenFOAMGPT: A retrieval-augmented large language model (LLM) agent for OpenFOAM-based computational fluid dynamics.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language OpenFOAMGPT: A retrieval-augmented large language model (LLM) agent for OpenFOAM-based computational fluid dynamics

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:e89d7359d2e5e7cea6c1fad266b1f78468800f05b101472185472d9bf1fc65df

Observation 26279d89-94a5-495a-8f10-c49953b0e576 · outbound

This paper cites MetaOpenFOAM: an LLM-based multi-agent framework for CFD.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language MetaOpenFOAM: an LLM-based multi-agent framework for CFD

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:13f921744def17a77009063159f7a6a69b6deef738494df0a38f970e61dbe8c2

Observation db51d044-e521-49b4-9d6e-8d9f7dd0fb6b · outbound

This paper cites CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:a2b6fe27d839d135fb66dbd809237d500e0ccf31a62c679ae2df4216de999ff9

Observation 47064579-345f-4d28-9dce-f5a7b2285cb4 · outbound

This paper cites MooseAgent: A LLM based multi-agent framework for automating MOOSE simulation ; 2025.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language MooseAgent: A LLM based multi-agent framework for automating MOOSE simulation ; 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:d8cf3fb1626848c589a0b1dd74b8fb97d36a83205278d62d2066379278837773

Observation ade3ced7-7bf8-49f6-8c9c-45359b431a77 · outbound

This paper cites AutoSAM: an agentic framework for automating input file generation for the SAM code with multi-modal retrieval-augmented generation.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language AutoSAM: an agentic framework for automating input file generation for the SAM code with multi-modal retrieval-augmented generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:90ca629716d0504a5851ed5c7e3ec18255d894f64263668506d163343b9991f3

Observation fcf21966-a402-4969-a9e8-e9c892463ce2 · outbound

This paper cites an unresolved cited work.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:9d970ad35cdd69e81af0e797188bc3955f35e52ac9da654f9671058e3070f3c5

Observation 2b838f9c-ea72-49c4-99a7-ddd74b777c1e · outbound

This paper cites SIMCODE: A Benchmark for Natural Language to ns-3 Network Simulation Code Generation.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language SIMCODE: A Benchmark for Natural Language to ns-3 Network Simulation Code Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:62062687bdc580855e616c59b063157e17799513ffd56ae760f3af171dff8b53

Observation 44b8a582-24c6-44dd-aa3c-ad2baa6257cd · outbound

This paper cites Evaluating LLM-generated code for domain-specific languages: molecular dynamics with LAMMPS.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Evaluating LLM-generated code for domain-specific languages: molecular dynamics with LAMMPS

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:3bfc7ddd8e9648d22956261907ecc6ff8321950efac02bae077093be93da5b7c

Observation 3d8dc4d3-787a-4b7f-9c6c-51e4f7f86467 · outbound

This paper cites GRACE: an agentic AI for particle physics experiment design and simulation ; 2026.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language GRACE: an agentic AI for particle physics experiment design and simulation ; 2026

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:aa4b6a167b2401e98a27a49504ab01ea8c6b2f863c7bc9cabed25a4c5a9c2a2d

Observation 5511e7fb-efc0-4a9e-8209-02e513c962db · outbound

This paper cites Exploring the Capabilities of the Frontier Large Language Models for Nuclear Energy Research.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Exploring the Capabilities of the Frontier Large Language Models for Nuclear Energy Research

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:753f053cd92bc5604873848ad308141c300074d698b63a2ecc7120abb47fa199

Observation 35839785-28fd-486b-9ecb-22dda7be9f28 · outbound

This paper cites DocPrompting: Generating Code by Retrieving the Docs.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language DocPrompting: Generating Code by Retrieving the Docs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:a777dce6f5c522c120db9cc2d55bc43c176cdda436478881e4d107f7a2b080c4

Observation ed989457-4c9a-47d9-8abd-7668b2dd119f · outbound

This paper cites DocCGen: Document-based Controlled Code Generation.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language DocCGen: Document-based Controlled Code Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:fd764ed56ec00e74e8aaa57c705c0f3f00a3e77c70e078e08eaeb348b0a47191

Observation 1e1b44e7-7e3e-4872-ad10-01ae6563dd02 · outbound

This paper cites MetaGPT: Meta programming for a multi-agent collabo- rative framework.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language MetaGPT: Meta programming for a multi-agent collabo- rative framework

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:268ddf38b3467286b15f22bd8126908afa2308ef2ba60934e36e07ea1de9689f

Observation 15c7861c-7add-4412-a054-9af9e2760ced · outbound

This paper cites AutoGen: Enabling next-gen LLM applications via multi-agent conversation.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language AutoGen: Enabling next-gen LLM applications via multi-agent conversation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:ca8a2b9e46512ef1fb0f488624803111ebf6a664a2437a251712feff07b37d58

Observation 6f42b1bd-8b7f-4ad2-8848-b9f1feb20171 · outbound

This paper cites SciAgents: Automating scientific discovery through bioinspired multi-agent intelligent graph reasoning.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language SciAgents: Automating scientific discovery through bioinspired multi-agent intelligent graph reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:83ba0bf98ea712ec305a27f9c901aff113f23c62322a332fe996d9e13d2834c2

Observation b28d99c9-199b-44fd-a2bd-2f5f8ad4f558 · outbound

This paper cites SWE-agent: Agent-computer interfaces enable automated software engineering.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language SWE-agent: Agent-computer interfaces enable automated software engineering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:3b18cc15b25bc33a44aa1ef9d17e4ca0ea37fc2e1af3e9accb1a9a8673949ae1

Observation aab0b677-b066-4aac-b79c-f88ed43ce24d · outbound

This paper cites ReAct: Synergizing reasoning and acting in language models.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language ReAct: Synergizing reasoning and acting in language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:e233cdfc990470490df7d4caa75373d3895acefe058c71a7b79d6eee3a69016f

Observation 77893e61-e976-4871-8ceb-b54da612f8a7 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Self-refine: Iterative refinement with self-feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:00903369fb82025e47c49f17f4ca17bfff0778a16ece396887d0f71878b91aac

Observation 3650f9e5-6de9-45a2-a462-2c478eb25457 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement 20 learning.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Reflexion: Language agents with verbal reinforcement 20 learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:7266b942282e1e0b5e1c487651b3155ba3161b2a60dbd7b0bb2bc0d5b0a00459

Observation bddc33ad-7626-4c93-b482-0c8266373cea · outbound

This paper cites Large language models cannot self-correct reasoning yet.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Large language models cannot self-correct reasoning yet

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:692c18f16a902fd7abba5457f61960cc237ef1e627a61c53fefc2344541604cd

Observation 40e6902e-b75a-44bd-ba15-268027f2cd00 · outbound

This paper cites Is self-repair a silver bullet for code generation? In: The Twelfth International Conference on Learning Representations (ICLR); 2024.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Is self-repair a silver bullet for code generation? In: The Twelfth International Conference on Learning Representations (ICLR); 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:98ce3afee3f39d6de4843d0196f79763e3ff1cc6a817f9b9477290705e426ff7

Observation 4bdb19ce-56ff-4bde-b79f-e1e2f2497bac · outbound

This paper cites When can LLMs actually correct their own mistakes? A critical survey of self-correction of LLMs.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language When can LLMs actually correct their own mistakes? A critical survey of self-correction of LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:3d49116fb940b5c9c979558fa40ec85727de94746b7f0c8852dbb4bea985b067

Observation 65e5bce6-6bc3-47e8-a30c-e4b7f8017a84 · outbound

This paper cites LiveCodeBench: Holistic and contamination free evaluation of large language models for code.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language LiveCodeBench: Holistic and contamination free evaluation of large language models for code

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:88a761304fb5fe7ec8526566a60baa22887357d0af6cf4b865b06bc1fafb9c9c

Observation bf91b133-af1e-49c7-ae73-c58328577741 · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:ebf2e88af4c89215e0ed24bd6799e72bef0597f94a7f43a4ef6c857da48a61fd

Observation 6ca96384-62a1-4cea-aa1d-2d3c332f1651 · outbound

This paper cites Codex CLI: Command-line coding agent [https://github.com/openai/codex]; 2025.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Codex CLI: Command-line coding agent [https://github.com/openai/codex]; 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:03882cd50a6b483d5a79e2a4611d3e3b905b2af9c9acb45a9e8fcbfc72713c58

Observation 8ffd9203-55d5-4113-bb45-fc810badda2f · outbound

This paper cites OpenAI Agents SDK [https://github.com/openai/openai-agents-python].

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language OpenAI Agents SDK [https://github.com/openai/openai-agents-python]

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:66d084daf9fb7fe0dbe09e68493faceaf742233fd94997f0fac871df75010326

Observation 968c8a8d-0454-4d0d-855b-0ecb0dbf56a2 · outbound

This paper cites an unresolved cited work.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:143b3f51ee467c1ca71b6106e36eecdefd500432f62ac672fa2da94f5f5c2716

Observation c7f85bdd-9ed9-4a06-94d9-dbf1870a0fce · outbound

This paper cites XCOM: Photon cross section database (version 1.5) [National institute of standards and technology, gaithersburg, md]; 2010.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language XCOM: Photon cross section database (version 1.5) [National institute of standards and technology, gaithersburg, md]; 2010

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:342e058f076d33e6250b18f3f4328dc7c125a1a5ac18963a1530218c1155025f

Observation 756cc65b-d048-4e7a-b306-b23f4059fe35 · outbound

This paper cites Nuclear data sheets for A = 137.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Nuclear data sheets for A = 137

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:635bdf7aa4ca76d238404c9c474c2d8e0e6bd34bacf3269da24289b42fe4fb72

Observation b50c2e4a-5bef-46a4-9afc-9b8305f862ef · outbound

This paper cites Benchmark study of particle and heavy-ion trans- port code system using shielding integral benchmark archive and database for accelerator- shielding experiments.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Benchmark study of particle and heavy-ion trans- port code system using shielding integral benchmark archive and database for accelerator- shielding experiments

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:ad392e7eeedb8e6f9ef895e1dd5f037ed37b15a07269891ce989c0dbcd3ef9a1

Observation 6803b39b-0930-40ce-b46f-27ca02a5b544 · outbound

This paper cites Validation of the physical and RBE-weighted dose estimator based on PHITS coupled with a microdosimetric kinetic model for proton therapy.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Validation of the physical and RBE-weighted dose estimator based on PHITS coupled with a microdosimetric kinetic model for proton therapy

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:acff71b3c1176110a4701d6ca1696bc5335afae899ece5305bc2a538a4d820e0

Observation 66efbf9a-2390-4974-932a-05a3afcc2523 · outbound

This paper cites Improvements in the particle and heavy-ion transport code system (PHITS) for simulating neutron-response functions and detection efficiencies of a liquid organic scintillator.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Improvements in the particle and heavy-ion transport code system (PHITS) for simulating neutron-response functions and detection efficiencies of a liquid organic scintillator

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:c542f4bbc880fc15c4105ba7e357162d60a191f2023d9498db5e8395fa7a12ca

Observation eb0acfbc-db60-42de-8e41-4b4f0a01d4b6 · outbound

This paper cites A PHITS-based computational model of a TRIGA-fueled subcritical reactor for gamma dose mapping.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language A PHITS-based computational model of a TRIGA-fueled subcritical reactor for gamma dose mapping

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:7a376c86ec425dc9fd77ebdaa84b21cb752d2eba2ffdf95dae8f05237e98ad92

Observation 75b2fa6e-a836-43e6-bfc1-d3c6ebcfca5d · outbound

This paper cites Dose estimation for astronauts using dose conversion coeffi- cients calculated with the PHITS code and the ICRP/ICRU adult reference computational phantoms.

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language Dose estimation for astronauts using dose conversion coeffi- cients calculated with the PHITS code and the ICRP/ICRU adult reference computational phantoms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T15:42:46.891891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:42:46.891891Z digest=sha256:49a57b28a5793737c870a921439b277ac20a6d68abc60e45b49e6933955795c0

Pith citing papers

No inbound Pith citation observations are available.