Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:29:10.888735Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 0 inbound Pith citation observations for arXiv:2502.07237.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:29:10.888735Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
88 of 88 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 23139c04-27bb-4f7e-813b-580216ed1d9f · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization De novo drug design using reinforcement learning with graph-based deep generative models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d0e9410-cafb-4473-8484-46ad1e2101fb · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization DrugCentral 2023 extends human clinical data and integrates veterinary drugs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 094e99f4-dcb2-4b13-bcd7-7b6683564a28 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization MolGPT: Molecular generation using a transformer-decoder model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccb5f472-fa15-445f-bb50-8fa447178299 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82d291ee-f4ce-4ba9-b9d1-b0baf5213678 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Constitutional AI: Harmlessness from AI Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb99ddb-7d6c-4982-8791-c187bf18704c · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Why is tanimoto index an appropriate choice for fingerprint-based similarity calculations? Journal of cheminformatics, 7:1–13, 2015
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cec44a0-b99d-494d-894c-40502357271f · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Molecular similarity: a key technique in molecular informatics
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 908a204e-7b65-4486-828e-6f67e9028956 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Better Rewards Yield Better Summaries: Learning to Summarise Without References
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36739dfb-d947-46c7-a108-9d1f8d560c50 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Paccmannrl: De novo generation of hit-like anticancer molecules from transcriptomic data via reinforcement learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a097645-0b34-49a0-99ab-973421e42fee · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Transformers and Large Language Models for Chemistry and Drug Discovery
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6a5f96de-cb61-450e-bbcb-094a52ef9aaf · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Offline rl without off-policy evaluation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5663653-79bf-4e7d-83ff-f771d6600718 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Safe learning in robotics: From learning-based control to safe reinforcement learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b206893b-d305-46af-9ba2-a6d27ec1fd6e · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Policy improvement via imitation of multiple oracles
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ffe4e10b-1f35-416f-9024-d5175ba07ca6 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Deep reinforcement learning from human preferences
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6d6b85ae-ff03-40d4-8545-30fe36030596 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Ai-accelerated protein-ligand docking for sars-cov-2 is 100-fold faster with no significant change in detection
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8011063-e22b-4fe0-b4c1-f863cc232dce · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization The cost of new drug discovery and development
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6408ef31-5451-4b82-9dd6-5bc38750038f · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbbe1424-7a55-4100-9f5e-7cd3bad2a6dc · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e225d3ef-bfcd-48bb-a2cc-4e3f05adb878 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Neural scaling of deep chemical models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99205759-fa1f-4eba-8e4d-fd3042074733 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Mimosa: Multi-constraint molecule sampling for molecule optimization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a7f36a5-e36b-4f42-b179-3ffa4f7b6085 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization A new algorithm for data compression
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b376e5c0-8b05-4431-ad25-18d3f51c0417 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Scaling laws for reward model overoptimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e21d29d2-23bd-4f8d-906e-1ae1ef185db7 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Learning to navigate the synthetically accessible chemical space using reinforcement learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d32b5402-a510-45ef-b87b-616781316bb6 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Objective-Reinforced Generative Adversarial Networks (ORGAN) for Sequence Generation Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7304e96-c6e4-4c95-91ca-e36e862cec1b · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Covid-19 vaccines and variants of concern: A review
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0e52c434-b836-469b-ab68-9ec432686a76 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Learning from Dialogue after Deployment: Feed Yourself, Chatbot!
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 114d79ca-e711-4826-a1ba-021da26e1722 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Molecular optimization by capturing chemist’s intuition using deep neural networks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dde8d5c8-2416-45be-b872-05cb4270c05b · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Transformer-based molecular optimization beyond matched molecular pairs
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5c1a2ea6-fde2-4a2b-9334-ec8c4e6932d3 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Reward learning from human preferences and demonstrations in atari
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ae7472bf-2b7d-46c6-8e9a-b17704c9eefa · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization The distribution of the flora in the alpine zone
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa297f5a-ade7-4a7d-b120-6cebc83fd29e · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fece7bd0-12ae-4d39-b813-073e94735335 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization A graph-based genetic algorithm and generative model/monte carlo tree search for the exploration of chemical space
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d3aabbc-ec55-47f2-a8b1-249e90dd010c · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Multi-objective molecule generation using interpretable substructures
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cb0cd3ed-101b-4b17-afec-2b76042e4c5b · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Posit: flexible shape-guided docking for pose prediction
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a001709c-5cf6-4e19-b714-4d43d8d6c97e · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Giraffe: Using Deep Reinforcement Learning to Play Chess
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7eed6fe-8191-457f-afc1-e099f4c25882 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization RDkit: Open-source cheminformatics software
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a663836-d1df-4baa-96fd-3b8c2d7ed932 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Rdkit: A software suite for cheminformatics, computational chemistry, and predictive modeling
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8dccded-8c51-48f6-b699-731479c7493b · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Scalable agent alignment via reward modeling: a research direction
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a7dd7a7-bef0-403e-a139-deb71922ae6e · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Drugimprover: Utilizing reinforcement learning for multi-objective alignment in drug optimization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 45814632-f5bc-4067-a44e-6ebf4a9eda45 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Blending Imitation and Reinforcement Learning for Robust Policy Improvement
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68f1ad5-5c2d-4a3e-9463-91497975427c · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Active policy improvement from multiple black-box oracles
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cbe878d8-589b-4a63-930d-8bb6f217e96f · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Entropy-reinforced planning with large language models for de novo drug discovery
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation abe3419c-4ce9-4b80-b31d-9602a5d27bf9 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Drugex v3: scaffold-constrained drug design with graph transformer-based reinforcement learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f23e3ed9-0fd9-432a-96a1-000f1f2be132 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Reinvent 4: Modern ai–driven generative molecule design
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 92dd2be8-1865-48ae-946d-fe2f337ae38d · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization The different mechanisms of cancer drug resistance: a brief review
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e480c6e2-d2ed-4c25-b705-2a498bc9e1ba · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Playing Atari with Deep Reinforcement Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 338e75be-78b2-4fd9-b61d-457709dea35e · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Exploring deep recurrent models with reinforcement learning for molecule design
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b4507f40-8c05-4b63-893c-b8f11b8e39df · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Molecular de-novo design through deep reinforcement learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 63ed8f5a-6c28-4931-880e-9b09de5fc5cb · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05bb62f3-110c-4816-afd9-60a17d77301a · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Combinatorial enumeration of groups, graphs, and chemical compounds
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 324daebc-83d2-460e-9377-5d455a92e951 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Alvinn: An autonomous land vehicle in a neural network
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 04c6d37d-e073-472f-8d33-a72063f62fd4 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Deep reinforcement learning for de novo drug design
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6360c346-a16e-4542-b7ff-68aeffca25f5 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Drug repurposing: Progress, challenges and recommendations
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 82924bfc-aaeb-4363-9c18-367187623f97 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6274477-7f12-4eb8-8b7f-b78cc1b06e12 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Learning by playing solving sparse reward tasks from scratch
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9cf7afa3-b90c-453f-9ab1-185431ee459f · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Extended-connectivity fingerprints
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 923c1ada-3d19-43c8-bf55-60a20962d984 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Reinforcement and Imitation Learning via Interactive No-Regret Learning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f3302b8-5f07-40e5-a7e3-74df4faaa6cd · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization A reduction of imitation learning and structured prediction to no-regret online learning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 08c26f04-fe45-446c-bb7e-c8d7bbe3b902 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization C5T5: Controllable Generation of Organic Molecules with Transformers
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e17aae1f-ad39-4e19-ad42-6afc84602047 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Neural Machine Translation of Rare Words with Subword Units
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5178c9a0-2e29-46b7-b83a-c5196241925f · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Preference Ranking Optimization for Human Alignment
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10823088-7042-4c39-b78c-4c9fde3ca91f · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Deep reinforcement learning for multiparameter optimization in de novo drug design
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 37ee144e-2f30-4b2f-b5af-21356c0f20be · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization ZINC15–ligand discovery for everyone
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 44a38208-e0ec-4566-a5d4-4d44c5fc3455 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Learning to summarize with human feedback
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 15067d09-f0c6-4d66-9de7-7883d3521f85 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Molsearch: search-based multi-objective molecular generation and property optimization
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 469feec5-13f5-47f2-8e77-8d19c8129579 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Policy gradient meth- ods for reinforcement learning with function approximation
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d3e42ab1-d33a-4709-affb-8a0690452915 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Reinforcement learning for systems pharmacology-oriented and personalized drug design
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 202f7111-55b2-412c-a7ae-549f1931acfe · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Drlinker: Deep reinforcement learning for optimization in fragment linking design
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e2ce34cb-fe22-4df6-a1f3-e20a89ae72cc · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization DeepMind Control Suite
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40b71e8d-b632-4ee5-9a72-464d617a31a7 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fec0ce4c-f3bf-4c3c-a04a-ce7dfaaf87a0 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Matched molecular pair analysis in short: algorithms, applications and limitations
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9a8ed734-2a4a-471a-84c4-28b01a56a7cc · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Benchmarking language-based docking models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8bf24568-87e4-47ae-9c4b-bf637c89ade6 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Attention is all you need
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52f6e24d-3e2b-4499-baeb-dc4bc2aab748 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edcf64a5-0810-4c24-a6e8-afa7a4f859d0 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization A reinforcement learning approach for protein–ligand binding pose prediction
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f994b61a-7269-49dd-8090-b8f1ec351a08 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Smiles, a chemical language and information system
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8b7c7211-8225-4e6a-b98a-6183fba0b0d5 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Simple statistical gradient-following algorithms for connectionist reinforce- ment learning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 58fa5c6a-c918-4eaa-9ef0-b8dc150444d3 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Recursively Summarizing Books with Human Feedback
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af742400-5b96-4b63-823e-c8c6b4b59aca · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Rlcg: When reinforce- ment learning meets coarse graining
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dba36a93-4910-49b2-9b99-285bbd402e00 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Towards Coherent and Engaging Spoken Dialog Response Generation Using Automatic Conversation Evaluators
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1231f94-5d75-4426-8925-59a0c3e6d2bc · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Population-based de novo molecule generation, using grammatical evolution
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 45f5e9b4-6aea-412f-9781-7ea6f37df6f9 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Graph convolutional policy network for goal-directed molecular graph generation
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3ebf7b5d-0ca2-4611-99a3-a929bb3607b9 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0959e6bb-b69f-4519-b139-d1afefd35b8d · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 910f057e-941d-4419-be53-bfbc0764eb8d · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Covid-19 pathophysiology: A review
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e9b7808-5ea1-409d-92de-313d0a645cb9 · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Universal approach to de novo drug design for target proteins using deep reinforcement learning
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eef72f93-4f95-40c2-8916-4eb13c4799fc · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Optimization of molecules via deep reinforcement learning
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b5fe6846-2656-4838-8bc3-70065bc8adff · outbound
DrugImproverGPT: A Large Language Model for Drug Optimization with Fine-Tuning via Structured Policy Optimization Fine-Tuning Language Models from Human Preferences
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.