Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T10:49:10.293120Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 100 of 229 outbound references and 0 inbound Pith citation observations for arXiv:2606.20295.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T10:49:10.293120Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 229 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 346564ca-59ff-4047-9b06-31b8fc02ed89 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models A Survey on Large Language Model Benchmarks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac064b73-c90d-441e-a1ce-8c7c424f3e38 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models A Survey on Evaluation of Large Language Models.ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3daf6302-7309-4785-b7df-097ee5615d15 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Measuring Massive Multitask Language Understanding,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4347beea-dc8c-4661-b59e-046758a29ef1 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a04cf945-606c-4710-8c70-40c9ac5b0c0b · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Holistic Evaluation of Language Models.Transactions on Machine Learning Research, 2023
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ed8bc40-233a-48ce-aad9-060219e30554 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Artificial Analysis intelligence benchmarking methodology, n.d
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce52ce6a-7e54-4785-a706-b41d3398dd93 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b2b587d-bb30-4052-b4b4-56a80ca39ed6 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Leaderboard, n.d
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a8ba6e-f885-435d-900c-89245db38013 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Quantifying Capability Boundaries: An Application-Driven Analysis for Large Language Model Selection
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d7513cb-f09b-4156-81c9-332fb0536e86 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Quantifying the Capability Boundary of DeepSeek Models: An Application-Driven Performance Analysis
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd96d426-2062-4af3-8eb5-3dcc48540390 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models What is the Best Model? Application-Driven Evaluation for Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 278cf955-18e8-4a38-846d-6f24c43754f3 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Pinchbench-upgraded, n.d
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b3b4dd-c93f-4784-af96-2b00e2aa1c2b · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models OrchestraLLM: Efficient Orchestration of Language Models for Dialogue State Tracking
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70059fd2-b878-4e06-b9f8-687ccd509a2d · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Semantic Router, n.d
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbcaa8fe-68aa-4ed4-9b7a-b8260444c222 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models LiteLLM: Python SDK & Proxy Server (AI Gateway) for 100+ LLM APIs, n.d
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b1f1821-007f-48bc-98d3-e62ba70c0699 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be4729ec-b1f0-4633-b381-2c3a8941c1b2 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models RouteLLM: Learning to Route LLMs with Preference Data,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c03313af-a61b-45f0-baf8-60dd239c29ea · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Large Language Model Routing with Benchmark Datasets
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e44ef3b8-fcb6-40c4-b8b0-6177b8974505 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d47765f-738e-477a-bdc4-20b7e223ca64 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Fusing Models with Complementary Expertise
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 036338fb-daa8-4603-937e-3c81b675f8a9 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Understanding intelligent prompt routing in Amazon Bedrock, n.d
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d304121e-8b6f-4772-a6b5-89628dc9af0f · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Introducing Martian - Better AI Tools Through Better Understanding, 2023
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 806ccf7f-fe27-4f20-baec-9020e33e7e34 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Model Routing for Agents, n.d
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d432253e-97b3-435f-a43b-a8444d0616d6 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models OpenRouter — One API for hundreds of models, n.d
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09d6aac1-ef02-4ce4-ae68-0c5331376dcf · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Introducing GPT-5, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01c7ef23-61b1-4fd3-a1ea-831c2e9ed319 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Using LLM intelligent routing to improve inference efficiency, n.d
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 042483ec-bf50-4bb0-997d-c5affd790eb6 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Intelligent model routing, n.d
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af0c6fba-9c15-41c4-8f03-c1d9793396ae · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79bf9a32-a2f2-425a-9005-ddfeea326fb8 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74bcfd1e-f01f-467c-946d-540790e449e1 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models AutoMix: Automatically Mixing Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88e93f93-4fcc-44c5-a66f-7c0003102685 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Tabi: An Efficient Multi-Level Inference System for Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6264491b-e95f-4a74-b743-0f06f69208e0 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models EcoAssistant: Using LLM Assistant More Affordably and Accurately
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10871e84-2b9b-4ac4-83e3-c3b2005c3bf8 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Fast Inference from Transformers via Speculative Decoding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9e35fe6-3962-42dd-9395-7a86a4d4607f · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Unity AI Gateway: Configure Fallbacks on Model Serving Endpoints, n.d
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 073a4f60-e68c-4104-86f6-e7a96754650e · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4118e3b-b6d9-4735-8b11-20dc79d87484 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models More Agents Is All You Need
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6682f2fa-11f2-4a3c-a781-50fe2fe608fb · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f90e1a6e-56d7-49f7-9abf-ceb3e3203cfc · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Knowledge Fusion of Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76875b0a-987d-4f9b-a086-c5f87fe6ce37 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Mixture-of-Agents Enhances Large Language Model Capabilities
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 085e036c-9f0c-40c6-9414-6594d336d71c · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Efficient Attention Mechanisms for Large Language Models: A Survey,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79f63b3b-284d-4d6e-8e1f-17d70b53b74c · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Qwen3 Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23dac4b8-f68d-487f-9393-c409c9a40f0a · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepSeek-V3 Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28e88532-8b0d-4dfb-98ae-1692befcf0fa · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 331b9e16-3160-419c-b748-f7530878bd32 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30edab7f-fbff-482e-af37-6115862bf3a6 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9857e8ce-7970-4278-8657-bb80e5bcc916 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepSeek-V4: towards highly efficient million-token context intelligence, 2026
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d1c1899-2a10-4a33-9645-b78c545f7387 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models GLM-5: from Vibe Coding to Agentic Engineering
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 796b09f8-23b8-48e5-9040-b08e1886d4fe · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b5f8fb-b3aa-4aa1-ad94-52d9983f47e5 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models MiniMax M3: Frontier Coding, 1M Context, Native Multimodality in One Model, 2026
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8792c76f-5bc4-4e7a-9eae-ce6a4bc2a184 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Qwen3.5-Omni Technical Report
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4005949f-3263-4849-836f-7cb216e74948 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21ec33ad-7dc9-4131-9cb8-3c9bdb8e2892 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models A Survey on Mixture of Experts in Large Language Models.IEEE Transactions on Knowledge and Data Engineering, pages 1–20, 2025
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 752887d1-532a-413d-8a6c-4370820297b3 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d4d7277-04dc-426f-a06f-fb2047358bb9 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fed5def-d20b-4922-93a4-9eceb84efe47 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Mixtral of Experts
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 823d9b52-e50c-4cfd-9809-119b100c5d2f · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 154f4660-b663-4a73-a7f5-aea7385bebf2 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5942b97a-fedc-4394-b856-be89cd3949c4 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models TensorRT-LLM expert parallelism documentation, n.d
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e6ae66b-3612-40d8-b652-57c992a9fed2 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09543c80-a541-4cf9-bdfb-03b6e6569b12 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Mixture of Heterogeneous Grouped Experts for Language Modeling
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f86d142c-53ba-4e73-8ddd-21e8f2d504c9 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Optimizing for the Shortest Path in Denoising Diffusion Model
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5044a919-65cd-406f-a2b2-37cc13d35210 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation, 2025
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46d8602d-7dfd-4bd5-bf8d-7762ec8a350e · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepCache: Accelerating Diffusion Models for Free
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ab818e2-6e84-40f3-9a27-5b79999ae9de · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d2edd56-123d-4481-b2aa-ba942ff026b5 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching, 2024
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b1704c-4eff-48d9-817d-57434c328938 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc6a258e-0061-4c31-9202-f35d94671705 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models URLhttps://doi.org/10.1109/cvpr52734.2025.01679
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4cea724-1380-47ac-a689-50e23834ee45 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33147e5a-3de8-44a2-bd67-d85895322ba7 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e64d0a0e-2dac-4d20-a294-4c6db1de7d35 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models STaR: Bootstrapping Reasoning With Reasoning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0deaae1-8abf-40a5-99ec-cd79efb0c0fe · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models ReAct: Synergizing Reasoning and Acting in Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c663ffe0-4920-44e0-8cc2-90ec60ad3b15 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b134d8b-b57b-4595-954f-0172c6b06cbf · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models MeanCache: From Instantaneous to Average Velocity for Accelerating Flow Matching Inference
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2dbebe4-e549-4409-80cd-5c3d59996158 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Introducing OpenAI o1-preview, 2024
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fc60caa-809f-40aa-82dd-0e771119ea81 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models OpenAI o1-mini: Advancing Cost-Efficient Reasoning, 2024
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 918cdb75-9f74-49c5-b883-8e0588268a54 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d12a6c4-a4f7-4709-a7ac-ac1098935bb7 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Claude 3.7 Sonnet and Claude Code, 2025
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92a2fd96-c260-4b28-b870-1c999ce109a8 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Qwen3: Think Deeper, Act Faster, 2025
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b778c97-dd94-4e3a-bcfb-e32edf682df5 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models PAL: Program-aided Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afebed7d-1074-402b-8971-541e6cb0e5d6 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Token-Budget-Aware LLM Reasoning
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e6b6886-b098-4cb7-a803-5960017cd5f4 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Training Language Models to Reason Efficiently, 2025
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e778285-b681-44bd-9652-4e0890b5ae65 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 539991cc-9c8d-41e7-8bd2-eaaf0a050631 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Chain of Draft: Thinking Faster by Writing Less
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31c1834d-7bdf-412a-af12-a1ac28ffe6f3 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning,
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7d3ed97-fc43-45b0-90a9-7be4101e2782 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Gemini Thinking / Thinking Budget Documentation, 2025
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a3c52c4-5d7a-4d58-b08b-52db7a206f91 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Training Large Language Models to Reason in a Continuous Latent Space
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61f2ca49-624e-4967-9372-58323ae51aaf · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models CoThink: Token-Efficient Reasoning via Instruct Models Guiding Reasoning Models, 2025
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d6f35f9-5d49-4796-a474-4cdbc43f7ded · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Not all tokens are needed(NAT): token efficient reinforcement learning, 2026
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 628fae5f-3e9e-44f2-bdc2-ca30acbcb9bd · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c063d0c-f510-4323-8a62-cb1e4fdb1acb · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models, 2025
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6967e08-c3c8-4e23-8297-c4dd9a0ee001 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 001e5436-9734-4b45-9811-0f72ac0f110d · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Wait, We Don’t Need to “Wait”! Removing Thinking Tokens Improves Reasoning Efficiency
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5a4b728-78aa-4493-b919-037f666fbf1f · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Reflexion: Language Agents with Verbal Reinforcement Learning
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac3354e6-79e0-49a3-b8c9-9f563ba7cb3f · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Voyager: An Open-Ended Embodied Agent with Large Language Models
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42e661ae-2a81-4114-bb30-8d421bd7cf8a · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models ExpeL: LLM Agents Are Experiential Learners
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73edc426-d1ec-4b53-95e4-de3c6fde6793 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory, n.d
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ce66f34-3b61-4416-89e5-960d6e1df9c1 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Zep: A Temporal Knowledge Graph Architecture for Agent Memory
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc653dad-8818-4263-bc61-a79a5d0343aa · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models MemoryBank: Enhancing Large Language Models with Long-Term Memory
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c65e98c0-5bb2-4827-8861-b5d27f3cb1c1 · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models MemGPT: Towards LLMs as Operating Systems
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a411dd21-0e1f-4e6a-9edb-796f4e9e51df · outbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.