Pith. sign in

Paper Citation Record · LEDGER

Token-Operations-Oriented Inference Optimization Techniques for Large Models

As of 5 August 2026, this Paper Citation Record lists 100 of 229 outbound references and 0 inbound Pith citation observations for arXiv:2606.20295.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.20295 v2

Coverage vector

measured 100 of 229 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:49:10.293120Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 229 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 346564ca-59ff-4047-9b06-31b8fc02ed89 · outbound

This paper cites A Survey on Large Language Model Benchmarks.

Token-Operations-Oriented Inference Optimization Techniques for Large Models A Survey on Large Language Model Benchmarks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:57.464446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:57.464446Z digest=sha256:3aaa7cd706c393ff3f8dc847c007ebf104a53b90028b10a73f7e023135ee8cf7

Observation ac064b73-c90d-441e-a1ce-8c7c424f3e38 · outbound

This paper cites A Survey on Evaluation of Large Language Models.ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024.

Token-Operations-Oriented Inference Optimization Techniques for Large Models A Survey on Evaluation of Large Language Models.ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:57.664807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:57.664807Z digest=sha256:7eac4e35ab906508f8802317e4f1f0bb911f0a1915eed0436e16afd630b93cec

Observation 3daf6302-7309-4785-b7df-097ee5615d15 · outbound

This paper cites Measuring Massive Multitask Language Understanding,.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Measuring Massive Multitask Language Understanding,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:57.734728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:57.734728Z digest=sha256:d1d8bb4e8a5b626666040d542794f29f7453f3d4af7ea2ab8b41431bcf3be3a4

Observation 4347beea-dc8c-4661-b59e-046758a29ef1 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:57.968427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:57.968427Z digest=sha256:82729294032eb39fa03bb8dd4ccef5a3ba2151a8ef9f23a8e34df8124d9314fa

Observation a04cf945-606c-4710-8c70-40c9ac5b0c0b · outbound

This paper cites Holistic Evaluation of Language Models.Transactions on Machine Learning Research, 2023.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Holistic Evaluation of Language Models.Transactions on Machine Learning Research, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:58.040701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:58.040701Z digest=sha256:96392911b0aaa5cde66d9ba7a5670cdbfdf62a6595e8dfcf2c36075970ebc65c

Observation 6ed8bc40-233a-48ce-aad9-060219e30554 · outbound

This paper cites Artificial Analysis intelligence benchmarking methodology, n.d.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Artificial Analysis intelligence benchmarking methodology, n.d

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:58.194754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:58.194754Z digest=sha256:289d678be482aeee2f088e67678287579ec75275608f8d87b4f14b3395bbf095

Observation ce52ce6a-7e54-4785-a706-b41d3398dd93 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:58.264838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:58.264838Z digest=sha256:272fc269eabcd8d95fd2ddff2bcce25b11bf2bac36e6bf0502e37aa0c10638b0

Observation 6b2b587d-bb30-4052-b4b4-56a80ca39ed6 · outbound

This paper cites Leaderboard, n.d.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Leaderboard, n.d

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:58.394754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:58.394754Z digest=sha256:a0c275d6eaa7086515d419a7c8e2be1a2aec3ea53ec8a7cdf35276f213860946

Observation 05a8ba6e-f885-435d-900c-89245db38013 · outbound

This paper cites Quantifying Capability Boundaries: An Application-Driven Analysis for Large Language Model Selection.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Quantifying Capability Boundaries: An Application-Driven Analysis for Large Language Model Selection

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:58.484750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:58.484750Z digest=sha256:8c305fd02c8a4fe2c9712ab3825c9b830ffa4acc6cbf3dc14aca0f282bea2905

Observation 9d7513cb-f09b-4156-81c9-332fb0536e86 · outbound

This paper cites Quantifying the Capability Boundary of DeepSeek Models: An Application-Driven Performance Analysis.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Quantifying the Capability Boundary of DeepSeek Models: An Application-Driven Performance Analysis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:58.544846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:58.544846Z digest=sha256:5cf38bc47b31f065a75509c010432183c31d7251f5c0eb6500d77862ff4164b1

Observation bd96d426-2062-4af3-8eb5-3dcc48540390 · outbound

This paper cites What is the Best Model? Application-Driven Evaluation for Large Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models What is the Best Model? Application-Driven Evaluation for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:58.657311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:58.657311Z digest=sha256:cc2dfa6ca51066e3139c5f48fb3c46967c16bc189d809bba0e245a63efc9d485

Observation 278cf955-18e8-4a38-846d-6f24c43754f3 · outbound

This paper cites Pinchbench-upgraded, n.d.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Pinchbench-upgraded, n.d

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:58.857605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:58.857605Z digest=sha256:f75bbec16732a0c6a6c33056751262a472ef6b384b01c7c249ce17c1d0dd737d

Observation 98b3b4dd-c93f-4784-af96-2b00e2aa1c2b · outbound

This paper cites OrchestraLLM: Efficient Orchestration of Language Models for Dialogue State Tracking.

Token-Operations-Oriented Inference Optimization Techniques for Large Models OrchestraLLM: Efficient Orchestration of Language Models for Dialogue State Tracking

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:59.054827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:59.054827Z digest=sha256:a5cff57edde0ec8a5515594b1ceee3fc1ee8ebc18a8922161379e59df969e659

Observation 70059fd2-b878-4e06-b9f8-687ccd509a2d · outbound

This paper cites Semantic Router, n.d.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Semantic Router, n.d

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:59.132644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:59.132644Z digest=sha256:8d13497232d88460811628137ddb2f8319ce8fc10d0f4f1b076209d6c202adb4

Observation bbcaa8fe-68aa-4ed4-9b7a-b8260444c222 · outbound

This paper cites LiteLLM: Python SDK & Proxy Server (AI Gateway) for 100+ LLM APIs, n.d.

Token-Operations-Oriented Inference Optimization Techniques for Large Models LiteLLM: Python SDK & Proxy Server (AI Gateway) for 100+ LLM APIs, n.d

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:59.184829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:59.184829Z digest=sha256:5fd44b825005eaff12df3b53ab4115f3f1a81fc36007f0584308007eb97f7b10

Observation 3b1f1821-007f-48bc-98d3-e62ba70c0699 · outbound

This paper cites Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:59.226127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:59.226127Z digest=sha256:3ebb005e46eaa88fdfe16d14865b2f9d183137c33eb219767fecb2b5b363718d

Observation be4729ec-b1f0-4633-b381-2c3a8941c1b2 · outbound

This paper cites RouteLLM: Learning to Route LLMs with Preference Data,.

Token-Operations-Oriented Inference Optimization Techniques for Large Models RouteLLM: Learning to Route LLMs with Preference Data,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:59.274749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:59.274749Z digest=sha256:7b794d058bb7fca9154fe487169b1bce8d20df2a12fc1c0d1215ba05037cd7ae

Observation c03313af-a61b-45f0-baf8-60dd239c29ea · outbound

This paper cites Large Language Model Routing with Benchmark Datasets.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Large Language Model Routing with Benchmark Datasets

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:59.419755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:59.419755Z digest=sha256:c8166b516511c5b1bb24ddcf24c71797f1108d37455f6a378382f0e41f08116d

Observation e44ef3b8-fcb6-40c4-b8b0-6177b8974505 · outbound

This paper cites RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language Models

Reference 19

Resolution
malformed identifier
no resolver link, observed 2026-08-02T10:48:59.534910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:59.534910Z digest=sha256:d9904989d31854e23454f8fccab99669c08f2d22e8985f91d52cfd6eb38284b0

Observation 3d47765f-738e-477a-bdc4-20b7e223ca64 · outbound

This paper cites Fusing Models with Complementary Expertise.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Fusing Models with Complementary Expertise

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:59.644900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:59.644900Z digest=sha256:7f3013b88ecd41b09408757939279ed49bff66adc8967d3be1d621ed7b7ef6dd

Observation 036338fb-daa8-4603-937e-3c81b675f8a9 · outbound

This paper cites Understanding intelligent prompt routing in Amazon Bedrock, n.d.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Understanding intelligent prompt routing in Amazon Bedrock, n.d

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:59.734847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:59.734847Z digest=sha256:05b8ef21d8725afafc77c883f323e6cdc1233bd5c3f8dbcda2de8656dfe76135

Observation d304121e-8b6f-4772-a6b5-89628dc9af0f · outbound

This paper cites Introducing Martian - Better AI Tools Through Better Understanding, 2023.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Introducing Martian - Better AI Tools Through Better Understanding, 2023

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:59.819557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:59.819557Z digest=sha256:f242d5a5010cfef102d419c98c42eaf240a4fd4eaffa39c11374b298ab646c00

Observation 806ccf7f-fe27-4f20-baec-9020e33e7e34 · outbound

This paper cites Model Routing for Agents, n.d.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Model Routing for Agents, n.d

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:59.904742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:59.904742Z digest=sha256:d3710001c6aaf4835183ef3051c2bf5bb2f8ec48b6d08daa49fbfdacd380dd23

Observation d432253e-97b3-435f-a43b-a8444d0616d6 · outbound

This paper cites OpenRouter — One API for hundreds of models, n.d.

Token-Operations-Oriented Inference Optimization Techniques for Large Models OpenRouter — One API for hundreds of models, n.d

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:00.000869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:00.000869Z digest=sha256:7a92f26a701d079e10e388bc2500a4b4e24b4a4ac0937830632f159fd96f733c

Observation 09d6aac1-ef02-4ce4-ae68-0c5331376dcf · outbound

This paper cites Introducing GPT-5, 2025.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Introducing GPT-5, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:00.084751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:00.084751Z digest=sha256:c63831ba1d3d1b955cda6ee34f930c2d8730cba94eaa6616cad0d9aa71b62ae3

Observation 01c7ef23-61b1-4fd3-a1ea-831c2e9ed319 · outbound

This paper cites Using LLM intelligent routing to improve inference efficiency, n.d.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Using LLM intelligent routing to improve inference efficiency, n.d

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:00.223510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:00.223510Z digest=sha256:df2f3d2229fc28ef1fa49c9a7158b3b29e6f212b2024768e48eead2c095206b7

Observation 042483ec-bf50-4bb0-997d-c5affd790eb6 · outbound

This paper cites Intelligent model routing, n.d.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Intelligent model routing, n.d

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:00.272052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:00.272052Z digest=sha256:c5cca32286acf88042c366f041b62c5100ad93d56777e5868e357ad880845558

Observation af0c6fba-9c15-41c4-8f03-c1d9793396ae · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

Token-Operations-Oriented Inference Optimization Techniques for Large Models FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:00.424815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:00.424815Z digest=sha256:05a69077724e2650ba920ce167006f8b2ebe2bbaf944a00f9a129af365ef8ab1

Observation 79bf9a32-a2f2-425a-9005-ddfeea326fb8 · outbound

This paper cites Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:00.720574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:00.720574Z digest=sha256:cfd5e8298e8181f945230e70a51c5e9f33ef6046b94b7534ceec546821cb2d8b

Observation 74bcfd1e-f01f-467c-946d-540790e449e1 · outbound

This paper cites AutoMix: Automatically Mixing Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models AutoMix: Automatically Mixing Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:00.800719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:00.800719Z digest=sha256:6ad53dcd68ad8f3995c16d5bfbb1ad8b3e4760e29fcd0d250cd5298aace0d026

Observation 88e93f93-4fcc-44c5-a66f-7c0003102685 · outbound

This paper cites Tabi: An Efficient Multi-Level Inference System for Large Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Tabi: An Efficient Multi-Level Inference System for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:00.928487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:00.928487Z digest=sha256:fa7836f23fb7ac378298655b1199d861f37c841ad09dd41e02610aca01b28718

Observation 6264491b-e95f-4a74-b743-0f06f69208e0 · outbound

This paper cites EcoAssistant: Using LLM Assistant More Affordably and Accurately.

Token-Operations-Oriented Inference Optimization Techniques for Large Models EcoAssistant: Using LLM Assistant More Affordably and Accurately

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:01.044830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:01.044830Z digest=sha256:f5444741daddea4cb47022332d4f871cf188acf945bc51cb840e1157245f431f

Observation 10871e84-2b9b-4ac4-83e3-c3b2005c3bf8 · outbound

This paper cites Fast Inference from Transformers via Speculative Decoding.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Fast Inference from Transformers via Speculative Decoding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:01.124963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:01.124963Z digest=sha256:4aef6eec97856ce3f26e56b367000405caf8133b0641d6dbcd7cf441aeb96351

Observation d9e35fe6-3962-42dd-9395-7a86a4d4607f · outbound

This paper cites Unity AI Gateway: Configure Fallbacks on Model Serving Endpoints, n.d.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Unity AI Gateway: Configure Fallbacks on Model Serving Endpoints, n.d

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:01.309687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:01.309687Z digest=sha256:a8c20e2588f271938c62322a0ce27d11d17ef0ef2f79a96f62bee29993e63337

Observation 073a4f60-e68c-4104-86f6-e7a96754650e · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:01.504761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:01.504761Z digest=sha256:4a8301e8fdd80acf76b7571f2931a48ff6c5514ca5bb90c1b975e44c77e35ccd

Observation f4118e3b-b6d9-4735-8b11-20dc79d87484 · outbound

This paper cites More Agents Is All You Need.

Token-Operations-Oriented Inference Optimization Techniques for Large Models More Agents Is All You Need

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:01.710633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:01.710633Z digest=sha256:cd69b46dce213eea03cd6ffca7369019b9b98087a8ca3ed15757aba5e152a882

Observation 6682f2fa-11f2-4a3c-a781-50fe2fe608fb · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

Token-Operations-Oriented Inference Optimization Techniques for Large Models LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:01.824067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:01.824067Z digest=sha256:57cc8fca0962e1c341da48d41e0e238bcddaba65d96e83aef3aaa1a2fe72171b

Observation f90e1a6e-56d7-49f7-9abf-ceb3e3203cfc · outbound

This paper cites Knowledge Fusion of Large Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Knowledge Fusion of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:01.931706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:01.931706Z digest=sha256:08ccb2efb24e39b97760fc4720411ec4a696a2230171f29f684e7a8c7c107697

Observation 76875b0a-987d-4f9b-a086-c5f87fe6ce37 · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:02.081223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:02.081223Z digest=sha256:35e2fe800b3c464b7950decad3ef4cb7ff7f6c990cdcddf51597aa9342fd1e0b

Observation 085e036c-9f0c-40c6-9414-6594d336d71c · outbound

This paper cites Efficient Attention Mechanisms for Large Language Models: A Survey,.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Efficient Attention Mechanisms for Large Language Models: A Survey,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:02.201957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:02.201957Z digest=sha256:2a59ceefd9cae0963ace519cb3dc83379945a7e9025b7d9b1df7ecaa74d7ea4e

Observation 79f63b3b-284d-4d6e-8e1f-17d70b53b74c · outbound

This paper cites Qwen3 Technical Report.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Qwen3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:02.346773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:02.346773Z digest=sha256:d8f8af0c6a7b75603045186b9c80ae80b7202f63236a4459f84eb78206b703c8

Observation 23dac4b8-f68d-487f-9393-c409c9a40f0a · outbound

This paper cites DeepSeek-V3 Technical Report.

Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepSeek-V3 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:02.492861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:02.492861Z digest=sha256:9d70394a1db97238e47c440871052b57782f3b16b1491c60126d4e8efc4c363a

Observation 28e88532-8b0d-4dfb-98ae-1692befcf0fa · outbound

This paper cites The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence.

Token-Operations-Oriented Inference Optimization Techniques for Large Models The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:02.690693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:02.690693Z digest=sha256:41965e320c73b33d9e988d95e0f733c922e2c77f7490cd529c97ba70e169a23a

Observation 331b9e16-3160-419c-b748-f7530878bd32 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Token-Operations-Oriented Inference Optimization Techniques for Large Models FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:02.767594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:02.767594Z digest=sha256:018225db34de09198341e258bd6349b5f33964984523a374802343da6f5cd615

Observation 30edab7f-fbff-482e-af37-6115862bf3a6 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Token-Operations-Oriented Inference Optimization Techniques for Large Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:02.964017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:02.964017Z digest=sha256:bf751a259c15a0bcea7bf296b9dce011c2b681866c9274e61dfdb468aee5ce8e

Observation 9857e8ce-7970-4278-8657-bb80e5bcc916 · outbound

This paper cites DeepSeek-V4: towards highly efficient million-token context intelligence, 2026.

Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepSeek-V4: towards highly efficient million-token context intelligence, 2026

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:03.045692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:03.045692Z digest=sha256:784dbb9ef7572f9bbe1f8760b392c3febe9141644beaf99a60b3d5edbfd47315

Observation 4d1c1899-2a10-4a33-9645-b78c545f7387 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Token-Operations-Oriented Inference Optimization Techniques for Large Models GLM-5: from Vibe Coding to Agentic Engineering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:03.114879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:03.114879Z digest=sha256:7f24ebec47b83d0d08359c2f7873ec10c93a17e19b2ec5d82127d56a40cca2b0

Observation 796b09f8-23b8-48e5-9040-b08e1886d4fe · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:03.204809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:03.204809Z digest=sha256:058d7155dc9404f1f0f1b0033e34e4db7e1d4f0c6fcf4c41cfd0c63b0fc337c1

Observation 46b5f8fb-b3aa-4aa1-ad94-52d9983f47e5 · outbound

This paper cites MiniMax M3: Frontier Coding, 1M Context, Native Multimodality in One Model, 2026.

Token-Operations-Oriented Inference Optimization Techniques for Large Models MiniMax M3: Frontier Coding, 1M Context, Native Multimodality in One Model, 2026

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:03.392996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:03.392996Z digest=sha256:ca8fbdbc2189dcc141822212f025344ebdc579b34c770c5b5d9d563e2ce5ef39

Observation 8792c76f-5bc4-4e7a-9eae-ce6a4bc2a184 · outbound

This paper cites Qwen3.5-Omni Technical Report.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Qwen3.5-Omni Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:03.511163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:03.511163Z digest=sha256:5627093ffe5ef65a5a69fa88244f9e4ef266fb023aba4a1740a4cf094b195966

Observation 4005949f-3263-4849-836f-7cb216e74948 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:03.636799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:03.636799Z digest=sha256:a68745a60479b87c3d9953961bdda8c6829fed1eb2abb214e28d384b0a36e10b

Observation 21ec33ad-7dc9-4131-9cb8-3c9bdb8e2892 · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.IEEE Transactions on Knowledge and Data Engineering, pages 1–20, 2025.

Token-Operations-Oriented Inference Optimization Techniques for Large Models A Survey on Mixture of Experts in Large Language Models.IEEE Transactions on Knowledge and Data Engineering, pages 1–20, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:03.769546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:03.769546Z digest=sha256:1cb8cbac4d79d0f36e2941ab19b6ede03893cb43877ecd791970ff1aabc1ed8d

Observation 752887d1-532a-413d-8a6c-4370820297b3 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

Token-Operations-Oriented Inference Optimization Techniques for Large Models GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:03.823350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:03.823350Z digest=sha256:243412823869f449c985ee93976e404df2d8f987c16616e3e5bb6571919fc85c

Observation 4d4d7277-04dc-426f-a06f-fb2047358bb9 · outbound

This paper cites Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:03.928187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:03.928187Z digest=sha256:ac74d79330672b94cc10662581535f87f5cb2d7a068af80bbd92991d84ea304c

Observation 6fed5def-d20b-4922-93a4-9eceb84efe47 · outbound

This paper cites Mixtral of Experts.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Mixtral of Experts

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:04.027295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:04.027295Z digest=sha256:3231234cfa4e9d938a3f4baf8c63d50741e123045c5bff47d5b8f8b128381269

Observation 823d9b52-e50c-4cfd-9809-119b100c5d2f · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:04.126215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:04.126215Z digest=sha256:2eb87a5100e5d9e3679d29931d5f0bcbc95b6c909f93eca9d0f56ddaa0a7058c

Observation 154f4660-b663-4a73-a7f5-aea7385bebf2 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:04.221071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:04.221071Z digest=sha256:25923b1ffa92546334aadbcee1ade9e1c28e8e5d6900d139821271519ab439bd

Observation 5942b97a-fedc-4394-b856-be89cd3949c4 · outbound

This paper cites TensorRT-LLM expert parallelism documentation, n.d.

Token-Operations-Oriented Inference Optimization Techniques for Large Models TensorRT-LLM expert parallelism documentation, n.d

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:04.312976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:04.312976Z digest=sha256:e6dfe87dfb40f065091c7b23748c462c844d293569a9d9a472713a72db2b1009

Observation 9e6ae66b-3612-40d8-b652-57c992a9fed2 · outbound

This paper cites Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:04.391670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:04.391670Z digest=sha256:11994cf4bf322c20d2396454e52275c3c8378de1a4b41979ed3bab8bbe07218c

Observation 09543c80-a541-4cf9-bdfb-03b6e6569b12 · outbound

This paper cites Mixture of Heterogeneous Grouped Experts for Language Modeling.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Mixture of Heterogeneous Grouped Experts for Language Modeling

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:04.554166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:04.554166Z digest=sha256:31b2b3c888e0f44952da99660ad0209834b13397d0f532cb4b32032520deca52

Observation f86d142c-53ba-4e73-8ddd-21e8f2d504c9 · outbound

This paper cites Optimizing for the Shortest Path in Denoising Diffusion Model.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Optimizing for the Shortest Path in Denoising Diffusion Model

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:04.626625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:04.626625Z digest=sha256:51d2da7b43139213dd1691a7ca4d916f7dbc64e5f84a3ceced9f91cf6f3c67c8

Observation 5044a919-65cd-406f-a2b2-37cc13d35210 · outbound

This paper cites A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation, 2025.

Token-Operations-Oriented Inference Optimization Techniques for Large Models A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:04.805876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:04.805876Z digest=sha256:a6bdfaa3b2d8a7d7e0f411c8d7203c5cbea21ef442d2083d3f9f5463d305387f

Observation 46d8602d-7dfd-4bd5-bf8d-7762ec8a350e · outbound

This paper cites DeepCache: Accelerating Diffusion Models for Free.

Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepCache: Accelerating Diffusion Models for Free

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:05.025293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:05.025293Z digest=sha256:616d2a9ce0aac3c44bfd68ce6ea18a2543415a92172bc0a2f4ee723264590c53

Observation 9ab818e2-6e84-40f3-9a27-5b79999ae9de · outbound

This paper cites Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:05.219314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:05.219314Z digest=sha256:5369d21e8b266a280189b4384d0569e0fa0be902817a885bc88e30c8e8a5dbac

Observation 5d2edd56-123d-4481-b2aa-ba942ff026b5 · outbound

This paper cites Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching, 2024.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching, 2024

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:05.420685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:05.420685Z digest=sha256:acbd599ef65fc489a18b7ad382547c04c0fbc7a541e15f0b3473e92cfa306311

Observation 25b1704c-4eff-48d9-817d-57434c328938 · outbound

This paper cites LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation.

Token-Operations-Oriented Inference Optimization Techniques for Large Models LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:05.554576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:05.554576Z digest=sha256:ca5f39153d876ba5d5498d47a4370f5c252e037002f04f224f176176757f02f6

Observation bc6a258e-0061-4c31-9202-f35d94671705 · outbound

This paper cites URLhttps://doi.org/10.1109/cvpr52734.2025.01679.

Token-Operations-Oriented Inference Optimization Techniques for Large Models URLhttps://doi.org/10.1109/cvpr52734.2025.01679

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:04.685324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:04.685324Z digest=sha256:d732c7bf352271b5199072f5b986268e5e899fa926b0a0cfc84e8068b9917bf9

Observation e4cea724-1380-47ac-a689-50e23834ee45 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:05.829494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:05.829494Z digest=sha256:7c704da0768f0f03bc25084ac178fcb854671a55555cf292fc69897c691fa95c

Observation 33147e5a-3de8-44a2-bd67-d85895322ba7 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:05.903146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:05.903146Z digest=sha256:36ea75dac204393bf452608cf69e743ab6dbc731466df303815a0421333640d9

Observation e64d0a0e-2dac-4d20-a294-4c6db1de7d35 · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning.

Token-Operations-Oriented Inference Optimization Techniques for Large Models STaR: Bootstrapping Reasoning With Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:06.089278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:06.089278Z digest=sha256:f18c7774d479608cd259fe83c11f2e5b8731b8a80c42a08e4367f2ec4cafadb3

Observation e0deaae1-8abf-40a5-99ec-cd79efb0c0fe · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models ReAct: Synergizing Reasoning and Acting in Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:06.290662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:06.290662Z digest=sha256:44e1fe5694759b4cb2fd0fd3d7a784346754727042ac8fe5d565c179e09c0ff0

Observation c663ffe0-4920-44e0-8cc2-90ec60ad3b15 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:06.444925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:06.444925Z digest=sha256:75400078f3739a30b979d1595aa35fb49a9c1176dc5d5ddcedd74b62aab457dd

Observation 6b134d8b-b57b-4595-954f-0172c6b06cbf · outbound

This paper cites MeanCache: From Instantaneous to Average Velocity for Accelerating Flow Matching Inference.

Token-Operations-Oriented Inference Optimization Techniques for Large Models MeanCache: From Instantaneous to Average Velocity for Accelerating Flow Matching Inference

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:05.696593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:05.696593Z digest=sha256:9314f9cf568356ae593e60abd887cb373ab3398ca12207f00521b69cfe487b22

Observation d2dbebe4-e549-4409-80cd-5c3d59996158 · outbound

This paper cites Introducing OpenAI o1-preview, 2024.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Introducing OpenAI o1-preview, 2024

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:06.767008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:06.767008Z digest=sha256:21ba47e77dd9e2b9a18f56f4ff18f01ffc2e76afb609f587c526f6f0e5b2527e

Observation 4fc60caa-809f-40aa-82dd-0e771119ea81 · outbound

This paper cites OpenAI o1-mini: Advancing Cost-Efficient Reasoning, 2024.

Token-Operations-Oriented Inference Optimization Techniques for Large Models OpenAI o1-mini: Advancing Cost-Efficient Reasoning, 2024

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:06.899051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:06.899051Z digest=sha256:1d19ba7fa95fa5f66a60b26276cecbc215712ddcda5c2878277b2362cd845104

Observation 918cdb75-9f74-49c5-b883-8e0588268a54 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Token-Operations-Oriented Inference Optimization Techniques for Large Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:07.038720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:07.038720Z digest=sha256:96b1de79ee2d408e9caabffb4a735f7f340f0a501209274228280663c0ebac54

Observation 7d12a6c4-a4f7-4709-a7ac-ac1098935bb7 · outbound

This paper cites Claude 3.7 Sonnet and Claude Code, 2025.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Claude 3.7 Sonnet and Claude Code, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:07.118728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:07.118728Z digest=sha256:d63967ca39b54e02b26802f7aba2abd06bb8f584bac5c904f981671281275abd

Observation 92a2fd96-c260-4b28-b870-1c999ce109a8 · outbound

This paper cites Qwen3: Think Deeper, Act Faster, 2025.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Qwen3: Think Deeper, Act Faster, 2025

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:07.261838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:07.261838Z digest=sha256:0fe3be8c6af434e29f2c4c218be80bad64b210a0d0339f4dce49b41a0aeb2f10

Observation 4b778c97-dd94-4e3a-bcfb-e32edf682df5 · outbound

This paper cites PAL: Program-aided Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models PAL: Program-aided Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:06.599436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:06.599436Z digest=sha256:8cf89302f6f1e48f1edb552226effecb07feb7fe525d59617557f6d247a2d3e6

Observation afebed7d-1074-402b-8971-541e6cb0e5d6 · outbound

This paper cites Token-Budget-Aware LLM Reasoning.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Token-Budget-Aware LLM Reasoning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:07.509981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:07.509981Z digest=sha256:1aa1cb0f25037512fe98211437f1968dda0c4089965e8914ec61d099d2949359

Observation 6e6b6886-b098-4cb7-a803-5960017cd5f4 · outbound

This paper cites Training Language Models to Reason Efficiently, 2025.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Training Language Models to Reason Efficiently, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:07.595312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:07.595312Z digest=sha256:3121123ee4af342153bc76c0498bc67217b7d2c492670ed182755d0958b11756

Observation 3e778285-b681-44bd-9652-4e0890b5ae65 · outbound

This paper cites Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:07.680058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:07.680058Z digest=sha256:e688d61187103726e4c2296692de7a76beffd2a9630d77324e8ec6c308140e05

Observation 539991cc-9c8d-41e7-8bd2-eaaf0a050631 · outbound

This paper cites Chain of Draft: Thinking Faster by Writing Less.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Chain of Draft: Thinking Faster by Writing Less

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:07.851036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:07.851036Z digest=sha256:a20e7cd19984090fe1a9b433ec9e11a61b4655781677b1fdf51deb2d65e51cd2

Observation 31c1834d-7bdf-412a-af12-a1ac28ffe6f3 · outbound

This paper cites Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning,.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:07.996094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:07.996094Z digest=sha256:d550a01e1b96859e791467f51ce55fa1e8e7e116b76849f0625ec689c3704f2e

Observation c7d3ed97-fc43-45b0-90a9-7be4101e2782 · outbound

This paper cites Gemini Thinking / Thinking Budget Documentation, 2025.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Gemini Thinking / Thinking Budget Documentation, 2025

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:07.436090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:07.436090Z digest=sha256:335061900520f0412163e6582ca5169c9e9d13f1b938c416b21bafdf104b1daa

Observation 4a3c52c4-5d7a-4d58-b08b-52db7a206f91 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Training Large Language Models to Reason in a Continuous Latent Space

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:08.460425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:08.460425Z digest=sha256:0fc6c5c3a28846ca2feffa3f290d033e656babaebf53775ac26c7fd6d4a8deb8

Observation 61f2ca49-624e-4967-9372-58323ae51aaf · outbound

This paper cites CoThink: Token-Efficient Reasoning via Instruct Models Guiding Reasoning Models, 2025.

Token-Operations-Oriented Inference Optimization Techniques for Large Models CoThink: Token-Efficient Reasoning via Instruct Models Guiding Reasoning Models, 2025

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:08.641974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:08.641974Z digest=sha256:1355b645e3afa5c7eb5edfb7e3502ab81ede99338d686e2a3065a5af78364c53

Observation 2d6f35f9-5d49-4796-a474-4cdbc43f7ded · outbound

This paper cites Not all tokens are needed(NAT): token efficient reinforcement learning, 2026.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Not all tokens are needed(NAT): token efficient reinforcement learning, 2026

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:08.741395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:08.741395Z digest=sha256:99da0a88109c113ad0fc39a8fe9b09a98ffd4406fb42a11abd648b0137abdcb8

Observation 628fae5f-3e9e-44f2-bdc2-ca30acbcb9bd · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:08.899979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:08.899979Z digest=sha256:6d4dda515cf872108c51d4d4666d01d476a0dcc48e42a4e23bd07b4ed6b59fe7

Observation 7c063d0c-f510-4323-8a62-cb1e4fdb1acb · outbound

This paper cites DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models, 2025.

Token-Operations-Oriented Inference Optimization Techniques for Large Models DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models, 2025

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:09.111767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:09.111767Z digest=sha256:865d27b7cf9bac44b15b430aecd56f0f7ce1465e0d1aa2ed7320c740aa599100

Observation e6967e08-c3c8-4e23-8297-c4dd9a0ee001 · outbound

This paper cites Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:08.087657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:08.087657Z digest=sha256:2b4f56d41cfb55d3f4f1ef158c7183d2ced0686eb30f61db8d140d4065e82ac2

Observation 001e5436-9734-4b45-9811-0f72ac0f110d · outbound

This paper cites Wait, We Don’t Need to “Wait”! Removing Thinking Tokens Improves Reasoning Efficiency.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Wait, We Don’t Need to “Wait”! Removing Thinking Tokens Improves Reasoning Efficiency

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:08.197707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:08.197707Z digest=sha256:7e96ee6d612d4bf3ec20a218bfaf224a2417ac17d88307cd6b1acafc2b0839bd

Observation e5a4b728-78aa-4493-b919-037f666fbf1f · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:09.474424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:09.474424Z digest=sha256:bcac1aea58d6585717bfc58d0ab22654d32637cab522e4975eab835e318dda28

Observation ac3354e6-79e0-49a3-b8c9-9f563ba7cb3f · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:09.539883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:09.539883Z digest=sha256:4895245be58388888263b159000ce2073e236f2fec8b1c2bb5def9166446b215

Observation 42e661ae-2a81-4114-bb30-8d421bd7cf8a · outbound

This paper cites ExpeL: LLM Agents Are Experiential Learners.

Token-Operations-Oriented Inference Optimization Techniques for Large Models ExpeL: LLM Agents Are Experiential Learners

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:09.625770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:09.625770Z digest=sha256:cfcd79a92194c01f109528452c532715fd877b686ec9c20c2fa45e2ca93af350

Observation 73edc426-d1ec-4b53-95e4-de3c6fde6793 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory, n.d.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory, n.d

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:09.743560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:09.743560Z digest=sha256:c7a543ae3e0387f7e19ed1c75e93209d7f724c50955f23603e7870c9fc1f7ab3

Observation 8ce66f34-3b61-4416-89e5-960d6e1df9c1 · outbound

This paper cites Zep: A Temporal Knowledge Graph Architecture for Agent Memory.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Zep: A Temporal Knowledge Graph Architecture for Agent Memory

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:09.841988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:09.841988Z digest=sha256:1cf4a69d4f57facccd25e2134c3068192de68522408ca2051d00094080629ee5

Observation cc653dad-8818-4263-bc61-a79a5d0343aa · outbound

This paper cites MemoryBank: Enhancing Large Language Models with Long-Term Memory.

Token-Operations-Oriented Inference Optimization Techniques for Large Models MemoryBank: Enhancing Large Language Models with Long-Term Memory

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:09.273065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:09.273065Z digest=sha256:43e757f56bff81deb54646886ffe2ea65b35509f7e91a9afa35262b68ec877a9

Observation c65e98c0-5bb2-4827-8861-b5d27f3cb1c1 · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

Token-Operations-Oriented Inference Optimization Techniques for Large Models MemGPT: Towards LLMs as Operating Systems

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:09.369986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:09.369986Z digest=sha256:5df295af4c7a6718cbb8d0cbeff1338112cd8f096b865a99301f20d08c68d2e9

Observation a411dd21-0e1f-4e6a-9edb-796f4e9e51df · outbound

This paper cites MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent.

Token-Operations-Oriented Inference Optimization Techniques for Large Models MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:10.293120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:10.293120Z digest=sha256:4dfb56d0f7ae8d5bc64db08768dac4f3443614683a013b782c2ccdd838b795dd

Pith citing papers

No inbound Pith citation observations are available.