Pith. sign in

Paper Citation Record · LEDGER

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs

As of 5 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 3 inbound Pith citation observations for arXiv:2605.20315.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.20315 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T07:49:07.128814Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T07:06:53.318182Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T08:55:35.141369Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact20
  • verified fuzzy24
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2a385ae5-409d-4994-a3d1-b1d0e5a84019 · outbound

This paper cites Artificial Analysis Long Context Reasoning Benchmark (LCR).

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Artificial Analysis Long Context Reasoning Benchmark (LCR)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.378024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:6548f97afcff4e744fbd28c1334c4d75d87879533278eb4f0687013bf5cbbb6b

Observation bfd33b8f-ffec-4025-a15d-5fc5cd592954 · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.972396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:f1b6dbe2088fdcccaf9a643ea4f7c33f0be699412b48c049f44c0c09f27dbeca

Observation d60d51c5-5ca8-4497-8691-803ab1ee0d02 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.969634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:7743927382c3d45b9c34dd680bc8a865ff467b58b4a9adba4f8990041efcbcdb

Observation 8d9f0956-f6e6-4d79-bdd7-042c0ceaf012 · outbound

This paper cites Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.966881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:4fe54fd9d94f8170f6248d5894b0069b8ffce3960348c6a5236a2eec47e5b4d9

Observation 7c328ebc-0313-47e2-8c9c-a44f3d222cbd · outbound

This paper cites LLM.int8(): 8-bit matrix multiplication for transformers at scale.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs LLM.int8(): 8-bit matrix multiplication for transformers at scale

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.375486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:05dd4ce13bb59c19c1599380de80ddaf89c21095c4de3c62b193a2e894fc3177

Observation 83876ccb-db30-41d5-8d14-038885cbb3dc · outbound

This paper cites Castro, Denis Kuznedelev, Andrei Panferov, Eldar Kurtic, Shubhra Pandit, Alexandre Noll Marques, Mark Kurtz, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Castro, Denis Kuznedelev, Andrei Panferov, Eldar Kurtic, Shubhra Pandit, Alexandre Noll Marques, Mark Kurtz, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:49:49.960736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:b111e802831d7050d2bf9d93cd0da47f210c7269d373229f5d0a681a1003ee79

Observation 5e6cc91b-ea7b-49d7-8ecd-c3bf82c89e4b · outbound

This paper cites Flashprefill: Instantaneous pattern discovery and thresholding for ultra-fast long-context prefilling.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Flashprefill: Instantaneous pattern discovery and thresholding for ultra-fast long-context prefilling

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:49:49.964034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:ff1f85c74fdf5c3a51908a7bf4641ab5c2529cb00244096e5cb0af919f12bc26

Observation f9fff8a0-bbce-40e1-b314-09dfad0454a7 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.954601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:4ec8f5a3fbfe748d640d4ea075d59516473f0d1db9f20a6e3c404ec0c969a01d

Observation 0a4e6f92-18fa-4e0f-bd71-822871563cd5 · outbound

This paper cites Gemma 4 model card.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Gemma 4 model card

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.369854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:af201f8c5043a9c8a2f8f456c396c841663edb9cae946a81e30212681b125d6b

Observation a7a730a8-ebf4-45e5-a14d-ea7649504a96 · outbound

This paper cites Minillm: Knowledge distillation of large language models.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Minillm: Knowledge distillation of large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.372980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:22568d34694b2d9dd58f2ed4d5aa48662b2f8a5a1bf3c3d7c3a51d1647b963bd

Observation ee3db034-e5b2-4b0a-b032-591650e2d55a · outbound

This paper cites Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.332349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:5135cd4580eb22696c403637f8faee258d4b0dc0bf33985aa037c6ea7e7eb9aa

Observation 8004bb3a-0bed-422c-a1d0-3bd81e3c2bda · outbound

This paper cites Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.326906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:76c5c5fdf325ce1143ae5ce382a9d67ae643c830a8255ee63800b31f9080d9e0

Observation 8cb856c8-3ed2-4fc2-ac3c-6fe52c7753b2 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.945471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:532fe107846d1566960f7fdf8b15b79c9661717b76787df55d0ac61dc35430ee

Observation abc28438-7dda-45df-b37f-5c2852826b12 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Gonzalez, Hao Zhang, and Ion Stoica

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.353536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:acf6648baaecfa89765089c9847816ef25a140bdf06c4ad6c1732223ea4c4865

Observation 90fb5a9d-bbc2-4524-9a97-b6bfe4c67954 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Gonzalez, Hao Zhang, and Ion Stoica

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.356126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:d1625921073f008e8f1a10b07eb2c849ebdae86403fff30fa71b86d3b8e23d40

Observation 158e9f64-0ce8-4f9e-8897-f3016a719d5b · outbound

This paper cites Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:49:49.951742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:f4c91a9870439700b7f761cbfe2ba4e9317bc1d98d7e0b426f39a6d810ec7625

Observation 966e5100-4c5b-4625-8d49-4e5b9fba6d3d · outbound

This paper cites Let’s verify step by step.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Let’s verify step by step

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.348042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:d4307991d86958ba90f574d2944e93f08f1b3ee51288470828db22e9996888a2

Observation 92b365b2-bff3-4f2f-ab30-d9ec586a32e3 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.350994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:967ade190142e12ddad7b72b2d506aa8f3d587f7ad1234f56a1f11274de6fbd1

Observation 56797825-8cc6-46a0-8dcd-f8dca6ba2c0c · outbound

This paper cites MixReasoning: Switching Modes to Think.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs MixReasoning: Switching Modes to Think

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-09T03:07:04.833097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:9abfc6a91d76c25db26c09560705682a9cb6fd5a2391c5b4204f6ffd63f5f108

Observation 758707e8-05c7-4f5a-8b56-31daaf3cf52c · outbound

This paper cites Large Language Model Agent: A Survey on Methodology, Applications and Challenges.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.938591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:4d21c2a267cbfb46e8df8fab795943c8b5b087aca41ad71776b0c72607580093

Observation 640f6f26-8c8d-4ce5-b575-2edd0f7c2c9f · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.Advances in neural information processing systems, 36:21702–21720.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Llm-pruner: On the structural pruning of large language models.Advances in neural information processing systems, 36:21702–21720

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.367035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:eed98bf6310379b14230748533c6ac88dfaba74987729a5277892c1b2674aabb

Observation 6415345b-ed10-4db1-9c3f-9387d04bcf93 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs WebGPT: Browser-assisted question-answering with human feedback

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.931375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:a7aca218ec37c73c5e9f7564cff5b8525bb17725cdb30c026a203c17e7109147

Observation 0b4853b3-0cfe-415d-9688-3c6b64e32a4a · outbound

This paper cites Introducing nvfp4 for efficient and accurate low- precision inference.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Introducing nvfp4 for efficient and accurate low- precision inference

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.341227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:c6bf1f4891c514b14533b92a30020b44f0d460f03c9689677d4db55e8714995b

Observation 58ddcb9f-a327-44fe-900d-144b6f22b243 · outbound

This paper cites Enhancing distributed inference performance with the nvidia inference transfer library.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Enhancing distributed inference performance with the nvidia inference transfer library

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.313218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:dea9297e7a534e83a5f35bb2ced56fa8632bbda9c53da526c5b7830766ac82a5

Observation a9885126-92b7-4362-bfbf-91e10aa528d8 · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs MemGPT: Towards LLMs as Operating Systems

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T07:49:49.957577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:fdf43a9aacafc0e526011ab90bd0726752a6136a5a9a3fb74f6d97f09653fe0b

Observation 64f81343-78c1-42a8-8efd-ee22672adb86 · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Splitwise: Efficient generative LLM inference using phase splitting

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:49:49.924669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:99bf259bea8b8ddc02a3bb9c4a82dd2d391e7b1826c759f5b0ff3873edde4671

Observation 03ff802a-6c80-47b0-b22c-0eafdd6b0eee · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Splitwise: Efficient generative llm inference using phase splitting

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.363770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:201f0fffe1a1d4f30d6eaed3bad43baa05c6acdba7aa508e0e6b322db8cfab45

Observation 0413313e-6ec5-4d57-89c4-a587c93ba989 · outbound

This paper cites The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.335327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:6e4cc76c591f3839e71d9516424d8b323b74965c362a213a4727982c5c8df717

Observation af2343a7-ab13-46d5-8779-0b74ef1a22a4 · outbound

This paper cites Yarn: Efficient context window extension of large language models.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Yarn: Efficient context window extension of large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.358780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:42a00d03ccaaf88339ba3af59e58798fa6a0bcdafced241ce28917ab7cf959ca

Observation 2d11e57d-a088-41c6-9ef7-db38a1ebaa7d · outbound

This paper cites Swiftkv: Fast prefill- optimized inference with knowledge-preserving model transformation.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Swiftkv: Fast prefill- optimized inference with knowledge-preserving model transformation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.360965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:f069b147303bdb664384b282793328699081242014e85d150b0475a865f72769

Observation 4f816f20-a4cf-4c42-a24a-c331a75ff4b1 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Qwen3.5: Towards native multimodal agents, February 2026

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.344987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:d6121adc8f903acbdc69a8b2683b81b43b25f31296840fc9f4ba80ad2fc144b0

Observation 8e1723f9-b021-43fb-8365-6503ae1224c6 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.948681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:c7234818a72488795e1ff488508a99378ed56e9cb3ddca3742e0449ce1f94b4e

Observation c33454c0-104a-41b1-84aa-e3a8e5d1aa77 · outbound

This paper cites Efficient llm serving for agentic workflows: A data systems perspective.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Efficient llm serving for agentic workflows: A data systems perspective

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:49:49.942271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:b3820e64c44f450c57fef2a59bab341e1d53ecf5af5c5d7a778fd593a96d3d61

Observation 214e712c-589e-4792-a522-7a3d145eb1a2 · outbound

This paper cites Long- memeval: Benchmarking chat assistants on long-term interactive memory.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Long- memeval: Benchmarking chat assistants on long-term interactive memory

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.338241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:f4a46a7242293f4de3e54e27be4dc4d1c01a32103aeb89263d74fce906073bc3

Observation dea70fbf-ce22-4e69-9d02-f5783188a110 · outbound

This paper cites Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.917837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:c7b0e1ced2ed066245d7f46889075624ce2dafac3c719d014c2ff090578bd831

Observation 93a4fce7-f10d-403b-ae33-47e4d9f178d4 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.316161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:85b35f30aa9b0ab7059cbe0319c29347e7a39b7a94a3dee2f1dba8ad21f112e4

Observation 8e16447e-4d6f-47e2-81ac-3273a059a3e1 · outbound

This paper cites A-MEM: Agentic Memory for LLM Agents.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs A-MEM: Agentic Memory for LLM Agents

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.920887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:6997f46ed9369f7bc304ddda713ba6424b7d81233f54551f298c1d77b0b17537

Observation 5b15b1b7-44ac-45a5-9a97-c39f280756c2 · outbound

This paper cites Qwen3 Technical Report.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Qwen3 Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.928093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:58ffa417901a1bf2f3efdc15fd2e19a8381c0b2940fdba037dbf3fab39296e54

Observation 0c304cd2-c41b-49a0-aa34-45a9e3b0478d · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.324073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:97df0a984bf24bb5bf2bbbdabdcbdb7e4d73f382eade5404c69ee0c2cfefe2cd

Observation 082ef578-7e54-4091-a4f7-222a2731af2d · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T07:49:49.905664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:8d85026281d4aeefe6aca4b43c070f11ef07b81eddea7b3369a2bde289679999

Observation 18f6fe46-d908-4334-9e9f-5d0516f1bd76 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs ReAct: Synergizing Reasoning and Acting in Language Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.908832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:c1a168e48f2be2cb653c79122b62c13bb0db58f42739f2feeee671594fa77e20

Observation 816b3edc-6414-470e-91e7-529ab802c486 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs React: Synergizing reasoning and acting in language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.318808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:813f8e7cc89feed267d79ccb7148d463851463539617c55d39a5ec80506f7be7

Observation 0c5816a9-3477-49f3-b916-015fa4e878ec · outbound

This paper cites Qspec: Speculative decoding with complementary quantization schemes.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Qspec: Speculative decoding with complementary quantization schemes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.321487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:379d1ff44fd34dcfa28d34ce3f73eba40d064ece225282d95fd3b8e70f443b74

Observation d27ba4df-1089-46e6-9ff2-471ec59aa2f9 · outbound

This paper cites Atom: Low-bit Quantization for Efficient and Accurate LLM Serving.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:49:49.914962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:588b9ab0afff8b56acb80d9b0d47e3ed180c5abd6c2e0740cd33caaa4fc063bc

Observation c85e2897-eba7-4a5e-903e-a2a0ff37827c · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T07:49:50.329713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:7326193eefd7ebd003e58ef54f3c967c7e39d99529af50dbd1c260daa7fcbd42

Observation 96d6637d-aae8-4ddc-a106-0be1a467de59 · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs A Survey on Efficient Inference for Large Language Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:49:49.911867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:5e976c2b0746917c4fb2720cc14135a6687efac759d7966a51883c8a0051db3f

Pith citing papers

Observation 25b1ed1e-ee39-4a35-97c2-0c3c676a14e3 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:04:45.308390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T05:37:13.211613Z digest=sha256:f97182b2851f7fed642064000860769180745dbdd782f894c69f3e4ce725b230

Observation 7ab5af06-553c-44c1-9326-0796ec6c9ade · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:55:35.142489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T07:06:53.318182Z digest=sha256:95bae7607fdf0d211cf46b55dfd906ebe81814c2fc6d565b5f6faa2996441024

Observation efe7403d-7d97-4a79-9285-2f95ab49a841 · inbound

HBM Is Not All You Need: Efficient Disaggregated LLM Serving across Memory-heterogeneous Accelerators cites this paper.

HBM Is Not All You Need: Efficient Disaggregated LLM Serving across Memory-heterogeneous Accelerators Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:14:57.498059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T04:26:36.990118Z digest=sha256:67cab1c4e64213bf546007f0edbc59d04c1b95a5a346ed3a1ecac1c119cd7494