{"work":{"id":"373baad2-322f-4573-8ba2-d7559e15d07b","openalex_id":"https://openalex.org/W4400024874","doi":"10.48550/arxiv.2406.16858","arxiv_id":"2406.16858","raw_key":null,"title":"EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees","authors":null,"authors_text":"Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang","year":2024,"venue":"cs.CL","abstract":"Inference with modern Large Language Models (LLMs) is expensive and time-consuming, and speculative sampling has proven to be an effective solution. Most speculative sampling methods such as EAGLE use a static draft tree, implicitly assuming that the acceptance rate of draft tokens depends only on their position. Interestingly, we found that the acceptance rate of draft tokens is also context-dependent. In this paper, building upon EAGLE, we propose EAGLE-2, which introduces a new technique of context-aware dynamic draft tree into drafting modeling. This improvement leverages the fact that the draft model of EAGLE is well-calibrated: the confidence scores from the draft model approximate acceptance rates with small errors. We conducted extensive evaluations on three series of LLMs and six tasks, with EAGLE-2 achieving speedup ratios 3.05x-4.26x, which is 20%-40% faster than EAGLE-1. EAGLE-2 also ensures that the distribution of the generated text remains unchanged, making it a lossless acceleration algorithm.","external_url":"https://arxiv.org/abs/2406.16858","cited_by_count":1,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2406.16858","created_at":"2026-05-10T06:01:13.608291+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Eagle-2: Faster inference of language models with dynamic draft trees","render_title":"Eagle-2: Faster inference of language models with dynamic draft trees"},"hub":{"state":{"work_id":"373baad2-322f-4573-8ba2-d7559e15d07b","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":22,"external_cited_by_count":1,"distinct_field_count":7,"first_pith_cited_at":"2025-01-31T15:10:29+00:00","last_pith_cited_at":"2026-07-09T16:16:35+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T12:49:42.384482+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":2},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":2},{"context_polarity":"use_method","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}