Pith. sign in

Paper Citation Record · LEDGER

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs

As of 4 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2510.18245.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.18245 v3

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T05:30:11.389756Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact44
  • verified fuzzy5
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1da1a75f-503e-45d4-8765-a892a0188f69 · outbound

This paper cites Phi-4 Technical Report.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Phi-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:54.969870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:9cc62fffbbb58f63e994bae471827e9427ad13da15b2f10a28781011ce337073

Observation 709620e1-3cdb-448e-ae3e-15838ae8e350 · outbound

This paper cites Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:54.974819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:805b89ed8785d264ca1b93011612a14fc3e4e5d8488a96980a950a4b3ad76e75

Observation bdc0747a-befe-4b43-9c4a-b213055b852e · outbound

This paper cites GPT-4 Technical Report.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs GPT-4 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:54.961428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:7ed0aea0ab00efe0cdfd1a04a5a73dac902fabb3fbf3498cbc26c7e4a085012a

Observation a6499279-3b04-4907-899c-77f49c282663 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:54.979456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:8f84e2ec5e78c2aac5307eb8cdfb3c83b7f778ff8ac99a372817fc5c21fac9af

Observation 145ea757-62b2-414d-bff1-47c0818dd646 · outbound

This paper cites Program Synthesis with Large Language Models.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Program Synthesis with Large Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:54.957133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:875c9608c722be4a93d32bf91193a208212593d8453abf787ae1b823b29220cf

Observation b109582b-e94e-4858-853b-129ad6b632fc · outbound

This paper cites Scaling Inference-Efficient Language Models.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Scaling Inference-Efficient Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:54.965606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:7830da4e62de39c02717e9b663951db2fc3acc97400cd6986de00b9ae9575898

Observation 3155b484-d3d6-4f13-9fd1-b30a2f9a2bf6 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.104414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:2d5f12f55f6d5417cf454d3c7e28ebb4b27e567d83990e97d52753b86abf1d31

Observation e5d40e9f-ebb8-4d07-be6e-17c8a44c3b17 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:30:55.442804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:74955e33351dcf1ec5ad42cec521b961a6a7ab55b332510ad3721b8aeed78c9b

Observation 4891e28d-626d-4305-8fa6-0820b87ffeac · outbound

This paper cites Exploring Diffusion Transformer Designs via Grafting.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Exploring Diffusion Transformer Designs via Grafting

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:54.987903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:0efb9b2ec2ccbf5ae70050487376dbd9262330f53720fb5157d8d9e697b6e4b3

Observation 83b1f712-d611-498f-9b56-7fe78a8936e6 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Evaluating Large Language Models Trained on Code

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.115874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:db49ffbddf4c4a9831e6f50b1a5851654bc8ae2ffe2eeefbf2be9cf6408b6eab

Observation 29230e97-6ab5-4f76-b4cd-0ae41a7c2b4f · outbound

This paper cites Scaling Law for Quantization-Aware Training.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Scaling Law for Quantization-Aware Training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.053160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:282bb6cba11603d84e3179d64acd901f0666cc05a9ce9a0fd468f775f3e4190f

Observation 34b87522-9e74-4c47-b8f5-bfbac89ef4f7 · outbound

This paper cites Reducing the carbon impact of generative ai inference (today and in 2035).

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Reducing the carbon impact of generative ai inference (today and in 2035)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:30:55.437460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:899c89f42805fb70cf0ff00896aac2aace7d9ea2e8605491f8867f52b5c190b1

Observation f971779c-a9fc-4fd9-a9fd-738e40178cc9 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.057630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:7cab6c60c6a1df68084041829c4e43952e99df522d2b02047de93ca40340f8c3

Observation ff32e271-e228-46d2-9303-efb746bd19f2 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.011142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:1433cade3bd16097a0e5a97f5f27621d0bcd4512b70115739fa3b1b646e9c4a2

Observation 6f7ef904-a7d1-46ba-b1c4-818fae591acc · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.147078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:374f0cc47e7bd84b4626f057a5f958f701bd0577f7bb1c46915d957f225e7613

Observation e85a2c3f-e98e-4305-be55-f6fefa8f9f0c · outbound

This paper cites The Llama 3 Herd of Models.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs The Llama 3 Herd of Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.031768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:a51a7aed2d2287a115b82a7e825a0c53bdae2b3f017dd1d8e6ac416e70523037

Observation e5b57a8b-8ff2-4038-9194-3c93348bc488 · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Language models scale reliably with over-training and on downstream tasks

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:30:55.093261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:37af6bbf3616fc4b06aabaf5cb6982983254010e2da1f7ac6a1ed03bf73e296e

Observation 0b9e6101-b575-4ecc-8611-f7a4b355e517 · outbound

This paper cites He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:30:55.069081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:7ee8a45e66b5ba9e4624aa2e4ab2e7f21e1158f6cb43e658615fbfdc27c0d2fd

Observation 2b12d6ce-7fe1-40c5-baf4-25f423b5945b · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.039488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:a32a9556ac8ac2d37f47408b4adbcffab12b4c115f11836c97c6f8cb2c79d5d0

Observation d14a1da3-bbd1-4d8c-b2d8-95c2ae267dc2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.134729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:f9402023ff0dfda58abfb3d97143f96811f834328f092a7a45bc7af78d2d7fb5

Observation 26c4dd41-d428-496f-94a5-2159be8a4e24 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.048040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:07ff2e5500961951539922148db4c6a069c664e5b14c78ec174900a706e8a77c

Observation 313ae315-8dad-482c-899a-14b5dbf82c82 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Training Compute-Optimal Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:54.997589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:958f800b302194de29ccfef0e16f1a383c2dd8f52aab626784eaf8f430db492f

Observation f662aaaf-2814-40c7-9a9f-294ddf7af60d · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.119433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:3c206b9a705a3086f1c983ff5c9c4bdd663f420b5c6570f3758027fb6c5123b5

Observation 87954259-ef6d-4d24-a4b2-7bf3539ab4e2 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:30:55.155139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:89828484cec78410df75a9a2d8505772e0c9617a88e2922393095c678c542d97

Observation 830e0122-ba17-4433-8bad-3fd466451386 · outbound

This paper cites Scaling Laws for Neural Language Models.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Scaling Laws for Neural Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.122777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:436850170029e1388d1787f2344e04a4b369ecc7dfda4e499032b867f29deb5d

Observation 8e35aa42-8db1-4f09-a7d3-54e90d3459ab · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Scaling Laws for Fine-Grained Mixture of Experts

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.127020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:b814bb1ee5510dd461eef36228e01ccc72b08ad3b1179a26db4d6f4037e5f5d1

Observation 77a1a9f5-b51f-4a18-b02d-d7fb973e58e6 · outbound

This paper cites Scaling Laws for Precision.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Scaling Laws for Precision

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:30:55.061410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:232a32925906e00d74172c1dc110de0c93de6f60187cf818a13b77e5808dca16

Observation e24042eb-c1cc-4cb0-bca9-681c1e1b0b19 · outbound

This paper cites DeepSeek-V3 Technical Report.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs DeepSeek-V3 Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.035438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:551ab1d13f02e74414bf3a6227a746f8e70109ed5155b8676765a2890ec77a34

Observation 21a2512e-a594-4953-aefc-4f2821587e21 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.002287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:4ee3256013fd37b691ed43df0291534593d8ebe44f3c3dbaa6361fa3fd435107

Observation b62a3903-ef66-4613-baba-02b1013818cf · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.043829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:e1b92e82324fbca159d16c17df6aec45acd4deb4dd40bf1a569f2afe92aa3cae

Observation cda07176-c96d-4846-9525-0d478dccbbda · outbound

This paper cites The Impact of Depth on Compositional Generalization in Transformer Language Models.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs The Impact of Depth on Compositional Generalization in Transformer Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.112323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:a2e21cb6c9e25932f327a4165995e881956b0687b5ba5e6a6584b3f6335c6d32

Observation e49369ee-2d24-4817-8bbe-d18c2ee685ec · outbound

This paper cites Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:30:54.993450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:611b6fe8b25a16bbb00a4aff66beab955a68cef0fd37b6e813ece465d9625a3f

Observation f272d83c-a6b2-4c07-b2f3-06e2045b63b0 · outbound

This paper cites Observational Scaling Laws and the Predictability of Language Model Performance.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Observational Scaling Laws and the Predictability of Language Model Performance

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.131188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:8a9673ca839dd042f850856b5c2867ffcd8f9e0d223c27a530f30844eb377f84

Observation 64973a38-ee69-4354-a109-9d52931c3803 · outbound

This paper cites Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.139126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:2ccd7fdf53efed8d5d7ceff54b16399d449ea94102a142eefc4043ba30a1cbf1

Observation 988fb2ff-4019-4f6b-9cd8-5ddf3e42bffa · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:54.983448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:f82f7afa25475c3754b56a6c1d73711f808ff50c095e8a4ab83034909da99f6f

Observation 02ad5651-0047-4b35-8e23-3585d10ce9f5 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.076637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:8cec1e6284204af1aa8f59a4458dd6e963ea51d564ac888a489645ee720db739

Observation e7ebcbc1-7631-4364-8326-91ff1f732d3c · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.019596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:872a5e052fb41c61eec298a877220193cbb2cc104db91395bab91da9e5d9e21d

Observation 4a349a68-e32e-4bc8-9752-804a537431a2 · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.101246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:1896ba5cf4d5d863cbea25a2e8c66bc0f577cb243f991156c23d724919fb20d9

Observation ad6a4eb1-b724-46c8-9c8d-59507364db17 · outbound

This paper cites Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.097385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:f5b10a74f34f52cfed086f68b239887683f247fbecfffe2fe403ff22432818fa

Observation b10b1dce-e4e7-4b07-b21d-a88466423d0d · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Gemma: Open Models Based on Gemini Research and Technology

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.088775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:8a27feae8635a6a43dfde5e7d350066a6f73c470830deee9f3ad05796e569302

Observation 693b4fc7-6cb5-4544-bafe-17a44bd934c4 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.142851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:0fa7ef1879f02e3610d8e99e5f5630be5642781b6e3a57eb269c65d756b43e14

Observation 5d0bb1cb-a0b8-4006-b2cc-121d4b7df4e1 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.064540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:c192346c68d6c2832a4177b0885b2e4291068be3b7b3c474014b699c33ec6c20

Observation a32cf55a-0359-4b57-b055-0f8711cce19d · outbound

This paper cites Emergent Abilities of Large Language Models.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Emergent Abilities of Large Language Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.072724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:a7f397dc875ec3983c90a9192be4ecadf0c8d0f46df06dad99b1539961c63f77

Observation 808db350-a29a-451b-8c6a-d6ad43d91182 · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Crowdsourcing Multiple Choice Science Questions

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.108024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:5fb2c8b8d260fa3d3b10546d651a8713a564d645093b0bac34fc5cdd3e3c9a69

Observation 3dc8d93e-5ac6-4414-b19f-b16be7165640 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Efficient Streaming Language Models with Attention Sinks

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.015207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:4476b7f2c7c0e6c938cc1090a7304f7d66a2fb1b794430adc5ebb1f2f53345dc

Observation 3da75dd6-ac2e-46ec-b23c-ecce9857d078 · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:49:16.947404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:c7a1c4068decdbb7d652b25a54198dc4573c4daf009bf3038d0328bc8c8c7367

Observation 72fca583-44f3-4719-980d-b715d56f87f3 · outbound

This paper cites Qwen3 Technical Report.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Qwen3 Technical Report

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.085018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:9c7e35b94cd67fa5a0deb9802aa154df4310e1592904f3539cf5811e869fa338

Observation 366e5d4d-bf24-4113-9445-f6f00f570e4d · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.023596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:8a4d36866bdc0be1a32a8ec5c6d62e3a31c97cc407b8183e3d41796e8abd1448

Observation 817612f6-1772-48df-8973-acaaf0cd85c5 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.027469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:4ecc3a46cb936b3ff0c9731255ea5f1f11d67b22edc35ab5950caa166891722a

Observation 3e18282a-a525-43b9-8ffc-343806a7f11b · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.007029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:28dc7858ded33719d7dc4d2dd87988f1b0a36377df7a91bf4519d21a06023d41

Observation c338478b-de5f-4260-961b-ead6fc97eb4e · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs TinyLlama: An Open-Source Small Language Model

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.150782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:2b66a343694ec3e638e45bef4cc1ed7055aa49159e33c85d5d82aa0eb245ba77

Observation 85388c2f-c1de-4033-879f-dc63fd0b9bcc · outbound

This paper cites It was not used to generate research ideas.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs It was not used to generate research ideas

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:30:55.447361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:608813b01445dc0934cd4b2a952dba7f7371988b8809af4f92a47eae3eab1328

Observation 5fc007a8-5807-40d6-89ee-8db1be66b5e0 · outbound

This paper cites Across varying batch sizes and model scales, larger hidden sizes yield higher inference throughput under a fixed parameter budget.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Across varying batch sizes and model scales, larger hidden sizes yield higher inference throughput under a fixed parameter budget

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:30:55.449858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:f30354d4353eef1fb952b1d234d617663f819b4d40aba5c087b84e260033753b

Observation 2b9ec17b-1782-4749-8eca-43101dfe8541 · outbound

This paper cites Moreover, while multiplicative and additive calibrations differ in formulation, their MSE and Spearman values remain nearly identical.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Moreover, while multiplicative and additive calibrations differ in formulation, their MSE and Spearman values remain nearly identical

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T05:30:55.445040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:a51dedb803336409234b1ce2593d0671cd2a3bdffd6da8a8615fe3940c339e3c

Observation 8cfb120f-2eff-4147-9a34-e2b24c757cf4 · outbound

This paper cites an unresolved cited work.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-18T05:30:55.439819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:8172e0c5fdd1d22133e8cfd4897640b90c0f243e86a8f2c5b845f4fd32f1c709

Pith citing papers

No inbound Pith citation observations are available.