REVIEW 3 major objections 3 minor 1 cited by
An Overview of Algorithms for Contactless Cardiac Feature Extraction from Radar Signals: Advances and Challenges
T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This review claims to be the first dedicated to signal-processing algorithms for extracting cardiac features from radar, supported by a new taxonomy and a public-dataset guide.
desk verdict Submission mismatch: the supplied PDF is a sparse-attention LLM paper, not the claimed radar cardiac-feature review, so the review's central claims are unverifiable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The working machinery is the review's classification scheme: a taxonomy of cardiac feature extraction algorithms organized by their core feature, a structured statement of the pros and cons of each algorithm class, and a table of public datasets with detailed configurations and ground-truth cardiac signals. The taxonomy does the argumentative work by making a scattered literature navigable, and the dataset inventory anchors algorithm comparisons in reproducible data. For a review, these organizing devices are what carry the claim that the field can be understood and advanced from the algorithm side.
What would settle it
Search scholarly databases for a review article published before this paper's submission whose stated focus is algorithms for extracting cardiac features from radar signals; finding one would falsify the first-review claim even if the rest of the review remains useful.
Extended reading notes
Core claim
On its own terms, this paper's central claim is a claim about the state of the literature: no earlier review has concentrated on algorithms for extracting cardiac features from received radar signals. To make that point useful, it introduces a new taxonomy designed to reveal the core feature of each algorithm, evaluates each algorithm's advantages and disadvantages in detail, and catalogues public datasets that contain both the received radar signal and the ground-truth cardiac feature signal, with configurations and evaluations meant to help readers choose among them. It closes by stating unsolved challenges and suggesting future research directions. The intended conclusion is that radar can give unobtrusive, accurate, and reliable contactless cardiac monitoring once the algorithm side of the field is systematically understood.
Load-bearing premise
The load-bearing premise is that the authors' literature search was comprehensive, so no earlier review focused on radar cardiac feature extraction algorithms was missed; if such a review exists, the claim of being first fails.
Editorial extensions
If this is right
- Researchers new to the area can use the taxonomy to compare algorithm families by their underlying principle instead of by the radar hardware used.
- The public-dataset list gives the field a shared reference point for benchmarking new cardiac feature extraction algorithms against recorded radar signals with ground truth.
- A clear statement of pros and cons for each algorithm class shows where current methods are mature and where they fall short, directing future effort to the limiting steps.
- The challenges and future directions listed in the paper supply a ready agenda for work aimed at making contactless radar cardiac monitoring practical in smart homes and in-cabin settings.
Reading between the lines
- A quantitative comparison of the catalogued algorithms run on the same public datasets would be a natural extension, since a taxonomy plus pros-and-cons discussion does not by itself rank methods by accuracy or reliability.
- The same taxonomy could plausibly be adapted to other contactless sensing modalities, such as cameras or Wi-Fi-based sensing, where the cardiac feature extraction problem has a similar structure.
- Because the supplied full text is a different paper, these extensions should be treated as inferences from the abstract; checking the review's actual taxonomy and dataset details against the published version is the first step before relying on them.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission is registered as arXiv:2508.02122 (eess.SP), with an abstract announcing a first review paper on algorithms for contactless cardiac feature extraction from radar signals, including a new algorithm taxonomy, a survey of public radar datasets for cardiac feature extraction, and a discussion of open challenges. The full text supplied for review is, however, an unrelated manuscript titled 'Trainable Dynamic Mask Sparse Attention' (arXiv:2508.02124v6, cs.AI), which addresses sparse attention mechanisms for large language models. The supplied full text contains no radar signal processing, no cardiac feature extraction algorithms, no taxonomy of such algorithms, and no list of radar cardiac datasets.
Significance. If the claimed radar cardiac-feature review existed as described, it could be a useful entry point for researchers in contactless cardiac monitoring: the proposed taxonomy would organize algorithm-level choices, the dataset survey would support reproducible benchmarking, and the challenge list would map open problems. None of these contributions can be assessed from the submitted artifact, because the text under review is a completely different paper. The open-source kernel code and experimental results in the supplied full text are strengths of that other paper, but they provide no evidence bearing on the radar-review claims and cannot substitute for the missing survey content.
major comments (3)
- [Manuscript header and full text] The submitted full text is not the manuscript advertised by the title and abstract. The header carries arXiv identifier 2508.02124v6 and the title 'Trainable Dynamic Mask Sparse Attention', whereas the claimed submission is arXiv:2508.02122, 'An Overview of Algorithms for Contactless Cardiac Feature Extraction from Radar Signals'. This is a load-bearing mismatch: the text under review contains no radar signal processing, no cardiac feature extraction, no algorithm taxonomy, and no cardiac radar dataset survey, so none of the abstract's central assertions can be verified from the submitted material.
- [Abstract, first-review claim] The abstract's claim that 'to the best of the author knowledge, this is the first review paper' cannot be evaluated because the review itself is absent. Even if the literature search were comprehensive, the claim of firstness requires a defined search strategy, inclusion criteria, and a comparison against prior surveys; the submitted text provides none of these, and for the same reason the proposed taxonomy and dataset tables promised in the abstract are not present.
- [Entire manuscript (claimed review content)] The paper promises 'pros and cons evaluated in detail' for cardiac feature extraction algorithms and 'public datasets containing the received radar signal and ground-truth cardiac feature signal' with 'detailed configurations'. No such evaluations or dataset tables appear anywhere in the supplied full text. The absence is structural rather than local: the supplied text's sections, equations, experiments, and references all concern sparse attention in transformers, making it impossible to fix the review content by minor revision.
minor comments (3)
- [Title/abstract vs. body] The title and abstract describe a radar cardiac feature extraction review, while the body is a sparse-attention methods paper; the inconsistency is visible already in the arXiv identifier on the first page, which does not match the claimed submission number.
- [Abstract wording] The abstract contains phrasing such as 'to the best of the author knowledge' and 'can be served as a guide'; these are presentation issues that would need correction in any resubmission of the actual review.
- [Survey methodology] The abstract does not mention the search strategy, inclusion criteria, or period covered by the literature review; a survey paper should state these explicitly, but this point is secondary to the identity mismatch documented above.
Circularity Check
No circular derivation found; the artifact is an unrelated paper, so the review's claims cannot be assessed but also cannot be shown circular.
full rationale
The claimed paper (arXiv:2508.02122) is a radar cardiac-feature-extraction review whose abstract makes no derivation claim: it asserts a first-review status, proposes a taxonomy, lists datasets, and discusses challenges. There are no equations, fitted parameters, or predictions in the abstract that could be equivalent by construction to its inputs. None of the circularity patterns (self-definitional, fitted-input-as-prediction, load-bearing self-citation, imported uniqueness, ansatz-via-citation, renaming) is present. I separately flag a structural completeness failure rather than circularity: the supplied full text is 'Trainable Dynamic Mask Sparse Attention' with header 'arXiv:2508.02124v6', containing no radar signal processing, cardiac feature extraction, taxonomy, or dataset survey. Therefore the abstract's first-review claim cannot be verified from the artifact. However, a mismatch between the claimed and supplied manuscript is not a circularity reduction, and I can quote no equation or fitted-parameter step that reduces the target result to its own inputs. Score is therefore 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The authors' literature search comprehensively covers prior work on radar cardiac feature extraction algorithms.
- domain assumption Radar is capable of high-accuracy, robust unobtrusive cardiac monitoring relative to other contactless sensors.
Cite this review
Pith. "Pith review of An Overview of Algorithms for Contactless Cardiac Feature Extraction from Radar Signals: Advances and Challenges." pith.science (2026). https://pith.science/paper/NLTXQR2M
@misc{pith2026250802122,
author = {Pith},
title = {Pith review of: An Overview of Algorithms for Contactless Cardiac Feature Extraction from Radar Signals: Advances and Challenges},
year = {2026},
howpublished = {\url{https://pith.science/paper/NLTXQR2M}},
note = {Machine review of arXiv:2508.02122}
}
read the original abstract
Contactless cardiac monitoring has vast potential to replace contact-based monitoring in various future scenarios such as smart home and in-cabin monitoring. Various contactless sensors can be potentially implemented for cardiac monitoring, such as cameras, acoustic sensors, Wi-Fi routers and radars. Among all these sensors, radar could achieve unobtrusive monitoring with high accuracy and robustness at the same time. The research about radar-based cardiac monitoring can be generally divided into the radar architecture design and signal-processing parts, where the former has been thoroughly reviewed in the literature but not the latter. To the best of the author knowledge, this is the first review paper that focuses on elaborating the algorithms for extracting cardiac features from the received radar signal. In addition, a new taxonomy is proposed to reveal the core feature of each algorithm, with the pros and cons evaluated in detail. Furthermore, the public datasets containing the received radar signal and ground-truth cardiac feature signal are listed with detailed configurations, and the corresponding evaluations may help the researchers select the suitable dataset. At last, several unsolved challenges and future directions are suggested and discussed in detail to encourage future research on solving the main obstacles in this field. In summary, this review can be served as a guide for researchers and practitioners to quickly understand the research trend and recent development of the cardiac feature extraction algorithms, and it is worth further investigating the relative area based on the proposed challenges and future directions.
Forward citations
Cited by 1 Pith paper
-
Transformer-Based Heartbeat Monitoring with FMCW Radar Under Random Body Motion
Hybrid model-based preprocessing plus CNN-Transformer reconstructs PPG signals from FMCW radar and achieves the highest score on the IEEE AESS Radar Challenge across stationary, deep-breathing, and random-body-motion ...
Reference graph
Works this paper leans on
-
[1]
Unitary Evolution Recurrent Neural Networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio. “Unitary Evolution Recurrent Neural Networks”. In:The Interna- tional Conference on Machine Learning (ICML). 2016, pp. 1120–1128
work page 2016
-
[2]
Zoology: Measuring and Improving Recall in Efficient Language Models
Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, and Christopher Ré. “Zoology: Measuring and Improving Recall in Efficient Language Models”. In:The International Conference on Learning Representations. 2024
work page 2024
-
[3]
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. “Program synthesis with large language models”. In:arXiv preprint arXiv:2108.07732(2021)
arXiv 2021
-
[4]
Longbench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al. “Longbench: A bilingual, multitask benchmark for long context understanding”. In:arXiv preprint arXiv:2308.14508(2023)
arXiv 2023
-
[5]
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. “Longformer: The long-document transformer”. In:arXiv preprint arXiv:2004.05150(2020)
arXiv 2020
- [6]
-
[7]
PIQA: Reasoning about Physical Commonsense in Natural Language
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. “PIQA: Reasoning about Physical Commonsense in Natural Language”. In:Proceedings of the AAAI conference on Artificial Intelligence. Vol. 34. 2020
work page 2020
-
[8]
Gpt-NeoX-20B: An Open-source Autoregressive Language Model
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. “Gpt-NeoX-20B: An Open-source Autoregressive Language Model”. In:arXiv preprint arXiv:2204.06745(2022)
arXiv 2022
Show all 58 references
-
[9]
Magicpig: Lsh sampling for efficient llm generation
Zhuoming Chen, Ranajoy Sadhukhan, Zihao Ye, Yang Zhou, Jianyu Zhang, Niklas Nolte, Yuandong Tian, Matthijs Douze, Leon Bottou, Zhihao Jia, et al. “Magicpig: Lsh sampling for efficient llm generation”. In:arXiv preprint arXiv:2410.16179(2024)
2024 arXiv
-
[10]
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. “Generating long sequences with sparse transformers”. In:arXiv preprint arXiv:1904.10509(2019)
2019 arXiv
-
[11]
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. “Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge”. In:arXiv preprint arXiv:1803.05457 (2018)
2018 arXiv
-
[12]
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. “Training verifiers to solve math word problems”. In:arXiv preprint arXiv:2110.14168(2021)
2021 arXiv
-
[13]
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. “Transformer-xl: Attentive language models beyond a fixed-length context”. In:arXiv preprint arXiv:1901.02860(2019)
2019 arXiv
-
[14]
FLASHATTENTION: fast and memory- efficient exact attention with IO-awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. “FLASHATTENTION: fast and memory- efficient exact attention with IO-awareness”. In:Proceedings of the 36th International Conference on Neural Informa- tion Processing Systems. 2022
2022
-
[15]
Google DeepMind.Gemini 2.5 Pro. Mar. 2025.url:https : / / blog . google / technology / google - deepmind / gemini-model-thinking-updates-march-2025
2025
-
[16]
HashAttention: Semantic Sparsity for Faster Inference
Aditya Desai, Shuo Yang, Alejandro Cuadron, Ana Klimovic, Matei Zaharia, Joseph E Gonzalez, and Ion Stoica. “HashAttention: Semantic Sparsity for Faster Inference”. In:arXiv preprint arXiv:2412.14468(2024)
2024 arXiv
-
[17]
Version 0.7.0
Clémentine Fourrier, Nathan Habib, Hynek Kydlíček, Thomas Wolf, and Lewis Tunstall.LightEval: A lightweight framework for LLM evaluation. Version 0.7.0. 2023.url:https://github.com/huggingface/lighteval
2023
-
[18]
Version v0.0.1
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou.A Framework for Few-shot Language Model ...
2021 doi
-
[19]
Seerattention: Learning intrinsic sparse attention in your llms
Yizhao Gao, Zhichen Zeng, Dayou Du, Shijie Cao, Peiyuan Zhou, Jiaxing Qi, Junjie Lai, Hayden Kwok-Hay So, Ting Cao, Fan Yang, et al. “Seerattention: Learning intrinsic sparse attention in your llms”. In:arXiv preprint arXiv:2410.13276(2024)
2024 arXiv
-
[20]
Model tells you what to discard: Adaptive kv cache compression for llms
Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao. “Model tells you what to discard: Adaptive kv cache compression for llms”. In:arXiv preprint arXiv:2310.01801(2023). 19
2023 arXiv
-
[21]
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning”. In:arXiv preprint arXiv:2501.12948(2025)
2025 arXiv
-
[22]
Scaling laws and compute- optimal training beyond fixed training durations
Alex Hägele, Elie Bakouch, Atli Kosson, Leandro Von Werra, Martin Jaggi, et al. “Scaling laws and compute- optimal training beyond fixed training durations”. In:Advances in Neural Information Processing Systems37 (2024), pp. 76232–76264
2024
-
[23]
Mea- suring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. “Mea- suring massive multitask language understanding”. In:arXiv preprint arXiv:2009.03300(2020)
2020 arXiv
-
[24]
Mea- suring Massive Multitask Language Understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. “Mea- suring Massive Multitask Language Understanding”. In:International Conference on Learning Representations. 2021
2021
-
[25]
An Empirical Analysis of Compute- Optimal Large Language Model Training
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. “An Empirical Analysis of Compute- Optimal Large Language Model Training”. In:Advances in Neural In...
2022
-
[26]
RULER: What’s the Real Context Size of Your Long-Context Language Models?
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. “RULER: What’s the Real Context Size of Your Long-Context Language Models?” In:arXiv preprint arXiv:2404.06654(2024)
2024 arXiv
-
[27]
2025.url:https://github.com/huggingface/ open-r1
HuggingFace.Open R1: A fully open reproduction of DeepSeek-R1. 2025.url:https://github.com/huggingface/ open-r1
2025
-
[28]
Weld, and Luke Zettlemoyer.TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer.TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. 2017. arXiv:1705.03551 [cs.CL]
2017 arXiv
-
[29]
2023.url:https://github.com/gkamradt/LLMTest_NeedleInAHaystack
G Kamradt.LLMTest NeedleInAHaystack. 2023.url:https://github.com/gkamradt/LLMTest_NeedleInAHaystack
2023
-
[30]
Houyi Li, Wenzhen Zheng, Qiufeng Wang, Hanshan Zhang, Zili Wang, Shijie Xuyang, Yuantao Fan, Shuigeng Zhou, Xiangyu Zhang, and Daxin Jiang.Predictable Scale: Part I – Optimal Hyperparameter Scaling Law in Large Language Model Pretraining. 2025. arXiv:2503.04715 [cs.LG].url:htt...
2025 arXiv
-
[31]
Snapkv: Llm knows what you are looking for before generation
Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen. “Snapkv: Llm knows what you are looking for before generation”. In:Advances in Neural Infor- mation Processing Systems37 (2024), pp. 22947–22970
2024
-
[32]
Deepseek-v3 technical report
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. “Deepseek-v3 technical report”. In:arXiv preprint arXiv:2412.19437(2024)
2024 arXiv
-
[33]
Clusterkv: Manipulating llm kv cache in semantic space for recallable compression
Guangda Liu, Chengwei Li, Jieru Zhao, Chenqi Zhang, and Minyi Guo. “Clusterkv: Manipulating llm kv cache in semantic space for recallable compression”. In:arXiv preprint arXiv:2412.03213(2024)
2024 arXiv
-
[34]
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. “Lost in the middle: How language models use long contexts”. In:Transactions of the Association for Computational Lin- guistics12 (2024), pp. 157–173
2024
-
[35]
Fixing Weight Decay Regularization in Adam
Ilya Loshchilov and Frank Hutter. “Fixing Weight Decay Regularization in Adam”. In:ArXivabs/1711.05101 (2017). url:https://api.semanticscholar.org/CorpusID:3312944
2017 arXiv
-
[36]
From softmax to sparsemax: A sparse model of attention and multi-label classification
Andre Martins and Ramon Astudillo. “From softmax to sparsemax: A sparse model of attention and multi-label classification”. In:International conference on machine learning. PMLR. 2016, pp. 1614–1623
2016
-
[37]
Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. “Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering”. In:arXiv preprint arXiv:1809.02789(2018)
2018 arXiv
-
[38]
Meta NVIDIA.PyTorch Container Image.https : / / catalog . ngc . nvidia . com / orgs / nvidia / containers / pytorch. 2022
2022
-
[39]
In-context Learning and Induction Heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kam...
2022
-
[40]
The LAMBADA Dataset: Word Prediction Requiring a Broad Discourse Context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. “The LAMBADA Dataset: Word Prediction Requiring a Broad Discourse Context”. In:Proceedings of the 54th Annual Meeting of...
2016
-
[41]
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. “Generative agents: Interactive simulacra of human behavior”. In:Proceedings of the 36th annual acm symposium on user interface software and technology. 2023, pp. 1–22
2023
-
[42]
CKConv: Continuous Kernel Convolution For Sequential Data
David W Romero, Anna Kuzina, Erik J Bekkers, Jakub M Tomczak, and Mark Hoogendoorn. “CKConv: Continuous Kernel Convolution For Sequential Data”. In:arXiv preprint arXiv:2102.02611(2021)
2021 arXiv
-
[43]
Winogrande: An Adversarial Winograd Schema Challenge at Scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. “Winogrande: An Adversarial Winograd Schema Challenge at Scale”. In:Communications of the ACM64.9 (2021), pp. 99–106
2021
-
[44]
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. “Scaling llm test-time compute optimally can be more effective than scaling model parameters”. In:arXiv preprint arXiv:2408.03314(2024)
2024 arXiv
-
[45]
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, et al. “Challenging big-bench tasks and whether chain-of-thought can solve them”. In:Findings of the Association for Computational Lin...
2023
-
[46]
Quest: Query-aware sparsity for efficient long-context llm inference
Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, and Song Han. “Quest: Query-aware sparsity for efficient long-context llm inference”. In:arXiv preprint arXiv:2406.10774(2024)
2024 arXiv
-
[47]
Qwen Team.Qwen3. Apr. 2025.url:https://qwenlm.github.io/blog/qwen3
2025
-
[48]
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. “Attention Is All You Need”. In:Advances in Neural Information Processing Systems. 2017
2017
-
[49]
Transformers: State-of-the-Art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020
-
[50]
Infllm: Training-free long-context extrapolation for llms with an efficient context memory
Chaojun Xiao, Pengle Zhang, Xu Han, Guangxuan Xiao, Yankai Lin, Zhengyan Zhang, Zhiyuan Liu, and Maosong Sun. “Infllm: Training-free long-context extrapolation for llms with an efficient context memory”. In:arXiv preprint arXiv:2402.04617(2024)
2024 arXiv
-
[51]
Effective long-context scaling of foundation models
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, et al. “Effective long-context scaling of foundation models”. In:arXiv preprint arXiv:2309.16039(2023)
2023 arXiv
-
[52]
Native sparse attention: Hardware-aligned and natively trainable sparse attention
Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, YX Wei, Lean Wang, Zhiping Xiao, et al. “Native sparse attention: Hardware-aligned and natively trainable sparse attention”. In: arXiv preprint arXiv:2502.11089(2025)
2025 arXiv
-
[53]
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. “Big bird: Transformers for longer sequences”. In:Advances in neural information processing systems33 (2020), pp. 17283–17297
2020
-
[54]
HellaSwag: Can a Machine Really Finish Your Sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. “HellaSwag: Can a Machine Really Finish Your Sentence?” In:Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019
2019
-
[55]
Hanzhi Zhang, Heng Fan, Kewei Sha, Yan Huang, and Yunhe Feng.DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration. 2025. arXiv:2506.11104 [cs.CL].url:https://arxiv.org/abs/ 2506.11104
2025 arXiv
-
[56]
Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges
Kechi Zhang, Jia Li, Ge Li, Xianjie Shi, and Zhi Jin. “Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges”. In:arXiv preprint arXiv:2401.07339(2024)
2024 arXiv
-
[57]
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al. “H2o: Heavy-hitter oracle for efficient generative inference of large language models”. In:Advances in Neural Information Processing ...
2023
-
[58]
LLM MapReduce: Simplified Long-Sequence Processing using Large Language Models
Zihan Zhou, Chong Li, Xinyi Chen, Shuo Wang, Yu Chao, Zhili Li, Haoyu Wang, Rongqiao An, Qi Shi, Zhixing Tan, et al. “LLM MapReduce: Simplified Long-Sequence Processing using Large Language Models”. In:CoRR(2024). 21 A Dynamic Mask Attention Implementation The following listin...
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.