{"work":{"id":"5b0449ac-92b0-41f2-8b4f-586c2b5a08b6","openalex_id":"https://openalex.org/W4384648484","doi":"10.48550/arxiv.2307.08621","arxiv_id":"2307.08621","raw_key":null,"title":"Retentive Network: A Successor to Transformer for Large Language Models","authors":null,"authors_text":"Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue","year":2023,"venue":"cs.CL","abstract":"In this work, we propose Retentive Network (RetNet) as a foundation architecture for large language models, simultaneously achieving training parallelism, low-cost inference, and good performance. We theoretically derive the connection between recurrence and attention. Then we propose the retention mechanism for sequence modeling, which supports three computation paradigms, i.e., parallel, recurrent, and chunkwise recurrent. Specifically, the parallel representation allows for training parallelism. The recurrent representation enables low-cost $O(1)$ inference, which improves decoding throughput, latency, and GPU memory without sacrificing performance. The chunkwise recurrent representation facilitates efficient long-sequence modeling with linear complexity, where each chunk is encoded parallelly while recurrently summarizing the chunks. Experimental results on language modeling show that RetNet achieves favorable scaling results, parallel training, low-cost deployment, and efficient inference. The intriguing properties make RetNet a strong successor to Transformer for large language models. Code will be available at https://aka.ms/retnet.","external_url":"https://arxiv.org/abs/2307.08621","cited_by_count":108,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2307.08621","created_at":"2026-05-10T00:54:48.519414+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Retentive Network: A Successor to Transformer for Large Language Models","render_title":"Retentive Network: A Successor to Transformer for Large Language Models"},"hub":{"state":{"work_id":"5b0449ac-92b0-41f2-8b4f-586c2b5a08b6","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":124,"external_cited_by_count":108,"distinct_field_count":10,"first_pith_cited_at":"2023-03-31T17:28:46+00:00","last_pith_cited_at":"2026-07-08T17:59:09+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T00:39:31.592512+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":21},{"context_role":"method","n":2}],"polarity_counts":[{"context_polarity":"background","n":17},{"context_polarity":"unclear","n":3},{"context_polarity":"use_method","n":2},{"context_polarity":"support","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Retentive Network: A Successor to Transformer for Large Language Models","claims":[{"claim_text":"In this work, we propose Retentive Network (RetNet) as a foundation architecture for large language models, simultaneously achieving training parallelism, low-cost inference, and good performance. We theoretically derive the connection between recurrence and attention. Then we propose the retention mechanism for sequence modeling, which supports three computation paradigms, i.e., parallel, recurrent, and chunkwise recurrent. Specifically, the parallel representation allows for training parallelism. The recurrent representation enables low-cost $O(1)$ inference, which improves decoding throughp","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"network: A successor to transformer for large language models\". In: arXiv preprint arXiv:2307.08621 (2023). [103] Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. \"Gemma: Open models based on gemini research and technology\". In: arXiv preprint arXiv:2403.08295 (2024). 22 [104] W Scott Terry. Learning and memory: Basic principles, processes, and procedures . Routledge, 2017. [105]","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"This makes long-sequence inference practical, but it also concentrates all prefix information into a limited matrix state, producing known weaknesses on associative recall and exact copying [2, 3, 60, 24]. Modern recurrent linear-attention layers improve this tradeoff through decay, data-dependent gating, or more expressive transition maps: RetNet [53], GLA [62], Gated DeltaNet [61], RWKV-7 [38], and KDA / Kimi Linear [55] all modify how the state is retained or overwritten. OSDN is orthogonal t","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"[72] Jimmy TH Smith, Andrew Warrington, and Scott W Linderman. Simplified state space layers for sequence modeling.arXiv preprint arXiv:2208.04933, 2022. [73] Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. Retentive network: A successor to transformer for large language models.arXiv preprint arXiv:2307.08621, 2023. [74] Ilya Sutskever, Oriol Vinyals, and Quoc V Le. Sequence to sequence learning with neural networks.Advances in neural informati","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"The FEP has inspired several ML architectures. Active inference agents [26], [27], [28] apply the FEP to reinforcement learning, replacing reward maximization with free energy minimization. These systems operate on episodic timescales and focus on action selection in environment interac- tion, a diﬀerent regime from token-by-token routing in language models. Predictive coding networks [29], [30] implement hierarchical prediction error minimization as an alternative to backpropagation, motivated ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. V on Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, ed- itors,Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper_files/paper/2017/file/ 3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf. [39] Roger Waleffe, Wonmin Byeon, Duncan Riach, Brandon Norick, Vijay Korthikanti, Tri Dao, Albert Gu, Ali Hatamizadeh, ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"attention score between a query and a key,r i−j denotes a learnable scalar based on the offset between the query and the key, andR Θ,t denotes a rotary matrix with rotation degreet·Θ. Configuration Method Equation Normalization position Post Norm [22] Norm(x+Sublayer(x)) Pre Norm [26] x+ Sublayer(Norm(x)) Sandwich Norm [274] x+ Norm(Sublayer(Norm(x))) Normalization method LayerNorm [275] x−µ σ ·γ+β, µ= 1 d Pd i=1 xi, σ= q 1 d Pd i=1(xi −µ)) 2 RMSNorm [276] x RMS(x) ·γ,RMS(x) = q 1 d Pd i=1 x2 i ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Retentive Network: A Successor to Transformer for Large Language Models because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (21 contexts).","role_counts":[{"n":21,"context_role":"background"},{"n":2,"context_role":"method"}]},"error":null,"updated_at":"2026-06-29T20:59:17.264071+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"32104b60-617c-49b7-bef0-900d62d14f88","orcid":null,"display_name":"Yutao Sun"},{"id":"e5d982b3-1c02-4828-8399-4f667dadefdb","orcid":null,"display_name":"Li Dong"},{"id":"019f6545-376f-4623-a727-dabd20ab9286","orcid":null,"display_name":"Shaohan Huang"},{"id":"6ed1f2fe-df78-4a3e-bdff-eebdc69382ff","orcid":null,"display_name":"Shuming Ma"},{"id":"1222d56d-5ab2-45b5-881d-9ce07af16216","orcid":null,"display_name":"Yuqing Xia"},{"id":"44dc4efb-f66b-47bd-97e7-48ab4f1143bc","orcid":null,"display_name":"Jilong Xue"}]},"error":null,"updated_at":"2026-06-29T20:59:17.671079+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T17:56:26.166579+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Mamba: Linear-Time Sequence Modeling with Selective State Spaces","work_id":"4ee75248-1199-492c-a52f-6661e0f4adff","shared_citers":21},{"title":"Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge","work_id":"28ea1282-d657-4c61-a83c-f1249be6d6b1","shared_citers":11},{"title":"Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality","work_id":"d8eba076-0449-4f6a-aae1-5a7260677f0f","shared_citers":11},{"title":"RWKV: Reinventing RNNs for the Transformer Era","work_id":"524dc80d-f4ef-4f89-bf1a-9a8c1e4b6a81","shared_citers":10},{"title":"Efficiently Modeling Long Sequences with Structured State Spaces","work_id":"4150b761-b8bf-4d9b-a2f8-cb2d1b73d378","shared_citers":9},{"title":"Gated linear attention transformers with hardware-efficient training","work_id":"65a18a30-6e80-4b64-a026-bb0368e38872","shared_citers":9},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":9},{"title":"RULER: What's the Real Context Size of Your Long-Context Language Models?","work_id":"c0bc4689-3ce8-4e3d-9442-bd74869445bb","shared_citers":9},{"title":"Griffin: Mixing gated linear recurrences with local attention for efficient language models","work_id":"546b02db-26f0-4dac-b50e-912e7f4e181c","shared_citers":7},{"title":"Eagle and finch: Rwkv with matrix- valued states and dynamic recurrence.arXiv preprint arXiv:2404.05892","work_id":"b993c2c1-c4d3-4ade-a380-dc5464bcb940","shared_citers":6},{"title":"Generating Long Sequences with Sparse Transformers","work_id":"c5b81688-45ee-4a9a-b095-e6290f45cb6c","shared_citers":6},{"title":"GLU Variants Improve Transformer","work_id":"17d0763c-1016-41ab-a478-478e890765eb","shared_citers":6},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":6},{"title":"Jamba: A Hybrid Transformer-Mamba Language Model","work_id":"129df0fe-8a66-4077-8991-3557cfa38274","shared_citers":6},{"title":"Linformer: Self-Attention with Linear Complexity","work_id":"4b717b51-6098-45d0-8e9e-b69bef651bc3","shared_citers":6},{"title":"Longformer: The Long-Document Transformer","work_id":"abea7a44-6668-4de7-aab6-f53a6e5aa088","shared_citers":6},{"title":"net/forum?id=HyUNwulC-","work_id":"137e6ed4-1c1d-40bd-801e-9d2baea0bd0c","shared_citers":6},{"title":"RWKV-7 “Goose” with expressive dynamic state evolution","work_id":"daf20308-c801-4745-a147-a1f1de4368e0","shared_citers":6},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":6},{"title":"arXiv preprint arXiv:2404.07904 , year=","work_id":"62f21939-fe4f-4159-bb1e-6a303fdc7b42","shared_citers":5},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":5},{"title":"Kimi Linear: An Expressive, Efficient Attention Architecture","work_id":"b2f9f1cd-c39c-4dbc-8637-f575681cdc01","shared_citers":5},{"title":"Native sparse attention: Hardware-aligned and natively trainable sparse attention.arXiv preprint arXiv:2502.11089","work_id":"c74d35ec-94d3-48f5-a291-20cfad7c04a3","shared_citers":5},{"title":"Rethinking Attention with Performers","work_id":"4c26d308-8b72-4a98-8e73-950617a75f50","shared_citers":5}],"time_series":[{"n":2,"year":2023},{"n":3,"year":2024},{"n":3,"year":2025},{"n":30,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T17:59:50.306256+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T17:56:22.262298+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Retentive Network: A Successor to Transformer for Large Language Models","claims":[{"claim_text":"In this work, we propose Retentive Network (RetNet) as a foundation architecture for large language models, simultaneously achieving training parallelism, low-cost inference, and good performance. We theoretically derive the connection between recurrence and attention. Then we propose the retention mechanism for sequence modeling, which supports three computation paradigms, i.e., parallel, recurrent, and chunkwise recurrent. Specifically, the parallel representation allows for training parallelism. The recurrent representation enables low-cost $O(1)$ inference, which improves decoding throughp","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"network: A successor to transformer for large language models\". In: arXiv preprint arXiv:2307.08621 (2023). [103] Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. \"Gemma: Open models based on gemini research and technology\". In: arXiv preprint arXiv:2403.08295 (2024). 22 [104] W Scott Terry. Learning and memory: Basic principles, processes, and procedures . Routledge, 2017. [105]","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"This makes long-sequence inference practical, but it also concentrates all prefix information into a limited matrix state, producing known weaknesses on associative recall and exact copying [2, 3, 60, 24]. Modern recurrent linear-attention layers improve this tradeoff through decay, data-dependent gating, or more expressive transition maps: RetNet [53], GLA [62], Gated DeltaNet [61], RWKV-7 [38], and KDA / Kimi Linear [55] all modify how the state is retained or overwritten. OSDN is orthogonal t","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"[72] Jimmy TH Smith, Andrew Warrington, and Scott W Linderman. Simplified state space layers for sequence modeling.arXiv preprint arXiv:2208.04933, 2022. [73] Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. Retentive network: A successor to transformer for large language models.arXiv preprint arXiv:2307.08621, 2023. [74] Ilya Sutskever, Oriol Vinyals, and Quoc V Le. Sequence to sequence learning with neural networks.Advances in neural informati","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"The FEP has inspired several ML architectures. Active inference agents [26], [27], [28] apply the FEP to reinforcement learning, replacing reward maximization with free energy minimization. These systems operate on episodic timescales and focus on action selection in environment interac- tion, a diﬀerent regime from token-by-token routing in language models. Predictive coding networks [29], [30] implement hierarchical prediction error minimization as an alternative to backpropagation, motivated ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. V on Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, ed- itors,Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper_files/paper/2017/file/ 3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf. [39] Roger Waleffe, Wonmin Byeon, Duncan Riach, Brandon Norick, Vijay Korthikanti, Tri Dao, Albert Gu, Ali Hatamizadeh, ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"attention score between a query and a key,r i−j denotes a learnable scalar based on the offset between the query and the key, andR Θ,t denotes a rotary matrix with rotation degreet·Θ. Configuration Method Equation Normalization position Post Norm [22] Norm(x+Sublayer(x)) Pre Norm [26] x+ Sublayer(Norm(x)) Sandwich Norm [274] x+ Norm(Sublayer(Norm(x))) Normalization method LayerNorm [275] x−µ σ ·γ+β, µ= 1 d Pd i=1 xi, σ= q 1 d Pd i=1(xi −µ)) 2 RMSNorm [276] x RMS(x) ·γ,RMS(x) = q 1 d Pd i=1 x2 i ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Retentive Network: A Successor to Transformer for Large Language Models because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (21 contexts).","role_counts":[{"n":21,"context_role":"background"},{"n":2,"context_role":"method"}]},"error":null,"updated_at":"2026-06-29T20:59:17.673667+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Retentive Network: A Successor to Transformer for Large Language Models","claims":[{"claim_text":"In this work, we propose Retentive Network (RetNet) as a foundation architecture for large language models, simultaneously achieving training parallelism, low-cost inference, and good performance. We theoretically derive the connection between recurrence and attention. Then we propose the retention mechanism for sequence modeling, which supports three computation paradigms, i.e., parallel, recurrent, and chunkwise recurrent. Specifically, the parallel representation allows for training parallelism. The recurrent representation enables low-cost $O(1)$ inference, which improves decoding throughp","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Retentive Network: A Successor to Transformer for Large Language Models because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T17:59:45.691731+00:00"}},"summary":{"title":"Retentive Network: A Successor to Transformer for Large Language Models","claims":[{"claim_text":"In this work, we propose Retentive Network (RetNet) as a foundation architecture for large language models, simultaneously achieving training parallelism, low-cost inference, and good performance. We theoretically derive the connection between recurrence and attention. Then we propose the retention mechanism for sequence modeling, which supports three computation paradigms, i.e., parallel, recurrent, and chunkwise recurrent. Specifically, the parallel representation allows for training parallelism. The recurrent representation enables low-cost $O(1)$ inference, which improves decoding throughp","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Retentive Network: A Successor to Transformer for Large Language Models because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Mamba: Linear-Time Sequence Modeling with Selective State Spaces","work_id":"4ee75248-1199-492c-a52f-6661e0f4adff","shared_citers":21},{"title":"Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge","work_id":"28ea1282-d657-4c61-a83c-f1249be6d6b1","shared_citers":11},{"title":"Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality","work_id":"d8eba076-0449-4f6a-aae1-5a7260677f0f","shared_citers":11},{"title":"RWKV: Reinventing RNNs for the Transformer Era","work_id":"524dc80d-f4ef-4f89-bf1a-9a8c1e4b6a81","shared_citers":10},{"title":"Efficiently Modeling Long Sequences with Structured State Spaces","work_id":"4150b761-b8bf-4d9b-a2f8-cb2d1b73d378","shared_citers":9},{"title":"Gated linear attention transformers with hardware-efficient training","work_id":"65a18a30-6e80-4b64-a026-bb0368e38872","shared_citers":9},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":9},{"title":"RULER: What's the Real Context Size of Your Long-Context Language Models?","work_id":"c0bc4689-3ce8-4e3d-9442-bd74869445bb","shared_citers":9},{"title":"Griffin: Mixing gated linear recurrences with local attention for efficient language models","work_id":"546b02db-26f0-4dac-b50e-912e7f4e181c","shared_citers":7},{"title":"Eagle and finch: Rwkv with matrix- valued states and dynamic recurrence.arXiv preprint arXiv:2404.05892","work_id":"b993c2c1-c4d3-4ade-a380-dc5464bcb940","shared_citers":6},{"title":"Generating Long Sequences with Sparse Transformers","work_id":"c5b81688-45ee-4a9a-b095-e6290f45cb6c","shared_citers":6},{"title":"GLU Variants Improve Transformer","work_id":"17d0763c-1016-41ab-a478-478e890765eb","shared_citers":6},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":6},{"title":"Jamba: A Hybrid Transformer-Mamba Language Model","work_id":"129df0fe-8a66-4077-8991-3557cfa38274","shared_citers":6},{"title":"Linformer: Self-Attention with Linear Complexity","work_id":"4b717b51-6098-45d0-8e9e-b69bef651bc3","shared_citers":6},{"title":"Longformer: The Long-Document Transformer","work_id":"abea7a44-6668-4de7-aab6-f53a6e5aa088","shared_citers":6},{"title":"net/forum?id=HyUNwulC-","work_id":"137e6ed4-1c1d-40bd-801e-9d2baea0bd0c","shared_citers":6},{"title":"RWKV-7 “Goose” with expressive dynamic state evolution","work_id":"daf20308-c801-4745-a147-a1f1de4368e0","shared_citers":6},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":6},{"title":"arXiv preprint arXiv:2404.07904 , year=","work_id":"62f21939-fe4f-4159-bb1e-6a303fdc7b42","shared_citers":5},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":5},{"title":"Kimi Linear: An Expressive, Efficient Attention Architecture","work_id":"b2f9f1cd-c39c-4dbc-8637-f575681cdc01","shared_citers":5},{"title":"Native sparse attention: Hardware-aligned and natively trainable sparse attention.arXiv preprint arXiv:2502.11089","work_id":"c74d35ec-94d3-48f5-a291-20cfad7c04a3","shared_citers":5},{"title":"Rethinking Attention with Performers","work_id":"4c26d308-8b72-4a98-8e73-950617a75f50","shared_citers":5}],"time_series":[{"n":2,"year":2023},{"n":3,"year":2024},{"n":3,"year":2025},{"n":30,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"44dc4efb-f66b-47bd-97e7-48ab4f1143bc","orcid":null,"display_name":"Jilong Xue","source":"manual","import_confidence":0.72},{"id":"e5d982b3-1c02-4828-8399-4f667dadefdb","orcid":null,"display_name":"Li Dong","source":"manual","import_confidence":0.72},{"id":"019f6545-376f-4623-a727-dabd20ab9286","orcid":null,"display_name":"Shaohan Huang","source":"manual","import_confidence":0.72},{"id":"6ed1f2fe-df78-4a3e-bdff-eebdc69382ff","orcid":null,"display_name":"Shuming Ma","source":"manual","import_confidence":0.72},{"id":"1222d56d-5ab2-45b5-881d-9ce07af16216","orcid":null,"display_name":"Yuqing Xia","source":"manual","import_confidence":0.72},{"id":"32104b60-617c-49b7-bef0-900d62d14f88","orcid":null,"display_name":"Yutao Sun","source":"manual","import_confidence":0.72}]}}