{"work":{"id":"aef70eae-f816-4598-84ec-429a2c09f5fc","openalex_id":null,"doi":null,"arxiv_id":null,"raw_key":"raw:06688f4f61bbc4727e68ca32","title":"Scalable training of","authors":null,"authors_text":"Andrew, Galen and Gao, Jianfeng , booktitle=","year":null,"venue":null,"abstract":null,"external_url":null,"cited_by_count":null,"metadata_source":"raw_reference","metadata_fetched_at":"2026-07-11T03:47:52.206662+00:00","pith_arxiv_id":null,"created_at":"2026-05-10T13:29:58.824445+00:00","updated_at":"2026-07-11T03:47:52.206662+00:00","title_quality_ok":false,"display_title":"Scalable training of","render_title":"Scalable training of"},"hub":{"state":{"work_id":"aef70eae-f816-4598-84ec-429a2c09f5fc","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":160,"external_cited_by_count":null,"distinct_field_count":16,"first_pith_cited_at":"2020-04-10T17:54:09+00:00","last_pith_cited_at":"2026-07-09T12:18:40+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T07:39:25.615068+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":9},{"context_role":"other","n":2}],"polarity_counts":[{"context_polarity":"unclear","n":7},{"context_polarity":"background","n":4}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Scalable training of","claims":[{"claim_text":"vey of graph meets large language model: progress and future directions. InProceedings of the Thirty- Third International Joint Conference on Artificial Intelligence, pages 8123-8131. Andrés Montoyo, Patricio Martínez-Barco, and Alexan- dra Balahur. 2012. Subjectivity and sentiment analy- sis: An overview of the current state of the area and envisaged developments.Decision Support Systems, 53(4):675-679. Bo Pang, Lillian Lee, and Shivakumar Vaithyanathan. 2002. Thumbs up? sentiment classificatio","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"whether u is semantically broader than k. The two samples are expressed as X={x i}n i=1 andY={y j}m j=1,(3) where xi =x u,i and yj =x k,j, with n, m fixed (typically, we subsample to a common size to con- trol the variance across words). A natural null hypothesis is that the two words have the same dispersion but different mean directions. H0 :disp(X) =disp(Y) withE[X]̸=E[Y]allowed.(4) This is because the mean direction is a strong nui- sance factor in contextual embedding spaces. Even if two wo","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"domain. Given Xv ∈R Sv×Dv and Xt ∈R St×Dt, the goal is to refine Xv by aggregating contextual information across scales. We define N scales with two adapter sets: G= {G1, . . . ,GN } (MGFA) and C={C 1, . . . ,CN } (MCFA). At each scale n, features are reshaped to a grid X (0) v ∈R H×W×D v and downsampled by Down(·,2 n−1): X (n) v = Down(X(0) v ,2 n−1).(4) Let Xv,n = Seq(X (n) v ) denote the flattened se- quence. We then refine and fuse: Gn =G n(Xv,n), C n =C n(Xv,n, Xt),(5) ˜Xv,n =G n +w C n,(6)","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"Question: Eukaryotic genes tend to consist of coding regions (exons) and non-coding regions (introns). The figure shows how such a gene leads to the production of a protein. Which of the following statements is true? A. Thymine content of (1) and (2) is approximately equal. B. The process occurring between (2) and (3) takes place in the cytosol. C. (4) can hybridise with (2). D. The number of amino acid residues in (5) must equal the number of nucleotide residues in (2). E. All processes occurri","claim_type":"other","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"Question: Eukaryotic genes tend to consist of coding regions (exons) and non-coding regions (introns). The figure shows how such a gene leads to the production of a protein. Which of the following statements is true? A. Thymine content of (1) and (2) is approximately equal. B. The process occurring between (2) and (3) takes place in the cytosol. C. (4) can hybridise with (2). D. The number of amino acid residues in (5) must equal the number of nucleotide residues in (2). E. All processes occurri","claim_type":"background","confidence":0.6,"evidence_strength":"citation_context"},{"claim_text":"sharing & image reaction functions are integrated to add a multi-modal dimension to the long-term dialogues.2 The image sharing function is called when the agent decides to send an image. This process includes: (1) Generate a caption c for the intended image using M; (2) Convert the caption c into relevant keywords w using M; (3) Use the keywords k to find an image through web search W EB(k)3; (4) Share the chosen image. Con- versely, the image reaction function is triggered upon receiving an im","claim_type":"other","confidence":0.6,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Scalable training of because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (8 contexts).","role_counts":[{"n":8,"context_role":"background"},{"n":2,"context_role":"other"}]},"error":null,"updated_at":"2026-05-23T15:14:40.195675+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"ab8fad20-ae27-45c2-9ff8-8fb5453b02a8","orcid":null,"display_name":"Andrew"},{"id":"60545d58-e79c-4643-8f4c-5d6d196a6227","orcid":null,"display_name":"Galen and Gao"},{"id":"a9a58439-ec7a-4653-83b9-cc0a3687866e","orcid":null,"display_name":"Jianfeng"},{"id":"e7d13929-bb54-4948-915d-228400764dda","orcid":null,"display_name":"booktitle="}]},"error":null,"updated_at":"2026-05-23T15:14:43.425094+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-21T17:23:01.468475+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":null,"work_id":"852d89f5-1e7b-4296-b4f2-71e578b5e9f6","shared_citers":62},{"title":"A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =","work_id":"6d196829-7173-4c45-aa5c-d0ee30947345","shared_citers":61},{"title":"Chandra and Dexter C","work_id":"c3270592-bd69-4213-95e1-4aaf8312be9b","shared_citers":61},{"title":"Tetreault , title =","work_id":"75f57f43-cb6d-44a9-9cc5-9b8cfc702ea3","shared_citers":61},{"title":null,"work_id":"aca2b566-99e0-4ebb-9c7a-a81219531259","shared_citers":59},{"title":"Aho and Jeffrey D","work_id":"b1f5cb43-a3c7-4ea0-85e7-9ccc9dfe1588","shared_citers":58},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":15},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":10},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":8},{"title":"2024 , eprint=","work_id":"94860f33-c1e9-46de-b7ef-cdfde74468a5","shared_citers":7},{"title":"Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback","work_id":"a1f2574b-a899-4713-be60-c87ba332656c","shared_citers":7},{"title":"2025 , eprint=","work_id":"26c7b6ed-f86e-4ed8-b9ed-b1783d90255b","shared_citers":6},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":6},{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","work_id":"41fe12c4-e538-4890-a244-480650ed3078","shared_citers":6},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":5},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":5},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":5},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":5},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":5},{"title":"Advances in neural information processing systems , volume=","work_id":"a1fd09f1-b62b-4aca-a5ef-dd2b50ad08b5","shared_citers":4},{"title":"Advances in Neural Information Processing Systems , volume=","work_id":"c2cc414b-17c6-4006-b389-f3ec0bf8141b","shared_citers":4},{"title":"Advances in Neural Information Processing Systems , volume=","work_id":"be2b69de-45c4-4db5-ab23-0bff300c6059","shared_citers":4},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":4},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":4}],"time_series":[{"n":1,"year":2020},{"n":1,"year":2021},{"n":3,"year":2023},{"n":4,"year":2024},{"n":1,"year":2025},{"n":52,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-21T17:23:06.381798+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-21T17:23:06.318973+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Scalable training of","claims":[{"claim_text":"vey of graph meets large language model: progress and future directions. InProceedings of the Thirty- Third International Joint Conference on Artificial Intelligence, pages 8123-8131. Andrés Montoyo, Patricio Martínez-Barco, and Alexan- dra Balahur. 2012. Subjectivity and sentiment analy- sis: An overview of the current state of the area and envisaged developments.Decision Support Systems, 53(4):675-679. Bo Pang, Lillian Lee, and Shivakumar Vaithyanathan. 2002. Thumbs up? sentiment classificatio","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"whether u is semantically broader than k. The two samples are expressed as X={x i}n i=1 andY={y j}m j=1,(3) where xi =x u,i and yj =x k,j, with n, m fixed (typically, we subsample to a common size to con- trol the variance across words). A natural null hypothesis is that the two words have the same dispersion but different mean directions. H0 :disp(X) =disp(Y) withE[X]̸=E[Y]allowed.(4) This is because the mean direction is a strong nui- sance factor in contextual embedding spaces. Even if two wo","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"domain. Given Xv ∈R Sv×Dv and Xt ∈R St×Dt, the goal is to refine Xv by aggregating contextual information across scales. We define N scales with two adapter sets: G= {G1, . . . ,GN } (MGFA) and C={C 1, . . . ,CN } (MCFA). At each scale n, features are reshaped to a grid X (0) v ∈R H×W×D v and downsampled by Down(·,2 n−1): X (n) v = Down(X(0) v ,2 n−1).(4) Let Xv,n = Seq(X (n) v ) denote the flattened se- quence. We then refine and fuse: Gn =G n(Xv,n), C n =C n(Xv,n, Xt),(5) ˜Xv,n =G n +w C n,(6)","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"Question: Eukaryotic genes tend to consist of coding regions (exons) and non-coding regions (introns). The figure shows how such a gene leads to the production of a protein. Which of the following statements is true? A. Thymine content of (1) and (2) is approximately equal. B. The process occurring between (2) and (3) takes place in the cytosol. C. (4) can hybridise with (2). D. The number of amino acid residues in (5) must equal the number of nucleotide residues in (2). E. All processes occurri","claim_type":"other","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"Question: Eukaryotic genes tend to consist of coding regions (exons) and non-coding regions (introns). The figure shows how such a gene leads to the production of a protein. Which of the following statements is true? A. Thymine content of (1) and (2) is approximately equal. B. The process occurring between (2) and (3) takes place in the cytosol. C. (4) can hybridise with (2). D. The number of amino acid residues in (5) must equal the number of nucleotide residues in (2). E. All processes occurri","claim_type":"background","confidence":0.6,"evidence_strength":"citation_context"},{"claim_text":"sharing & image reaction functions are integrated to add a multi-modal dimension to the long-term dialogues.2 The image sharing function is called when the agent decides to send an image. This process includes: (1) Generate a caption c for the intended image using M; (2) Convert the caption c into relevant keywords w using M; (3) Use the keywords k to find an image through web search W EB(k)3; (4) Share the chosen image. Con- versely, the image reaction function is triggered upon receiving an im","claim_type":"other","confidence":0.6,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Scalable training of because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (8 contexts).","role_counts":[{"n":8,"context_role":"background"},{"n":2,"context_role":"other"}]},"error":null,"updated_at":"2026-05-23T15:14:40.190764+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Scalable training of","claims":[{"claim_text":"vey of graph meets large language model: progress and future directions. InProceedings of the Thirty- Third International Joint Conference on Artificial Intelligence, pages 8123-8131. Andrés Montoyo, Patricio Martínez-Barco, and Alexan- dra Balahur. 2012. Subjectivity and sentiment analy- sis: An overview of the current state of the area and envisaged developments.Decision Support Systems, 53(4):675-679. Bo Pang, Lillian Lee, and Shivakumar Vaithyanathan. 2002. Thumbs up? sentiment classificatio","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"whether u is semantically broader than k. The two samples are expressed as X={x i}n i=1 andY={y j}m j=1,(3) where xi =x u,i and yj =x k,j, with n, m fixed (typically, we subsample to a common size to con- trol the variance across words). A natural null hypothesis is that the two words have the same dispersion but different mean directions. H0 :disp(X) =disp(Y) withE[X]̸=E[Y]allowed.(4) This is because the mean direction is a strong nui- sance factor in contextual embedding spaces. Even if two wo","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"domain. Given Xv ∈R Sv×Dv and Xt ∈R St×Dt, the goal is to refine Xv by aggregating contextual information across scales. We define N scales with two adapter sets: G= {G1, . . . ,GN } (MGFA) and C={C 1, . . . ,CN } (MCFA). At each scale n, features are reshaped to a grid X (0) v ∈R H×W×D v and downsampled by Down(·,2 n−1): X (n) v = Down(X(0) v ,2 n−1).(4) Let Xv,n = Seq(X (n) v ) denote the flattened se- quence. We then refine and fuse: Gn =G n(Xv,n), C n =C n(Xv,n, Xt),(5) ˜Xv,n =G n +w C n,(6)","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"Question: Eukaryotic genes tend to consist of coding regions (exons) and non-coding regions (introns). The figure shows how such a gene leads to the production of a protein. Which of the following statements is true? A. Thymine content of (1) and (2) is approximately equal. B. The process occurring between (2) and (3) takes place in the cytosol. C. (4) can hybridise with (2). D. The number of amino acid residues in (5) must equal the number of nucleotide residues in (2). E. All processes occurri","claim_type":"other","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"Question: Eukaryotic genes tend to consist of coding regions (exons) and non-coding regions (introns). The figure shows how such a gene leads to the production of a protein. Which of the following statements is true? A. Thymine content of (1) and (2) is approximately equal. B. The process occurring between (2) and (3) takes place in the cytosol. C. (4) can hybridise with (2). D. The number of amino acid residues in (5) must equal the number of nucleotide residues in (2). E. All processes occurri","claim_type":"background","confidence":0.6,"evidence_strength":"citation_context"},{"claim_text":"sharing & image reaction functions are integrated to add a multi-modal dimension to the long-term dialogues.2 The image sharing function is called when the agent decides to send an image. This process includes: (1) Generate a caption c for the intended image using M; (2) Convert the caption c into relevant keywords w using M; (3) Use the keywords k to find an image through web search W EB(k)3; (4) Share the chosen image. Con- versely, the image reaction function is triggered upon receiving an im","claim_type":"other","confidence":0.6,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Scalable training of because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (8 contexts).","role_counts":[{"n":8,"context_role":"background"},{"n":2,"context_role":"other"}]},"error":null,"updated_at":"2026-05-21T17:23:01.473345+00:00"}},"summary":{"title":"Scalable training of","claims":[{"claim_text":"vey of graph meets large language model: progress and future directions. InProceedings of the Thirty- Third International Joint Conference on Artificial Intelligence, pages 8123-8131. Andrés Montoyo, Patricio Martínez-Barco, and Alexan- dra Balahur. 2012. Subjectivity and sentiment analy- sis: An overview of the current state of the area and envisaged developments.Decision Support Systems, 53(4):675-679. Bo Pang, Lillian Lee, and Shivakumar Vaithyanathan. 2002. Thumbs up? sentiment classificatio","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"whether u is semantically broader than k. The two samples are expressed as X={x i}n i=1 andY={y j}m j=1,(3) where xi =x u,i and yj =x k,j, with n, m fixed (typically, we subsample to a common size to con- trol the variance across words). A natural null hypothesis is that the two words have the same dispersion but different mean directions. H0 :disp(X) =disp(Y) withE[X]̸=E[Y]allowed.(4) This is because the mean direction is a strong nui- sance factor in contextual embedding spaces. Even if two wo","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"domain. Given Xv ∈R Sv×Dv and Xt ∈R St×Dt, the goal is to refine Xv by aggregating contextual information across scales. We define N scales with two adapter sets: G= {G1, . . . ,GN } (MGFA) and C={C 1, . . . ,CN } (MCFA). At each scale n, features are reshaped to a grid X (0) v ∈R H×W×D v and downsampled by Down(·,2 n−1): X (n) v = Down(X(0) v ,2 n−1).(4) Let Xv,n = Seq(X (n) v ) denote the flattened se- quence. We then refine and fuse: Gn =G n(Xv,n), C n =C n(Xv,n, Xt),(5) ˜Xv,n =G n +w C n,(6)","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"Question: Eukaryotic genes tend to consist of coding regions (exons) and non-coding regions (introns). The figure shows how such a gene leads to the production of a protein. Which of the following statements is true? A. Thymine content of (1) and (2) is approximately equal. B. The process occurring between (2) and (3) takes place in the cytosol. C. (4) can hybridise with (2). D. The number of amino acid residues in (5) must equal the number of nucleotide residues in (2). E. All processes occurri","claim_type":"other","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"Question: Eukaryotic genes tend to consist of coding regions (exons) and non-coding regions (introns). The figure shows how such a gene leads to the production of a protein. Which of the following statements is true? A. Thymine content of (1) and (2) is approximately equal. B. The process occurring between (2) and (3) takes place in the cytosol. C. (4) can hybridise with (2). D. The number of amino acid residues in (5) must equal the number of nucleotide residues in (2). E. All processes occurri","claim_type":"background","confidence":0.6,"evidence_strength":"citation_context"},{"claim_text":"sharing & image reaction functions are integrated to add a multi-modal dimension to the long-term dialogues.2 The image sharing function is called when the agent decides to send an image. This process includes: (1) Generate a caption c for the intended image using M; (2) Convert the caption c into relevant keywords w using M; (3) Use the keywords k to find an image through web search W EB(k)3; (4) Share the chosen image. Con- versely, the image reaction function is triggered upon receiving an im","claim_type":"other","confidence":0.6,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Scalable training of because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (8 contexts).","role_counts":[{"n":8,"context_role":"background"},{"n":2,"context_role":"other"}]},"graph":{"co_cited":[{"title":null,"work_id":"852d89f5-1e7b-4296-b4f2-71e578b5e9f6","shared_citers":62},{"title":"A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =","work_id":"6d196829-7173-4c45-aa5c-d0ee30947345","shared_citers":61},{"title":"Chandra and Dexter C","work_id":"c3270592-bd69-4213-95e1-4aaf8312be9b","shared_citers":61},{"title":"Tetreault , title =","work_id":"75f57f43-cb6d-44a9-9cc5-9b8cfc702ea3","shared_citers":61},{"title":null,"work_id":"aca2b566-99e0-4ebb-9c7a-a81219531259","shared_citers":59},{"title":"Aho and Jeffrey D","work_id":"b1f5cb43-a3c7-4ea0-85e7-9ccc9dfe1588","shared_citers":58},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":15},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":10},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":8},{"title":"2024 , eprint=","work_id":"94860f33-c1e9-46de-b7ef-cdfde74468a5","shared_citers":7},{"title":"Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback","work_id":"a1f2574b-a899-4713-be60-c87ba332656c","shared_citers":7},{"title":"2025 , eprint=","work_id":"26c7b6ed-f86e-4ed8-b9ed-b1783d90255b","shared_citers":6},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":6},{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","work_id":"41fe12c4-e538-4890-a244-480650ed3078","shared_citers":6},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":5},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":5},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":5},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":5},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":5},{"title":"Advances in neural information processing systems , volume=","work_id":"a1fd09f1-b62b-4aca-a5ef-dd2b50ad08b5","shared_citers":4},{"title":"Advances in Neural Information Processing Systems , volume=","work_id":"c2cc414b-17c6-4006-b389-f3ec0bf8141b","shared_citers":4},{"title":"Advances in Neural Information Processing Systems , volume=","work_id":"be2b69de-45c4-4db5-ab23-0bff300c6059","shared_citers":4},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":4},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":4}],"time_series":[{"n":1,"year":2020},{"n":1,"year":2021},{"n":3,"year":2023},{"n":4,"year":2024},{"n":1,"year":2025},{"n":52,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"ab8fad20-ae27-45c2-9ff8-8fb5453b02a8","orcid":null,"display_name":"Andrew","source":"manual","import_confidence":0.72},{"id":"e7d13929-bb54-4948-915d-228400764dda","orcid":null,"display_name":"booktitle=","source":"manual","import_confidence":0.72},{"id":"60545d58-e79c-4643-8f4c-5d6d196a6227","orcid":null,"display_name":"Galen and Gao","source":"manual","import_confidence":0.72},{"id":"a9a58439-ec7a-4653-83b9-cc0a3687866e","orcid":null,"display_name":"Jianfeng","source":"manual","import_confidence":0.72}]}}