{"work":{"id":"b6274271-7af9-4ee8-993b-ba1ba4205ba8","openalex_id":"https://openalex.org/W4405354744","doi":"10.48550/arxiv.2412.08905","arxiv_id":"2412.08905","raw_key":null,"title":"Phi-4 Technical Report","authors":null,"authors_text":"Marah Abdin, Jyoti Aneja, Harkirat Behl, S\\'ebastien Bubeck, Ronen Eldan, Suriya Gunasekar","year":2024,"venue":"cs.CL","abstract":"We present phi-4, a 14-billion parameter language model developed with a training recipe that is centrally focused on data quality. Unlike most language models, where pre-training is based primarily on organic data sources such as web content or code, phi-4 strategically incorporates synthetic data throughout the training process. While previous models in the Phi family largely distill the capabilities of a teacher model (specifically GPT-4), phi-4 substantially surpasses its teacher model on STEM-focused QA capabilities, giving evidence that our data-generation and post-training techniques go beyond distillation. Despite minimal changes to the phi-3 architecture, phi-4 achieves strong performance relative to its size -- especially on reasoning-focused benchmarks -- due to improved data, training curriculum, and innovations in the post-training scheme.","external_url":"https://arxiv.org/abs/2412.08905","cited_by_count":24,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2412.08905","created_at":"2026-05-09T05:50:28.059621+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":false,"display_title":"Phi-4 Technical Report","render_title":"Phi-4 Technical Report"},"hub":{"state":{"work_id":"b6274271-7af9-4ee8-993b-ba1ba4205ba8","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":182,"external_cited_by_count":24,"distinct_field_count":21,"first_pith_cited_at":"2025-01-16T17:37:58+00:00","last_pith_cited_at":"2026-07-07T22:36:30+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T17:09:23.281336+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":19},{"context_role":"baseline","n":9},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"background","n":19},{"context_polarity":"baseline","n":9},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Phi-4 Technical Report","claims":[{"claim_text":"We present phi-4, a 14-billion parameter language model developed with a training recipe that is centrally focused on data quality. Unlike most language models, where pre-training is based primarily on organic data sources such as web content or code, phi-4 strategically incorporates synthetic data throughout the training process. While previous models in the Phi family largely distill the capabilities of a teacher model (specifically GPT-4), phi-4 substantially surpasses its teacher model on STEM-focused QA capabilities, giving evidence that our data-generation and post-training techniques go","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Code Attribute Quality Issues Correctness LLMs Meet Library Evolution[125], Less Is More[146], SStuBs[50], Quality In, Quality Out[45], SwallowCode[28], Every Sample Matters[19], Synthetic Data Generation[86], Cracks in The Stack[48], RTL-Breaker[84], MG-Verilog[150], Code Generation Survey[51], DataRecipe[55], AiXcoder-7B[52], RustEvo2 [67], AATK Benchmark[95] Security Phi-4[2], Quality In, Quality Out[45], StarCoder 2 and The Stack v2[78], Cracks in The Stack[48], RTL-Breaker[84], AATK Benchma","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"scores 10T tokens on educational value (scale 1-10), (2) keep only score≥7(1.5T tokens remaining), (3) cluster by topic and balance distribution, (4) train with progres- sive curriculum. The quality filtering formula: Quality(x) =α·Clarity(x)+β·Reasoning(x)+γ·Factuality(x) (34) where GPT-4 estimates each component. Only top 15% of tokens used for training. Phi-4[1] (2024, 14B parameters) achieved 84.8% on MATH-competitive with GPT-4.What:The culmina- tion of the Phi paradigm, demonstrating that ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"models (MLLMs), which are divided according to their reasoning mechanisms into general instruction- following models and reasoning-enhanced models with long chain-of-thought (Long-CoT) capa- bilities. The former category includes the Qwen3-VL-Instruct series (2B, 8B, 30B-A3B, 32B) [5], Qwen2.5-VL-Instruct (7B) [6], Phi-3.5-Vision-Instruct (4B) [1], Phi-4-Multimodal-Instruct (6B) [2], InternVL3-Instruct (8B) [106], GPT-4o [35], GPT-4.1 [49], and Gemini-3-Flash-Preview [26]. The latter category in","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"RemoteAgent RemoteReasoner GeoChat Falcon Figure 4. Intent recognition performance across diverse EO tasks on our VagueEO. RemoteAgent eclipses all baselines. 5 Table 1. Comparison of scene classification results. Methods Publication AID [61] WHU-RS19 [3] Acc Acc InternVL3.5 [56] arXiv'25 73.80 91.50 Qwen2.5-VL [2] arXiv'25 63.07 76.60 Phi3.5-Vision [1] arXiv'24 56.57 68.90 GeoChat [20] CVPR'24 73.17 84.80 EarthDial [50] CVPR'25 87.5795.80 GeoMag [37] MM'25 83.03 77.62 VHM [40] AAAI'25 91.7095.8","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"formance evaluation in building the sports insights dataset T able 2.Comparative performance of open-source LLMs on 996 manually labeled articles for match relevance validation Model Precision (%) Recall (%) F1-score (%) Falcon 10B [2] 81.90 80.10 81.00 Qwen 2.5 14B [48] 84.90 91.50 88.07 Mistral Nemo-12.2B [22] 82.30 88.10 85.09 Llama 3.1 8B [15] 79.10 91.30 84.75 Phi-4 14B [1] 86.00 80.80 83.30 Llama 3.3 70B (4-bit quant.) [15] 81.90 95.10 88.02 Llama 3.3 70B [15] 83.10 93.50 87.95 Qwen 2.5 32","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"feedback to filter them, and Pareto-front selection to form balanced training targets. Experiments on AgentDojo, AgentHarm, and ATBench show consistent safety improvements across backbone families, model scales, and evolution rounds while preserving useful task behavior. We hope this work motivates using trajectory-level failures as structured supervision for safer self-evolving agents. 9 References [1] Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Mic","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Phi-4 Technical Report because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (16 contexts).","role_counts":[{"n":16,"context_role":"background"},{"n":8,"context_role":"baseline"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-05-20T07:21:52.469276+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"2df2d991-a259-4c92-9e30-f428e986de0c","orcid":null,"display_name":"Marah Abdin"},{"id":"af8698dc-c88d-447f-91a0-e07de89bdd20","orcid":null,"display_name":"Jyoti Aneja"},{"id":"954c9994-5c2c-46cc-96cd-33b1e58e7ad5","orcid":null,"display_name":"Harkirat Behl"},{"id":"1bfa2323-ac8e-448e-9c18-9ac61eb17d31","orcid":null,"display_name":"S\\'ebastien Bubeck"},{"id":"4819041d-4322-431a-aac2-8cc6d048411d","orcid":null,"display_name":"Ronen Eldan"},{"id":"2a2951ee-cf7e-41c5-8967-50a50c92f2b7","orcid":null,"display_name":"Suriya Gunasekar"}]},"error":null,"updated_at":"2026-05-20T07:21:53.227145+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T12:20:30.557815+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":29},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":26},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":12},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":12},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":10},{"title":"Gemma 3 Technical Report","work_id":"f93e08bf-9e96-409b-8ac6-b8385fd17fd7","shared_citers":9},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":9},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":8},{"title":"Gemma 2: Improving Open Language Models at a Practical Size","work_id":"4dd94e2f-2b27-4cbf-88a0-4910f0772a57","shared_citers":8},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":7},{"title":"gpt-oss-120b & gpt-oss-20b Model Card","work_id":"178c1f7e-4f19-4392-a45d-45a6dfa88ead","shared_citers":7},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":7},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":7},{"title":"Qwen Technical Report","work_id":"bb1fd52f-6b2f-437c-9516-37bdf6eb9be8","shared_citers":7},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":6},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":6},{"title":"Scaling Laws for Neural Language Models","work_id":"b7dd8749-9c45-4977-ab9b-64478dce1ae8","shared_citers":6},{"title":"Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge","work_id":"28ea1282-d657-4c61-a83c-f1249be6d6b1","shared_citers":6},{"title":"Constitutional AI: Harmlessness from AI Feedback","work_id":"faaaa4e0-2676-4fac-a0b4-99aef10d2095","shared_citers":5},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":5},{"title":"Mixtral of Experts","work_id":"0de8c352-9daa-4e1e-8c7b-3d0dec69f369","shared_citers":5},{"title":"Program Synthesis with Large Language Models","work_id":"fd241a05-03b9-4de2-9588-9d77ce176125","shared_citers":5},{"title":"Qwen2 Technical Report","work_id":"a1857881-ab9b-4b80-9b5f-9ae4b5c2566d","shared_citers":5},{"title":"Gemma: Open Models Based on Gemini Research and Technology","work_id":"a9ea2870-df28-40b8-a9e0-a7e9a116f793","shared_citers":4}],"time_series":[{"n":4,"year":2025},{"n":54,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T12:30:34.196697+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T12:20:38.373816+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Phi-4 Technical Report","claims":[{"claim_text":"We present phi-4, a 14-billion parameter language model developed with a training recipe that is centrally focused on data quality. Unlike most language models, where pre-training is based primarily on organic data sources such as web content or code, phi-4 strategically incorporates synthetic data throughout the training process. While previous models in the Phi family largely distill the capabilities of a teacher model (specifically GPT-4), phi-4 substantially surpasses its teacher model on STEM-focused QA capabilities, giving evidence that our data-generation and post-training techniques go","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Code Attribute Quality Issues Correctness LLMs Meet Library Evolution[125], Less Is More[146], SStuBs[50], Quality In, Quality Out[45], SwallowCode[28], Every Sample Matters[19], Synthetic Data Generation[86], Cracks in The Stack[48], RTL-Breaker[84], MG-Verilog[150], Code Generation Survey[51], DataRecipe[55], AiXcoder-7B[52], RustEvo2 [67], AATK Benchmark[95] Security Phi-4[2], Quality In, Quality Out[45], StarCoder 2 and The Stack v2[78], Cracks in The Stack[48], RTL-Breaker[84], AATK Benchma","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"scores 10T tokens on educational value (scale 1-10), (2) keep only score≥7(1.5T tokens remaining), (3) cluster by topic and balance distribution, (4) train with progres- sive curriculum. The quality filtering formula: Quality(x) =α·Clarity(x)+β·Reasoning(x)+γ·Factuality(x) (34) where GPT-4 estimates each component. Only top 15% of tokens used for training. Phi-4[1] (2024, 14B parameters) achieved 84.8% on MATH-competitive with GPT-4.What:The culmina- tion of the Phi paradigm, demonstrating that ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"models (MLLMs), which are divided according to their reasoning mechanisms into general instruction- following models and reasoning-enhanced models with long chain-of-thought (Long-CoT) capa- bilities. The former category includes the Qwen3-VL-Instruct series (2B, 8B, 30B-A3B, 32B) [5], Qwen2.5-VL-Instruct (7B) [6], Phi-3.5-Vision-Instruct (4B) [1], Phi-4-Multimodal-Instruct (6B) [2], InternVL3-Instruct (8B) [106], GPT-4o [35], GPT-4.1 [49], and Gemini-3-Flash-Preview [26]. The latter category in","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"RemoteAgent RemoteReasoner GeoChat Falcon Figure 4. Intent recognition performance across diverse EO tasks on our VagueEO. RemoteAgent eclipses all baselines. 5 Table 1. Comparison of scene classification results. Methods Publication AID [61] WHU-RS19 [3] Acc Acc InternVL3.5 [56] arXiv'25 73.80 91.50 Qwen2.5-VL [2] arXiv'25 63.07 76.60 Phi3.5-Vision [1] arXiv'24 56.57 68.90 GeoChat [20] CVPR'24 73.17 84.80 EarthDial [50] CVPR'25 87.5795.80 GeoMag [37] MM'25 83.03 77.62 VHM [40] AAAI'25 91.7095.8","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"formance evaluation in building the sports insights dataset T able 2.Comparative performance of open-source LLMs on 996 manually labeled articles for match relevance validation Model Precision (%) Recall (%) F1-score (%) Falcon 10B [2] 81.90 80.10 81.00 Qwen 2.5 14B [48] 84.90 91.50 88.07 Mistral Nemo-12.2B [22] 82.30 88.10 85.09 Llama 3.1 8B [15] 79.10 91.30 84.75 Phi-4 14B [1] 86.00 80.80 83.30 Llama 3.3 70B (4-bit quant.) [15] 81.90 95.10 88.02 Llama 3.3 70B [15] 83.10 93.50 87.95 Qwen 2.5 32","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"feedback to filter them, and Pareto-front selection to form balanced training targets. Experiments on AgentDojo, AgentHarm, and ATBench show consistent safety improvements across backbone families, model scales, and evolution rounds while preserving useful task behavior. We hope this work motivates using trajectory-level failures as structured supervision for safer self-evolving agents. 9 References [1] Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Mic","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Phi-4 Technical Report because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (16 contexts).","role_counts":[{"n":16,"context_role":"background"},{"n":8,"context_role":"baseline"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-05-20T07:21:52.473794+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Phi-4 Technical Report","claims":[{"claim_text":"We present phi-4, a 14-billion parameter language model developed with a training recipe that is centrally focused on data quality. Unlike most language models, where pre-training is based primarily on organic data sources such as web content or code, phi-4 strategically incorporates synthetic data throughout the training process. While previous models in the Phi family largely distill the capabilities of a teacher model (specifically GPT-4), phi-4 substantially surpasses its teacher model on STEM-focused QA capabilities, giving evidence that our data-generation and post-training techniques go","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Phi-4 Technical Report because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T12:30:38.049334+00:00"}},"summary":{"title":"Phi-4 Technical Report","claims":[{"claim_text":"We present phi-4, a 14-billion parameter language model developed with a training recipe that is centrally focused on data quality. Unlike most language models, where pre-training is based primarily on organic data sources such as web content or code, phi-4 strategically incorporates synthetic data throughout the training process. While previous models in the Phi family largely distill the capabilities of a teacher model (specifically GPT-4), phi-4 substantially surpasses its teacher model on STEM-focused QA capabilities, giving evidence that our data-generation and post-training techniques go","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Phi-4 Technical Report because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":29},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":26},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":12},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":12},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":10},{"title":"Gemma 3 Technical Report","work_id":"f93e08bf-9e96-409b-8ac6-b8385fd17fd7","shared_citers":9},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":9},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":8},{"title":"Gemma 2: Improving Open Language Models at a Practical Size","work_id":"4dd94e2f-2b27-4cbf-88a0-4910f0772a57","shared_citers":8},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":7},{"title":"gpt-oss-120b & gpt-oss-20b Model Card","work_id":"178c1f7e-4f19-4392-a45d-45a6dfa88ead","shared_citers":7},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":7},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":7},{"title":"Qwen Technical Report","work_id":"bb1fd52f-6b2f-437c-9516-37bdf6eb9be8","shared_citers":7},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":6},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":6},{"title":"Scaling Laws for Neural Language Models","work_id":"b7dd8749-9c45-4977-ab9b-64478dce1ae8","shared_citers":6},{"title":"Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge","work_id":"28ea1282-d657-4c61-a83c-f1249be6d6b1","shared_citers":6},{"title":"Constitutional AI: Harmlessness from AI Feedback","work_id":"faaaa4e0-2676-4fac-a0b4-99aef10d2095","shared_citers":5},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":5},{"title":"Mixtral of Experts","work_id":"0de8c352-9daa-4e1e-8c7b-3d0dec69f369","shared_citers":5},{"title":"Program Synthesis with Large Language Models","work_id":"fd241a05-03b9-4de2-9588-9d77ce176125","shared_citers":5},{"title":"Qwen2 Technical Report","work_id":"a1857881-ab9b-4b80-9b5f-9ae4b5c2566d","shared_citers":5},{"title":"Gemma: Open Models Based on Gemini Research and Technology","work_id":"a9ea2870-df28-40b8-a9e0-a7e9a116f793","shared_citers":4}],"time_series":[{"n":4,"year":2025},{"n":54,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"954c9994-5c2c-46cc-96cd-33b1e58e7ad5","orcid":null,"display_name":"Harkirat Behl","source":"manual","import_confidence":0.72},{"id":"af8698dc-c88d-447f-91a0-e07de89bdd20","orcid":null,"display_name":"Jyoti Aneja","source":"manual","import_confidence":0.72},{"id":"2df2d991-a259-4c92-9e30-f428e986de0c","orcid":null,"display_name":"Marah Abdin","source":"manual","import_confidence":0.72},{"id":"4819041d-4322-431a-aac2-8cc6d048411d","orcid":null,"display_name":"Ronen Eldan","source":"manual","import_confidence":0.72},{"id":"1bfa2323-ac8e-448e-9c18-9ac61eb17d31","orcid":null,"display_name":"S\\'ebastien Bubeck","source":"manual","import_confidence":0.72},{"id":"2a2951ee-cf7e-41c5-8967-50a50c92f2b7","orcid":null,"display_name":"Suriya Gunasekar","source":"manual","import_confidence":0.72}]}}