{"work":{"id":"9e2a976b-f5ad-4aee-af5c-243fe0fe75d2","openalex_id":"https://openalex.org/W4388926373","doi":"10.48550/arxiv.2311.12022","arxiv_id":"2311.12022","raw_key":null,"title":"GPQA: A Graduate-Level Google-Proof Q&A Benchmark","authors":null,"authors_text":"David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani","year":2023,"venue":"cs.AI","abstract":"We present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. We ensure that the questions are high-quality and extremely difficult: experts who have or are pursuing PhDs in the corresponding domains reach 65% accuracy (74% when discounting clear mistakes the experts identified in retrospect), while highly skilled non-expert validators only reach 34% accuracy, despite spending on average over 30 minutes with unrestricted access to the web (i.e., the questions are \"Google-proof\"). The questions are also difficult for state-of-the-art AI systems, with our strongest GPT-4 based baseline achieving 39% accuracy. If we are to use future AI systems to help us answer very hard questions, for example, when developing new scientific knowledge, we need to develop scalable oversight methods that enable humans to supervise their outputs, which may be difficult even if the supervisors are themselves skilled and knowledgeable. The difficulty of GPQA both for skilled non-experts and frontier AI systems should enable realistic scalable oversight experiments, which we hope can help devise ways for human experts to reliably get truthful information from AI systems that surpass human capabilities.","external_url":"https://arxiv.org/abs/2311.12022","cited_by_count":23,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2311.12022","created_at":"2026-05-09T05:45:21.558103+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"GPQA: A Graduate-Level Google-Proof Q&A Benchmark","render_title":"GPQA: A Graduate-Level Google-Proof Q&A Benchmark"},"hub":{"state":{"work_id":"9e2a976b-f5ad-4aee-af5c-243fe0fe75d2","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":236,"external_cited_by_count":23,"distinct_field_count":18,"first_pith_cited_at":"2024-06-11T17:32:21+00:00","last_pith_cited_at":"2026-07-09T09:56:50+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T23:09:20.973985+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"dataset","n":18},{"context_role":"background","n":14},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"use_dataset","n":18},{"context_polarity":"background","n":10},{"context_polarity":"unclear","n":4},{"context_polarity":"use_method","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"GPQA: A Graduate-Level Google-Proof Q&A Benchmark","claims":[{"claim_text":"We present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. We ensure that the questions are high-quality and extremely difficult: experts who have or are pursuing PhDs in the corresponding domains reach 65% accuracy (74% when discounting clear mistakes the experts identified in retrospect), while highly skilled non-expert validators only reach 34% accuracy, despite spending on average over 30 minutes with unrestricted access to the web (i.e., the questions are \"Google-proof\"). The questions are also difficult for state-","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"These benchmarks cover different instruction- following scenarios, including constraint satisfaction, multi-turn dialogue, writing-oriented tasks, agentic tool-use settings, and multilingual instruction following. To assess the general capabil- ities of the trained models, we also evaluate on several general-purpose benchmarks, including GPQA-Diamond [27], MMLU-Pro [35], BBEH [16], and AIME. Details are provided in Appendix E. 3.2 Main Results Self-Evolving Enhances Instruction-Following.Table 1","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"example, theVerifiershould blindly score both teacher and student rollouts together to calibrate bias arise from varying question difficulties. These findings suggest that rubric-based OPD is not merely a heuristic replacement for logit-based OPD, but a principled and robust distillation framework. We extensively validate ROPD across diverse benchmarks (e.g., AIME24/25 [1, 2], HMMT25 [3], GPQA-Diamond [11], HealthBench [12], and IFEval [13]) and model configurations (e.g., Qwen3- 4B [5] and Gemm","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"rollouts during policy optimization, learning rate1e −6, and maximum response length 8192. Evaluation Setting.Evaluation is conducted on a diverse set of benchmarks spanning math reasoning and broader reasoning domains. For in-domain evaluation, we use AIME2024 [ 42], AIME2025, MATH500 [6], Minerva Math [ 17], and Olympiad [ 11]. For out-of-domain generalization, we evaluate on ARC-Challenge [5], GPQA [25], and MMLU-Pro [34], covering multi-domain reasoning. For smaller test sets(AIME 24/25), we","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"↑Tokens↓Acc.↑Tokens↓ Ours 65.2 483.4K 57.6 592.5K 36.5 650.7K 53.1 575.5K 39.9 Ablation on beta-controlled search space w/o Beta Parameterization 60.781.2K54.493.9K31.9104.7K49.093.3K46.4 Ablation on history design w/o Execution Traces 64.0 703.1K 56.7 823.7K 34.1 946.1K 51.6 824.3K30.9 from DeepSeek-R1 [22] on HMMT25, and a non-math benchmark (GPQA-Diamond [ 23]) with Qwen3-1.7B. On DeepSeek-R1-Distill-Llama-8B with HMMT25, our discovered controller with β= 1 achieves the best accuracy among al","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"the model to solve real-life situational problems using mathematical concepts. We use chain-of-thought prompting [46] for this benchmark. • MATH: 12,500 challenging competition-level mathematics problems (5,000 in the test set). We use chain-of-thought prompting [46] for this benchmark. 7 • BBH [39]: A suite of 23 challenging BIG-Bench [ 37] tasks. We use chain-of-thought prompting [46] for this benchmark. • GPQA [32]: A graduate-level multi-choice benchmark in biology, chemistry, and physics. •","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"Our evaluation covers two categories of benchmarks. The first category is in-domain search-agent benchmarks, including four representative datasets,GAIA[ 39],WebWalkerQA[ 63], XBench[ 5], andBrowseComp-ZH (BC-ZH)[ 83], which span diverse difficulty levels, multiple languages, and real-world multi-step reasoning scenarios. The second category is out-of-domain benchmarks, includingGPQA[ 47],TruthfulQA[ 34], andIFEval[ 82], which are used to evaluate the out-of-domain generalization ability of mode","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks GPQA: A Graduate-Level Google-Proof Q&A Benchmark because it crossed a citation-hub threshold. Current citing contexts most often use it as dataset evidence (16 contexts).","role_counts":[{"n":16,"context_role":"dataset"},{"n":11,"context_role":"background"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-05-19T01:11:26.298382+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"ae3eb554-8661-4827-92bd-c9329fff916a","orcid":null,"display_name":"David Rein"},{"id":"3ec6a7cf-88d5-400f-a72f-472722d0c607","orcid":null,"display_name":"Betty Li Hou"},{"id":"ef098013-28a2-47fe-a46c-8255675e4a7c","orcid":null,"display_name":"Asa Cooper Stickland"},{"id":"d4880cc7-e458-4172-a53b-58a97c5a2ab0","orcid":null,"display_name":"Jackson Petty"},{"id":"6a7a9cd6-3845-4ebc-b188-257bafaac758","orcid":null,"display_name":"Richard Yuanzhe Pang"},{"id":"661f5328-fa83-46b3-ad95-c7bb6a531d02","orcid":null,"display_name":"Julien Dirani"}]},"error":null,"updated_at":"2026-05-19T01:11:27.343912+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T07:57:47.131130+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":27},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":23},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":22},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":21},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":20},{"title":"Instruction-Following Evaluation for Large Language Models","work_id":"3aa06177-125a-4f5a-8f4a-8070c5986c26","shared_citers":18},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":17},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":16},{"title":"Measuring Mathematical Problem Solving With the MATH Dataset","work_id":"50652ac6-fb7c-4675-a2c2-159c241feb17","shared_citers":14},{"title":"MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark","work_id":"3c028052-035a-4c22-b80e-3046edb44adc","shared_citers":14},{"title":"LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code","work_id":"ea9e51ce-1e75-4182-92d8-4d25f70d2ee4","shared_citers":13},{"title":"Measuring Massive Multitask Language Understanding","work_id":"e87ec49a-544b-4ec8-8991-75298c64ff5e","shared_citers":13},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":11},{"title":"Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters","work_id":"a8d50b24-bdf5-46ed-bc4f-2927dfd81f1d","shared_citers":11},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":10},{"title":"Let's Verify Step by Step","work_id":"6d05b790-04c5-4fd2-91b2-ba1dfdd5770f","shared_citers":9},{"title":"Qwen2 Technical Report","work_id":"a1857881-ab9b-4b80-9b5f-9ae4b5c2566d","shared_citers":9},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":8},{"title":"Humanity's Last Exam","work_id":"59ea00d4-16a8-45e1-aafc-290a6f91d9f4","shared_citers":8},{"title":"OpenAI o1 System Card","work_id":"68d3c334-0fc9-49e3-b7b0-a69afae933e2","shared_citers":8},{"title":"Program Synthesis with Large Language Models","work_id":"fd241a05-03b9-4de2-9588-9d77ce176125","shared_citers":8},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":8},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":7},{"title":"Group Sequence Policy Optimization","work_id":"3a98b53b-9f52-4d95-adf7-89353c0a9a65","shared_citers":7}],"time_series":[{"n":3,"year":2024},{"n":7,"year":2025},{"n":64,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T08:07:41.232518+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T07:57:51.494024+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"GPQA: A Graduate-Level Google-Proof Q&A Benchmark","claims":[{"claim_text":"We present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. We ensure that the questions are high-quality and extremely difficult: experts who have or are pursuing PhDs in the corresponding domains reach 65% accuracy (74% when discounting clear mistakes the experts identified in retrospect), while highly skilled non-expert validators only reach 34% accuracy, despite spending on average over 30 minutes with unrestricted access to the web (i.e., the questions are \"Google-proof\"). The questions are also difficult for state-","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"These benchmarks cover different instruction- following scenarios, including constraint satisfaction, multi-turn dialogue, writing-oriented tasks, agentic tool-use settings, and multilingual instruction following. To assess the general capabil- ities of the trained models, we also evaluate on several general-purpose benchmarks, including GPQA-Diamond [27], MMLU-Pro [35], BBEH [16], and AIME. Details are provided in Appendix E. 3.2 Main Results Self-Evolving Enhances Instruction-Following.Table 1","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"example, theVerifiershould blindly score both teacher and student rollouts together to calibrate bias arise from varying question difficulties. These findings suggest that rubric-based OPD is not merely a heuristic replacement for logit-based OPD, but a principled and robust distillation framework. We extensively validate ROPD across diverse benchmarks (e.g., AIME24/25 [1, 2], HMMT25 [3], GPQA-Diamond [11], HealthBench [12], and IFEval [13]) and model configurations (e.g., Qwen3- 4B [5] and Gemm","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"rollouts during policy optimization, learning rate1e −6, and maximum response length 8192. Evaluation Setting.Evaluation is conducted on a diverse set of benchmarks spanning math reasoning and broader reasoning domains. For in-domain evaluation, we use AIME2024 [ 42], AIME2025, MATH500 [6], Minerva Math [ 17], and Olympiad [ 11]. For out-of-domain generalization, we evaluate on ARC-Challenge [5], GPQA [25], and MMLU-Pro [34], covering multi-domain reasoning. For smaller test sets(AIME 24/25), we","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"↑Tokens↓Acc.↑Tokens↓ Ours 65.2 483.4K 57.6 592.5K 36.5 650.7K 53.1 575.5K 39.9 Ablation on beta-controlled search space w/o Beta Parameterization 60.781.2K54.493.9K31.9104.7K49.093.3K46.4 Ablation on history design w/o Execution Traces 64.0 703.1K 56.7 823.7K 34.1 946.1K 51.6 824.3K30.9 from DeepSeek-R1 [22] on HMMT25, and a non-math benchmark (GPQA-Diamond [ 23]) with Qwen3-1.7B. On DeepSeek-R1-Distill-Llama-8B with HMMT25, our discovered controller with β= 1 achieves the best accuracy among al","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"the model to solve real-life situational problems using mathematical concepts. We use chain-of-thought prompting [46] for this benchmark. • MATH: 12,500 challenging competition-level mathematics problems (5,000 in the test set). We use chain-of-thought prompting [46] for this benchmark. 7 • BBH [39]: A suite of 23 challenging BIG-Bench [ 37] tasks. We use chain-of-thought prompting [46] for this benchmark. • GPQA [32]: A graduate-level multi-choice benchmark in biology, chemistry, and physics. •","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"Our evaluation covers two categories of benchmarks. The first category is in-domain search-agent benchmarks, including four representative datasets,GAIA[ 39],WebWalkerQA[ 63], XBench[ 5], andBrowseComp-ZH (BC-ZH)[ 83], which span diverse difficulty levels, multiple languages, and real-world multi-step reasoning scenarios. The second category is out-of-domain benchmarks, includingGPQA[ 47],TruthfulQA[ 34], andIFEval[ 82], which are used to evaluate the out-of-domain generalization ability of mode","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks GPQA: A Graduate-Level Google-Proof Q&A Benchmark because it crossed a citation-hub threshold. Current citing contexts most often use it as dataset evidence (16 contexts).","role_counts":[{"n":16,"context_role":"dataset"},{"n":11,"context_role":"background"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-05-19T01:11:27.348048+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"GPQA: A Graduate-Level Google-Proof Q&A Benchmark","claims":[{"claim_text":"We present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. We ensure that the questions are high-quality and extremely difficult: experts who have or are pursuing PhDs in the corresponding domains reach 65% accuracy (74% when discounting clear mistakes the experts identified in retrospect), while highly skilled non-expert validators only reach 34% accuracy, despite spending on average over 30 minutes with unrestricted access to the web (i.e., the questions are \"Google-proof\"). The questions are also difficult for state-","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks GPQA: A Graduate-Level Google-Proof Q&A Benchmark because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T08:07:54.639293+00:00"}},"summary":{"title":"GPQA: A Graduate-Level Google-Proof Q&A Benchmark","claims":[{"claim_text":"We present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. We ensure that the questions are high-quality and extremely difficult: experts who have or are pursuing PhDs in the corresponding domains reach 65% accuracy (74% when discounting clear mistakes the experts identified in retrospect), while highly skilled non-expert validators only reach 34% accuracy, despite spending on average over 30 minutes with unrestricted access to the web (i.e., the questions are \"Google-proof\"). The questions are also difficult for state-","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks GPQA: A Graduate-Level Google-Proof Q&A Benchmark because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":27},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":23},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":22},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":21},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":20},{"title":"Instruction-Following Evaluation for Large Language Models","work_id":"3aa06177-125a-4f5a-8f4a-8070c5986c26","shared_citers":18},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":17},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":16},{"title":"Measuring Mathematical Problem Solving With the MATH Dataset","work_id":"50652ac6-fb7c-4675-a2c2-159c241feb17","shared_citers":14},{"title":"MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark","work_id":"3c028052-035a-4c22-b80e-3046edb44adc","shared_citers":14},{"title":"LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code","work_id":"ea9e51ce-1e75-4182-92d8-4d25f70d2ee4","shared_citers":13},{"title":"Measuring Massive Multitask Language Understanding","work_id":"e87ec49a-544b-4ec8-8991-75298c64ff5e","shared_citers":13},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":11},{"title":"Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters","work_id":"a8d50b24-bdf5-46ed-bc4f-2927dfd81f1d","shared_citers":11},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":10},{"title":"Let's Verify Step by Step","work_id":"6d05b790-04c5-4fd2-91b2-ba1dfdd5770f","shared_citers":9},{"title":"Qwen2 Technical Report","work_id":"a1857881-ab9b-4b80-9b5f-9ae4b5c2566d","shared_citers":9},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":8},{"title":"Humanity's Last Exam","work_id":"59ea00d4-16a8-45e1-aafc-290a6f91d9f4","shared_citers":8},{"title":"OpenAI o1 System Card","work_id":"68d3c334-0fc9-49e3-b7b0-a69afae933e2","shared_citers":8},{"title":"Program Synthesis with Large Language Models","work_id":"fd241a05-03b9-4de2-9588-9d77ce176125","shared_citers":8},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":8},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":7},{"title":"Group Sequence Policy Optimization","work_id":"3a98b53b-9f52-4d95-adf7-89353c0a9a65","shared_citers":7}],"time_series":[{"n":3,"year":2024},{"n":7,"year":2025},{"n":64,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"ef098013-28a2-47fe-a46c-8255675e4a7c","orcid":null,"display_name":"Asa Cooper Stickland","source":"manual","import_confidence":0.72},{"id":"3ec6a7cf-88d5-400f-a72f-472722d0c607","orcid":null,"display_name":"Betty Li Hou","source":"manual","import_confidence":0.72},{"id":"ae3eb554-8661-4827-92bd-c9329fff916a","orcid":null,"display_name":"David Rein","source":"manual","import_confidence":0.72},{"id":"d4880cc7-e458-4172-a53b-58a97c5a2ab0","orcid":null,"display_name":"Jackson Petty","source":"manual","import_confidence":0.72},{"id":"661f5328-fa83-46b3-ad95-c7bb6a531d02","orcid":null,"display_name":"Julien Dirani","source":"manual","import_confidence":0.72},{"id":"6a7a9cd6-3845-4ebc-b188-257bafaac758","orcid":null,"display_name":"Richard Yuanzhe Pang","source":"manual","import_confidence":0.72}]}}