{"work":{"id":"42c46ece-c4e8-4d5e-abae-f8d5b4208995","openalex_id":"https://openalex.org/W4391272787","doi":"10.48550/arxiv.2401.14159","arxiv_id":"2401.14159","raw_key":null,"title":"Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks","authors":null,"authors_text":"Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao","year":2024,"venue":"cs.CV","abstract":"We introduce Grounded SAM, which uses Grounding DINO as an open-set object detector to combine with the segment anything model (SAM). This integration enables the detection and segmentation of any regions based on arbitrary text inputs and opens a door to connecting various vision models. As shown in Fig.1, a wide range of vision tasks can be achieved by using the versatile Grounded SAM pipeline. For example, an automatic annotation pipeline based solely on input images can be realized by incorporating models such as BLIP and Recognize Anything. Additionally, incorporating Stable-Diffusion allows for controllable image editing, while the integration of OSX facilitates promptable 3D human motion analysis. Grounded SAM also shows superior performance on open-vocabulary benchmarks, achieving 48.7 mean AP on SegInW (Segmentation in the wild) zero-shot benchmark with the combination of Grounding DINO-Base and SAM-Huge models.","external_url":"https://arxiv.org/abs/2401.14159","cited_by_count":92,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2401.14159","created_at":"2026-05-09T06:00:35.790483+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks","render_title":"Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks"},"hub":{"state":{"work_id":"42c46ece-c4e8-4d5e-abae-f8d5b4208995","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":133,"external_cited_by_count":92,"distinct_field_count":5,"first_pith_cited_at":"2025-02-04T16:19:20+00:00","last_pith_cited_at":"2026-07-09T12:21:06+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T14:29:25.717421+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":9},{"context_role":"method","n":4},{"context_role":"baseline","n":2},{"context_role":"other","n":2}],"polarity_counts":[{"context_polarity":"background","n":9},{"context_polarity":"use_method","n":4},{"context_polarity":"baseline","n":2},{"context_polarity":"unclear","n":2}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks","claims":[{"claim_text":"We introduce Grounded SAM, which uses Grounding DINO as an open-set object detector to combine with the segment anything model (SAM). This integration enables the detection and segmentation of any regions based on arbitrary text inputs and opens a door to connecting various vision models. As shown in Fig.1, a wide range of vision tasks can be achieved by using the versatile Grounded SAM pipeline. For example, an automatic annotation pipeline based solely on input images can be realized by incorporating models such as BLIP and Recognize Anything. Additionally, incorporating Stable-Diffusion all","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Ren, Justin Lidard, Lars Lien Ankile, Anthony Simeonov, Pulkit Agrawal, Anirudha Majumdar, Benjamin Burchfiel, Hongkai Dai, and Max Simchowitz. Diffusion policy policy optimization. InICLR, 2025. 22 [65] Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, et al. Grounded sam: Assembling open-world models for diverse visual tasks.arXiv preprint arXiv:2401.14159, 2024. 22 [66] Xuanchi Ren, Yifan Lu, Tianshi Cao, Ruiyuan Gao, Shengyu ","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"cording to someunknowntime-invariant process(s t+1)∼f(s t, at). We assume the state can be decomposed into a pose elementpt, defined relative to an in- ertial frame, and a body-velocity elementvt, so thats t := (p t, vt). We further assume the robot is equipped with a depth sensor and an algorithm providing 4 L. Marques, and D. Berenson semantically segmented images (e.g., [23,25]) where each pixel has an associated class. We consider as observationot := (odepth t , osemantics t ), whereo t =h(s","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"stereo image pair(l i, ri)at framei, we first estimate the depth mapd i using FoundationStereo [36]. The depth map is then back-projected to reconstruct an RGB point-cloud in the camera framep cam i . In subsequent steps, this point cloud will be transformed into a unified world frame. To remove human-specific visual artifacts, we segment the operator's hands using Grounded-SAM [37] to obtain the hand maskm hand i , and remove the corresponding points from the 4 reconstructed point cloud. We als","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"•Each part_idx from 0 to N-1 (where N is total parts) must appear exactly once •The part_idx=x description must correspond STRICTLY to the x-th normalized image shown in the input sequence Keep captions concise but descriptive. Focus on visual appearance, object identification, and assembly functionality. Here is an example of the expected format: 1{ 2\" c a p t i o n s \" : [ 3{ 4\" p a r t _ i d x \" : 0 , 5\" c a p t i o n \": \" A round side t a b l e with a dark wood top and a black metal splayed ","claim_type":"other","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"suggests that even with complete candidates,instance-level alignmentunder di- verse cues and multi-targets remain challenging, and it motivates improving the second-stage visual prompting and reasoning strategy. 5.3 Analysis Chart Instance segmentation resultsIn this section, we compare our fine- tunedMask2Formerwith two off-the-shelf instance segmentation baselines, GroundedSAM2[14] andSAM3[2], onChartREG++images. To run these baselines, for each target category we manually design multiple shor","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"lenging multi-arm manipulation benchmarks demonstrate CoEnv's ef- fectiveness in achieving high task success rates and execution efficiency, establishing a new paradigm for multi-agent embodied AI. Keywords:Embodied AI·Multi-Agent Systems·Robotic Manipula- tion·Vision-Language Models 1 Introduction The rapid evolution of foundation models, particularly multimodal large lan- guage models [46,47] and vision-language-action architectures [7,21,59], has unlocked unprecedented capabilities in embodie","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (9 contexts).","role_counts":[{"n":9,"context_role":"background"},{"n":4,"context_role":"method"},{"n":2,"context_role":"baseline"},{"n":2,"context_role":"other"}]},"error":null,"updated_at":"2026-07-01T22:41:53.170251+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"0ab620d7-a875-49b7-90e6-f022a0e27ad4","orcid":null,"display_name":"Tianhe Ren"},{"id":"58750f5c-a26b-48e5-919d-1ea442a1077d","orcid":null,"display_name":"Shilong Liu"},{"id":"09a63137-256f-4d79-b884-90df342dead2","orcid":null,"display_name":"Ailing Zeng"},{"id":"a336ca11-66c7-4b4e-9cab-39dbe45c82ae","orcid":null,"display_name":"Jing Lin"},{"id":"7d9fea06-97b4-4683-b934-cacd7dc6aa35","orcid":null,"display_name":"Kunchang Li"},{"id":"20334f7b-fb10-428f-84fe-0fd4a3028dea","orcid":null,"display_name":"He Cao"}]},"error":null,"updated_at":"2026-07-01T22:41:53.165163+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T15:32:01.357946+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"SAM 2: Segment Anything in Images and Videos","work_id":"acc13f66-d814-44f9-9688-375688bf2d4a","shared_citers":8},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":7},{"title":"SAM 3: Segment Anything with Concepts","work_id":"4a72a006-2592-4554-aad0-a9c41a9f952d","shared_citers":7},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":5},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":5},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":4},{"title":"DINOv3","work_id":"c8b07deb-8fe7-4e18-9620-f3569d3529ce","shared_citers":4},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":4},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":4},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":3},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":3},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":3},{"title":"Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection","work_id":"3757dc8f-79d5-4beb-a03b-eb4c9a33427d","shared_citers":3},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":3},{"title":"Matter- port3d: Learning from rgb-d data in indoor environments","work_id":"a6675134-1bd7-4d3f-9344-d7072e7449e9","shared_citers":3},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","work_id":"3e7e65c5-5aed-4fe9-8414-2092bcb31cc7","shared_citers":3},{"title":"RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots","work_id":"11232b35-bd17-402a-9234-951c46015815","shared_citers":3},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":2},{"title":"An open and comprehensive pipeline for unified object grounding and de- tection","work_id":"e2d251e7-a46b-4b55-9ddb-9364d020f415","shared_citers":2},{"title":"Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting","work_id":"f2f765af-f0ba-41f1-8dd1-37bf1d134d31","shared_citers":2},{"title":"arXiv preprint arXiv:2507.23134 (2025) 9, 12, 13","work_id":"90e42fa9-770a-442d-a6e6-4df3d73fc3da","shared_citers":2},{"title":"Depth pro: Sharp monocular metric depth in less than a second","work_id":"0b67883b-1901-45f1-9d58-1ef7a928df23","shared_citers":2},{"title":"Distilling the Knowledge in a Neural Network","work_id":"d927ab1f-17b8-4002-9d09-c3d55764fbad","shared_citers":2},{"title":"Egovla: Learning vision-language-action models from egocentric human videos","work_id":"14ae45e8-e465-4230-99bf-0a2a191dabfd","shared_citers":2}],"time_series":[{"n":1,"year":2025},{"n":43,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T15:42:10.146261+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T15:32:05.538452+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks","claims":[{"claim_text":"We introduce Grounded SAM, which uses Grounding DINO as an open-set object detector to combine with the segment anything model (SAM). This integration enables the detection and segmentation of any regions based on arbitrary text inputs and opens a door to connecting various vision models. As shown in Fig.1, a wide range of vision tasks can be achieved by using the versatile Grounded SAM pipeline. For example, an automatic annotation pipeline based solely on input images can be realized by incorporating models such as BLIP and Recognize Anything. Additionally, incorporating Stable-Diffusion all","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Ren, Justin Lidard, Lars Lien Ankile, Anthony Simeonov, Pulkit Agrawal, Anirudha Majumdar, Benjamin Burchfiel, Hongkai Dai, and Max Simchowitz. Diffusion policy policy optimization. InICLR, 2025. 22 [65] Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, et al. Grounded sam: Assembling open-world models for diverse visual tasks.arXiv preprint arXiv:2401.14159, 2024. 22 [66] Xuanchi Ren, Yifan Lu, Tianshi Cao, Ruiyuan Gao, Shengyu ","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"cording to someunknowntime-invariant process(s t+1)∼f(s t, at). We assume the state can be decomposed into a pose elementpt, defined relative to an in- ertial frame, and a body-velocity elementvt, so thats t := (p t, vt). We further assume the robot is equipped with a depth sensor and an algorithm providing 4 L. Marques, and D. Berenson semantically segmented images (e.g., [23,25]) where each pixel has an associated class. We consider as observationot := (odepth t , osemantics t ), whereo t =h(s","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"stereo image pair(l i, ri)at framei, we first estimate the depth mapd i using FoundationStereo [36]. The depth map is then back-projected to reconstruct an RGB point-cloud in the camera framep cam i . In subsequent steps, this point cloud will be transformed into a unified world frame. To remove human-specific visual artifacts, we segment the operator's hands using Grounded-SAM [37] to obtain the hand maskm hand i , and remove the corresponding points from the 4 reconstructed point cloud. We als","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"•Each part_idx from 0 to N-1 (where N is total parts) must appear exactly once •The part_idx=x description must correspond STRICTLY to the x-th normalized image shown in the input sequence Keep captions concise but descriptive. Focus on visual appearance, object identification, and assembly functionality. Here is an example of the expected format: 1{ 2\" c a p t i o n s \" : [ 3{ 4\" p a r t _ i d x \" : 0 , 5\" c a p t i o n \": \" A round side t a b l e with a dark wood top and a black metal splayed ","claim_type":"other","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"suggests that even with complete candidates,instance-level alignmentunder di- verse cues and multi-targets remain challenging, and it motivates improving the second-stage visual prompting and reasoning strategy. 5.3 Analysis Chart Instance segmentation resultsIn this section, we compare our fine- tunedMask2Formerwith two off-the-shelf instance segmentation baselines, GroundedSAM2[14] andSAM3[2], onChartREG++images. To run these baselines, for each target category we manually design multiple shor","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"lenging multi-arm manipulation benchmarks demonstrate CoEnv's ef- fectiveness in achieving high task success rates and execution efficiency, establishing a new paradigm for multi-agent embodied AI. Keywords:Embodied AI·Multi-Agent Systems·Robotic Manipula- tion·Vision-Language Models 1 Introduction The rapid evolution of foundation models, particularly multimodal large lan- guage models [46,47] and vision-language-action architectures [7,21,59], has unlocked unprecedented capabilities in embodie","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (9 contexts).","role_counts":[{"n":9,"context_role":"background"},{"n":4,"context_role":"method"},{"n":2,"context_role":"baseline"},{"n":2,"context_role":"other"}]},"error":null,"updated_at":"2026-07-01T22:41:53.173827+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks","claims":[{"claim_text":"We introduce Grounded SAM, which uses Grounding DINO as an open-set object detector to combine with the segment anything model (SAM). This integration enables the detection and segmentation of any regions based on arbitrary text inputs and opens a door to connecting various vision models. As shown in Fig.1, a wide range of vision tasks can be achieved by using the versatile Grounded SAM pipeline. For example, an automatic annotation pipeline based solely on input images can be realized by incorporating models such as BLIP and Recognize Anything. Additionally, incorporating Stable-Diffusion all","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T15:41:57.504680+00:00"}},"summary":{"title":"Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks","claims":[{"claim_text":"We introduce Grounded SAM, which uses Grounding DINO as an open-set object detector to combine with the segment anything model (SAM). This integration enables the detection and segmentation of any regions based on arbitrary text inputs and opens a door to connecting various vision models. As shown in Fig.1, a wide range of vision tasks can be achieved by using the versatile Grounded SAM pipeline. For example, an automatic annotation pipeline based solely on input images can be realized by incorporating models such as BLIP and Recognize Anything. Additionally, incorporating Stable-Diffusion all","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"SAM 2: Segment Anything in Images and Videos","work_id":"acc13f66-d814-44f9-9688-375688bf2d4a","shared_citers":8},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":7},{"title":"SAM 3: Segment Anything with Concepts","work_id":"4a72a006-2592-4554-aad0-a9c41a9f952d","shared_citers":7},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":5},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":5},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":4},{"title":"DINOv3","work_id":"c8b07deb-8fe7-4e18-9620-f3569d3529ce","shared_citers":4},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":4},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":4},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":3},{"title":"CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer","work_id":"f38fc088-12aa-4bf4-9ecd-08d3e797ccb7","shared_citers":3},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":3},{"title":"Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection","work_id":"3757dc8f-79d5-4beb-a03b-eb4c9a33427d","shared_citers":3},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":3},{"title":"Matter- port3d: Learning from rgb-d data in indoor environments","work_id":"a6675134-1bd7-4d3f-9344-d7072e7449e9","shared_citers":3},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","work_id":"3e7e65c5-5aed-4fe9-8414-2092bcb31cc7","shared_citers":3},{"title":"RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots","work_id":"11232b35-bd17-402a-9234-951c46015815","shared_citers":3},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":2},{"title":"An open and comprehensive pipeline for unified object grounding and de- tection","work_id":"e2d251e7-a46b-4b55-9ddb-9364d020f415","shared_citers":2},{"title":"Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting","work_id":"f2f765af-f0ba-41f1-8dd1-37bf1d134d31","shared_citers":2},{"title":"arXiv preprint arXiv:2507.23134 (2025) 9, 12, 13","work_id":"90e42fa9-770a-442d-a6e6-4df3d73fc3da","shared_citers":2},{"title":"Depth pro: Sharp monocular metric depth in less than a second","work_id":"0b67883b-1901-45f1-9d58-1ef7a928df23","shared_citers":2},{"title":"Distilling the Knowledge in a Neural Network","work_id":"d927ab1f-17b8-4002-9d09-c3d55764fbad","shared_citers":2},{"title":"Egovla: Learning vision-language-action models from egocentric human videos","work_id":"14ae45e8-e465-4230-99bf-0a2a191dabfd","shared_citers":2}],"time_series":[{"n":1,"year":2025},{"n":43,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"09a63137-256f-4d79-b884-90df342dead2","orcid":null,"display_name":"Ailing Zeng","source":"manual","import_confidence":0.72},{"id":"20334f7b-fb10-428f-84fe-0fd4a3028dea","orcid":null,"display_name":"He Cao","source":"manual","import_confidence":0.72},{"id":"a336ca11-66c7-4b4e-9cab-39dbe45c82ae","orcid":null,"display_name":"Jing Lin","source":"manual","import_confidence":0.72},{"id":"7d9fea06-97b4-4683-b934-cacd7dc6aa35","orcid":null,"display_name":"Kunchang Li","source":"manual","import_confidence":0.72},{"id":"58750f5c-a26b-48e5-919d-1ea442a1077d","orcid":null,"display_name":"Shilong Liu","source":"manual","import_confidence":0.72},{"id":"0ab620d7-a875-49b7-90e6-f022a0e27ad4","orcid":null,"display_name":"Tianhe Ren","source":"manual","import_confidence":0.72}]}}