{"work":{"id":"bf1e8e81-ff31-401a-a5dc-d9c49df168ab","openalex_id":null,"doi":null,"arxiv_id":"2505.05470","raw_key":null,"title":"Flow-GRPO: Training Flow Matching Models via Online RL","authors":null,"authors_text":"Jie Liu, Gongye Liu, Jiajun Liang, Yangguang Li, Jiaheng Liu, Xintao Wang","year":2025,"venue":"cs.CV","abstract":"We propose Flow-GRPO, the first method to integrate online policy gradient reinforcement learning (RL) into flow matching models. Our approach uses two key strategies: (1) an ODE-to-SDE conversion that transforms a deterministic Ordinary Differential Equation (ODE) into an equivalent Stochastic Differential Equation (SDE) that matches the original model's marginal distribution at all timesteps, enabling statistical sampling for RL exploration; and (2) a Denoising Reduction strategy that reduces training denoising steps while retaining the original number of inference steps, significantly improving sampling efficiency without sacrificing performance. Empirically, Flow-GRPO is effective across multiple text-to-image tasks. For compositional generation, RL-tuned SD3.5-M generates nearly perfect object counts, spatial relations, and fine-grained attributes, increasing GenEval accuracy from $63\\%$ to $95\\%$. In visual text rendering, accuracy improves from $59\\%$ to $92\\%$, greatly enhancing text generation. Flow-GRPO also achieves substantial gains in human preference alignment. Notably, very little reward hacking occurred, meaning rewards did not increase at the cost of appreciable image quality or diversity degradation.","external_url":"https://arxiv.org/abs/2505.05470","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-09T02:25:55.898592+00:00","pith_arxiv_id":"2505.05470","created_at":"2026-05-09T06:10:42.496560+00:00","updated_at":"2026-07-09T02:25:55.898592+00:00","title_quality_ok":true,"display_title":"Flow-GRPO: Training Flow Matching Models via Online RL","render_title":"Flow-GRPO: Training Flow Matching Models via Online RL"},"hub":{"state":{"work_id":"bf1e8e81-ff31-401a-a5dc-d9c49df168ab","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":138,"external_cited_by_count":null,"distinct_field_count":7,"first_pith_cited_at":"2025-06-02T17:54:39+00:00","last_pith_cited_at":"2026-07-08T17:49:49+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T10:59:25.028715+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":20},{"context_role":"method","n":11},{"context_role":"baseline","n":5}],"polarity_counts":[{"context_polarity":"background","n":18},{"context_polarity":"use_method","n":10},{"context_polarity":"baseline","n":5},{"context_polarity":"unclear","n":2},{"context_polarity":"support","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Flow-GRPO: Training Flow Matching Models via Online RL","claims":[{"claim_text":"We propose Flow-GRPO, the first method to integrate online policy gradient reinforcement learning (RL) into flow matching models. Our approach uses two key strategies: (1) an ODE-to-SDE conversion that transforms a deterministic Ordinary Differential Equation (ODE) into an equivalent Stochastic Differential Equation (SDE) that matches the original model's marginal distribution at all timesteps, enabling statistical sampling for RL exploration; and (2) a Denoising Reduction strategy that reduces training denoising steps while retaining the original number of inference steps, significantly impro","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Early methods such as DDPO [23], DPOK [24], and ImageReward/ReFL [25] formulate diffusion generation as policy optimization with rewards for aesthetics, human preference, or text-image alignment, while Diffusion-DPO [26] aligns diffusion models using preference pairs. More recent GRPO-style methods extend RL to modern visual generators, including those for flow models [27, 17], and AR paradigms [28, 29, 30, 31, 32] . However, T2I generation requires multiple rewards to cover aesthetics, alignmen","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"reasoning and the final edited result in existing models. To address this, we propose the CoT-Editing Consistency Reward. Specifically, we employ a VLM (e.g., Qwen2.5- VL [2]) to assess the consistency between the CoT and the edited image from both the task and object perspectives, and provide corresponding rewards. Using this reward, we per- form Flow-GRPO [37] and effectively improve the align- ment between the CoT reasoning and editing outcomes. To validate our approach, we construct a benchm","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"arrangements, and long text to improve rendering accuracy and readability. Second, for multi-view generation, we introduce data that contains consistent subjects across different viewpoints. These samples are designed to improve the model's ability to maintain identity, structure, and spatial consistency under viewpoint changes. 4.2.4 Reinforcement Learning We mainly follow Flow-GRPO [49] as our reinforcement learning framework for text-to-image generation. To construct diverse and informative o","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"further improves training efficiency while achieving com- parable performance. MixGRPO exhibits substantial gains across multiple dimensions of human preference alignment, outperforming DanceGRPO in both effectiveness and ef- ficiency, with nearly 50% lower training time. Notably, MixGRPO-Flash further reduces training time by 71%.1 1. Introduction Recent advances [18, 19, 45, 46, 52] in Text-to-Image (T2I) tasks have demonstrated that probability flow models can *Equal contribution. (lijunzhe10","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"For flow models, the policy in Eq. (2) is defined by the reverse-time SDE transition [19, 40]. Under the flow parameterization, each step follows a Gaussian transition xt−1 =µ θ(xt, t, c)+σ t √ dtϵ,ϵ∼ N(0,I)⇒p θ(xt−1 |x t, c) =N xt−1;µ θ(xt, t, c), σ2 t dtI \u0001 ,(4) where covariance σ2 t dtI is fixed by noise schedule, while mean µθ is given by the velocity field[40]: µθ(xt, t, c) =x t + h vθ(xt, t, c) + σ2 t 2t xt + (1−t)v θ(xt, t, c) \u0001i dt.(5) Flow-Based GRPO Objective.Flow-GRPO optimizes the po","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"representations and the trajectory planner decodes these representations into trajectories. (a) For stage one, to build the strong and comprehensive future reasoning ability of EponaV2, we supervise the future image, depth and semantic maps decoded from the inferred future representations with visual foundation models. (b) For stage two, EponaV2 finetunes the predicted trajectory by flow matching GRPO [45] with several simple reward functions. to supply the driving model with feature representat","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Flow-GRPO: Training Flow Matching Models via Online RL because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (19 contexts).","role_counts":[{"n":19,"context_role":"background"},{"n":11,"context_role":"method"},{"n":5,"context_role":"baseline"}]},"error":null,"updated_at":"2026-06-29T09:08:39.860895+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"bb7c03e0-5597-473a-b71c-5f598c466bb5","orcid":null,"display_name":"Jie Liu"},{"id":"30abb676-35b8-4398-8b81-0e1bc8df6133","orcid":null,"display_name":"Gongye Liu"},{"id":"9d50a9ef-2175-4f00-8e1d-e6a3ec8cf22f","orcid":null,"display_name":"Jiajun Liang"},{"id":"6d0090ac-764a-4c18-88c2-47219e588855","orcid":null,"display_name":"Yangguang Li"},{"id":"142aa718-6eb3-4b74-a953-618e4fbf4a75","orcid":null,"display_name":"Jiaheng Liu"},{"id":"87647f8b-2b87-4a20-8b52-0209c3dc7d0b","orcid":null,"display_name":"Xintao Wang"}]},"error":null,"updated_at":"2026-06-29T09:08:39.858614+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T14:11:13.405263+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"DanceGRPO: Unleashing GRPO on Visual Generation","work_id":"7404dd36-8f9c-478f-b089-ef9f8189c711","shared_citers":31},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":22},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":21},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":20},{"title":"Training Diffusion Models with Reinforcement Learning","work_id":"67684dda-3930-452a-b91a-36cbb8e2e219","shared_citers":18},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":18},{"title":"Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis","work_id":"40702548-f094-4c67-a5db-a62f426f852e","shared_citers":17},{"title":"MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE","work_id":"8b0ab84a-b7ea-46ea-a6ba-bf490a84d251","shared_citers":16},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":14},{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow","work_id":"a1989e1b-d66d-4533-be3a-fb9c5fd62290","shared_citers":14},{"title":"DiffusionNFT: Online Diffusion Reinforcement with Forward Process","work_id":"0ed3cf57-36ba-4962-847e-7a8f5f99901d","shared_citers":13},{"title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","work_id":"5dfe19d5-3541-4803-8fe9-3c8b9e29b281","shared_citers":13},{"title":"Qwen-Image Technical Report","work_id":"d06d7ecc-7579-4f89-a60b-4278a0f3c562","shared_citers":13},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":13},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":10},{"title":"Score-Based Generative Modeling through Stochastic Differential Equations","work_id":"d9110e53-a5d4-4794-a4c5-a575e91c31ad","shared_citers":10},{"title":"Tempflow-grpo: When timing matters for grpo in flow models.arXiv preprint arXiv:2508.04324","work_id":"fecc731c-f8a2-4f00-ab1b-51b87a133726","shared_citers":10},{"title":"Unified Reward Model for Multimodal Understanding and Generation","work_id":"bf9fcf9a-1781-4008-960e-2bec1a717e4e","shared_citers":10},{"title":"Emerging Properties in Unified Multimodal Pretraining","work_id":"e0cfd82c-f5d4-44fd-b531-ec73ab0a805b","shared_citers":9},{"title":"Improving Video Generation with Human Feedback","work_id":"cfe4c01d-1cf7-4a00-ba86-06d583ca2cff","shared_citers":9},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":9},{"title":"Aligning text-to-image models using human feedback.arXiv preprint arXiv:2302.12192","work_id":"39a08cdc-3986-4994-baf0-91e8ddf1b855","shared_citers":8},{"title":"arXiv preprint arXiv:2509.06040 (2025) 2, 3","work_id":"4b5ca02b-b12f-4ca1-88b1-99007b8f9528","shared_citers":8},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":8}],"time_series":[{"n":7,"year":2025},{"n":39,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T14:21:34.301565+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T14:11:22.340306+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Flow-GRPO: Training Flow Matching Models via Online RL","claims":[{"claim_text":"We propose Flow-GRPO, the first method to integrate online policy gradient reinforcement learning (RL) into flow matching models. Our approach uses two key strategies: (1) an ODE-to-SDE conversion that transforms a deterministic Ordinary Differential Equation (ODE) into an equivalent Stochastic Differential Equation (SDE) that matches the original model's marginal distribution at all timesteps, enabling statistical sampling for RL exploration; and (2) a Denoising Reduction strategy that reduces training denoising steps while retaining the original number of inference steps, significantly impro","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Early methods such as DDPO [23], DPOK [24], and ImageReward/ReFL [25] formulate diffusion generation as policy optimization with rewards for aesthetics, human preference, or text-image alignment, while Diffusion-DPO [26] aligns diffusion models using preference pairs. More recent GRPO-style methods extend RL to modern visual generators, including those for flow models [27, 17], and AR paradigms [28, 29, 30, 31, 32] . However, T2I generation requires multiple rewards to cover aesthetics, alignmen","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"reasoning and the final edited result in existing models. To address this, we propose the CoT-Editing Consistency Reward. Specifically, we employ a VLM (e.g., Qwen2.5- VL [2]) to assess the consistency between the CoT and the edited image from both the task and object perspectives, and provide corresponding rewards. Using this reward, we per- form Flow-GRPO [37] and effectively improve the align- ment between the CoT reasoning and editing outcomes. To validate our approach, we construct a benchm","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"arrangements, and long text to improve rendering accuracy and readability. Second, for multi-view generation, we introduce data that contains consistent subjects across different viewpoints. These samples are designed to improve the model's ability to maintain identity, structure, and spatial consistency under viewpoint changes. 4.2.4 Reinforcement Learning We mainly follow Flow-GRPO [49] as our reinforcement learning framework for text-to-image generation. To construct diverse and informative o","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"further improves training efficiency while achieving com- parable performance. MixGRPO exhibits substantial gains across multiple dimensions of human preference alignment, outperforming DanceGRPO in both effectiveness and ef- ficiency, with nearly 50% lower training time. Notably, MixGRPO-Flash further reduces training time by 71%.1 1. Introduction Recent advances [18, 19, 45, 46, 52] in Text-to-Image (T2I) tasks have demonstrated that probability flow models can *Equal contribution. (lijunzhe10","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"For flow models, the policy in Eq. (2) is defined by the reverse-time SDE transition [19, 40]. Under the flow parameterization, each step follows a Gaussian transition xt−1 =µ θ(xt, t, c)+σ t √ dtϵ,ϵ∼ N(0,I)⇒p θ(xt−1 |x t, c) =N xt−1;µ θ(xt, t, c), σ2 t dtI \u0001 ,(4) where covariance σ2 t dtI is fixed by noise schedule, while mean µθ is given by the velocity field[40]: µθ(xt, t, c) =x t + h vθ(xt, t, c) + σ2 t 2t xt + (1−t)v θ(xt, t, c) \u0001i dt.(5) Flow-Based GRPO Objective.Flow-GRPO optimizes the po","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"representations and the trajectory planner decodes these representations into trajectories. (a) For stage one, to build the strong and comprehensive future reasoning ability of EponaV2, we supervise the future image, depth and semantic maps decoded from the inferred future representations with visual foundation models. (b) For stage two, EponaV2 finetunes the predicted trajectory by flow matching GRPO [45] with several simple reward functions. to supply the driving model with feature representat","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Flow-GRPO: Training Flow Matching Models via Online RL because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (19 contexts).","role_counts":[{"n":19,"context_role":"background"},{"n":11,"context_role":"method"},{"n":5,"context_role":"baseline"}]},"error":null,"updated_at":"2026-06-29T09:08:39.518835+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Flow-GRPO: Training Flow Matching Models via Online RL","claims":[{"claim_text":"We propose Flow-GRPO, the first method to integrate online policy gradient reinforcement learning (RL) into flow matching models. Our approach uses two key strategies: (1) an ODE-to-SDE conversion that transforms a deterministic Ordinary Differential Equation (ODE) into an equivalent Stochastic Differential Equation (SDE) that matches the original model's marginal distribution at all timesteps, enabling statistical sampling for RL exploration; and (2) a Denoising Reduction strategy that reduces training denoising steps while retaining the original number of inference steps, significantly impro","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Flow-GRPO: Training Flow Matching Models via Online RL because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T14:21:37.802788+00:00"}},"summary":{"title":"Flow-GRPO: Training Flow Matching Models via Online RL","claims":[{"claim_text":"We propose Flow-GRPO, the first method to integrate online policy gradient reinforcement learning (RL) into flow matching models. Our approach uses two key strategies: (1) an ODE-to-SDE conversion that transforms a deterministic Ordinary Differential Equation (ODE) into an equivalent Stochastic Differential Equation (SDE) that matches the original model's marginal distribution at all timesteps, enabling statistical sampling for RL exploration; and (2) a Denoising Reduction strategy that reduces training denoising steps while retaining the original number of inference steps, significantly impro","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Flow-GRPO: Training Flow Matching Models via Online RL because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"DanceGRPO: Unleashing GRPO on Visual Generation","work_id":"7404dd36-8f9c-478f-b089-ef9f8189c711","shared_citers":31},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":22},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":21},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":20},{"title":"Training Diffusion Models with Reinforcement Learning","work_id":"67684dda-3930-452a-b91a-36cbb8e2e219","shared_citers":18},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":18},{"title":"Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis","work_id":"40702548-f094-4c67-a5db-a62f426f852e","shared_citers":17},{"title":"MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE","work_id":"8b0ab84a-b7ea-46ea-a6ba-bf490a84d251","shared_citers":16},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":14},{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow","work_id":"a1989e1b-d66d-4533-be3a-fb9c5fd62290","shared_citers":14},{"title":"DiffusionNFT: Online Diffusion Reinforcement with Forward Process","work_id":"0ed3cf57-36ba-4962-847e-7a8f5f99901d","shared_citers":13},{"title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","work_id":"5dfe19d5-3541-4803-8fe9-3c8b9e29b281","shared_citers":13},{"title":"Qwen-Image Technical Report","work_id":"d06d7ecc-7579-4f89-a60b-4278a0f3c562","shared_citers":13},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":13},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":10},{"title":"Score-Based Generative Modeling through Stochastic Differential Equations","work_id":"d9110e53-a5d4-4794-a4c5-a575e91c31ad","shared_citers":10},{"title":"Tempflow-grpo: When timing matters for grpo in flow models.arXiv preprint arXiv:2508.04324","work_id":"fecc731c-f8a2-4f00-ab1b-51b87a133726","shared_citers":10},{"title":"Unified Reward Model for Multimodal Understanding and Generation","work_id":"bf9fcf9a-1781-4008-960e-2bec1a717e4e","shared_citers":10},{"title":"Emerging Properties in Unified Multimodal Pretraining","work_id":"e0cfd82c-f5d4-44fd-b531-ec73ab0a805b","shared_citers":9},{"title":"Improving Video Generation with Human Feedback","work_id":"cfe4c01d-1cf7-4a00-ba86-06d583ca2cff","shared_citers":9},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":9},{"title":"Aligning text-to-image models using human feedback.arXiv preprint arXiv:2302.12192","work_id":"39a08cdc-3986-4994-baf0-91e8ddf1b855","shared_citers":8},{"title":"arXiv preprint arXiv:2509.06040 (2025) 2, 3","work_id":"4b5ca02b-b12f-4ca1-88b1-99007b8f9528","shared_citers":8},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":8}],"time_series":[{"n":7,"year":2025},{"n":39,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"30abb676-35b8-4398-8b81-0e1bc8df6133","orcid":null,"display_name":"Gongye Liu","source":"manual","import_confidence":0.72},{"id":"142aa718-6eb3-4b74-a953-618e4fbf4a75","orcid":null,"display_name":"Jiaheng Liu","source":"manual","import_confidence":0.72},{"id":"9d50a9ef-2175-4f00-8e1d-e6a3ec8cf22f","orcid":null,"display_name":"Jiajun Liang","source":"manual","import_confidence":0.72},{"id":"bb7c03e0-5597-473a-b71c-5f598c466bb5","orcid":null,"display_name":"Jie Liu","source":"manual","import_confidence":0.72},{"id":"87647f8b-2b87-4a20-8b52-0209c3dc7d0b","orcid":null,"display_name":"Xintao Wang","source":"manual","import_confidence":0.72},{"id":"6d0090ac-764a-4c18-88c2-47219e588855","orcid":null,"display_name":"Yangguang Li","source":"manual","import_confidence":0.72}]}}