{"id":"a9399c7c-4cbd-43c9-a304-be7a8a48c494","arxiv_id":"2508.17449","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that taxonomizes robotic manipulation policies trained by imitation learning, traces their evolution, and compiles benchmark comparisons.","lead":"This paper is a survey of imitation learning for robotic manipulation, organizing about 80 to 120 methods into a taxonomy and comparing them on standard benchmarks. It aims to be a reference for newcomers and a guide for researchers, but contains several citation and counting errors that reduce its reliability.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Quantitative benchmark tables contain verifiable arithmetic and citation errors; the survey's core claim as a trustworthy comparison resource is undermined.","rationale":"The reader flagged inaccurate transcriptions as the weakest assumption and gave a CONDITIONAL verdict. This stress-test confirms that concern and sharpens it: the problem is not merely possible transcription sloppiness but actual, checkable arithmetic errors in a headline comparison table and multiple citation collisions in the main taxonomy narrative. These errors directly violate the survey's stated purpose as a resource for comparing methods. However, the taxonomic structure and qualitative summaries may still be salvageable if the tables and references are corrected, so a conditional accept rather than a reject is appropriate. I selected 'partial' agreement because the reader's general concern matches, but I found a more specific and less ambiguous failure mode (Table VI arithmetic) than the examples the reader cited.","tokens_in":75,"tokens_out":5595,"duration_ms":67914,"concrete_test":"For every row in Table VI, recompute the Avg. column from the four difficulty values and compare with the paper's reported number for Lift3D and DP3. Also resolve Table III's duplicate [23] by looking up the CALVIN leaderboard entries for SIE and DeeR, and verify Section III.A's [48] citations against the reference list. If any mismatch persists, the tables and in-text attributions must be corrected before the survey can serve as a reliable benchmark reference.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the survey lets researchers locate and compare IL policies via accurate tables and citations. The most load-bearing weakness is that the benchmark aggregation in Section V is demonstrably unreliable. In Table VI (MetaWorld), the reported 'Avg.' for Lift3D is 84.5, but the four listed difficulty scores (93.1, 82.4, 88.0, 28.0) average to 72.9; for DP3 the correct mean is 52.6, not 65.3. Two of six rows are arithmetically wrong, so the table's quantitative ranking cannot be used as is. This is compounded by bibliographic errors: Table III labels two different methods (SIE and DeeR) with the same reference [23], and Section III.A cites CLIPort, ACT, and Diffusion Policy all as [48], while [48] is actually Reactive Diffusion Policy. The recurring mismatch between references and methods means the 'reproducible resource' premise fails: a reader cannot trace a table row to its source or trust the aggregated numbers. These are not stylistic issues; they are factual errors in the core deliverable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper surveys imitation learning (IL) for robotic manipulation (RM). It proposes a hierarchical taxonomy based on control strategy (action generation vs. task planning; diffusion, flow matching, regression, autoregression, classification, affordance), provides structured summaries of representative works, a technological timeline, a discussion of pretraining strategies, benchmark comparisons (CALVIN, RLBench, LIBERO, MetaWorld, SimplerEnv, COLOSSEUM), application categories, and open challenges. The authors claim to identify 82 representative papers and to provide a reproducible resource for locating and comparing IL-based manipulation policies.","tokens_in":35343,"tokens_out":4723,"duration_ms":45822,"significance":"If accurate, this survey would be a valuable entry point and reference: the taxonomy is reasonable, the structured per-paper summaries are useful, and the timeline and benchmark aggregation address a real community need. The paper does not perform circular reasoning; it is a literature survey. However, the resource value depends critically on the correctness of its citations and benchmark tables, and the current version contains several concrete errors in exactly those deliverables. These errors undermine the central claim that readers can trust the tables and references as a reliable comparison resource. The strengths—a first systematic IL-for-RM survey, a sensible taxonomy, and structured paper summaries—are real, but they require a careful fact-checking revision.","major_comments":[{"comment":"The paragraph on single-task policies cites CLIPort, ACT, and Diffusion Policy all as [48]; reference [48] is Reactive Diffusion Policy. The correct citations are CLIPort [96], ACT [63], and Diffusion Policy [22]. This breaks the paper's stated goal of letting readers trace methods to their sources. The same problem appears in Section III.C, where RT-1 is cited as [103] (which is RT-2) instead of [14]; RT-1 and RT-2 then both appear as [103] in adjacent sentences. Table III also labels two distinct methods, SIE and DeeR, both as [23], when [23] is SuSIE.","section":"Section III.A"},{"comment":"The MetaWorld averages in Table VI are arithmetically incorrect for two of six rows. For Lift3D, the reported Avg. of 84.5 does not match the four difficulty scores (93.1, 82.4, 88.0, 28.0), whose mean is 72.9. For DP3, the reported Avg. of 65.3 does not match (85.7, 49.6, 57.0, 18.0), whose mean is 52.6. Since Table VI is one of the paper's principal quantitative comparisons, these errors mislead readers and must be corrected.","section":"Table VI"},{"comment":"The paper's scope statement is internally inconsistent: Section I.C says 'we identified 82 of the most representative papers,' while Section VII.C concludes 'We analyzed 120 research papers.' The discrepancy is not explained, and Table I appears to contain roughly 80 entries. Additionally, Table II lists 'GeminiRob [102]' twice, with different strengths/limitations and different average monthly citations (7.0 and 2.32); one row appears to describe OmniManip [101], not GeminiRob. This duplication and citation mismatch further impair the survey's reliability.","section":"Section I.C vs. Section VII.C"}],"minor_comments":[{"comment":"In the definition of SPL (Eq. (1)), the symbol M is used but not defined; it should be the number of episodes. Also, in Table II, the CARP row contains the typo 'losed-loop feedback' (should be 'closed-loop').","section":"Section V.B"},{"comment":"Table III is introduced as reporting 'success rate' but the columns and metric are 'Avg. Len' and 'Task completed in a row'; the text should clarify the metric and state the source/version of the 'publicly available leaderboard' used, so readers can verify the numbers.","section":"Section V.C"},{"comment":"Minor wording issue: 'Reactive diffusion policy (RDP) [48], a novel slow-fast visual-tactile imitation learning algorithm, allowing robots...' is a sentence fragment; consider restructuring.","section":"Section II.A.b"}],"recommendation":"major_revision","confidential_remarks":"The taxonomy and survey structure are a reasonable contribution, but the number and nature of citation and arithmetic errors in the core deliverables (Tables III, VI, and the paper count) are above what I would expect for a TPAMI-level survey. The authors should perform a systematic audit of all references and benchmark numbers before resubmission. I would support publication if these errors are definitively fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on arXiv:2508.17449. It's a survey of imitation learning for robotic manipulation, organized by control strategy (diffusion, flow matching, regression, autoregressive, affordance, etc.) with a timeline, taxonomy tables, and benchmark comparisons. The organizational work is genuinely useful: the hierarchical taxonomy and the timeline give newcomers a quick map of the field, and the structured per-paper summaries cover a lot of recent work. The ambition to compile quantitative comparisons across CALVIN, RLBench, LIBERO, MetaWorld, SimplerEnv, and COLOSSEUM is the right kind of service for the community.\n\nThe problem is that the core deliverable—the part that makes this a \"reproducible resource\"—has several verifiable errors, and they land right on the load-bearing claims. In Table VI (MetaWorld), the listed Avg. for Lift3D is 84.5, but the four difficulty scores (93.1, 82.4, 88.0, 28.0) average to 72.9. For DP3, the correct mean over 85.7, 49.6, 57.0, 18.0 is 52.6, not 65.3. Two of six rows are arithmetically wrong. That alone means readers can't trust the rankings. On top of that, Section III.A cites CLIPort, ACT, and Diffusion Policy all as [48], when [48] is actually Reactive Diffusion Policy. Table III labels two different methods (SIE and DeeR) with the same reference [23]. And Section I.C says 82 papers while Section VII.C says 120. These are not formatting nits; they undermine the promise that a reader can trace a table row to its source and trust the aggregated numbers.\n\nThere is no new method or measurement here, which is fine for a survey; the paper does not need to be a new result to be valuable. The taxonomy and timeline are reasonable, and the coverage is up to date through 2025. The soft spots are entirely in the accuracy of the aggregation and citation. A careful revision—recomputing every average, fixing the reference keys, and reconciling the paper counts—could make this a genuinely useful entry point. As it stands, I wouldn't cite the benchmark tables in my own work, and I'd caution students against quoting the MetaWorld numbers.\n\nMy recommendation: send it to peer review, but with the expectation that the referees will demand a thorough correction pass before publication. The field could use a good IL-manipulation survey; this one just needs to earn the trust it asks for.","headline":"Useful survey structure, but the benchmark tables and reference list have too many verifiable errors to trust as a reference without heavy revision.","tokens_in":35796,"tokens_out":2369,"would_cite":false,"duration_ms":24604,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey organizes 82 robot imitation-learning methods into one taxonomy.","keywords":["imitation learning","robotic manipulation","vision-language-action models","diffusion policy","flow matching","affordance prediction","robot learning benchmarks","generalization"],"falsifier":"Re-run the survey's literature search for 2021–2025 using an independent citation index, then measure how many of the most-cited imitation-learning manipulation papers appear in the survey's 82-paper selection; if major works are missing or the benchmark numbers in Tables III–VIII disagree with the original papers, the survey's claim to be a comprehensive reference fails.","tokens_in":35026,"feed_emoji":"🤖","tokens_out":6485,"duration_ms":70956,"temperature":0.7,"pith_summary":"This survey tries to give the field of imitation learning for robotic manipulation a single organizing structure. It identifies 82 representative papers and classifies each by how the robot policy produces actions — diffusion models, flow matching, plain regression, autoregressive generation, classification, or affordance-based planning — then traces how these techniques evolved from 2021 to 2025. It also collects benchmark results across six evaluation suites so that methods can be compared on the same tasks. If the survey's selection and transcriptions are trustworthy, it becomes a reference that newcomers and active researchers can use to locate the main approach families, see what each is good at, and find the open problems that remain.","feed_headline":"One taxonomy maps 82 robot imitation-learning methods","feed_subtitle":"Benchmark tables let researchers compare diffusion, flow-matching, and VLA policies side by side.","key_machinery":"The carrying object is the control-strategy taxonomy of Table I: every surveyed method is placed in a grid whose rows distinguish action generation from task planning and whose columns distinguish seven output mechanisms — diffusion, flow matching, Gaussian mixture, naive regression, autoregression, naive classification, and affordance. The survey uses this grid as the spine for its narrative, and pairs it with a chronological timeline (Fig. 2) and six benchmark tables (CALVIN, RLBench, LIBERO, MetaWorld, SimplerEnv, COLOSSEUM) to make the comparison empirical rather than purely taxonomic.","core_discovery":"The paper's central claim is that imitation-learning-based robotic manipulation policies can be systematically organized by a two-level taxonomy: first, whether the policy directly generates actions or plans via task-level outputs such as key poses, affordance maps, or heatmaps, and second, which underlying mechanism — diffusion, flow matching, naive regression or classification, autoregression, or affordance prediction — produces the output. On this basis the survey positions each of its 82 representative papers, arranges them on a technology timeline, identifies the dominant trend toward large-scale pretrained vision-language-action models, and compiles benchmark tables from CALVIN, RLBenc","pith_inferences":["A reader could test the taxonomy's completeness by applying it to the works the survey explicitly excludes — reinforcement-learning policies and grasp-only papers — and checking whether every new method still falls into one of the existing cells.","The survey's numerical inconsistencies (82 analyzed papers in Section I.C versus 120 in Section VII.C, and several different methods sharing the same reference number) suggest that any coverage or citation-count claims should be re-verified from the original sources before being repeated.","The juxtaposition of benchmark tables implies an implicit ranking; a natural extension would be to publish per-task numbers and seeds so the tables can be updated as new policies appear.","The suggestion that bio-inspired structural priors can compensate for scarce data predicts a testable comparison: whether equivariance-, affordance-, or chain-of-thought-based policies reach the same success as data scaling at a fraction of the data."],"forward_implications":["A newcomer can use the taxonomy to identify the main policy families and the canonical paper in each, shortening the entry path into the field.","The benchmark tables provide a current snapshot of which methods lead on long-horizon language-conditioned tasks (CALVIN), keypose prediction (RLBench), lifelong learning (LIBERO), and robustness to perturbation (COLOSSEUM).","The timeline documents a shift from single-task diffusion policies toward generalist, large-scale vision-language-action policies, suggesting that scaling data and model size is the field's main current trajectory.","The survey's open-challenge list — generalization, embodiment diversity, data efficiency, expert-data dependence, and benchmark standardization — defines where near-term research efforts are most needed.","Pretraining on actionless video is highlighted as a way to reduce reliance on expensive action-labeled demonstrations."],"supporting_citations":[{"why":"Diffusion Policy is the foundational method of the diffusion action-generation family that anchors the taxonomy's first category.","marker":"[22]"},{"why":"CLIPort anchors the affordance-prediction task-planner family and the language-conditioned pick-and-place comparison.","marker":"[96]"},{"why":"ACT introduces action chunking, a core mechanism for autoregressive policies and sample-efficient learning.","marker":"[63]"},{"why":"RT-1 establishes the vision-language-action recipe for large-scale real-world robot control.","marker":"[14]"},{"why":"RT-2 extends the VLA paradigm by transferring web-scale knowledge into robotic control.","marker":"[103]"},{"why":"OpenVLA provides the open-source VLA reference point used for fine-tuning and benchmark comparisons.","marker":"[67]"},{"why":"π0 anchors the flow-matching-plus-pretrained-VLM family and the generalist policy evaluation.","marker":"[50]"},{"why":"The Open-X Embodiment dataset is the large-scale multi-embodiment resource that enables pretraining claims for generalist policies.","marker":"[106]"},{"why":"CALVIN is the long-horizon language-conditioned benchmark behind Table III's comparison.","marker":"[117]"},{"why":"RLBench is the simulation benchmark behind Table IV and the COLOSSEUM generalization evaluation.","marker":"[108]"}],"fun_headline_variants":["82 robot policies, two levels, one survey","Direct vs planned: taxonomy of 82 robot policies","How robots learn from demos: 82 methods compared","Survey maps robot imitation learning by mechanism","From demos to dexterity: a taxonomy of 82 methods"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"Everything the survey offers depends on its hand-curated selection of 82 representative papers and on the benchmark numbers transcribed from those papers; the paper itself reports 82 analyzed papers in one section and 120 in another, and some table entries cite the same reference for different methods, so the selection and transcriptions are the load-bearing parts.","fun_headline_variants_meta":{"raw":{"variants":["82 robot policies, two levels, one survey","Direct vs planned: taxonomy of 82 robot policies","How robots learn from demos: 82 methods compared","Survey maps robot imitation learning by mechanism","From demos to dexterity: a taxonomy of 82 methods"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000421,"raw_usage":{"total_tokens":1969,"prompt_tokens":680,"completion_tokens":1289,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":424,"completion_tokens_details":{"reasoning_tokens":1222}},"tokens_in":424,"tokens_out":1289,"duration_ms":16052,"temperature":1.0,"reasoning_tokens":1222,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T16:51:32.121802+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the survey's literature search for 2021–2025 using an independent citation index, then measure how many of the most-cited imitation-learning manipulation papers appear in the survey's 82-paper selection; if major works are missing or the benchmark numbers in Tables III–VIII disagree with the original papers, the survey's claim to be a comprehensive reference fails.","supporting_citations":[],"review_version":1}