{"id":"0563b7c9-35e4-4c62-ad05-bfd46b3d015b","arxiv_id":"2505.11282","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces MTevent, an event camera dataset with 75 scenes, 16 objects, MoCap-based 6D pose and 3D box annotations, plus a FoundationPose RGB baseline with Average Recall 0.22.","lead":"MTevent is a new event camera dataset with stereo event cameras and an RGB camera recording 75 indoor scenes of large moving objects with ground-truth 6D poses. It targets high-speed robot perception tasks like pose estimation and moving object detection, where ordinary RGB cameras suffer from motion blur.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Event-camera ground truth is generated at RGB frame rate and its accuracy is asserted without measurement; temporal misalignment is the load-bearing risk.","rationale":"The paper is a dataset contribution, and its central claim is that MTevent fills a gap by providing multi-task event-camera ground truth under large distances, occlusions, and high motion. The strongest claim in the abstract is about combining high-speed motion, long-range perception, and real-world interactions; the supporting Sections 3.4 and 3.5 promise ground-truth 6D poses and 3D boxes. The reader's verdict is CONDITIONAL with moderate confidence, and the weakest assumption is that the MoCap-to-camera calibration chain and the RGB-rate event annotations are accurate. I agree that this is the soft spot, but I narrow it to one specific, testable mechanism: the event annotations are generated at the RGB camera's rate, and the statement that they remain accurate is a bare assertion. The dataset may still be valuable if the annotations are actually aligned to event times, but the manuscript does not demonstrate this. The proposed test would settle whether the released labels are trustworthy for event-based methods. This concern does not amount to an internal inconsistency or fraud; it is a missing validation of a condition necessary for the central claim. Therefore the verdict should remain CONDITIONAL, pending a concrete check of temporal alignment accuracy. I do not recommend REJECT because the dataset, toolkit, and raw MoCap data are being released, making the check feasible and the concern corrigible.","tokens_in":9458,"tokens_out":2577,"duration_ms":28750,"concrete_test":"Select one 25 FPS scene with fast camera/object motion. Using the raw 200 Hz MoCap data and the provided extrinsic calibration, recompute object 6D poses and 3D bounding boxes at the start and end of each 10 ms event accumulation window. Compare these per-event interval poses with the released annotations assigned to the corresponding RGB frame. If the median positional difference in fast-motion segments exceeds 2 cm (or the rotation difference exceeds 2 degrees), the event annotations are temporally misaligned and the dataset's ground truth for event cameras is not reliable as released. A complementary check is to compare the 25 FPS-derived event annotations with annotations derived from the 100 FPS RGB data over the same event intervals; a large discrepancy would confirm the issue.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central value of MTevent is that it provides reliable event-camera ground truth for 6D pose estimation and moving-object detection in fast, dynamic scenes. The load-bearing assumption is that the annotations transferred to the event cameras remain accurate, despite being generated at the RGB camera's recording rate. Section 3.5 states: 'Since eye-in-hand calibration is performed using the RGB camera, all annotations for the three cameras are generated at the RGB camera's recording rate.' It then asserts: 'However, event camera annotations remain accurate. This is evident in Fig. 7.' This is unsupported. Events are asynchronous and are accumulated over 10 ms windows; during fast motion, object and camera positions change within that window. At the reported camera speed of up to 2 m/s and with objects moving at comparable speeds, a 10 ms window allows displacements of 1-2 cm, and the 40 ms gap at 25 FPS corresponds to 4-8 cm of motion. The provided 200 Hz MoCap data could support per-event-interval interpolation, but the paper describes a pipeline that labels all three cameras at RGB frame times. If the event annotations are simply copied from RGB times, they contain temporal misalignment that directly corrupts ground truth for event-based methods. No quantitative validation (e.g., comparison against per-event MoCap interpolation or against the 100 FPS RGB annotations) is reported. The reader's weakest assumption correctly identifies the calibration chain; the most critical element is this temporal transfer, because it is asserted as accurate in the text while being unmeasured.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MTevent, a multi-task event-camera dataset for 6D pose estimation and moving-object detection in dynamic indoor scenes. The recording platform is a stereo pair of DVXplorer event cameras plus an RGB camera (either 25 FPS or 100 FPS), tracked by a Vicon motion-capture system at 200 Hz. The dataset contains 75 scenes with 16 rigid industrial and household objects; annotations include 6D poses and segmentation masks for rigid objects, and 3D bounding boxes for humans and a forklift. All annotations are generated at the RGB camera's frame rate and transferred to the event cameras. The authors report a FoundationPose baseline on RGB images using ground-truth masks and synthetic depth, achieving an Average Recall of 0.2207, and argue this highlights the limitations of RGB-based approaches in high-speed scenarios.","tokens_in":9836,"tokens_out":4317,"duration_ms":44707,"significance":"If the annotation quality is established, MTevent fills a genuine gap: existing event-camera datasets with independently moving objects are mostly limited to small household items or short ranges, while MTevent provides larger objects, distances of 2–7 m, occlusions, varying lighting, and MoCap-based ground truth for multiple tasks. The public release of data and toolkit, the use of standard calibration tools (Kalibr, e2vid), the BOP-compatible evaluation format, and the inclusion of a FoundationPose baseline are concrete strengths. The main weakness is that the accuracy of the event-camera ground truth is asserted rather than measured; because the dataset's core value is reliable event-frame labels, this missing validation is load-bearing.","major_comments":[{"comment":"The claim that event-camera annotations remain accurate is not supported by quantitative evidence. The paper states that all annotations are generated at the RGB camera's recording rate, which is 25 FPS in the lower-rate setup, and that event images are accumulated over 10 ms windows. With camera translation speeds up to 2 m/s (Fig. 8), an object or camera can move 4–8 cm between consecutive 25 FPS annotation times, and 1–2 cm within a 10 ms accumulation window. This is not negligible relative to the object sizes reported in Fig. 9 and to the 10 cm tolerance used in the BOP evaluation. The cited evidence, Fig. 7, shows RGB bounding boxes lagging behind the object, but it does not quantitatively demonstrate that the event-frame annotations are correct. Because the MoCap system provides 200 Hz poses, the authors should either (a) report a quantitative temporal-alignment analysis (e.g., reprojection error of meshes at event timestamps versus at RGB frame times, or a comparison against per-event-interval MoCap interpolation), or (b) release annotations interpolated to event timestamps. Without this, the central resource of the dataset is not verified.","section":"3.5, Fig. 7"},{"comment":"The accuracy of all 6D pose annotations depends on an unvalidated calibration chain: Vicon tracking, eye-in-hand calibration between the MoCap-tracked camera system and the RGB optical frame, and object calibration aligning the MoCap-tracked frame with the mesh geometric center. The paper does not report calibration residuals or reprojection errors for any of these transformations. I request a quantitative validation, for example the mean/median reprojection error of projected object meshes into RGB and event frames on a held-out set of frames, or a comparison of MoCap-derived poses with an independent pose estimator on sampled frames. Without such numbers, the ground-truth poses cannot be assessed.","section":"3.1, 3.3, 3.5"},{"comment":"The 3D bounding box annotations for non-rigid objects (humans and forklift) are computed from individual MoCap marker positions, but the manuscript does not specify the marker placement protocol or the rule by which markers are converted to a 3D bounding box. Since moving-object detection is one of the two headline tasks, the accuracy of these boxes is load-bearing. The authors should provide a precise definition of the bounding box (e.g., axes from markers, extents, and coordinate frame), and report a quantitative evaluation, such as 3D IoU against manual annotations on a sample of frames.","section":"3.5"}],"minor_comments":[{"comment":"The table header repeats 'AR' and is confusing: 0.2207 is the mean of the three component scores, so the columns should be labeled 'AR (mean)', 'AR-VSD', 'AR-MSSD', and 'AR-MSPD'.","section":"Table 2"},{"comment":"References [7] and [8] appear to be the same paper, and Section 2 cites both; one duplicate should be removed.","section":"2, References"},{"comment":"The sentence 'We collected the dataset in 2 research hall both are hangar buildings' is grammatically incomplete; it should state the number of halls and the building type more clearly.","section":"3.3"},{"comment":"In Table 1, the row for EED lists 'UA Vs' under Environment, which appears to be a typo; also 'MoCAP' is inconsistently capitalized across the table.","section":"Table 1"},{"comment":"The conclusion that RGB-based methods struggle in this setting is based on a single method, FoundationPose, evaluated with synthetic depth and ground-truth masks; I recommend framing this as a proof-of-concept baseline rather than a general statement about RGB approaches.","section":"4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a useful dataset contribution, and the collection pipeline is standard. The decisive issue is the lack of quantitative validation of the event-camera annotations and the calibration chain. If the authors supply the requested measurements and clarify the non-rigid annotation protocol, I would support acceptance; without them, the dataset's ground-truth trustworthiness remains unestablished. The 'first dataset' claim should also be checked against very recent event-camera datasets not cited here, though I did not find a clear counterexample in the manuscript's own comparison table."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper gives the event-based vision community something it genuinely lacks: a dataset with larger objects, 2–7 m distances, moving cameras, and independently moving objects with 6D pose ground truth. The comparison table in Section 2 checks out—no prior indoor event dataset combines all of those properties. The recording setup is described in detail, the calibration chain is standard, and the objects are decently varied. The FoundationPose baseline is a reasonable way to show that RGB-based methods struggle in these scenes; the low AR is credible, not a red flag.\n\nThe soft spot is the temporal transfer of annotations to the event cameras. Section 3.5 says all annotations are generated at the RGB camera's recording rate, then asserts that event camera annotations remain accurate, citing Figure 7. That is not measured. At the reported camera speed of up to 2 m/s and similar object speeds, a 10 ms event accumulation window allows 1–2 cm of displacement, and the 40 ms gap at 25 FPS allows 4–8 cm. For objects that range from 19 cm to 80 cm, that is not negligible. The 200 Hz MoCap data could support per-event-interval interpolation, but the authors do not report doing that, nor do they validate against an independent method or against the 100 FPS RGB annotations. This is a load-bearing issue: if the event annotations are simply copied from RGB times, they contain temporal misalignment that directly corrupts ground truth for event-based methods.\n\nTwo minor things: the abstract's \"high-speed motion\" framing outruns the actual 2 m/s camera speed, though the conclusion does acknowledge this. And there is no event-based baseline, which is fine for showing RGB limitations but leaves the dataset's utility for event methods unproven.\n\nWho is this for? Researchers working on event-based 6D pose estimation, motion segmentation, or object tracking in robotics. It could become a useful benchmark, but only after the annotation timing is addressed. I'd send this to peer review—the dataset is worth referee time—but the authors should be required to quantify the temporal accuracy of the event annotations, ideally by interpolating the MoCap data to event timestamps or at least by providing a comparison against the 100 FPS annotations. Right now, the claim that event annotations are accurate is asserted, not demonstrated, and that is the gap a serious referee should press on.","headline":"Useful new dataset for event-based robotics, but the event-camera ground truth rests on an unmeasured temporal transfer that needs to be fixed before the annotations can be trusted.","tokens_in":10272,"tokens_out":1703,"would_cite":false,"duration_ms":18880,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MTevent is a new event-camera benchmark that provides ground-truth 6D poses and 3D bounding boxes for fast-moving objects, and its RGB-only baseline scores only 0.22 average recall.","keywords":["event camera dataset","6D pose estimation","moving object detection","3D bounding boxes","motion capture ground truth","high-speed robotics","stereo event camera","RGB baseline evaluation"],"falsifier":"Check the calibration by reprojecting the tracked 6D poses into high-contrast frames and measuring pixel error against manually clicked object corners; if the error exceeds a few pixels, the ground truth is biased. Second, in a fast-motion scene recorded with both 25 FPS and 100 FPS RGB, compare the 25 FPS labels against the 100 FPS labels: systematic disagreement during fast motion would disprove the paper's assertion that annotations generated at the RGB rate remain accurate for the event cameras.","tokens_in":9287,"feed_emoji":"📷","tokens_out":11037,"duration_ms":96646,"temperature":0.7,"pith_summary":"MTevent is a 75-scene dataset recorded with a stereo event camera pair and a single RGB camera in two large hangar halls, built for robotics perception in fast-moving, cluttered scenes. The paper's central claim is that MTevent is the first event-camera dataset to combine high-speed motion, long detection distances, and realistic object interactions, with ground-truth 6D poses for 16 rigid objects and 3D bounding boxes for every moving object. Existing event datasets, the authors argue, are limited to small household objects or short ranges, leaving a gap for high-speed mobile robots that must perceive larger objects at distances mostly between 2 and 7 meters. To demonstrate the gap, the authors run FoundationPose on the RGB frames with ground-truth masks and report an Average Recall of only 0.22. If the dataset delivers what it promises, it gives the event-vision community a benchmark in which RGB-based pose estimation is the clear bottleneck, motivating event-based perception for high-speed robots.","feed_headline":"RGB pose estimation scores only 0.22 on new event-camera benchmark","feed_subtitle":"75 scenes, 16 objects, and 200 Hz motion-capture labels put event and RGB cameras in the same high-speed test.","key_machinery":"The load-bearing mechanism is the annotation pipeline that converts 200 Hz motion-capture tracks into camera-frame ground truth. Eye-in-hand calibration maps the tracked camera rig to the RGB camera's optical frame, while a per-object calibration, performed with the BOP manual annotation tool, aligns each object's tracked frame with its mesh geometric center. Object 6D poses are then transformed from the world frame into the optical frames of all three cameras, and segmentation masks for rigid objects come from projecting the 3D meshes; for humans and forklifts, masks come from Segment-Anything-2 and 3D boxes from the tracked markers. All annotations are generated at the RGB camera's frame rate, which means the 25 FPS scenes—with 25 ms exposure time—can carry motion blur that the paper acknowledges degrades RGB annotation accuracy, while asserting that the event-camera annotations remain accurate. An event camera is defined by asynchronous per-pixel brightness-change reporting, which is the property the dataset is designed to exploit.","core_discovery":"The paper's discovery is the dataset itself, presented as the missing combination of properties in event-camera benchmarks. MTevent provides 75 scenes averaging 16 seconds each, recorded with two synchronized DVXplorer event cameras (640×480, 10.2 cm baseline) and an RGB camera running at either 25 FPS or 100 FPS. The 16 rigid objects are larger than typical household items—ranging from 16 × 19 × 42 cm to 80 × 120 × 144 cm—and each has a 3D mesh (CAD for the Euro pallet, BundleSDF reconstructions for the rest) plus motion-capture-based 6D pose annotations. For non-rigid moving entities such as humans and a forklift, the dataset supplies 3D bounding boxes derived from marker tracks, and segmentation masks for those entities come from Segment-Anything-2. The authors evaluate 6D pose estimation on the 25 FPS RGB subset using FoundationPose with ground-truth masks and synthetic depth, obtaining an Average Recall of 0.2207 (VSD 0.1874, MSSD 0.1716, MSPD 0.3031), which they interpret as evidence that RGB-only methods are limited in these dynamic conditions.","pith_inferences":["A straightforward test of the paper's motivation would be to run FoundationPose on the 100 FPS RGB scenes and compare average recall with the 25 FPS result; a large gap would confirm motion blur, rather than background clutter, as the main reason RGB pose estimation underperforms.","Because the motion-capture system tracks at 200 Hz but labels are published at RGB frame rate, re-rendering the same annotations at motion-capture frequency would give event-based methods a higher-rate evaluation without any additional hardware.","The claim that event annotations remain accurate even when RGB annotations lag could be quantified by measuring whether projected mesh edges align with event-image edges in fast-motion slices; the paper shows a visual example but no such metric.","The same motion-capture-based annotation pipeline could be ported to outdoor high-speed platforms, replacing the indoor tracking system with GPS/RTK, to cover the 5-10 m/s regime the introduction motivates; the authors note the camera rig in MTevent moved only up to 2 m/s."],"forward_implications":["With ground-truth masks and synthetic depth supplied, FoundationPose reaches only 0.22 average recall on MTevent's 25 FPS RGB subset, making the dataset a stress test in which RGB-only pose estimators visibly struggle under motion blur and clutter.","Each scene bundles synchronized event streams, RGB frames, motion-capture tracks, 6D poses for 16 rigid objects, and 3D bounding boxes for all moving objects, so one dataset supports direct comparison across six perception tasks.","The 100 FPS RGB scenes provide higher-temporal-resolution annotations, enabling controlled experiments on how frame rate and motion blur affect pose-estimation accuracy and annotation reliability.","Because synthetic depth is rendered from the object meshes, depth-requiring pose pipelines can be evaluated on MTevent without a physical depth sensor.","The overlap of the 16 objects with an existing RGB pose-tracking benchmark means models already tested on those objects can be transferred directly to the event-based setting."],"supporting_citations":[{"why":"Prior event-camera dataset with motion-capture ground truth but small household objects; serves as the closest comparison that MTevent extends.","marker":"[1]"},{"why":"Multi-sensor event dataset with independently moving objects and large detection distances; establishes the baseline comparison in the dataset table.","marker":"[2]"},{"why":"BundleSDF: reconstructs the object meshes used to project rigid-object masks and to generate synthetic depth for pose estimation.","marker":"[25]"},{"why":"BOP manual annotation tool: aligns each object's tracked frame with its mesh geometric center during object calibration.","marker":"[12]"},{"why":"Kalibr toolbox: performs the intrinsic and extrinsic calibration of the three-camera system.","marker":"[24]"},{"why":"FoundationPose: the RGB-based pose estimator whose 0.22 average recall on ground-truth masks is the paper's reported baseline.","marker":"[26]"},{"why":"BOP challenge metrics: define VSD, MSSD, MSPD and the average-recall evaluation used in the baseline experiment.","marker":"[13]"},{"why":"Segment-Anything-2: generates segmentation masks for humans and the forklift, from which their 3D bounding boxes are derived.","marker":"[22]"}],"fun_headline_variants":["Event-camera benchmark exposes RGB pose limits at speed","New dataset pairs event and RGB for high-speed 6D pose","MTevent: first dataset for fast, long-range robot vision","RGB-only pose scores 0.22 on new event-camera dataset","First event-camera dataset for long-range, high-speed 6D pose"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The accuracy of every pose and bounding-box label rests on the motion-capture-to-camera and object-to-mesh calibration chain, and the paper asserts rather than measures that event annotations stay accurate even when RGB annotations are unreliable.","fun_headline_variants_meta":{"raw":{"variants":["Event-camera benchmark exposes RGB pose limits at speed","New dataset pairs event and RGB for high-speed 6D pose","MTevent: first dataset for fast, long-range robot vision","RGB-only pose scores 0.22 on new event-camera dataset","First event-camera dataset for long-range, high-speed 6D pose"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000879,"raw_usage":{"total_tokens":3877,"prompt_tokens":1100,"completion_tokens":2777,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":716,"completion_tokens_details":{"reasoning_tokens":2686}},"tokens_in":716,"tokens_out":2777,"duration_ms":20119,"temperature":1.0,"reasoning_tokens":2686,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:53:40.659713+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the calibration by reprojecting the tracked 6D poses into high-contrast frames and measuring pixel error against manually clicked object corners; if the error exceeds a few pixels, the ground truth is biased. Second, in a fast-motion scene recorded with both 25 FPS and 100 FPS RGB, compare the 25 FPS labels against the 100 FPS labels: systematic disagreement during fast motion would disprove the paper's assertion that annotations generated at the RGB rate remain accurate for the event cameras.","supporting_citations":[{"cited_title":"M3ed: Multi-robot, multi-sensor, multi-environment event dataset","cited_arxiv_id":null,"evidence_quote":"Multi-sensor event dataset with independently moving objects and large detection distances; establishes the baseline comparison in the dataset table."},{"cited_title":"Extending kalibr: Cali- brating the extrinsics of multiple imus and of individual axes","cited_arxiv_id":null,"evidence_quote":"Kalibr toolbox: performs the intrinsic and extrinsic calibration of the three-camera system."},{"cited_title":"FoundationPose: Unified 6d pose estimation and tracking of novel objects","cited_arxiv_id":null,"evidence_quote":"FoundationPose: the RGB-based pose estimator whose 0.22 average recall on ground-truth masks is the paper's reported baseline."},{"cited_title":"BOP challenge 2020 on 6D object localization","cited_arxiv_id":null,"evidence_quote":"BOP challenge metrics: define VSD, MSSD, MSPD and the average-recall evaluation used in the baseline experiment."}],"review_version":1}