{"id":"fad72a2c-b549-4e84-a8f8-4722bd1489ae","arxiv_id":"2505.11970","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured review of real-time scheduling techniques for CPU-plus-accelerator systems, grouped by deadline type, task model, and application area.","lead":"This paper surveys a decade of real-time scheduling methods for computers that combine CPUs with GPUs, FPGAs, TPUs, or other accelerators. It organizes the field into soft and hard deadline approaches, common task models, and application-driven designs, and points out open problems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's exclusion of latency/QoS-driven scheduling (Sec. II.B.2) is contradicted by the inclusion of works whose stated objectives are explicitly QoS/SLO or latency-based, such as Pegasus [79], Llumnix [112], and BOXR [114]; this weakens the coherence of the 'comprehensive survey' claim.","rationale":"The reader's ACCEPT is reasonable given the survey's coherent internal structure, but the scope inconsistency I found is concrete and directly bears on the central comprehensiveness claim. The paper explicitly excludes latency/QoS-driven scheduling in Sec. II.B.2, yet several included entries are described in the paper's own text as QoS-, SLO-, or latency-driven. This does not invalidate the useful taxonomy, but it does mean the inclusion boundary is not being applied as stated. Before final publication the authors should either revise Sec. II.B.2 to acknowledge that latency/QoS/SLO approaches with time budgets are included, or remove or relabel the contradictory entries. Because the issue is correctable and the survey otherwise has a well-structured organization, acceptance should be conditional on resolving this scope inconsistency rather than unconditional.","tokens_in":29270,"tokens_out":4043,"duration_ms":39674,"concrete_test":"Audit Tables II-IV and Section VI against the Sec. II.B.2 criterion by reading the cited papers' stated scheduling objective. For each entry, classify it as deadline-based (explicit periods/deadlines/WCRT) or latency/QoS/SLO-based. Start with the five clearest test cases: Pegasus [79], Llumnix [112], BOXR [114], Heimdall [115], and R-TOD [106]. If even one is primarily latency/QoS/SLO-driven with no explicit deadline model, the scope statement is contradicted. The test passes only if all retained entries have hard/soft deadlines as their primary scheduling metric.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that this is 'a comprehensive survey of real-time scheduling techniques' (Abstract). That claim depends on the inclusion criterion stated in Sec. II.B.2: a task qualifies as real-time if it has hard or soft deadlines, and 'latency- or quality-of-service (QoS)-driven scheduling approaches' are explicitly not covered. The text does not consistently apply this criterion. In Sec. IV.B.2, Pegasus is described as coordinating CPU/GPU scheduling in a hypervisor with 'feedback-driven adjustments for QoS' and a policy named SLAF; this is QoS-driven scheduling by the survey's own definition. In Sec. VI.C, Llumnix is included for 'differentiated Service Level Objectives (SLOs)' and lower latency/cost, with no deadline model. In Sec. VI.E, BOXR and Heimdall are included based on latency budgets, 'M2D/C2D latencies', and 'responsiveness', not on hard/soft deadlines. Similar latency-driven language appears for R-TOD and other perception entries. If the exclusion is real, these entries should be removed; if it is not, the survey's boundary for 'real-time' is broader than stated, and the claim of comprehensive coverage is underspecified because it silently absorbs works whose primary metric is SLO/latency rather than deadlines. This is an internal consistency problem, not an external completeness judgment.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper surveys real-time scheduling techniques for accelerator-based heterogeneous architectures, covering CPU-GPU, CPU-TPU, and CPU-FPGA systems. It organizes the material around task execution models (accelerator-only, self-suspending segmented model, DAG, and task chain), soft versus hard deadline classes, vendor and research-designed schedulers, multi-objective approaches (energy and thermal), and application-driven scheduling for autonomous systems, perception, language models, satellites, and extended reality. The survey claims to be comprehensive for the past ten years and concludes with open challenges such as multi-core heterogeneity, memory-copy overheads, response-time analysis, and standardized evaluation.","tokens_in":29564,"tokens_out":3094,"duration_ms":33090,"significance":"If the scope were applied consistently, this would be a valuable reference map for researchers entering real-time scheduling on accelerator-based heterogeneous platforms. The paper's strengths are its broad collection of recent works, the explicit task-model taxonomy, the structured tables summarizing soft/hard real-time and application-driven schedulers, and the honest enumeration of open problems. It does not make new formal claims, so correctness hinges on accurate representation and consistent selection criteria; the tables and descriptions are generally faithful to the cited works.","major_comments":[{"comment":"The stated inclusion criterion is inconsistent with several included works. Section II.B.2 says a task qualifies as real-time only if it has hard or soft deadlines, and that latency- or QoS-driven scheduling approaches are not covered. However, Pegasus (Section IV.B.2, Table II) is included with a policy named SLAF described as 'feedback-driven adjustments for QoS'; Llumnix (Section VI.C, Table IV) is included for 'differentiated Service Level Objectives (SLOs)' without any deadline model; and BOXR (Section VI.E, Table IV) is included on the basis of M2D/C2D latency budgets. These works are latency/QoS-driven by the survey's own definition, so the 'comprehensive survey' claim is underspecified: either the exclusion should be enforced or the scope should be broadened and explicitly justified.","section":"II.B.2 vs IV.B.2, VI.C, VI.E, Table IV"},{"comment":"The same scope paragraph excludes cloud-based virtualized accelerator servers, yet Pegasus is a hypervisor-level scheduling system for virtualized GPU access and Llumnix is a datacenter-scale LLM serving system. This contradicts the stated emphasis on single-machine embedded and mobile real-time platforms. The authors should either remove these entries or revise the scope statement to acknowledge that some cloud/virtualized or latency-driven works are included because of their deadline-related contributions.","section":"II.B.2 vs IV.B.2 and VI.C"}],"minor_comments":[{"comment":"The name 'Elloit' should be 'Elliott'; the same misspelling appears in Table II as 'Elliott [76] et al.'.","section":"IV.B.1"},{"comment":"The phrase 'zhDNN stages' appears to be a typo and should be corrected to 'DNN stages' or the intended technical term.","section":"V.C, PredJoule paragraph"},{"comment":"The word 'espeacially' in the DAG model description should be 'especially'.","section":"III.B.3"},{"comment":"The entry 'Kernal' under the FRED row should be 'Kernel'.","section":"Table I"},{"comment":"The scheduling algorithm label 'N P F Pf lex' has a formatting artifact and should read 'NPFPflex' consistently with the text in Section VI.B.","section":"Table IV"},{"comment":"The phrase 'V oltage and Frequency Scaling' contains an unintended space and should be 'Voltage and Frequency Scaling'.","section":"V.C"}],"recommendation":"major_revision","confidential_remarks":"The scope inconsistency is the main substantive issue and is fixable by either tightening the selection or openly broadening the definition of real-time scheduling to include latency/SLO-driven work. The paper is otherwise a competent survey; the number of self-citations is noticeable but not inappropriate for a survey from an active group in the area."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this survey does what a survey should do. It gives a readable, structured map of a decade of real-time scheduling work on CPU-GPU, CPU-FPGA, and CPU-TPU platforms, organized by soft/hard deadlines, task models, and application domains. The taxonomies are sensible, and the summaries of the primary literature are mostly accurate — I spot-checked several and they held up. If you are new to this subfield, this is a genuinely useful entry point.\n\nThe stress-test concern about the scope definition is real, but I would grade it as a moderate internal inconsistency rather than a fatal flaw. The paper explicitly says it does not cover latency- or QoS-driven scheduling (Sec. II.B.2), yet the application section includes Pegasus (feedback-driven QoS), Llumnix (differentiated SLOs), BOXR (M2D/C2D latencies), and R-TOD (end-to-end delay). That is a genuine boundary-rule violation. The fix is easy: either broaden the stated scope to include latency/QoS-driven approaches that still respect soft or hard timing constraints, or move those entries out. As written, the inclusion criterion is contradictory and should be resolved before publication.\n\nOther soft spots are minor. There are typos like “Elloit” and “zhDNN” that suggest a rushed proofread. Table I lists Elliott [32] under thread-block preemption, but the cited GPUSync paper is a scheduling framework rather than a preemption mechanism — that is an imprecise summary. The survey also leans on the authors’ own prior work in several places, but those citations are to real, peer-reviewed results, so I do not count that against it.\n\nWhat the paper does not do is also important: it is a review, not a research contribution. No new algorithm, proof, or measurement. That is fine — not every useful paper needs to be a new result — and the significance is in the consolidation of a scattered literature.\n\nWho is this for? A graduate student or a researcher entering the area who needs a map of the field, its models, and its open problems. The discussion of evaluation methodology and the challenge of response-time pessimism with mixed processor types is well done.\n\nRecommendation: send it to peer review. It deserves referee time. A careful reviewer will flag the scope contradiction and the imprecise entries, and after a light revision it will be a solid reference that people will cite.","headline":"A genuinely useful, well-organized survey that makes a few internal-consistency mistakes in its stated scope; worth refereeing, but the authors should fix the latency/QoS inclusion contradiction before publication.","tokens_in":30072,"tokens_out":1928,"would_cite":false,"duration_ms":22119,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey maps ten years of real-time scheduling on accelerator-based heterogeneous platforms, organizing the field by soft versus hard deadlines, task execution models, and application domains.","keywords":["real-time scheduling","heterogeneous computing","GPU","FPGA","TPU","self-suspending tasks","DAG scheduling","task chain"],"falsifier":"Run a systematic literature search over 2015-2025 for deadline-constrained scheduling on GPU, FPGA, TPU, or NPU platforms; if the majority of retrieved works are latency-, QoS-, or virtualization-driven, the survey's scope decision would place the field's center of gravity outside the map it claims to draw.","tokens_in":1688,"feed_emoji":"⏱️","tokens_out":2398,"duration_ms":69083,"temperature":0.7,"pith_summary":"The paper tries to establish that real-time scheduling on accelerator-based heterogeneous systems--chips that combine CPU cores with GPUs, TPUs, or FPGAs--has matured into a coherent research area that can be mapped by deadline type, task model, and application. It argues that the last decade divides naturally into soft-real-time work, which improves schedulability without strict response-time analysis, and hard-real-time work, which carries formal worst-case response-time guarantees, with a third stream driven by applications such as autonomous driving, perception, language models, satellites, and extended reality. A sympathetic reader would care because the survey supplies the categories and reference points needed to locate where a new scheduling problem fits and which existing techniques are candidates to extend.","feed_headline":"A decade of real-time scheduling on AI accelerators, organized by deadline","feed_subtitle":"Soft and hard deadline techniques for CPU-GPU, CPU-TPU, and CPU-FPGA systems from vendor to DNN schedulers.","key_machinery":"The organizing machinery is a two-axis taxonomy: deadline class (soft versus hard) crossed with accelerator count (single versus multiple), task execution model, and objective (timing only versus energy or thermal). The load-bearing task models are the self-suspending segmented model $\\tau_i = (C^1_i, A^1_i, C^2_i, \\ldots, A^{M_i-1}_i, C^{M_i}_i), D_i, T_i$, the directed acyclic graph (DAG) model, and the task-chain model, each mapped onto a real workload such as CNN inference, transformer attention, or ROS pipelines. The taxonomy does the work: it converts a list of papers into a decision tree that a researcher can enter with their deadline type, task shape, and accelerator count.","core_discovery":"On the paper's own terms, the discovery is organizational: the seemingly scattered literature on scheduling real-time tasks on CPU-plus-accelerator platforms can be read as a progression from architectural features through task models to scheduling objectives. It shows that three task models--the segmented self-suspension model, the directed acyclic graph (DAG) model, and the task-chain model--capture most existing analyses, and that soft-real-time designs (vendor schedulers, priority queues, spatial partitioning, energy and thermal heuristics) and hard-real-time designs (fixed-priority and EDF analyses, response-time bounds, multi-objective thermal and energy control) form distinct lineages. It also claims that application-driven work is where the field is currently expanding, and it closes by identifying open problems: platforms with more than two processor types, memory-bus copy overhead, co-designed CPU-plus-accelerator metrics, and standardized evaluation.","pith_inferences":["Editorial extension: the survey's observation that two different utilization definitions coexist, one CPU-dominant and one treating CPU and accelerator equally, suggests the field lacks a shared unit of load across processor types; a testable next step is a benchmark suite that reports both definitions and shows when scheduling conclusions flip.","Editorial extension: the task-chain model is likely to grow in importance as ROS 2 becomes the standard integration layer for robots using accelerators, so scheduling work that ignores chain-level data dependencies may increasingly miss the true end-to-end latency.","Editorial extension: because memory copies are modeled as non-preemptive stages, unified-memory platforms where CPU and accelerator share addresses may eventually blur the CPU-segment versus accelerator-segment distinction, which would strain the segmented model that much of the hard-real-time analysis relies on.","Editorial extension: a reader could test the taxonomy's completeness by applying it to papers outside the survey's scope, such as latency- or QoS-driven accelerator systems; if those papers resist the soft-versus-hard deadline distinction, the taxonomy would need a third axis."],"forward_implications":["A researcher with a new hard-deadline single-accelerator problem can locate the relevant analysis lineage, such as fixed-priority or EDF-like response-time bounds over self-suspending segments, without re-deriving the field.","Multi-accelerator systems shift the focus from response-time analysis to task-to-accelerator allocation, so new work there should expect to compare against allocation and load-balancing approaches rather than purely analytical schedulability tests.","Soft-real-time designs that tolerate occasional deadline misses can use vendor-provided mechanisms as baselines before adopting heuristic frameworks that add preemption, spatial partitioning, or energy and thermal control.","Application-driven scheduling, covering autonomous systems, perception, language models, satellites, and extended reality, is presented as the current frontier where deadline type, task model, and platform constraints are chosen together.","The survey's open problems imply that the next advances are likely to come from platforms with more than two processor types, explicit memory-copy modeling in schedules, and standardized evaluation metrics that allow fair comparison."],"supporting_citations":[{"why":"GPUSync, the framework for real-time GPU management with fixed- and dynamic-priority scheduling, anchors the multi-accelerator soft-real-time lineage.","marker":"[32]"},{"why":"Defines fixed-relative-deadline scheduling for self-suspending tasks and supplies the segmented model used throughout the hard-real-time analysis.","marker":"[56]"},{"why":"EDF-like scheduling for self-suspending tasks provides the adaptive scheduling and response-time analysis framework for single-accelerator hard deadlines.","marker":"[57]"},{"why":"RTGPU introduces spatial partitioning and explicit memory-copy modeling for hard-deadline single-GPU scheduling, a central reference for the survey's hard-real-time category.","marker":"[58]"},{"why":"SHAPE extends fixed-priority analysis to multiple CPUs and many processing elements with a resource-pool viewpoint, grounding the multi-CPU single-accelerator case.","marker":"[59]"},{"why":"Enhanced-MPCP contributes blocking-time analysis for self-suspending tasks under multiprocessor priority ceiling protocols, loading the multi-accelerator response-time analysis branch.","marker":"[10]"},{"why":"The survey of real-time DAG scheduling supplies the DAG task model and its analytical context used in the survey's task-model taxonomy.","marker":"[62]"},{"why":"Response-time analysis of ROS 2 processing chains anchors the task-chain model, which the survey applies to robot and data-dependent workloads.","marker":"[66]"},{"why":"sBEET demonstrates energy-and-timing GPU scheduling with spatial multitasking, serving as the representative multi-objective soft-real-time design.","marker":"[12]"}],"fun_headline_variants":["Real-time scheduling on AI accelerators: a decade surveyed","Soft vs hard deadlines in accelerator real-time scheduling","Scheduling time-critical tasks on CPU-GPU/TPU/FPGA","Survey: real-time scheduling for heterogeneous accelerators","A decade of accelerator scheduling across deadline types"],"cache_read_input_tokens":32256,"weakest_assumption_plain":"The survey's claim to be comprehensive depends on its explicit decision to exclude latency- and QoS-driven scheduling and cloud-based virtualized accelerator servers, so if a large share of time-critical accelerator scheduling research falls in those excluded areas, the map is incomplete.","fun_headline_variants_meta":{"raw":{"variants":["Real-time scheduling on AI accelerators: a decade surveyed","Soft vs hard deadlines in accelerator real-time scheduling","Scheduling time-critical tasks on CPU-GPU/TPU/FPGA","Survey: real-time scheduling for heterogeneous accelerators","A decade of accelerator scheduling across deadline types"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000655,"raw_usage":{"total_tokens":3013,"prompt_tokens":968,"completion_tokens":2045,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":584,"completion_tokens_details":{"reasoning_tokens":1967}},"tokens_in":584,"tokens_out":2045,"duration_ms":15382,"temperature":1.0,"reasoning_tokens":1967,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:43:05.326695+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a systematic literature search over 2015-2025 for deadline-constrained scheduling on GPU, FPGA, TPU, or NPU platforms; if the majority of retrieved works are latency-, QoS-, or virtualization-driven, the survey's scope decision would place the field's center of gravity outside the map it claims to draw.","supporting_citations":[{"cited_title":"Fixed-relative-deadline scheduling of hard real-time tasks with self-suspensions","cited_arxiv_id":null,"evidence_quote":"Defines fixed-relative-deadline scheduling for self-suspending tasks and supplies the segmented model used throughout the hard-real-time analysis."},{"cited_title":"Edf-like scheduling for self-suspending real-time tasks","cited_arxiv_id":null,"evidence_quote":"EDF-like scheduling for self-suspending tasks provides the adaptive scheduling and response-time analysis framework for single-accelerator hard deadlines."},{"cited_title":"RTGPU: Real-Time GPU Scheduling of Hard Deadline Parallel Tasks with Fine-Grain Utilization","cited_arxiv_id":"2101.10463","evidence_quote":"RTGPU introduces spatial partitioning and explicit memory-copy modeling for hard-deadline single-GPU scheduling, a central reference for the survey's hard-real-time category."},{"cited_title":"Shape: Scheduling of fixed-priority tasks on heterogeneous architectures with multiple cpus and many pes","cited_arxiv_id":null,"evidence_quote":"SHAPE extends fixed-priority analysis to multiple CPUs and many processing elements with a resource-pool viewpoint, grounding the multi-CPU single-accelerator case."},{"cited_title":"A survey on real-time dag scheduling, revisiting the global-partitioned infinity war.Real-Time Systems, 59(3):479–530, 2023","cited_arxiv_id":null,"evidence_quote":"The survey of real-time DAG scheduling supplies the DAG task model and its analytical context used in the survey's task-model taxonomy."},{"cited_title":"Response-time analysis of ros 2 processing chains under reservation- based scheduling","cited_arxiv_id":null,"evidence_quote":"Response-time analysis of ROS 2 processing chains anchors the task-chain model, which the survey applies to robot and data-dependent workloads."}],"review_version":1}