Pith. sign in

REVIEW 3 major objections 4 minor 145 references

The Open-Source BlackParrot-BedRock Cache Coherence System

T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Programmable cache coherence can match fixed-function engines

desk verdict A solid engineering dissertation that delivers an open-source programmable coherence system with real measurements, though the headline performance claim is not tested at the protocol's serialization bottleneck. read the letter →

arxiv 2505.00962 v1 pith:TPJWPJD5 submitted 2025-05-02 cs.AR

classification cs.AR
keywords cachecoherenceprogrammableengineMOESIFprotocoldirectory-basedRISC-VmulticoreBlackParrot-BedRockmicrocode-programmabledirectoryopen-sourcehardware
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This dissertation argues that cache coherence controllers can be made programmable—able to change protocol behavior after fabrication or adapt to applications—without sacrificing the performance or area efficiency of fixed-function hardware. To show this, it presents the BedRock coherence protocol, which keeps all cache blocks in stable MOESIF states by letting the coherence directory serialize every transaction per way group. It then describes three full directory implementations inside the open-source BlackParrot multicore: a fixed-function finite-state machine, a microcode-programmable engine, and a hybrid of the two. The programmable and hybrid engines match the fixed-function design's request-processing occupancy and Splash-3 benchmark results, with a small reported area overhead relative to the whole multicore. The thesis is that programmability in the coherence system is feasible today and can be studied in an open-source setting.

What carries the argument

The central objects are the way group and the pending bit. A way group collects one tag set from every cache in the system for a single cache set, and the pending bit allows only one active coherence transaction per way group. Because the CCE is the only component that changes coherence state and it always waits for the coherence acknowledgment before clearing the pending bit, the protocol never exposes transient states, which dramatically simplifies verification and the cache controller logic. The microcode-programmable CCE uses a small base ISA plus coherence-specific instructions for flag, directory, and queue operations to process BedRock's messages, while the hybrid CCE runs a fixed-function request pipe and a programmable pipe side by side. This machinery is what lets the dissertation compare protocol-processing occupancy directly across the three designs.

What would settle it

Run a many-core workload whose address stream is deliberately hashed so that a small number of way groups receive all the coherence traffic, then measure request throughput against the same workload with addresses hashed evenly. If the programmable or hybrid CCE's throughput collapses relative to the fixed-function design under way-group contention, the programmability-without-overhead claim fails for that regime.

Watch

Extended reading notes

Core claim

The central claim is that adding programmability to the cache coherence directory does not inherently require significant performance or area cost. The evidence is the BP-BedRock system, in which a microcode-programmable CCE and a hybrid CCE both achieve request-processing occupancy and benchmark performance comparable to the fixed-function FSM CCE, while occupying minimal additional area relative to the full multicore. BedRock itself is the enabling protocol design: it is a directory-based invalidate protocol using MOESIF states with no transient coherence states, because the directory is the sole arbiter of state changes and one pending bit per way group serializes all transactions touching that way group. The directory's full-duplicate tag sets give it exact knowledge of every cached block, so cache-to-cache transfers and upgrades can be scheduled without races. Together the protocol and the three engines demonstrate that programmable coherence is implementable and competitive, not just theoretically attractive.

Load-bearing premise

The load-bearing premise is that serializing all coherence transactions touching the same way group with a single pending bit does not become a performance bottleneck when many cores miss to addresses that map to the same group; if that serialization ever becomes the bottleneck, the programmable engines' competitive standing is lost.

Editorial extensions

If this is right

  • A protocol without transient states verifies dramatically faster: CMurphi checks the BedRock MESI variant up to 66x faster than a traditional MESI protocol with six caches, completing in hours where the traditional model would take months.
  • Because the duplicate-tag directory has a constant 6.25% storage overhead relative to L1 capacity regardless of core count, the directory can be tiled across cores with a fixed per-tile size.
  • A microcode-programmable coherence engine can match fixed-function request occupancy on common MOESIF protocol paths, so a chip could ship one programmable directory and later change the protocol without new hardware.
  • The hybrid CCE offers a path to fixed-function performance on common operations and programmable flexibility for uncommon or system-specific operations in the same directory.
  • All three designs are open source, giving other researchers a working multicore platform on which to prototype and evaluate coherence features.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the way-group serialization premise holds under adversarial access patterns, the same pending-bit mechanism could be reused for lightweight ordered or speculative transactions in the coherence system, an extension the dissertation does not explore.
  • Because the duplicate-tag directory overhead is constant while complete-directory overhead grows with core count, programmable duplicate-tag directories may become comparatively cheaper in many-core designs than traditional directory organizations.
  • A direct test of the thesis would be to program the ucode CCE with a non-MOESIF protocol, such as a coherence scheme tailored to accelerators, and measure whether the occupancy advantage persists; this is an extension beyond the paper's claims.
  • The performance-competitiveness result is demonstrated for Splash-3 and microbenchmark workloads, not for worst-case hashing, so the claim should be read as holding for typical rather than adversarial way-group contention.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The dissertation presents the BedRock directory-based MOESIF cache coherence protocol and three open-source implementations within the BlackParrot RISC-V multicore: a fixed-function FSM coherence directory engine, a microcode-programmable engine, and a hybrid of the two. The central claim, stated in Section 1.2 and revisited in Chapters 4 and 5, is that programmability can be added to the coherence system without significant performance or area overhead. Evidence includes exhaustive protocol tables and state-transition diagrams, occupancy tables for each engine, FPGA resource measurements, and Splash-3 and microbenchmark execution comparisons. The work also provides a protocol-complexity comparison against a canonical directory protocol, including CMurphi verification results for the MESI variant and a mathematical latency model for both protocols.

Significance. If the central claim holds, this is a solid, reproducible systems contribution: it ships an open-source, well-documented programmable coherence engine inside a real multicore, with machine-readable protocol artifacts and clear occupancy/area accounting. The three-engine comparison is a useful data point for a long-standing architectural question. The work is explicit about its parameter-free methodology, and the open-source release (Section 4) makes the measurements independently checkable. The main significance is therefore as an engineering demonstration and an open research platform, not as a new protocol concept; the protocol itself is a simplified directory protocol whose novelty lies in the elimination of exposed transient states and directory-controlled replacements.

major comments (3)
  1. [§4.3.1, §4.6, Tables 4.10/4.18/4.19] The performance-competitiveness claim is not tested where the protocol's serialization point is exposed. Section 4.3.1 states that only one coherence transaction per way group may be active, enforced by a pending bit, and that all requests to that way group stall at the directory. The occupancy comparison in Tables 4.10, 4.18, and 4.19 shows the ucode and hybrid engines have higher per-request occupancy than the FSM engine, but the evaluation in Section 4.6 and Chapter 5 reports only aggregate Splash-3 and microbenchmark execution times, without way-group conflict counts, pending-bit stall cycles, or access-pattern variance. Under realistic or adversarial contention on a single way group, the higher occupancy directly becomes lower throughput. This is a load-bearing gap because the thesis's central claim is parity with the fixed-function design; the manuscript should either measure contention directly or qualify the claim to no-contention and low-contention regimes.
  2. [§3.4, Table 3.10; §4.4–4.5] The verification evidence covers only the MESI variant, while the implemented engines run the MOESIF protocol. Table 3.10 reports CMurphi verification for BedRock MESI, and the text in Section 3.4 says only the MESI protocol has been verified. The BP-BedRock implementations described in Chapter 4 and 5 execute the MOESIF protocol with the O and F states (Tables 3.7, 4.18, and related state tables). Since O/F state interactions are precisely where directory-controlled transfers and writebacks are most complex, the correctness claim for the shipped design is not established by the reported verification. The authors should either verify the MOESIF variant or provide a rigorous argument for why MESI verification transfers to MOESIF, and should state this limitation explicitly where correctness is claimed.
  3. [§5.3, Figure 5.11] The hybrid CCE performance comparison in Chapter 5 is presented as a single aggregate result without reporting variance across runs or the configuration parameters (e.g., number of cores, cache sizes, network widths) used. Since the hybrid design's programmable pipe is the key new contribution, the evaluation should show the occupancy and execution-time relationship for the specific requests that use the programmable pipe, not only the blended result. This would make the claimed 'performance parity' of the hybrid design falsifiable and would address the way-group serialization concern raised above.
minor comments (4)
  1. [§1.2] The sentence 'programmability can be be introduced to the cache coherence system' contains a duplicated 'be'; please fix.
  2. [§3.5.5, Tables 3.14–3.15] The mathematical models use abbreviations such as 'M em', 'F ill', and 'Ack' without a legend; add a short note defining each symbol and explaining why the 'Ack' term is counted as latency for BedRock but the requester can overlap it with execution.
  3. [§4.3.3, Figure 4.15] The caption of Figure 4.15 contains 'T able' instead of 'Table'; also, the y-axis label 'Directory Storage Overhead' should specify the normalization basis (L1 cache capacity) directly in the axis title.
  4. [§4.2.2, Table 4.2] The description of the 'Set State & Transfer & Writeback (dirty)' occupancy of '2 + (2*N)' cycles is clear, but the table would benefit from an explicit note that N is the cache block width divided by the fill width, as defined in Section 4.2.2.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central performance and area claims rest on direct measurements of open-source implementations, on internal occupancy tables, and on model-checking runs, not on fitted parameters or load-bearing self-citations.

full rationale

Every load-bearing result in this dissertation is evaluated empirically or by deterministic analysis within the document itself. The central comparison among the fixed-function FSM CCE, the microcode-programmable ucode CCE, and the hybrid CCE is supported by request-occupancy tables (Tables 4.10, 4.18, 4.19, 5.2, 5.5), Splash-3 normalized execution time (Figure 4.22), FPGA resource utilization (Appendix E), and microbenchmark measurements (Section 5.3). These are direct measurements of the described designs; no parameter is fitted to force performance parity, and no 'prediction' is obtained by renaming a fitted constant. The BedRock protocol description in Chapter 3 is a specification rather than a derived result: the comparisons to a canonical directory protocol are based on explicit state-transition tables, message-equivalency analysis (Table 3.13), and CMurphi model-checking runs (Section 3.4), all self-contained or machine-checked. Self-citations to BlackParrot and the BedRock protocol specification identify the prior platform and protocol, but the new contribution--three coherence-engine implementations and their measured trade-offs--does not reduce to those citations. The weakest point noted by a skeptical reader is the way-group serialization assumption (Section 4.3.1), which may limit the generality of the performance-competitiveness claim under adversarial contention; that is an evaluation-coverage concern, not a circular dependency. No circular step meeting the quoted-evidence standard was found.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No numeric parameters are fitted to data in this dissertation. Design choices such as L1 cache capacity, associativity, and block size are configuration parameters documented in Chapter 2, not fitted to make results match. The axioms listed are the load-bearing assumptions that the protocol and implementation rest on; they are clearly stated in the text.

assumptions (4)
  • domain assumption Coherence networks deliver messages error-free and may be unordered.
    Section 3.1.1 states this as a requirement; the protocol's race-free design depends on this assumption.
  • domain assumption Cache controllers process coherence commands atomically as indivisible operations.
    Section 3.2.5 lists this as a protocol assumption; it is load-bearing for the elimination of transient states.
  • ad hoc to paper Only one coherence transaction per way group is active at a time, enforced by pending bits.
    Section 4.3.1 introduces way groups and pending bits to serialize transactions; this implementation choice is specific to BP-BedRock and directly enables the claim of race-free operation.
  • standard math The CMurphi verification model with a single cache block and a single directory is sufficient for coherence correctness.
    Section 3.4 justifies this because coherence is per-address and directories are independent; this mathematical reduction is standard for this kind of verification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Open-Source BlackParrot-BedRock Cache Coherence System." pith.science (2026). https://pith.science/paper/TPJWPJD5

@misc{pith2026250500962,
  author       = {Pith},
  title        = {Pith review of: The Open-Source BlackParrot-BedRock Cache Coherence System},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TPJWPJD5}},
  note         = {Machine review of arXiv:2505.00962}
}
read the original abstract

This dissertation revisits the topic of programmable cache coherence engines in the context of modern shared-memory multicore processors. First, the open-source BedRock cache coherence protocol is described. BedRock employs the canonical MOESIF coherence states and reduces implementation burden by eliminating transient coherence states from the protocol. The protocol's design complexity, concurrency, and verification effort are analyzed and compared to a canonical directory-based invalidate coherence protocol. Second, the architecture and microarchitecture of three separate cache coherence directories implementing the BedRock protocol within the BlackParrot 64-bit RISC-V multicore processor, collectively called BlackParrot-BedRock (BP-BedRock), are described. A fixed-function coherence directory engine implementation provides a baseline design for performance and area comparisons. A microcode-programmable coherence directory implementation demonstrates the feasibility of implementing a programmable coherence engine capable of maintaining sufficient protocol processing performance. A hybrid fixed-function and programmable coherence directory blends the protocol processing performance of the fixed-function design with the programmable flexibility of the microcode-programmable design. Collectively, the BedRock coherence protocol and its three BP-BedRock implementations demonstrate the feasibility and challenges of including programmable logic within the coherence system of modern shared-memory multicore processors, paving the way for future research into the application- and system-level benefits of programmable coherence engines.

Figures

Figures reproduced from arXiv: 2505.00962 by the authors.

Figure 2
Figure 2. [PITH_FULL_IMAGE:figures/full_fig_p019_2.png] view at source ↗
Figure 2.1
Figure 2.1. Canonical Shared-Memory Multicore Processor Architecture [PITH_FULL_IMAGE:figures/full_fig_p020_2_1.png] view at source ↗
Figure 2.2
Figure 2.2. The Cache Coherence Problem in the system. A detailed discussion of memory consistency is beyond the scope of this dissertation. Cache coherence, on the other hand, defines the ordering and outcomes of read and write operations made by more than one processing element to a single memory location in a system employing private data caches that hold local copies of memory data. In many systems, cache coherence defines … view at source ↗
Figures from the paper (93 more)
Figure 2
Figure 2. Figure 2: a [PITH_FULL_IMAGE:figures/full_fig_p021_2.png]
Figure 2.3
Figure 2.3. Figure 2.3: Cache Coherence Single Writer Multiple Reader (SWMR) Invariant [PITH_FULL_IMAGE:figures/full_fig_p022_2_3.png]
Figure 2
Figure 2. Figure 2 [PITH_FULL_IMAGE:figures/full_fig_p022_2.png]
Figure 2.4
Figure 2.4. Figure 2.4: Cache Coherence Data Value Invariant location is guaranteed to be identical. Data Value Invariant The Data Value Invariant, depicted in [PITH_FULL_IMAGE:figures/full_fig_p023_2_4.png]
Figure 2.5
Figure 2.5. Figure 2.5: BlackParrot Tiled Multicore Processor core complex is configurable and supports most common organizations. The remaining tile types surround the core complex in 1-dimensional complexes. Along the North side of the core complex are the I/O tiles, which connect off-chi…
Figure 2
Figure 2. Figure 2 [PITH_FULL_IMAGE:figures/full_fig_p024_2.png]
Figure 2.6
Figure 2.6. Figure 2.6: BlackParrot Multicore Tiles I/O and streaming accelerator tiles. These tiles only support uncached load and store operations, and therefore do not utilize the Fill or Response BedRock networks. DRAM Network The DRAM network supports both cacheable and uncacheable mem…
Figure 2
Figure 2. Figure 2: a [PITH_FULL_IMAGE:figures/full_fig_p025_2.png]
Figure 2
Figure 2. Figure 2: b [PITH_FULL_IMAGE:figures/full_fig_p026_2.png]
Figure 2.7
Figure 2.7. Figure 2.7: BlackParrot Physical Address Space maintaining coherence between the off-chip and DRAM domains. In practice, off-chip I/O access to cacheable DRAM addresses is used to load software or data for execution into BlackParrot’s DRAM prior to the processor cores beginning …
Figure 2
Figure 2. Figure 2: c [PITH_FULL_IMAGE:figures/full_fig_p027_2.png]
Figure 2.8
Figure 2.8. Figure 2.8: BlackParrot Cache Engine Interface is not managed by the BedRock coherence protocol. The BlackParrot Platform Guide found in the BlackParrot documentation provides a more detailed breakdown of the local memory region. Cacheable Global Memory Cacheable global memory i…
Figure 2
Figure 2. Figure 2 [PITH_FULL_IMAGE:figures/full_fig_p028_2.png]
Figure 2
Figure 2. Figure 2: a [PITH_FULL_IMAGE:figures/full_fig_p029_2.png]
Figure 3.1
Figure 3.1. Figure 3.1: Canonical BedRock Coherence System Organization [PITH_FULL_IMAGE:figures/full_fig_p031_3_1.png]
Figure 3
Figure 3. Figure 3 [PITH_FULL_IMAGE:figures/full_fig_p031_3.png]
Figure 3.2
Figure 3.2. Figure 3.2: BedRock Coherence Networks Request Network The Request network carries messages from a cache controller to a coherence directory. Coherence requests are initiated when a cache miss occurs due to the cache having insufficient permissions to complete the requested oper…
Figure 3.3
Figure 3.3. Figure 3.3: Canonical BedRock Coherence Directory Request Processing Flow [PITH_FULL_IMAGE:figures/full_fig_p034_3_3.png]
Figure 3
Figure 3. Figure 3 [PITH_FULL_IMAGE:figures/full_fig_p034_3.png]
Figure 3.4
Figure 3.4. Figure 3.4: Canonical Address Space Layout and Cacheability Properties [PITH_FULL_IMAGE:figures/full_fig_p036_3_4.png]
Figure 3
Figure 3. Figure 3 [PITH_FULL_IMAGE:figures/full_fig_p036_3.png]
Figure 3.5
Figure 3.5. Figure 3.5: BedRock I→S Transitions (1) ReqRd (3) CohAck (2) Data Req IE Dir IE [PITH_FULL_IMAGE:figures/full_fig_p052_3_5.png]
Figure 3.6
Figure 3.6. Figure 3.6: BedRock I→E Transitions 38 [PITH_FULL_IMAGE:figures/full_fig_p052_3_6.png]
Figure 3.7
Figure 3.7. Figure 3.7: BedRock I/S→M Transitions (1) ReqWr (3) CohAck (2) STW Req FM OM Dir FM OM Req FM OM Dir FM OM Sharer SI (3) InvAck (2) Inv (1) ReqWr (3) CohAck Sharer SI (2) Inv (2) STW (3) InvAck [PITH_FULL_IMAGE:figures/full_fig_p053_3_7.png]
Figure 3.8
Figure 3.8. Figure 3.8: BedRock F/O→M Transitions 39 [PITH_FULL_IMAGE:figures/full_fig_p053_3_8.png]
Figure 3.9
Figure 3.9. Figure 3.9: Canonical I→S Transitions (1) GetS (2) Data Req IE Dir IE [PITH_FULL_IMAGE:figures/full_fig_p054_3_9.png]
Figure 3.10
Figure 3.10. Figure 3.10: Canonical I→E Transitions 40 [PITH_FULL_IMAGE:figures/full_fig_p054_3_10.png]
Figure 3.11
Figure 3.11. Figure 3.11: Canonical I/S→M Transitions (1) GetM (2) AckCount[ack=0] Req FM OM Dir FM OM Req FM OM Dir FM OM Sharer SI (3) Inv-Ack (2) Inv (1) GetM Sharer SI (2) Inv (2) AckCount[ack>0] (3) Inv-Ack [PITH_FULL_IMAGE:figures/full_fig_p055_3_11.png]
Figure 3.12
Figure 3.12. Figure 3.12: Canonical F/O→M Transitions 41 [PITH_FULL_IMAGE:figures/full_fig_p055_3_12.png]
Figure 3
Figure 3. Figure 3 [PITH_FULL_IMAGE:figures/full_fig_p056_3.png]
Figure 3
Figure 3. Figure 3 [PITH_FULL_IMAGE:figures/full_fig_p057_3.png]
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p061_4.png]
Figure 4.1
Figure 4.1. Figure 4.1: BP-Bedrock Stream Protocol ‘ define declare_bp_bedrock_header_s ( addr_width_mp , payload_mp , name_mp ) \ typedef struct packed \ { \ payload_mp payload ; \ bp_bedrock_msg_size_e size ; \ logic [ addr_width_mp -1:0] addr ; \ bp_bedrock_wr_subop_e subop ; \ bp_bedroc…
Figure 4.2
Figure 4.2. Figure 4.2: BP-BedRock LCE Block Diagram header and data signals that can be examined by the protocol processing logic to easily take action at important points in messaging processing. The new signal is asserted on the first beat of a message, the last signal is asserted on the…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p063_4.png]
Figure 4.3
Figure 4.3. Figure 4.3: BP-BedRock LCE Request Block Diagram Crossbar before being processed by the Command FSM. The BP-BedRock LCE supports cacheable, uncacheable, and atomic read-modify-write operations to lower levels of the memory hierarchy. As discussed in Section 2.3, the BP-BedRock L…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p064_4.png]
Figure 4.4
Figure 4.4. Figure 4.4: BP-BedRock LCE Request FSM Cache Request Occupancy (cycles) Initiation Interval (cycles) Description Cacheable Load, Cacheable Store 2 + N – Block-based cache miss, send cacheable miss request Uncacheable Load 2 + N – 1, 2, 4, or 8-byte load, send uncached load reque…
Figure 4.5
Figure 4.5. Figure 4.5: BP-BedRock LCE Command Block Diagram Reset Clear Ready Send Ack Transfer Stat Clear Writeback [PITH_FULL_IMAGE:figures/full_fig_p066_4_5.png]
Figure 4.6
Figure 4.6. Figure 4.6: BP-BedRock LCE Command State Machine request in the Ready state and then both receiving the valid metadata and sending the outbound coherence requst in the Send Request state in the same cycle. Thus, the extra cycle of occupancy is a consequence of the BP-BedRock L1 …
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p066_4.png]
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p069_4.png]
Figure 4.7
Figure 4.7. Figure 4.7: BP-BedRock Coherence Directory Architecture [PITH_FULL_IMAGE:figures/full_fig_p070_4_7.png]
Figure 4.8
Figure 4.8. Figure 4.8: BP-BedRock Tag Set modification, and each segment’s outputs are either combined or multiplexed to the output ports of the coherence directory. If no coherent accelerators are present in the design, the accelerator cache directory segment is not instantiated. 4.3.1 Ta…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p070_4.png]
Figure 4.9
Figure 4.9. Figure 4.9: BP-BedRock Way Group Pending Bit Cache 1 - Tag Set 1 Way Group 1 Cache 2 - Tag Set 1 Cache N - Tag Set 1 Pending Bit Cache 1 - Tag Set 2 Way Group 2 Cache 2 - Tag Set 2 Cache N - Tag Set 2 Pending Bit Cache 1 - Tag Set 3 Way Group 3 Cache 2 - Tag Set 3 Cache N - Tag …
Figure 4.10
Figure 4.10. Figure 4.10: BP-BedRock Way Groups At all times, the tag sets tracked at the coherence directory are considered to be the golden copies of the tag sets, which hold the current state of the coherence system across all cached blocks. The directory updates its tag sets during reque…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p071_4.png]
Figure 4.11
Figure 4.11. Figure 4.11: BP-BedRock Address Breakdown block that maps to cache set X (equivalently, tag set X), is a member of way group X [PITH_FULL_IMAGE:figures/full_fig_p072_4_11.png]
Figure 4.12
Figure 4.12. Figure 4.12: illustrates how cache blocks in two caches with different organizations may be related. Cache A has S sets and Cache B has 2*S sets. A cache block that maps to set N in Cache A may map to either set N or S+N in Cache B. Conversely, a block that maps to either set N …
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p073_4.png]
Figure 4.13
Figure 4.13. Figure 4.13: BP-BedRock Coherence Directory Segment 4.3.2 Coherence Directory Segment Architecture Directory Segment Storage [PITH_FULL_IMAGE:figures/full_fig_p074_4_13.png]
Figure 4.14
Figure 4.14. Figure 4.14: BP-BedRock Sharers Vectors ID bits. The least significant bits determine the cache’s position within a row, and the remaining bits index a lookup table that provides the row within the directory memory. The lookup table is computed at compilation and provides fast-a…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p075_4.png]
Figure 4.15
Figure 4.15. Figure 4.15: Coherence Directory Storage Overhead Comparison [PITH_FULL_IMAGE:figures/full_fig_p076_4_15.png]
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p077_4.png]
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p078_4.png]
Figure 4.16
Figure 4.16. Figure 4.16: BP-Bedrock FSM CCE Block Diagram track speculative memory reads issued during request processing, Pending Bits to enforce coherence transaction ordering for each way group, and a Flow Counter to provide network flow control on the memory network. The LCE Request sta…
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p083_4.png]
Figure 4.17
Figure 4.17. Figure 4.17: BP-Bedrock FSM CCE Memory Response Abstract State Machine [PITH_FULL_IMAGE:figures/full_fig_p084_4_17.png]
Figure 4.18
Figure 4.18. Figure 4.18: BP-Bedrock FSM LCE Request State Machine [PITH_FULL_IMAGE:figures/full_fig_p085_4_18.png]
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p087_4.png]
Figure 4.19
Figure 4.19. Figure 4.19: BP-BedRock Microcode-Programmable CCE Block Diagram [PITH_FULL_IMAGE:figures/full_fig_p089_4_19.png]
Figure 4.20
Figure 4.20. Figure 4.20: BP-BedRock MOESIF Microcode Processing Flow - Initial and Fast Path [PITH_FULL_IMAGE:figures/full_fig_p097_4_20.png]
Figure 4.21
Figure 4.21. Figure 4.21: BP-BedRock MOESIF Microcode Processing Flow - Slow Path [PITH_FULL_IMAGE:figures/full_fig_p097_4_21.png]
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p097_4.png]
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p098_4.png]
Figure 4.22
Figure 4.22. Figure 4.22: BP-BedRock Splash-3 Normalized Execution Time - 8 core [PITH_FULL_IMAGE:figures/full_fig_p102_4_22.png]
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p102_4.png]
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p104_5.png]
Figure 5.1
Figure 5.1. Figure 5.1: Hybrid CCE Block Diagram processed by either the Uncacheable Request Pipe or the Coherent Request Pipe after first being classified by the Request Arbiter block. The rest of this section describes the functionality of each processing pipe and the other major modules …
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p105_5.png]
Figure 5.2
Figure 5.2. Figure 5.2: Hybrid CCE Coherence State and Pipe Interaction [PITH_FULL_IMAGE:figures/full_fig_p106_5_2.png]
Figure 5.3
Figure 5.3. Figure 5.3: Hybrid CCE Control State Machine The Memory Command crossbar arbitrates access to the outbound memory command interface from the uncacheable request, coherent request, and LCE response pipelines. During normal op￾eration, the majority of memory commands are issued by…
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p106_5.png]
Figure 5.4
Figure 5.4. Figure 5.4: Hybrid CCE Coherent Request Pipe Block Diagram [PITH_FULL_IMAGE:figures/full_fig_p107_5_4.png]
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p107_5.png]
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p108_5.png]
Figure 5.5
Figure 5.5. Figure 5.5: Hybrid CCE Coherent Request Pipe State Machine [PITH_FULL_IMAGE:figures/full_fig_p109_5_5.png]
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p109_5.png]
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p111_5.png]
Figure 5.6
Figure 5.6. Figure 5.6: BP-Bedrock Hybrid CCE Memory Response Pipe Block Diagram [PITH_FULL_IMAGE:figures/full_fig_p112_5_6.png]
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p112_5.png]
Figure 5.8
Figure 5.8. Figure 5.8: Hybrid CCE LCE Response Pipe Block Diagram [PITH_FULL_IMAGE:figures/full_fig_p113_5_8.png]
Figure 5.9
Figure 5.9. Figure 5.9: Hybrid CCE LCE Response Pipe State Machine [PITH_FULL_IMAGE:figures/full_fig_p114_5_9.png]
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p114_5.png]
Figure 5.10
Figure 5.10. Figure 5.10: Hybrid CCE Programmable Pipe Block Diagram [PITH_FULL_IMAGE:figures/full_fig_p115_5_10.png]
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p119_5.png]
Figure 5.11
Figure 5.11. Figure 5.11: Hybrid CCE Performance Comparison [PITH_FULL_IMAGE:figures/full_fig_p120_5_11.png]
Figure 5.12
Figure 5.12. Figure 5.12: FPGA Resource Utilization Overheads - Full Design [PITH_FULL_IMAGE:figures/full_fig_p121_5_12.png]
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p121_5.png]
Figure 5.13
Figure 5.13. Figure 5.13: FPGA Resource Utilization Overheads - BlackParrot Multicore [PITH_FULL_IMAGE:figures/full_fig_p122_5_13.png]
Figure 5
Figure 5. Figure 5 [PITH_FULL_IMAGE:figures/full_fig_p122_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

145 extracted references · 45 canonical work pages

  1. [1]

    Comparison of hardware and software cache coherence schemes,

    S. V. Adve, V. S. Adve, M. D. Hill, and M. K. Vernon, “Comparison of hardware and software cache coherence schemes,” in Proceedings of the 18th annual international symposium on Computer architecture - ISCA ’91 , Toronto, Ontario, Canada: ACM Press, 1991, pp. 298– 308, isbn: 978-0-89791-394-2. doi: 10.1145/115952.115982

  2. [2]

    The MIT alewife machine,

    A. Agarwal, R. Bianchini, D. Chaiken, F. Chong, K. Johnson, D. Kranz, J. Kubiatowicz, Beng-Hong Lim, K. Mackenzie, and D. Yeung, “The MIT alewife machine,” Proceedings of the IEEE, vol. 87, no. 3, pp. 430–444, Mar. 1999, issn: 00189219. doi: 10.1109/5.747864

  3. [3]

    An evaluation of directory schemes for cache coherence,

    A. Agarwal, R. Simoni, J. Hennessy, and M. Horowitz, “An evaluation of directory schemes for cache coherence,” in [1988] The 15th Annual International Symposium on Computer Architecture. Conference Proceedings, Honolulu, HI, USA: IEEE Comput. Soc. Press, 1988, pp. 280–289, isbn: 978-0-8186-0861-2. doi: 10.1109/ISCA.1988.5238

  4. [4]

    The MIT alewife machine: Architecture and performance,

    A. Agarwal, R. Bianchini, D. Chaiken, K. L. Johnson, D. Kranz, J. Kubiatowicz, B.-H. Lim, K. Mackenzie, and D. Yeung, “The MIT alewife machine: Architecture and performance,” in Proceedings of the 22nd annual international symposium on Computer architecture - ISCA ’95, S. Margherita Ligure, Italy: ACM Press, 1995, pp. 2–13, isbn: 978-0-89791-698-1. doi: 1...

  5. [5]

    An event-triggered programmable prefetcher for irregular workloads,

    S. Ainsworth and T. M. Jones, “An event-triggered programmable prefetcher for irregular workloads,” in Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and Operating Systems, Williamsburg VA USA: ACM, Mar. 19, 2018, pp. 578–592, isbn: 978-1-4503-4911-6. doi: 10.1145/3173162.3173189

  6. [6]

    Coherence protocol for transparent management of scratchpad memories in shared memory manycore architectures,

    L. Alvarez, L. Vilanova, M. Moreto, M. Casas, M. Gonz` alez, X. Martorell, N. Navarro, E. Ayguad´ e, and M. Valero, “Coherence protocol for transparent management of scratchpad memories in shared memory manycore architectures,” in Proceedings of the 42nd Annual International Symposium on Computer Architecture, Portland Oregon: ACM, Jun. 13, 2015, pp. 720–...

  7. [7]

    AMBA AXI and ACE protocol specification, version ARM IHI 0022H.c, ARM Limited, 2021

  8. [8]

    Chipyard: Integrated design, simulation, and implementation framework for custom SoCs,

    A. Amid, D. Biancolin, A. Gonzalez, D. Grubb, S. Karandikar, H. Liew, A. Magyar, H. Mao, A. Ou, N. Pemberton, P. Rigge, C. Schmidt, J. Wright, J. Zhao, Y. S. Shao, K. Asanovic, and B. Nikolic, “Chipyard: Integrated design, simulation, and implementation framework for custom SoCs,” IEEE Micro , vol. 40, no. 4, pp. 10–21, Jul. 1, 2020, issn: 0272-1732, 1937...

Show all 145 references
  1. [9]

    Instruction sets should be free: The case for risc-v,

    K. Asanovic and D. A. Patterson, “Instruction sets should be free: The case for risc-v,” University of California at Berkeley, Tech. Rep., 2014

  2. [10]

    The rocket chip generator,

    Asanovic et al, “The rocket chip generator,” UC Berkeley EECS Tech Report UCB/EECS- 2016-17, 2016. 163

  3. [11]

    Software-based cache coherence with hardware-assisted selective self-invalidations using bloom filters,

    T. J. Ashby, P. Diaz, and M. Cintra, “Software-based cache coherence with hardware-assisted selective self-invalidations using bloom filters,” IEEE Transactions on Computers , vol. 60, no. 4, pp. 472–483, Apr. 2011, issn: 0018-9340. doi: 10.1109/TC.2010.155

  4. [12]

    Chisel: Constructing hardware in a scala embedded language,

    J. Bachrach, H. Vo, B. C. Richards, Y. Lee, A. Waterman, R. Avizienis, J. Wawrzynek, and K. Asanovic, “Chisel: Constructing hardware in a scala embedded language,” in The 49th Annual Design Automation Conference 2012, DAC ’12, San Francisco, CA, USA, June 3- 7, 2012, P. Groene...

  5. [13]

    Open source platforms for enabling full-stack hardware-software research,

    J. Balkind, “Open source platforms for enabling full-stack hardware-software research,” Ph.D. dissertation, Princeton University, USA, 2022

  6. [14]

    BYOC: A

    J. Balkind, K. Lim, M. Schaffner, F. Gao, G. Chirkov, A. Li, A. Lavrov, T. M. Nguyen, Y. Fu, F. Zaruba, K. Gulati, L. Benini, and D. Wentzlaff, “BYOC: A ”bring your own core” framework for heterogeneous-ISA research,” inProceedings of the Twenty-Fifth International Conference ...

  7. [15]

    OpenPiton: An open source manycore research framework,

    J. Balkind, M. McKeown, Y. Fu, T. Nguyen, Y. Zhou, A. Lavrov, M. Shahrad, A. Fuchs, S. Payne, X. Liang, M. Matl, and D. Wentzlaff, “OpenPiton: An open source manycore research framework,” in Proceedings of the Twenty-First International Conference on Architectural Support for ...

  8. [16]

    Piranha: A scalable architecture based on single-chip multipro- cessing,

    L. A. Barroso, K. Gharachorloo, R. McNamara, A. Nowatzyk, S. Qadeer, B. Sano, S. Smith, R. Stets, and B. Verghese, “Piranha: A scalable architecture based on single-chip multipro- cessing,” in Proceedings of the 27th annual international symposium on Computer archi- tecture - ...

  9. [17]

    Analysis and op- timization of the memory hierarchy for graph processing workloads,

    A. Basak, S. Li, X. Hu, S. M. Oh, X. Xie, L. Zhao, X. Jiang, and Y. Xie, “Analysis and op- timization of the memory hierarchy for graph processing workloads,” in 25th IEEE Interna- tional Symposium on High Performance Computer Architecture, HPCA 2019, Washington, DC, USA, Febr...

  10. [18]

    Buildroot

    Buildroot. “Buildroot.” (), [Online]. Available: https://buildroot.org/

  11. [19]

    Busybox

    BusyBox. “Busybox.” (), [Online]. Available: https://busybox.net/

  12. [20]

    Invited - the case for embedded scalable platforms,

    L. P. Carloni, “Invited - the case for embedded scalable platforms,” in Proceedings of the 53rd Annual Design Automation Conference , ser. DAC ’16, Austin, Texas: Association for Computing Machinery, 2016, isbn: 9781450342360. doi: 10.1145/2897937.2905018

  13. [21]

    A highly productive implementation of an out-of-order processor generator,

    C. Celio, “A highly productive implementation of an out-of-order processor generator,” Ph.D. dissertation, University of California, Berkeley, USA, 2017

  14. [22]

    Boom v2: An open- source out-of-order risc-v core,

    C. Celio, P.-F. Chiu, B. Nikolic, D. A. Patterson, and K. Asanovi´ c, “Boom v2: An open- source out-of-order risc-v core,” University of California, Berkeley, Tech. Rep. UCB/EECS- 2017-157, Sep. 2017

  15. [23]

    The berkeley out-of-order machine (boom): An industry-competitive, synthesiz- able, parameterized risc-v processor,

    Celio et al, “The berkeley out-of-order machine (boom): An industry-competitive, synthesiz- able, parameterized risc-v processor,” UC Berkeley EECS Tech. Rep. UCB/EECS-2015-167, 2015

  16. [24]

    A new solution to coherence problems in multicache systems,

    Censier and Feautrier, “A new solution to coherence problems in multicache systems,” IEEE Transactions on Computers , vol. C-27, no. 12, pp. 1112–1118, Dec. 1978, issn: 0018-9340. doi: 10.1109/TC.1978.1675013. 164

  17. [25]

    Software-extended coherent shared memory: Performance and cost,

    D. Chaiken and A. Agarwal, “Software-extended coherent shared memory: Performance and cost,” ACM SIGARCH Computer Architecture News, vol. 22, no. 2, pp. 314–324, Apr. 1994, issn: 0163-5964. doi: 10.1145/192007.192060

  18. [26]

    LimitLESS directories: A scalable cache coherence scheme,

    D. Chaiken, J. Kubiatowicz, and A. Agarwal, “LimitLESS directories: A scalable cache coherence scheme,” ACM SIGOPS Operating Systems Review , vol. 25, pp. 224–234, Special Issue Apr. 2, 1991, issn: 0163-5980. doi: 10.1145/106974.106995

  19. [28]

    SMTp: An architecture for next-generation scalable multi- threading,

    M. Chaudhuri and M. Heinrich, “SMTp: An architecture for next-generation scalable multi- threading,” ACM SIGARCH Computer Architecture News , vol. 32, no. 2, p. 124, Mar. 2, 2004, issn: 0163-5964. doi: 10.1145/1028176.1006712

  20. [29]

    Integrated memory controllers with parallel coherence streams,

    M. Chaudhuri and M. Heinrich, “Integrated memory controllers with parallel coherence streams,” IEEE Transactions on Parallel and Distributed Systems , vol. 18, no. 8, pp. 1159– 1173, Aug. 2007, issn: 1045-9219. doi: 10.1109/TPDS.2007.1044

  21. [30]

    Software-controlled caches in the VMP multiprocessor,

    D. R. Cheriton, G. Slavenburg, and P. D. Boyle, “Software-controlled caches in the VMP multiprocessor,” in Proceedings of the 13th Annual Symposium on Computer Architecture, Tokyo, Japan, June 1986 , H. Aiso, Ed., IEEE Computer Society, 1986, pp. 366–374. doi: 10.1145/17356.17399

  22. [31]

    DeNovo: Rethinking the memory hierarchy for disciplined parallelism,

    B. Choi, R. Komuravelli, H. Sung, R. Smolinski, N. Honarmand, S. V. Adve, V. S. Adve, N. P. Carter, and C.-T. Chou, “DeNovo: Rethinking the memory hierarchy for disciplined parallelism,” in 2011 International Conference on Parallel Architectures and Compilation Techniques, Gal...

  23. [32]

    Application performance on the MIT alewife machine,

    F. Chong, Beng-Hong Lim, R. Bianchini, J. Kubiatowicz, and A. Agarwal, “Application performance on the MIT alewife machine,” Computer, vol. 29, no. 12, pp. 57–64, Dec. 1996, issn: 00189162. doi: 10.1109/2.546610

  24. [33]

    Lightweight hardware support for selective co- herence in heterogeneous manycore accelerators,

    A. Cilardo, M. Gagliardi, and V. Scotti, “Lightweight hardware support for selective co- herence in heterogeneous manycore accelerators,” in 2019 Design, Automation & Test in Europe Conference & Exhibition (DATE) , Florence, Italy: IEEE, Mar. 2019, pp. 932–935, isbn: 978-3-981...

  25. [34]

    The AMD opteron northbridge architecture,

    P. Conway and B. Hughes, “The AMD opteron northbridge architecture,” IEEE Micro , vol. 27, no. 2, pp. 10–21, Mar. 2007, issn: 0272-1732. doi: 10.1109/MM.2007.43

  26. [35]

    Cache hierarchy and memory subsystem of the AMD opteron processor,

    P. Conway, N. Kalyanasundharam, G. Donley, K. Lepak, and B. Hughes, “Cache hierarchy and memory subsystem of the AMD opteron processor,” IEEE Micro, vol. 30, no. 2, pp. 16– 29, Mar. 2010, issn: 0272-1732. doi: 10.1109/MM.2010.31

  27. [36]

    Corporation, An introduction to the Intel quickpath interconnect, january 2009 , 2009

    I. Corporation, An introduction to the Intel quickpath interconnect, january 2009 , 2009

  28. [37]

    The celerity open-source 511-core RISC-v tiered accelerator fabric: Fast architectures and design methodologies for fast chips,

    S. Davidson, S. Xie, C. Torng, K. Al-Hawai, A. Rovinski, T. Ajayi, L. Vega, C. Zhao, R. Zhao, S. Dai, A. Amarnath, B. Veluri, P. Gao, A. Rao, G. Liu, R. K. Gupta, Z. Zhang, R. Dreslinski, C. Batten, and M. B. Taylor, “The celerity open-source 511-core RISC-v tiered accelerator...

  29. [38]

    Design of ion-implanted MOSFET’s with very small physical dimensions,

    R. H. Dennard, F. H. Gaensslen, H.-N. Yu, V. L. Rideout, E. Bassous, and A. R. LeBlanc, “Design of ion-implanted MOSFET’s with very small physical dimensions,” IEEE Journal of Solid-State Circuits , vol. SC-9, no. 5, pp. 256–268, Oct. 1974. 165

  30. [39]

    Coherency traffic reduction in manycore sys- tems,

    E. Derebasoglu, I. Kadayif, and O. Ozturk, “Coherency traffic reduction in manycore sys- tems,” in 2022 25th Euromicro Conference on Digital System Design (DSD) , Maspalomas, Spain: IEEE, Aug. 2022, pp. 262–267, isbn: 978-1-66547-404-7. doi: 10.1109/DSD57027. 2022.00043

  31. [40]

    Protocol verification as a hardware design aid,

    D. Dill, A. Drexler, A. Hu, and C. Yang, “Protocol verification as a hardware design aid,” in Proceedings 1992 IEEE International Conference on Computer Design: VLSI in Computers & Processors, IEEE Comput. Soc. Press, 2004. doi: 10.1109/iccd.1992.276232

  32. [41]

    Ds890 ultrascale architecture and product data sheet: Overview, v4.1.1 , Xilinx, 2022

  33. [43]

    Design verification of the s3.mp cache-coherent shared-memory system,

    Fong Pong, M. Browne, A. Nowatzyk, and M. Dubois, “Design verification of the s3.mp cache-coherent shared-memory system,” IEEE Transactions on Computers , vol. 47, no. 1, pp. 135–140, Jan. 1998, issn: 00189340. doi: 10.1109/12.656100

  34. [45]

    Coherence domain restriction on large scale sys- tems,

    Y. Fu, T. M. Nguyen, and D. Wentzlaff, “Coherence domain restriction on large scale sys- tems,” in Proceedings of the 48th International Symposium on Microarchitecture , Waikiki Hawaii: ACM, Dec. 5, 2015, pp. 686–698, isbn: 978-1-4503-4034-2. doi: 10.1145/2830772. 2830832

  35. [46]

    F. Gao, T. Chang, A. Li, M. Orenes-Vera, D. Giri, P. J. Jackson, A. Ning, G. Tziantzioulis, J. Zuckerman, J. Tu, K. Xu, G. Chirkov, G. Tombesi, J. Balkind, M. Martonosi, L. P. Carloni, and D. Wentzlaff, “DECADES: A 67mm 2, 1.46tops, 55 giga cache-coherent 64-bit RISC-V instruc...

  36. [47]

    Accelerators and coherence: An soc perspective,

    D. Giri, P. Mantovani, and L. P. Carloni, “Accelerators and coherence: An soc perspective,” IEEE Micro, vol. 38, no. 6, pp. 36–45, 2018. doi: 10.1109/MM.2018.2877288

  37. [48]

    Runtime reconfigurable memory hierarchy in embedded scalable platforms,

    D. Giri, P. Mantovani, and L. P. Carloni, “Runtime reconfigurable memory hierarchy in embedded scalable platforms,” in Proceedings of the 24th Asia and South Pacific Design Automation Conference, ASPDAC 2019, Tokyo, Japan, January 21-24, 2019 , T. Shibuya, Ed., ACM, 2019, pp. ...

  38. [49]

    Using cache memory to reduce processor-memory traffic,

    J. R. Goodman, “Using cache memory to reduce processor-memory traffic,” in Proceedings of the 10th annual international symposium on Computer architecture - ISCA ’83 , Stockholm, Sweden: ACM Press, 1983, pp. 124–131, isbn: 978-0-89791-101-6. doi: 10.1145/800046. 801647

  39. [50]

    Dynamically specialized datapaths for en- ergy efficient computing,

    V. Govindaraju, C. Ho, and K. Sankaralingam, “Dynamically specialized datapaths for en- ergy efficient computing,” in 17th International Conference on High-Performance Computer Architecture (HPCA-17 2011), February 12-16 2011, San Antonio, Texas, USA, IEEE Com- puter Society, ...

  40. [51]

    Efficient strategies for software-only protocols in shared- memory multiprocessors,

    H. Grahn and P. Stenstr¨ om, “Efficient strategies for software-only protocols in shared- memory multiprocessors,” in Proceedings of the 22nd Annual International Symposium on Computer Architecture, ISCA ’95, Santa Margherita Ligure, Italy, June 22-24, 1995 , D. A. Patterson, ...

  41. [52]

    A comparative evaluation of hardware-only and software-only directory protocols in shared-memory multiprocessors,

    H. Grahn and P. Stenstr¨ om, “A comparative evaluation of hardware-only and software-only directory protocols in shared-memory multiprocessors,” Journal of Systems Architecture , vol. 50, no. 9, pp. 537–561, Sep. 2004, issn: 13837621. doi: 10.1016/j.sysarc.2003.08. 014

  42. [53]

    SCC: A flexible architecture for many- core platform research,

    M. Gries, U. Hoffmann, M. Konow, and M. Riepen, “SCC: A flexible architecture for many- core platform research,” Computing in Science & Engineering, vol. 13, no. 6, pp. 79–83, Nov. 2011, issn: 1521-9615. doi: 10.1109/MCSE.2011.109

  43. [54]

    W. P. R. Group, Openpiton microarchitecture specification, Princeton University, 2016

  44. [55]

    Performance analysis and benchmarking of the intel SCC,

    P. Gschwandtner, T. Fahringer, and R. Prodan, “Performance analysis and benchmarking of the intel SCC,” in 2011 IEEE International Conference on Cluster Computing , Austin, TX, USA: IEEE, Sep. 2011, pp. 139–149, isbn: 978-1-4577-1355-2. doi: 10.1109/CLUSTER. 2011.24

  45. [56]

    Reducing memory and traffic requirements for scalable directory-based cache coherence schemes,

    A. Gupta, W. Weber, and T. C. Mowry, “Reducing memory and traffic requirements for scalable directory-based cache coherence schemes,” in Proceedings of the 1990 International Conference on Parallel Processing, Urbana-Champaign, IL, USA, August 1990. Volume 1: Architecture, B. ...

  46. [57]

    Integration of message passing and shared memory in the stanford FLASH multiprocessor,

    J. Heinlein, K. Gharachorloo, S. Dresser, and A. Gupta, “Integration of message passing and shared memory in the stanford FLASH multiprocessor,” in Proceedings of the sixth international conference on Architectural support for programming languages and operating systems - ASPL...

  47. [58]

    Hardware/software co-design of the stanford FLASH multiprocessor,

    M. Heinrich, D. Ofelt, M. Horowitz, and J. Hennessy, “Hardware/software co-design of the stanford FLASH multiprocessor,” Proceedings of the IEEE, vol. 85, no. 3, pp. 455–466, Mar. 1997, issn: 00189219. doi: 10.1109/5.558720

  48. [59]

    A quantitative analysis of the performance and scalability of distributed shared memory cache coherence protocols,

    M. Heinrich, V. Soundararajan, J. Hennessy, and A. Gupta, “A quantitative analysis of the performance and scalability of distributed shared memory cache coherence protocols,” IEEE Transactions on Computers , vol. 48, no. 2, pp. 205–217, Feb. 1999, issn: 00189340. doi: 10.1109/...

  49. [60]

    The performance and scalability of distributed shared-memory cache coher- ence protocols,

    M. Heinrich, “The performance and scalability of distributed shared-memory cache coher- ence protocols,” Ph.D. dissertation, 1998

  50. [61]

    The performance impact of flexibility in the stanford FLASH multiprocessor,

    M. Heinrich, M. Horowitz, A. Gupta, M. Rosenblum, J. Hennessy, J. Kuskin, D. Ofelt, J. Heinlein, J. Baxter, J. P. Singh, R. Simoni, K. Gharachorloo, and D. Nakahira, “The performance impact of flexibility in the stanford FLASH multiprocessor,” in Proceedings of the sixth inter...

  51. [62]

    Cooperative shared memory: Software and hardware support for scalable multiprocesors,

    M. D. Hill, J. R. Larus, S. K. Reinhardt, and D. A. Wood, “Cooperative shared memory: Software and hardware support for scalable multiprocesors,” in ASPLOS-V Proceedings - Fifth International Conference on Architectural Support for Programming Languages and Operating Systems, ...

  52. [63]

    Cooperative shared memory: Software and hardware for scalable multiprocessors,

    M. D. Hill, J. R. Larus, S. K. Reinhardt, and D. A. Wood, “Cooperative shared memory: Software and hardware for scalable multiprocessors,” ACM Transactions on Computer Sys- tems, vol. 11, no. 4, pp. 300–318, Nov. 1993, issn: 0734-2071, 1557-7333. doi: 10.1145/ 161541.161544

  53. [64]

    A 48-core IA-32 message-passing processor with DVFS in 45nm CMOS,

    J. Howard, S. Dighe, Y. Hoskote, S. Vangal, D. Finan, G. Ruhl, D. Jenkins, H. Wilson, N. Borkar, G. Schrom, F. Pailet, S. Jain, T. Jacob, S. Yada, S. Marella, P. Salihundam, V. Erraguntla, M. Konow, M. Riepen, G. Droege, J. Lindemann, M. Gries, T. Apel, K. 167 Henriss, T. Lund...

  54. [65]

    Ieee standard for systemverilog–unified hardware design, specification, and verification lan- guage,

    “Ieee standard for systemverilog–unified hardware design, specification, and verification lan- guage,” IEEE Std 1800-2017 (Revision of IEEE Std 1800-2012) , pp. 1–1315, 2018. doi: 10.1109/IEEESTD.2018.8299595

  55. [66]

    Inclusive cache

    “Inclusive cache.” (), [Online]. Available: https://github.com/chipsalliance/rocket- chip-inclusive-cache

  56. [67]

    IntelSCC

    Intel Corporation. “IntelSCC.” (), [Online]. Available: https://www.intel.cn/content/ dam/www/public/us/en/documents/technology- briefs/intel- labs- single- chip- platform-overview-paper.pdf

  57. [68]

    Distributed-directory scheme: Scalable coherent interface,

    D. James, A. Laundrie, S. Gjessing, and G. Sohi, “Distributed-directory scheme: Scalable coherent interface,” Computer, vol. 23, no. 6, pp. 74–77, Jun. 1990, issn: 0018-9162. doi: 10.1109/2.55503

  58. [69]

    Scalable, programmable and dense: The hammerblade open-source RISC-V manycore,

    D. C. Jung, M. Ruttenberg, P. Gao, S. Davidson, D. Petrisko, K. Li, A. K. Kamath, L. Cheng, S. Xie, P. Pan, Z. Zhao, Z. Yue, B. Veluri, S. Muralitharan, A. Sampson, A. Lumsdaine, Z. Zhang, C. Batten, M. Oskin, D. Richmond, and M. B. Taylor, “Scalable, programmable and dense: T...

  59. [70]

    Implementing a cache consistency protocol,

    R. H. Katz, S. J. Eggers, D. A. Wood, C. L. Perkins, and R. G. Sheldon, “Implementing a cache consistency protocol,” ACM SIGARCH Computer Architecture News , vol. 13, no. 3, pp. 276–283, Jun. 1985, issn: 0163-5964. doi: 10.1145/327070.327237

  60. [71]

    Optimizing co- herence traffic in manycore processors using closed-form caching/home agent mappings,

    S. Kommrusch, M. Horro, L.-N. Pouchet, G. Rodriguez, and J. Tourino, “Optimizing co- herence traffic in manycore processors using closed-form caching/home agent mappings,” IEEE Access, vol. 9, pp. 28 930–28 945, 2021, issn: 2169-3536. doi: 10.1109/ACCESS.2021. 3058280

  61. [72]

    Revisiting the complexity of hardware cache coherence and some implications,

    R. Komuravelli, S. V. Adve, and C.-T. Chou, “Revisiting the complexity of hardware cache coherence and some implications,” ACM Transactions on Architecture and Code Optimiza- tion, vol. 11, no. 4, pp. 1–22, Jan. 9, 2015, issn: 1544-3566, 1544-3973. doi: 10 . 1145 / 2663345

  62. [73]

    Post-silicon microarchitecture,

    C. Kumar, A. Chaudhary, S. Bhawalkar, U. Mathur, S. Jain, A. Vastrad, and E. Rotenberg, “Post-silicon microarchitecture,” IEEE Computer Architecture Letters, vol. 19, no. 1, pp. 26– 29, Jan. 1, 2020, issn: 1556-6056, 1556-6064, 2473-2575. doi: 10.1109/LCA.2020.2978841

  63. [74]

    Post- fabrication microarchitecture,

    C. Kumar, A. Seshadri, A. Chaudhary, S. Bhawalkar, R. Singh, and E. Rotenberg, “Post- fabrication microarchitecture,” in MICRO-54: 54th Annual IEEE/ACM International Sym- posium on Microarchitecture, Virtual Event Greece: ACM, Oct. 18, 2021, pp. 1270–1281, isbn: 978-1-4503-855...

  64. [75]

    Herov2: Full-stack open-source research platform for heterogeneous computing,

    A. Kurth, B. Forsberg, and L. Benini, “Herov2: Full-stack open-source research platform for heterogeneous computing,” IEEE Trans. Parallel Distributed Syst., vol. 33, no. 10, pp. 4368– 4382, 2022. doi: 10.1109/TPDS.2022.3189390

  65. [76]

    HERO: heterogeneous embedded research platform for exploring RISC-V manycore accelerators on FPGA,

    A. Kurth, P. Vogel, A. Capotondi, A. Marongiu, and L. Benini, “HERO: heterogeneous embedded research platform for exploring RISC-V manycore accelerators on FPGA,”CoRR, vol. abs/1712.06497, 2017. arXiv: 1712.06497

  66. [77]

    The stan- ford FLASH multiprocessor,

    J. Kuskin, D. Ofelt, M. Heinrich, J. Heinlein, R. Simoni, K. Gharachorloo, J. Chapin, D. Nakahira, J. Baxter, M. Horowitz, A. Gupta, M. Rosenblum, and J. Hennessy, “The stan- ford FLASH multiprocessor,” in Proceedings of 21 International Symposium on Computer 168 Architecture,...

  67. [78]

    THE FLASH MULTIPROCESSOR: DESIGNING a FLEXIBLE AND SCAL- ABLE SYSTEM,

    J. Kuskin, “THE FLASH MULTIPROCESSOR: DESIGNING a FLEXIBLE AND SCAL- ABLE SYSTEM,” Ph.D. dissertation, 1997

  68. [79]

    COMIC: A coherent shared memory interface for cell be,

    J. Lee, S. Seo, C. Kim, J. Kim, P. Chun, Z. Sura, J. Kim, and S. Han, “COMIC: A coherent shared memory interface for cell be,” in Proceedings of the 17th international conference on Parallel architectures and compilation techniques , Toronto Ontario Canada: ACM, Oct. 25, 2008,...

  69. [80]

    An OpenCL framework for homogeneous manycores with no hardware cache coherence,

    J. Lee, J. Kim, J. Kim, S. Seo, and J. Lee, “An OpenCL framework for homogeneous manycores with no hardware cache coherence,” in 2011 International Conference on Parallel Architectures and Compilation Techniques, Galveston, TX, USA: IEEE, Oct. 2011, pp. 56–

  70. [81]

    doi: 10.1109/PACT.2011.12

  71. [82]

    The stanford dash multiprocessor,

    D. Lenoski, J. Laudon, K. Gharachorloo, W.-D. Weber, A. Gupta, J. Hennessy, M. Horowitz, and M. Lam, “The stanford dash multiprocessor,” Computer, vol. 25, no. 3, pp. 63–79, Mar. 1992, issn: 0018-9162. doi: 10.1109/2.121510

  72. [83]

    Cifer: A cache-coherent 12-nm 16-mm2 soc with four 64-bit risc-v application cores, 18 32-bit risc-v compute cores, and a 1541 lut6/mm2 synthesizable efpga,

    A. Li, T.-J. Chang, F. Gao, T. Ta, G. Tziantzioulis, Y. Ou, M. Wang, J. Tu, K. Xu, P. Jackson, A. Ning, G. Chirkov, M. Orenes-Vera, S. Agwa, X. Yan, E. Tang, J. Balkind, C. Batten, and D. Wentzlaff, “Cifer: A cache-coherent 12-nm 16-mm2 soc with four 64-bit risc-v application ...

  73. [84]

    Sting: A CC-NUMA computer system for the commercial marketplace,

    T. Lovett and R. M. Clapp, “Sting: A CC-NUMA computer system for the commercial marketplace,” in Proceedings of the 23rd Annual International Symposium on Computer Architecture, Philadelphia, PA, USA, May 22-24, 1996 , J. Baer, Ed., ACM, 1996, pp. 308–

  74. [85]

    Token coherence: Decoupling performance and correct- ness,

    M. Martin, M. Hill, and D. Wood, “Token coherence: Decoupling performance and correct- ness,” in 30th Annual International Symposium on Computer Architecture, 2003. Proceed- ings., San Diego, CA, USA: IEEE Comput. Soc, 2003, pp. 182–193, isbn: 978-0-7695-1945-6. doi: 10.1109/I...

  75. [86]

    Agile soc development with open ESP : Invited paper,

    P. Mantovani, D. Giri, G. D. Guglielmo, L. Piccolboni, J. Zuckerman, E. G. Cota, M. Petracca, C. Pilato, and L. P. Carloni, “Agile soc development with open ESP : Invited paper,” in IEEE/ACM International Conference On Computer Aided Design, ICCAD 2020, San Diego, CA, USA, Nov...

  76. [87]

    Integrating performance monitoring and com- munication in parallel computers,

    M. Martonosi, D. Ofelt, and M. Heinrich, “Integrating performance monitoring and com- munication in parallel computers,” in Proceedings of the 1996 ACM SIGMETRICS inter- national conference on Measurement and modeling of computer systems - SIGMETRICS ’96, Philadelphia, Pennsyl...

  77. [88]

    Why on-chip cache coherence is here to stay,

    M. M. K. Martin, M. D. Hill, and D. J. Sorin, “Why on-chip cache coherence is here to stay,” Communications of the ACM , vol. 55, no. 7, pp. 78–89, Jul. 2012, issn: 0001-0782, 1557-7317. doi: 10.1145/2209249.2209269

  78. [89]

    Flask coherence: A morphable hybrid coher- ence protocol to balance energy, performance and scalability,

    L. G. Menezo, V. Puente, and J.-A. Gregorio, “Flask coherence: A morphable hybrid coher- ence protocol to balance energy, performance and scalability,” in 2015 IEEE 21st Interna- 169 tional Symposium on High Performance Computer Architecture (HPCA) , Burlingame, CA, USA: IEEE,...

  79. [90]

    Rainbow: A composable coherence proto- col for ¡span style=

    L. G. Menezo, V. Puente, and J. A. Gregorio, “Rainbow: A composable coherence proto- col for ¡span style=”font-variant:small-caps;”¿multi-chip¡/span¿ servers,” Concurrency and Computation: Practice and Experience, vol. 32, no. 24, Dec. 25, 2020, issn: 1532-0626, 1532-

  80. [91]

    Coherence controller architectures for scalable shared-memory multiprocessors,

    M. Michael, A. Nanda, and Beng-Hong Lim, “Coherence controller architectures for scalable shared-memory multiprocessors,” IEEE Transactions on Computers, vol. 48, no. 2, pp. 245– 255, Feb. 1999, issn: 00189340. doi: 10.1109/12.752666

  81. [92]

    Coherence controller architectures for SMP-based CC-NUMA multiprocessors,

    M. M. Michael, A. K. Nanda, B.-H. Lim, and M. L. Scott, “Coherence controller architectures for SMP-based CC-NUMA multiprocessors,” ACM SIGARCH Computer Architecture News, vol. 25, no. 2, pp. 219–228, May 1997, issn: 0163-5964. doi: 10.1145/384286.264203

  82. [93]

    A coherence-capable write-back l1 data cache for ariane,

    M. Miceli, “A coherence-capable write-back l1 data cache for ariane,” Politecnico di Torino, 2023

  83. [94]

    Cramming more components onto integrated circuits,

    G. E. Moore, “Cramming more components onto integrated circuits,” Electronics, vol. 38, no. 8, pp. 114–117, 1965

  84. [95]

    Cramming more components onto integrated circuits,

    G. E. Moore, “Cramming more components onto integrated circuits,” Proc. IEEE, vol. 86, no. 1, pp. 82–85, 1998. doi: 10.1109/JPROC.1998.658762

  85. [96]

    A timestamp-based cache coherence scheme,

    S. L. Min and J. Baer, “A timestamp-based cache coherence scheme,” in Proceedings of the International Conference on Parallel Processing, ICPP ’89, The Pennsylvania State University, University Park, PA, USA, August 1989. Volume 1: Architecture , Pennsylvania State University ...

  86. [97]

    First draft of a report on the EDVAC,

    J. von Neumann, “First draft of a report on the EDVAC,” IEEE Ann. Hist. Comput., vol. 15, no. 4, pp. 27–75, 1993. doi: 10.1109/85.238389

  87. [98]

    The s3.mp scalable shared memory multiprocessor,

    A. Nowatzyk, G. Aybay, M. Browne, E. Kelly, D. Lee, and M. Parkin, “The s3.mp scalable shared memory multiprocessor,” in Proceedings of the Twenty-Seventh Hawaii International Conference on System Sciences HICSS-94 , Wailea, HI, USA: IEEE Comput. Soc. Press, 1994, pp. 144–153,...

  88. [99]

    Nagarajan, D

    V. Nagarajan, D. J. Sorin, M. D. Hill, and D. A. Wood, A Primer on Memory Consistency and Cache Coherence (Synthesis Lectures on Computer Architecture). Springer Interna- tional Publishing, 2020. doi: 10.1007/978-3-031-01764-3

  89. [100]

    Exploiting parallelism in cache coherency protocol engines,

    A. Nowatzyk, G. Aybay, M. C. Browne, E. J. Kelly, M. Parkin, B. Radke, and S. Vishin, “Exploiting parallelism in cache coherency protocol engines,” in Euro-Par ’95 Parallel Pro- cessing, First International Euro-Par Conference, Stockholm, Sweden, August 29-31, 1995, Proceeding...

  90. [101]

    Opensbi

    OpenSBI. “Opensbi.” (), [Online]. Available: https : / / github . com / riscv - software - src/opensbi

  91. [102]

    The s3.mp architecture: A local area multiprocessor,

    A. Nowatzyk, M. Monger, M. Parkin, E. Kelly, M. Browne, G. Aybay, and D. Lee, “The s3.mp architecture: A local area multiprocessor,” in Proceedings of the fifth annual ACM symposium on Parallel algorithms and architectures - SPAA ’93 , Velen, Germany: ACM Press, 1993, pp. 140–...

  92. [103]

    Fast and efficient automatic memory management for GPUs using compiler-assisted runtime coherence scheme,

    S. Pai, R. Govindarajan, and M. J. Thazhuthaveetil, “Fast and efficient automatic memory management for GPUs using compiler-assisted runtime coherence scheme,” in Proceedings of the 21st international conference on Parallel architectures and compilation techniques , Minneapoli...

  93. [104]

    A low-overhead coherence solution for multiprocessors with private cache memories,

    M. S. Papamarcos and J. H. Patel, “A low-overhead coherence solution for multiprocessors with private cache memories,” in Proceedings of the 11th annual international symposium 170 on Computer architecture - ISCA ’84 , Not Known: ACM Press, 1984, pp. 348–354, isbn: 978-0-8186-...

  94. [105]

    OpenSPARC t1 micro architecture specification, 2008

  95. [106]

    Exploiting transition locality in automatic verification of finite-state concurrent systems,

    G. D. Penna, B. Intrigila, I. Melatti, E. Tronci, and M. V. Zilli, “Exploiting transition locality in automatic verification of finite-state concurrent systems,” International Journal on Software Tools for Technology Transfer , vol. 6, no. 4, pp. 320–341, Jul. 2004. doi: 10. 1...

  96. [107]

    Blackparrot: An agile open-source RISC-V multicore for accelerator socs,

    D. Petrisko, F. Gilani, M. Wyse, D. C. Jung, S. Davidson, P. Gao, C. Zhao, Z. Azad, S. Canakci, B. Veluri, T. Guarino, A. Joshi, M. Oskin, and M. B. Taylor, “Blackparrot: An agile open-source RISC-V multicore for accelerator socs,” IEEE Micro, vol. 40, no. 4, pp. 93–102,

  97. [108]

    Paulin, P

    G. Paulin, P. Scheffler, T. Benz, M. A. Cavalcante, T. Fischer, M. Eggimann, Y. Zhang, N. Wistoff, L. Bertaccini, L. Colagrande, G. Ottavi, F. K. G¨ urkaynak, D. Rossi, and L. Benini, “Occamy: A 432-core 28.1 dp-gflop/s/w 83% FPU utilization dual-chiplet, dual- hbm2e risc-v-ba...

  98. [109]

    Tempest and typhoon: User-level shared memory,

    S. Reinhardt, J. Larus, and D. Wood, “Tempest and typhoon: User-level shared memory,” in Proceedings of 21 International Symposium on Computer Architecture, Chicago, IL, USA: IEEE Comput. Soc. Press, 1994, pp. 325–336, isbn: 978-0-8186-5510-4. doi: 10.1109/ISCA. 1994.288138

  99. [110]

    Splash-3: A properly synchronized benchmark suite for contemporary research,

    C. Sakalis, C. Leonardsson, S. Kaxiras, and A. Ros, “Splash-3: A properly synchronized benchmark suite for contemporary research,” in 2016 IEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2016, Uppsala, Sweden, April 17- 19, 2016, IEEE Compu...

  100. [111]

    SCD: A scalable coherence directory with flexible sharer set encoding,

    D. Sanchez and C. Kozyrakis, “SCD: A scalable coherence directory with flexible sharer set encoding,” in IEEE International Symposium on High-Performance Comp Architecture , New Orleans, LA, USA: IEEE, Feb. 2012, pp. 1–12. doi: 10.1109/HPCA.2012.6168950

  101. [112]

    SWEL: Hardware cache coherence protocols to map shared data onto shared caches,

    S. H. Pugsley, J. B. Spjut, D. W. Nellans, and R. Balasubramonian, “SWEL: Hardware cache coherence protocols to map shared data onto shared caches,” in Proceedings of the 19th international conference on Parallel architectures and compilation techniques , Vienna Austria: ACM, ...

  102. [113]

    T¨ ak¯ o: A polymorphic cache hierarchy for general-purpose optimization of data movement,

    B. C. Schwedock, P. Yoovidhya, J. Seibert, and N. Beckmann, “T¨ ak¯ o: A polymorphic cache hierarchy for general-purpose optimization of data movement,” in Proceedings of the 49th Annual International Symposium on Computer Architecture , New York New York: ACM, Jun. 18, 2022, ...

  103. [114]

    A fully programmable FSM-based process- ing engine for gigabytes/s header parsing,

    K. Septinus, P. Pirsch, H. Blume, and U. Mayer, “A fully programmable FSM-based process- ing engine for gigabytes/s header parsing,” in 2010 International Conference on Embedded Computer Systems: Architectures, Modeling and Simulation, Samos, Greece: IEEE, Jul. 2010, pp. 45–54...

  104. [115]

    SiFive, SiFive TileLink Specification, Version 1.8.1 , 2020. 171

  105. [116]

    Fine- grain access control for distributed shared memory,

    I. Schoinas, B. Falsafi, A. R. Lebeck, S. K. Reinhardt, J. R. Larus, and D. A. Wood, “Fine- grain access control for distributed shared memory,” in Proceedings of the sixth interna- tional conference on Architectural support for programming languages and operating systems - AS...

  106. [117]

    A class of compatible cache consistency protocols and their support by the IEEE futurebus,

    P. Sweazey and A. J. Smith, “A class of compatible cache consistency protocols and their support by the IEEE futurebus,” ACM SIGARCH Computer Architecture News , vol. 14, no. 2, pp. 414–423, May 1986, issn: 0163-5964. doi: 10.1145/17356.17404

  107. [118]

    Prodigy: Improving the memory latency of data-indirect irregular workloads using hardware-software co-design,

    N. Talati, K. May, A. Behroozi, Y. Yang, K. Kaszyk, C. Vasiladiotis, T. Verma, L. Li, B. Nguyen, J. Sun, J. M. Morton, A. Ahmadi, T. M. Austin, M. F. P. O’Boyle, S. A. Mahlke, T. N. Mudge, and R. G. Dreslinski, “Prodigy: Improving the memory latency of data-indirect irregular ...

  108. [119]

    Cache system design in the tightly coupled multiprocessor system,

    C. K. Tang, “Cache system design in the tightly coupled multiprocessor system,” in Pro- ceedings of the June 7-10, 1976, national computer conference and exposition on - AFIPS ’76, New York, New York: ACM Press, 1976, p. 749. doi: 10.1145/1499799.1499901

  109. [120]

    Specifying and verifying a broadcast and a multicast snooping cache coherence protocol,

    D. Sorin, M. Plakal, A. Condon, M. Hill, M. Martin, and D. Wood, “Specifying and verifying a broadcast and a multicast snooping cache coherence protocol,” IEEE Transactions on Parallel and Distributed Systems , vol. 13, no. 6, pp. 556–578, Jun. 2002, issn: 1045-9219. doi: 10.1...

  110. [121]

    Basejump STL: systemverilog needs a standard template library for hardware design,

    M. B. Taylor, “Basejump STL: systemverilog needs a standard template library for hardware design,” in Proceedings of the 55th Annual Design Automation Conference, DAC 2018, San Francisco, CA, USA, June 24-29, 2018 , ACM, 2018, 73:1–73:6. doi: 10.1145/3195970. 3199848

  111. [122]

    Tedeschi, L

    R. Tedeschi, L. Valente, G. Ottavi, E. Zelioli, N. Wistoff, M. Giacometti, A. B. Sajjad, L. Benini, and D. Rossi, Culsans: An efficient snoop-based coherency unit for the cva6 open source risc-v application processor, 2024. arXiv: 2407.19895 [eess.SY]

  112. [123]

    Firefly: A multiprocessor workstation,

    C. P. Thacker and L. C. Stewart, “Firefly: A multiprocessor workstation,” in Proceedings of the second international conference on Architectual support for programming languages and operating systems, Palo Alto California USA: ACM, Oct. 1987, pp. 164–172. doi: 10.1145/ 36206.36199

  113. [124]

    ADir/sub p/NB: A cost-effective way to implement full map directory- based cache coherence protocols,

    Tao Li and L. John, “ADir/sub p/NB: A cost-effective way to implement full map directory- based cache coherence protocols,” IEEE Transactions on Computers, vol. 50, no. 9, pp. 921– 934, Sep. 2001, issn: 00189340. doi: 10.1109/12.954507

  114. [125]

    A case for second-level software cache coherency on many-core accelerators,

    A. Vianes, F. Petrot, and F. Rousseau, “A case for second-level software cache coherency on many-core accelerators,” in 2022 IEEE International Workshop on Rapid System Proto- typing (RSP), Shanghai, China: IEEE, Oct. 13, 2022, pp. 29–35. doi: 10.1109/RSP57251. 2022.10038999

  115. [126]

    Efficiently supporting dynamic task parallelism on heterogeneous cache-coherent systems,

    M. Wang, T. Ta, L. Cheng, and C. Batten, “Efficiently supporting dynamic task parallelism on heterogeneous cache-coherent systems,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA) , Valencia, Spain: IEEE, May 2020, pp. 173– 186, isbn: 978-1...

  116. [127]

    A programmable state machine architecture for packet processing,

    Wangyang Lai and Chin-Tau Lea, “A programmable state machine architecture for packet processing,” IEEE Micro , vol. 23, no. 4, pp. 32–42, Jul. 2003. doi: 10 . 1109 / MM . 2003 . 1225965

  117. [128]

    Torvalds

    L. Torvalds. “Linux.” (), [Online]. Available: https://github.com/torvalds/linux

  118. [129]

    The RISC-V instruc- tion set manual, volume i: User-level ISA, version 2.0,

    A. Waterman, Y. Lee, R. Avizienis, D. A. Patterson, and K. Asanovi´ c, “The RISC-V instruc- tion set manual, volume i: User-level ISA, version 2.0,” University of California, Berkeley, Tech. Rep. UCB/EECS-2016-118, May 2016. 172

  119. [130]

    Scalable directories for cache-coherent shared-memory multiprocessors,

    W.-D. Weber, “Scalable directories for cache-coherent shared-memory multiprocessors,” Ph.D. dissertation, Stanford University, 1993

  120. [131]

    The SPLASH-2 programs: Characterization and methodological considerations,

    S. C. Woo, M. Ohara, E. Torrie, J. P. Singh, and A. Gupta, “The SPLASH-2 programs: Characterization and methodological considerations,” in Proceedings of the 22nd Annual International Symposium on Computer Architecture, ISCA ’95, Santa Margherita Ligure, Italy, June 22-24, 199...

  121. [132]

    Design of the RISC-V instruction set architecture,

    A. Waterman, “Design of the RISC-V instruction set architecture,” Ph.D. dissertation, Uni- versity of California, Berkeley, USA, 2016

  122. [133]

    The bedrock cache coherence protocol and system v1.1

    M. Wyse. “The bedrock cache coherence protocol and system v1.1.” (2022), [Online]. Avail- able: https://github.com/black-parrot/black-parrot/blob/master/docs/bedrock_ protocol_specification.pdf

  123. [134]

    M. Wyse, D. Petrisko, F. Gilani, Y.-M. Chueh, P. Gao, D. C. Jung, S. Muralitharan, S. V. Ranga, M. Oskin, and M. Taylor, The blackparrot bedrock cache coherence system , 2022. arXiv: 2211.06390 [cs.AR]

  124. [135]

    Virtual channels and multiple physical networks: Two alternatives to improve noc performance,

    Y. Yoon, N. Concer, M. Petracca, and L. P. Carloni, “Virtual channels and multiple physical networks: Two alternatives to improve noc performance,” IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. , vol. 32, no. 12, pp. 1906–1919, 2013. doi: 10 . 1109 / TCAD . 2013 . 2276399

  125. [136]

    Mechanisms for cooperative shared mem- ory,

    D. A. Wood, S. K. Reinhardt, S. Chandra, B. Falsafi, M. D. Hill, J. R. Larus, A. R. Lebeck, J. C. Lewis, S. S. Mukherjee, and S. Palacharla, “Mechanisms for cooperative shared mem- ory,” in Proceedings of the 20th annual international symposium on Computer architecture - ISCA ...

  126. [137]

    Manticore: A 4096-core RISC-v chiplet architecture for ultraefficient floating-point computing,

    F. Zaruba, F. Schuiki, and L. Benini, “Manticore: A 4096-core RISC-v chiplet architecture for ultraefficient floating-point computing,” IEEE Micro, vol. 41, no. 2, pp. 36–42, Mar. 1, 2021, issn: 0272-1732, 1937-4143. doi: 10.1109/MM.2020.3045564

  127. [138]

    RnR: A software-assisted record-and-replay hard- ware prefetcher,

    C. Zhang, Y. Zeng, J. Shalf, and X. Guo, “RnR: A software-assisted record-and-replay hard- ware prefetcher,” in 2020 53rd Annual IEEE/ACM International Symposium on Microarchi- tecture (MICRO), Athens, Greece: IEEE, Oct. 2020, pp. 609–621, isbn: 978-1-72817-383-2. doi: 10.1109...

  128. [139]

    Fractal coherence: Scalably verifiable cache co- herence,

    M. Zhang, A. R. Lebeck, and D. J. Sorin, “Fractal coherence: Scalably verifiable cache co- herence,” in 2010 43rd Annual IEEE/ACM International Symposium on Microarchitecture , Atlanta, GA, USA: IEEE, Dec. 2010, pp. 471–482, isbn: 978-1-4244-9071-4. doi: 10.1109/ MICRO.2010.11

  129. [140]

    The cost of application-class processing: Energy and performance analysis of a linux-ready 1.7-GHz 64-bit RISC-v core in 22-nm FDSOI technology,

    F. Zaruba and L. Benini, “The cost of application-class processing: Energy and performance analysis of a linux-ready 1.7-GHz 64-bit RISC-v core in 22-nm FDSOI technology,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , vol. 27, no. 11, pp. 2629– 2640, Nov. ...

  130. [141]

    J. Zhao, B. Korpan, A. Gonzalez, and K. Asanovic, Sonicboom: The 3rd generation berkeley out-of-order machine, May 2020

  131. [142]

    Software assistance for directory-based caches,

    Zhiyuan Li, “Software assistance for directory-based caches,” in Proceedings of 8th Inter- national Parallel Processing Symposium, Cancun, Mexico: IEEE Comput. Soc. Press, 1994, pp. 151–157, isbn: 978-0-8186-5602-6. doi: 10.1109/IPPS.1994.288307

  132. [143]

    A programmable co-processor for profiling,

    C. Zilles and G. Sohi, “A programmable co-processor for profiling,” in Proceedings HPCA Seventh International Symposium on High-Performance Computer Architecture , Monterrey, 173 Mexico: IEEE Comput. Soc, 2001, pp. 241–252, isbn: 978-0-7695-1019-4. doi: 10 . 1109 / HPCA.2001.903267

  133. [144]

    Constellation: An open-source soc- capable noc generator,

    J. Zhao, A. Agrawal, B. Nikolic, and K. Asanovi´ c, “Constellation: An open-source soc- capable noc generator,” in 2022 15th IEEE/ACM International Workshop on Network on Chip Architectures (NoCArc), IEEE, 2022, pp. 1–7

  134. [148]

    PULP Platform

    E. Z¨ urich. “PULP Platform.” (2013), [Online]. Available: https://pulp-platform.org/ index.html. 174

  135. [317]

    doi: 10.1145/232973.233006

  136. [634]

    doi: 10.1002/cpe.5947

  137. [2020]

    doi: 10.1109/MM.2020.2996145

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.