REVIEW 3 major objections 4 minor 145 references
The Open-Source BlackParrot-BedRock Cache Coherence System
T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Programmable cache coherence can match fixed-function engines
desk verdict A solid engineering dissertation that delivers an open-source programmable coherence system with real measurements, though the headline performance claim is not tested at the protocol's serialization bottleneck. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central objects are the way group and the pending bit. A way group collects one tag set from every cache in the system for a single cache set, and the pending bit allows only one active coherence transaction per way group. Because the CCE is the only component that changes coherence state and it always waits for the coherence acknowledgment before clearing the pending bit, the protocol never exposes transient states, which dramatically simplifies verification and the cache controller logic. The microcode-programmable CCE uses a small base ISA plus coherence-specific instructions for flag, directory, and queue operations to process BedRock's messages, while the hybrid CCE runs a fixed-function request pipe and a programmable pipe side by side. This machinery is what lets the dissertation compare protocol-processing occupancy directly across the three designs.
What would settle it
Run a many-core workload whose address stream is deliberately hashed so that a small number of way groups receive all the coherence traffic, then measure request throughput against the same workload with addresses hashed evenly. If the programmable or hybrid CCE's throughput collapses relative to the fixed-function design under way-group contention, the programmability-without-overhead claim fails for that regime.
Extended reading notes
Core claim
The central claim is that adding programmability to the cache coherence directory does not inherently require significant performance or area cost. The evidence is the BP-BedRock system, in which a microcode-programmable CCE and a hybrid CCE both achieve request-processing occupancy and benchmark performance comparable to the fixed-function FSM CCE, while occupying minimal additional area relative to the full multicore. BedRock itself is the enabling protocol design: it is a directory-based invalidate protocol using MOESIF states with no transient coherence states, because the directory is the sole arbiter of state changes and one pending bit per way group serializes all transactions touching that way group. The directory's full-duplicate tag sets give it exact knowledge of every cached block, so cache-to-cache transfers and upgrades can be scheduled without races. Together the protocol and the three engines demonstrate that programmable coherence is implementable and competitive, not just theoretically attractive.
Load-bearing premise
The load-bearing premise is that serializing all coherence transactions touching the same way group with a single pending bit does not become a performance bottleneck when many cores miss to addresses that map to the same group; if that serialization ever becomes the bottleneck, the programmable engines' competitive standing is lost.
Editorial extensions
If this is right
- A protocol without transient states verifies dramatically faster: CMurphi checks the BedRock MESI variant up to 66x faster than a traditional MESI protocol with six caches, completing in hours where the traditional model would take months.
- Because the duplicate-tag directory has a constant 6.25% storage overhead relative to L1 capacity regardless of core count, the directory can be tiled across cores with a fixed per-tile size.
- A microcode-programmable coherence engine can match fixed-function request occupancy on common MOESIF protocol paths, so a chip could ship one programmable directory and later change the protocol without new hardware.
- The hybrid CCE offers a path to fixed-function performance on common operations and programmable flexibility for uncommon or system-specific operations in the same directory.
- All three designs are open source, giving other researchers a working multicore platform on which to prototype and evaluate coherence features.
Reading between the lines
- If the way-group serialization premise holds under adversarial access patterns, the same pending-bit mechanism could be reused for lightweight ordered or speculative transactions in the coherence system, an extension the dissertation does not explore.
- Because the duplicate-tag directory overhead is constant while complete-directory overhead grows with core count, programmable duplicate-tag directories may become comparatively cheaper in many-core designs than traditional directory organizations.
- A direct test of the thesis would be to program the ucode CCE with a non-MOESIF protocol, such as a coherence scheme tailored to accelerators, and measure whether the occupancy advantage persists; this is an extension beyond the paper's claims.
- The performance-competitiveness result is demonstrated for Splash-3 and microbenchmark workloads, not for worst-case hashing, so the claim should be read as holding for typical rather than adversarial way-group contention.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The dissertation presents the BedRock directory-based MOESIF cache coherence protocol and three open-source implementations within the BlackParrot RISC-V multicore: a fixed-function FSM coherence directory engine, a microcode-programmable engine, and a hybrid of the two. The central claim, stated in Section 1.2 and revisited in Chapters 4 and 5, is that programmability can be added to the coherence system without significant performance or area overhead. Evidence includes exhaustive protocol tables and state-transition diagrams, occupancy tables for each engine, FPGA resource measurements, and Splash-3 and microbenchmark execution comparisons. The work also provides a protocol-complexity comparison against a canonical directory protocol, including CMurphi verification results for the MESI variant and a mathematical latency model for both protocols.
Significance. If the central claim holds, this is a solid, reproducible systems contribution: it ships an open-source, well-documented programmable coherence engine inside a real multicore, with machine-readable protocol artifacts and clear occupancy/area accounting. The three-engine comparison is a useful data point for a long-standing architectural question. The work is explicit about its parameter-free methodology, and the open-source release (Section 4) makes the measurements independently checkable. The main significance is therefore as an engineering demonstration and an open research platform, not as a new protocol concept; the protocol itself is a simplified directory protocol whose novelty lies in the elimination of exposed transient states and directory-controlled replacements.
major comments (3)
- [§4.3.1, §4.6, Tables 4.10/4.18/4.19] The performance-competitiveness claim is not tested where the protocol's serialization point is exposed. Section 4.3.1 states that only one coherence transaction per way group may be active, enforced by a pending bit, and that all requests to that way group stall at the directory. The occupancy comparison in Tables 4.10, 4.18, and 4.19 shows the ucode and hybrid engines have higher per-request occupancy than the FSM engine, but the evaluation in Section 4.6 and Chapter 5 reports only aggregate Splash-3 and microbenchmark execution times, without way-group conflict counts, pending-bit stall cycles, or access-pattern variance. Under realistic or adversarial contention on a single way group, the higher occupancy directly becomes lower throughput. This is a load-bearing gap because the thesis's central claim is parity with the fixed-function design; the manuscript should either measure contention directly or qualify the claim to no-contention and low-contention regimes.
- [§3.4, Table 3.10; §4.4–4.5] The verification evidence covers only the MESI variant, while the implemented engines run the MOESIF protocol. Table 3.10 reports CMurphi verification for BedRock MESI, and the text in Section 3.4 says only the MESI protocol has been verified. The BP-BedRock implementations described in Chapter 4 and 5 execute the MOESIF protocol with the O and F states (Tables 3.7, 4.18, and related state tables). Since O/F state interactions are precisely where directory-controlled transfers and writebacks are most complex, the correctness claim for the shipped design is not established by the reported verification. The authors should either verify the MOESIF variant or provide a rigorous argument for why MESI verification transfers to MOESIF, and should state this limitation explicitly where correctness is claimed.
- [§5.3, Figure 5.11] The hybrid CCE performance comparison in Chapter 5 is presented as a single aggregate result without reporting variance across runs or the configuration parameters (e.g., number of cores, cache sizes, network widths) used. Since the hybrid design's programmable pipe is the key new contribution, the evaluation should show the occupancy and execution-time relationship for the specific requests that use the programmable pipe, not only the blended result. This would make the claimed 'performance parity' of the hybrid design falsifiable and would address the way-group serialization concern raised above.
minor comments (4)
- [§1.2] The sentence 'programmability can be be introduced to the cache coherence system' contains a duplicated 'be'; please fix.
- [§3.5.5, Tables 3.14–3.15] The mathematical models use abbreviations such as 'M em', 'F ill', and 'Ack' without a legend; add a short note defining each symbol and explaining why the 'Ack' term is counted as latency for BedRock but the requester can overlap it with execution.
- [§4.3.3, Figure 4.15] The caption of Figure 4.15 contains 'T able' instead of 'Table'; also, the y-axis label 'Directory Storage Overhead' should specify the normalization basis (L1 cache capacity) directly in the axis title.
- [§4.2.2, Table 4.2] The description of the 'Set State & Transfer & Writeback (dirty)' occupancy of '2 + (2*N)' cycles is clear, but the table would benefit from an explicit note that N is the cache block width divided by the fill width, as defined in Section 4.2.2.
Circularity Check
No significant circularity: the central performance and area claims rest on direct measurements of open-source implementations, on internal occupancy tables, and on model-checking runs, not on fitted parameters or load-bearing self-citations.
full rationale
Every load-bearing result in this dissertation is evaluated empirically or by deterministic analysis within the document itself. The central comparison among the fixed-function FSM CCE, the microcode-programmable ucode CCE, and the hybrid CCE is supported by request-occupancy tables (Tables 4.10, 4.18, 4.19, 5.2, 5.5), Splash-3 normalized execution time (Figure 4.22), FPGA resource utilization (Appendix E), and microbenchmark measurements (Section 5.3). These are direct measurements of the described designs; no parameter is fitted to force performance parity, and no 'prediction' is obtained by renaming a fitted constant. The BedRock protocol description in Chapter 3 is a specification rather than a derived result: the comparisons to a canonical directory protocol are based on explicit state-transition tables, message-equivalency analysis (Table 3.13), and CMurphi model-checking runs (Section 3.4), all self-contained or machine-checked. Self-citations to BlackParrot and the BedRock protocol specification identify the prior platform and protocol, but the new contribution--three coherence-engine implementations and their measured trade-offs--does not reduce to those citations. The weakest point noted by a skeptical reader is the way-group serialization assumption (Section 4.3.1), which may limit the generality of the performance-competitiveness claim under adversarial contention; that is an evaluation-coverage concern, not a circular dependency. No circular step meeting the quoted-evidence standard was found.
Assumptions & free parameters
assumptions (4)
- domain assumption Coherence networks deliver messages error-free and may be unordered.
- domain assumption Cache controllers process coherence commands atomically as indivisible operations.
- ad hoc to paper Only one coherence transaction per way group is active at a time, enforced by pending bits.
- standard math The CMurphi verification model with a single cache block and a single directory is sufficient for coherence correctness.
Cite this review
Pith. "Pith review of The Open-Source BlackParrot-BedRock Cache Coherence System." pith.science (2026). https://pith.science/paper/TPJWPJD5
@misc{pith2026250500962,
author = {Pith},
title = {Pith review of: The Open-Source BlackParrot-BedRock Cache Coherence System},
year = {2026},
howpublished = {\url{https://pith.science/paper/TPJWPJD5}},
note = {Machine review of arXiv:2505.00962}
}
read the original abstract
This dissertation revisits the topic of programmable cache coherence engines in the context of modern shared-memory multicore processors. First, the open-source BedRock cache coherence protocol is described. BedRock employs the canonical MOESIF coherence states and reduces implementation burden by eliminating transient coherence states from the protocol. The protocol's design complexity, concurrency, and verification effort are analyzed and compared to a canonical directory-based invalidate coherence protocol. Second, the architecture and microarchitecture of three separate cache coherence directories implementing the BedRock protocol within the BlackParrot 64-bit RISC-V multicore processor, collectively called BlackParrot-BedRock (BP-BedRock), are described. A fixed-function coherence directory engine implementation provides a baseline design for performance and area comparisons. A microcode-programmable coherence directory implementation demonstrates the feasibility of implementing a programmable coherence engine capable of maintaining sufficient protocol processing performance. A hybrid fixed-function and programmable coherence directory blends the protocol processing performance of the fixed-function design with the programmable flexibility of the microcode-programmable design. Collectively, the BedRock coherence protocol and its three BP-BedRock implementations demonstrate the feasibility and challenges of including programmable logic within the coherence system of modern shared-memory multicore processors, paving the way for future research into the application- and system-level benefits of programmable coherence engines.
Figures
Figures from the paper (93 more)
Reference graph
Works this paper leans on
-
[1]
Comparison of hardware and software cache coherence schemes,
S. V. Adve, V. S. Adve, M. D. Hill, and M. K. Vernon, “Comparison of hardware and software cache coherence schemes,” in Proceedings of the 18th annual international symposium on Computer architecture - ISCA ’91 , Toronto, Ontario, Canada: ACM Press, 1991, pp. 298– 308, isbn: 978-0-89791-394-2. doi: 10.1145/115952.115982
arXiv 1991
-
[2]
A. Agarwal, R. Bianchini, D. Chaiken, F. Chong, K. Johnson, D. Kranz, J. Kubiatowicz, Beng-Hong Lim, K. Mackenzie, and D. Yeung, “The MIT alewife machine,” Proceedings of the IEEE, vol. 87, no. 3, pp. 430–444, Mar. 1999, issn: 00189219. doi: 10.1109/5.747864
-
[3]
An evaluation of directory schemes for cache coherence,
A. Agarwal, R. Simoni, J. Hennessy, and M. Horowitz, “An evaluation of directory schemes for cache coherence,” in [1988] The 15th Annual International Symposium on Computer Architecture. Conference Proceedings, Honolulu, HI, USA: IEEE Comput. Soc. Press, 1988, pp. 280–289, isbn: 978-0-8186-0861-2. doi: 10.1109/ISCA.1988.5238
arXiv 1988
-
[4]
The MIT alewife machine: Architecture and performance,
A. Agarwal, R. Bianchini, D. Chaiken, K. L. Johnson, D. Kranz, J. Kubiatowicz, B.-H. Lim, K. Mackenzie, and D. Yeung, “The MIT alewife machine: Architecture and performance,” in Proceedings of the 22nd annual international symposium on Computer architecture - ISCA ’95, S. Margherita Ligure, Italy: ACM Press, 1995, pp. 2–13, isbn: 978-0-89791-698-1. doi: 1...
arXiv 1995
-
[5]
An event-triggered programmable prefetcher for irregular workloads,
S. Ainsworth and T. M. Jones, “An event-triggered programmable prefetcher for irregular workloads,” in Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and Operating Systems, Williamsburg VA USA: ACM, Mar. 19, 2018, pp. 578–592, isbn: 978-1-4503-4911-6. doi: 10.1145/3173162.3173189
arXiv 2018
-
[6]
L. Alvarez, L. Vilanova, M. Moreto, M. Casas, M. Gonz` alez, X. Martorell, N. Navarro, E. Ayguad´ e, and M. Valero, “Coherence protocol for transparent management of scratchpad memories in shared memory manycore architectures,” in Proceedings of the 42nd Annual International Symposium on Computer Architecture, Portland Oregon: ACM, Jun. 13, 2015, pp. 720–...
arXiv 2015
-
[7]
AMBA AXI and ACE protocol specification, version ARM IHI 0022H.c, ARM Limited, 2021
2021
-
[8]
Chipyard: Integrated design, simulation, and implementation framework for custom SoCs,
A. Amid, D. Biancolin, A. Gonzalez, D. Grubb, S. Karandikar, H. Liew, A. Magyar, H. Mao, A. Ou, N. Pemberton, P. Rigge, C. Schmidt, J. Wright, J. Zhao, Y. S. Shao, K. Asanovic, and B. Nikolic, “Chipyard: Integrated design, simulation, and implementation framework for custom SoCs,” IEEE Micro , vol. 40, no. 4, pp. 10–21, Jul. 1, 2020, issn: 0272-1732, 1937...
arXiv 2020
Show all 145 references
-
[9]
Instruction sets should be free: The case for risc-v,
K. Asanovic and D. A. Patterson, “Instruction sets should be free: The case for risc-v,” University of California at Berkeley, Tech. Rep., 2014
2014
-
[10]
The rocket chip generator,
Asanovic et al, “The rocket chip generator,” UC Berkeley EECS Tech Report UCB/EECS- 2016-17, 2016. 163
2016
-
[11]
Software-based cache coherence with hardware-assisted selective self-invalidations using bloom filters,
T. J. Ashby, P. Diaz, and M. Cintra, “Software-based cache coherence with hardware-assisted selective self-invalidations using bloom filters,” IEEE Transactions on Computers , vol. 60, no. 4, pp. 472–483, Apr. 2011, issn: 0018-9340. doi: 10.1109/TC.2010.155
2011 doi
-
[12]
Chisel: Constructing hardware in a scala embedded language,
J. Bachrach, H. Vo, B. C. Richards, Y. Lee, A. Waterman, R. Avizienis, J. Wawrzynek, and K. Asanovic, “Chisel: Constructing hardware in a scala embedded language,” in The 49th Annual Design Automation Conference 2012, DAC ’12, San Francisco, CA, USA, June 3- 7, 2012, P. Groene...
2012
-
[13]
Open source platforms for enabling full-stack hardware-software research,
J. Balkind, “Open source platforms for enabling full-stack hardware-software research,” Ph.D. dissertation, Princeton University, USA, 2022
2022
-
[14]
BYOC: A
J. Balkind, K. Lim, M. Schaffner, F. Gao, G. Chirkov, A. Li, A. Lavrov, T. M. Nguyen, Y. Fu, F. Zaruba, K. Gulati, L. Benini, and D. Wentzlaff, “BYOC: A ”bring your own core” framework for heterogeneous-ISA research,” inProceedings of the Twenty-Fifth International Conference ...
2020
-
[15]
OpenPiton: An open source manycore research framework,
J. Balkind, M. McKeown, Y. Fu, T. Nguyen, Y. Zhou, A. Lavrov, M. Shahrad, A. Fuchs, S. Payne, X. Liang, M. Matl, and D. Wentzlaff, “OpenPiton: An open source manycore research framework,” in Proceedings of the Twenty-First International Conference on Architectural Support for ...
2016
-
[16]
Piranha: A scalable architecture based on single-chip multipro- cessing,
L. A. Barroso, K. Gharachorloo, R. McNamara, A. Nowatzyk, S. Qadeer, B. Sano, S. Smith, R. Stets, and B. Verghese, “Piranha: A scalable architecture based on single-chip multipro- cessing,” in Proceedings of the 27th annual international symposium on Computer archi- tecture - ...
-
[17]
Analysis and op- timization of the memory hierarchy for graph processing workloads,
A. Basak, S. Li, X. Hu, S. M. Oh, X. Xie, L. Zhao, X. Jiang, and Y. Xie, “Analysis and op- timization of the memory hierarchy for graph processing workloads,” in 25th IEEE Interna- tional Symposium on High Performance Computer Architecture, HPCA 2019, Washington, DC, USA, Febr...
2019
-
[18]
Buildroot
Buildroot. “Buildroot.” (), [Online]. Available: https://buildroot.org/
-
[19]
Busybox
BusyBox. “Busybox.” (), [Online]. Available: https://busybox.net/
-
[20]
Invited - the case for embedded scalable platforms,
L. P. Carloni, “Invited - the case for embedded scalable platforms,” in Proceedings of the 53rd Annual Design Automation Conference , ser. DAC ’16, Austin, Texas: Association for Computing Machinery, 2016, isbn: 9781450342360. doi: 10.1145/2897937.2905018
2016
-
[21]
A highly productive implementation of an out-of-order processor generator,
C. Celio, “A highly productive implementation of an out-of-order processor generator,” Ph.D. dissertation, University of California, Berkeley, USA, 2017
2017
-
[22]
Boom v2: An open- source out-of-order risc-v core,
C. Celio, P.-F. Chiu, B. Nikolic, D. A. Patterson, and K. Asanovi´ c, “Boom v2: An open- source out-of-order risc-v core,” University of California, Berkeley, Tech. Rep. UCB/EECS- 2017-157, Sep. 2017
2017
-
[23]
The berkeley out-of-order machine (boom): An industry-competitive, synthesiz- able, parameterized risc-v processor,
Celio et al, “The berkeley out-of-order machine (boom): An industry-competitive, synthesiz- able, parameterized risc-v processor,” UC Berkeley EECS Tech. Rep. UCB/EECS-2015-167, 2015
2015
-
[24]
A new solution to coherence problems in multicache systems,
Censier and Feautrier, “A new solution to coherence problems in multicache systems,” IEEE Transactions on Computers , vol. C-27, no. 12, pp. 1112–1118, Dec. 1978, issn: 0018-9340. doi: 10.1109/TC.1978.1675013. 164
1978
-
[25]
Software-extended coherent shared memory: Performance and cost,
D. Chaiken and A. Agarwal, “Software-extended coherent shared memory: Performance and cost,” ACM SIGARCH Computer Architecture News, vol. 22, no. 2, pp. 314–324, Apr. 1994, issn: 0163-5964. doi: 10.1145/192007.192060
1994
-
[26]
LimitLESS directories: A scalable cache coherence scheme,
D. Chaiken, J. Kubiatowicz, and A. Agarwal, “LimitLESS directories: A scalable cache coherence scheme,” ACM SIGOPS Operating Systems Review , vol. 25, pp. 224–234, Special Issue Apr. 2, 1991, issn: 0163-5980. doi: 10.1145/106974.106995
1991
-
[28]
SMTp: An architecture for next-generation scalable multi- threading,
M. Chaudhuri and M. Heinrich, “SMTp: An architecture for next-generation scalable multi- threading,” ACM SIGARCH Computer Architecture News , vol. 32, no. 2, p. 124, Mar. 2, 2004, issn: 0163-5964. doi: 10.1145/1028176.1006712
2004
-
[29]
Integrated memory controllers with parallel coherence streams,
M. Chaudhuri and M. Heinrich, “Integrated memory controllers with parallel coherence streams,” IEEE Transactions on Parallel and Distributed Systems , vol. 18, no. 8, pp. 1159– 1173, Aug. 2007, issn: 1045-9219. doi: 10.1109/TPDS.2007.1044
2007
-
[30]
Software-controlled caches in the VMP multiprocessor,
D. R. Cheriton, G. Slavenburg, and P. D. Boyle, “Software-controlled caches in the VMP multiprocessor,” in Proceedings of the 13th Annual Symposium on Computer Architecture, Tokyo, Japan, June 1986 , H. Aiso, Ed., IEEE Computer Society, 1986, pp. 366–374. doi: 10.1145/17356.17399
1986
-
[31]
DeNovo: Rethinking the memory hierarchy for disciplined parallelism,
B. Choi, R. Komuravelli, H. Sung, R. Smolinski, N. Honarmand, S. V. Adve, V. S. Adve, N. P. Carter, and C.-T. Chou, “DeNovo: Rethinking the memory hierarchy for disciplined parallelism,” in 2011 International Conference on Parallel Architectures and Compilation Techniques, Gal...
2011 doi
-
[32]
Application performance on the MIT alewife machine,
F. Chong, Beng-Hong Lim, R. Bianchini, J. Kubiatowicz, and A. Agarwal, “Application performance on the MIT alewife machine,” Computer, vol. 29, no. 12, pp. 57–64, Dec. 1996, issn: 00189162. doi: 10.1109/2.546610
1996 doi
-
[33]
Lightweight hardware support for selective co- herence in heterogeneous manycore accelerators,
A. Cilardo, M. Gagliardi, and V. Scotti, “Lightweight hardware support for selective co- herence in heterogeneous manycore accelerators,” in 2019 Design, Automation & Test in Europe Conference & Exhibition (DATE) , Florence, Italy: IEEE, Mar. 2019, pp. 932–935, isbn: 978-3-981...
2019
-
[34]
The AMD opteron northbridge architecture,
P. Conway and B. Hughes, “The AMD opteron northbridge architecture,” IEEE Micro , vol. 27, no. 2, pp. 10–21, Mar. 2007, issn: 0272-1732. doi: 10.1109/MM.2007.43
2007 doi
-
[35]
Cache hierarchy and memory subsystem of the AMD opteron processor,
P. Conway, N. Kalyanasundharam, G. Donley, K. Lepak, and B. Hughes, “Cache hierarchy and memory subsystem of the AMD opteron processor,” IEEE Micro, vol. 30, no. 2, pp. 16– 29, Mar. 2010, issn: 0272-1732. doi: 10.1109/MM.2010.31
2010 doi
-
[36]
Corporation, An introduction to the Intel quickpath interconnect, january 2009 , 2009
I. Corporation, An introduction to the Intel quickpath interconnect, january 2009 , 2009
2009
-
[37]
The celerity open-source 511-core RISC-v tiered accelerator fabric: Fast architectures and design methodologies for fast chips,
S. Davidson, S. Xie, C. Torng, K. Al-Hawai, A. Rovinski, T. Ajayi, L. Vega, C. Zhao, R. Zhao, S. Dai, A. Amarnath, B. Veluri, P. Gao, A. Rao, G. Liu, R. K. Gupta, Z. Zhang, R. Dreslinski, C. Batten, and M. B. Taylor, “The celerity open-source 511-core RISC-v tiered accelerator...
2018
-
[38]
Design of ion-implanted MOSFET’s with very small physical dimensions,
R. H. Dennard, F. H. Gaensslen, H.-N. Yu, V. L. Rideout, E. Bassous, and A. R. LeBlanc, “Design of ion-implanted MOSFET’s with very small physical dimensions,” IEEE Journal of Solid-State Circuits , vol. SC-9, no. 5, pp. 256–268, Oct. 1974. 165
1974
-
[39]
Coherency traffic reduction in manycore sys- tems,
E. Derebasoglu, I. Kadayif, and O. Ozturk, “Coherency traffic reduction in manycore sys- tems,” in 2022 25th Euromicro Conference on Digital System Design (DSD) , Maspalomas, Spain: IEEE, Aug. 2022, pp. 262–267, isbn: 978-1-66547-404-7. doi: 10.1109/DSD57027. 2022.00043
2022
-
[40]
Protocol verification as a hardware design aid,
D. Dill, A. Drexler, A. Hu, and C. Yang, “Protocol verification as a hardware design aid,” in Proceedings 1992 IEEE International Conference on Computer Design: VLSI in Computers & Processors, IEEE Comput. Soc. Press, 2004. doi: 10.1109/iccd.1992.276232
1992
-
[41]
Ds890 ultrascale architecture and product data sheet: Overview, v4.1.1 , Xilinx, 2022
2022
-
[43]
Design verification of the s3.mp cache-coherent shared-memory system,
Fong Pong, M. Browne, A. Nowatzyk, and M. Dubois, “Design verification of the s3.mp cache-coherent shared-memory system,” IEEE Transactions on Computers , vol. 47, no. 1, pp. 135–140, Jan. 1998, issn: 00189340. doi: 10.1109/12.656100
1998 doi
-
[45]
Coherence domain restriction on large scale sys- tems,
Y. Fu, T. M. Nguyen, and D. Wentzlaff, “Coherence domain restriction on large scale sys- tems,” in Proceedings of the 48th International Symposium on Microarchitecture , Waikiki Hawaii: ACM, Dec. 5, 2015, pp. 686–698, isbn: 978-1-4503-4034-2. doi: 10.1145/2830772. 2830832
2015 doi
-
[46]
F. Gao, T. Chang, A. Li, M. Orenes-Vera, D. Giri, P. J. Jackson, A. Ning, G. Tziantzioulis, J. Zuckerman, J. Tu, K. Xu, G. Chirkov, G. Tombesi, J. Balkind, M. Martonosi, L. P. Carloni, and D. Wentzlaff, “DECADES: A 67mm 2, 1.46tops, 55 giga cache-coherent 64-bit RISC-V instruc...
2023
-
[47]
Accelerators and coherence: An soc perspective,
D. Giri, P. Mantovani, and L. P. Carloni, “Accelerators and coherence: An soc perspective,” IEEE Micro, vol. 38, no. 6, pp. 36–45, 2018. doi: 10.1109/MM.2018.2877288
2018
-
[48]
Runtime reconfigurable memory hierarchy in embedded scalable platforms,
D. Giri, P. Mantovani, and L. P. Carloni, “Runtime reconfigurable memory hierarchy in embedded scalable platforms,” in Proceedings of the 24th Asia and South Pacific Design Automation Conference, ASPDAC 2019, Tokyo, Japan, January 21-24, 2019 , T. Shibuya, Ed., ACM, 2019, pp. ...
2019
-
[49]
Using cache memory to reduce processor-memory traffic,
J. R. Goodman, “Using cache memory to reduce processor-memory traffic,” in Proceedings of the 10th annual international symposium on Computer architecture - ISCA ’83 , Stockholm, Sweden: ACM Press, 1983, pp. 124–131, isbn: 978-0-89791-101-6. doi: 10.1145/800046. 801647
1983 doi
-
[50]
Dynamically specialized datapaths for en- ergy efficient computing,
V. Govindaraju, C. Ho, and K. Sankaralingam, “Dynamically specialized datapaths for en- ergy efficient computing,” in 17th International Conference on High-Performance Computer Architecture (HPCA-17 2011), February 12-16 2011, San Antonio, Texas, USA, IEEE Com- puter Society, ...
2011
-
[51]
Efficient strategies for software-only protocols in shared- memory multiprocessors,
H. Grahn and P. Stenstr¨ om, “Efficient strategies for software-only protocols in shared- memory multiprocessors,” in Proceedings of the 22nd Annual International Symposium on Computer Architecture, ISCA ’95, Santa Margherita Ligure, Italy, June 22-24, 1995 , D. A. Patterson, ...
1995
-
[52]
A comparative evaluation of hardware-only and software-only directory protocols in shared-memory multiprocessors,
H. Grahn and P. Stenstr¨ om, “A comparative evaluation of hardware-only and software-only directory protocols in shared-memory multiprocessors,” Journal of Systems Architecture , vol. 50, no. 9, pp. 537–561, Sep. 2004, issn: 13837621. doi: 10.1016/j.sysarc.2003.08. 014
2004 doi
-
[53]
SCC: A flexible architecture for many- core platform research,
M. Gries, U. Hoffmann, M. Konow, and M. Riepen, “SCC: A flexible architecture for many- core platform research,” Computing in Science & Engineering, vol. 13, no. 6, pp. 79–83, Nov. 2011, issn: 1521-9615. doi: 10.1109/MCSE.2011.109
2011 doi
-
[54]
W. P. R. Group, Openpiton microarchitecture specification, Princeton University, 2016
2016
-
[55]
Performance analysis and benchmarking of the intel SCC,
P. Gschwandtner, T. Fahringer, and R. Prodan, “Performance analysis and benchmarking of the intel SCC,” in 2011 IEEE International Conference on Cluster Computing , Austin, TX, USA: IEEE, Sep. 2011, pp. 139–149, isbn: 978-1-4577-1355-2. doi: 10.1109/CLUSTER. 2011.24
2011 doi
-
[56]
Reducing memory and traffic requirements for scalable directory-based cache coherence schemes,
A. Gupta, W. Weber, and T. C. Mowry, “Reducing memory and traffic requirements for scalable directory-based cache coherence schemes,” in Proceedings of the 1990 International Conference on Parallel Processing, Urbana-Champaign, IL, USA, August 1990. Volume 1: Architecture, B. ...
1990
-
[57]
Integration of message passing and shared memory in the stanford FLASH multiprocessor,
J. Heinlein, K. Gharachorloo, S. Dresser, and A. Gupta, “Integration of message passing and shared memory in the stanford FLASH multiprocessor,” in Proceedings of the sixth international conference on Architectural support for programming languages and operating systems - ASPL...
1994
-
[58]
Hardware/software co-design of the stanford FLASH multiprocessor,
M. Heinrich, D. Ofelt, M. Horowitz, and J. Hennessy, “Hardware/software co-design of the stanford FLASH multiprocessor,” Proceedings of the IEEE, vol. 85, no. 3, pp. 455–466, Mar. 1997, issn: 00189219. doi: 10.1109/5.558720
1997 doi
-
[59]
A quantitative analysis of the performance and scalability of distributed shared memory cache coherence protocols,
M. Heinrich, V. Soundararajan, J. Hennessy, and A. Gupta, “A quantitative analysis of the performance and scalability of distributed shared memory cache coherence protocols,” IEEE Transactions on Computers , vol. 48, no. 2, pp. 205–217, Feb. 1999, issn: 00189340. doi: 10.1109/...
1999 doi
-
[60]
The performance and scalability of distributed shared-memory cache coher- ence protocols,
M. Heinrich, “The performance and scalability of distributed shared-memory cache coher- ence protocols,” Ph.D. dissertation, 1998
1998
-
[61]
The performance impact of flexibility in the stanford FLASH multiprocessor,
M. Heinrich, M. Horowitz, A. Gupta, M. Rosenblum, J. Hennessy, J. Kuskin, D. Ofelt, J. Heinlein, J. Baxter, J. P. Singh, R. Simoni, K. Gharachorloo, and D. Nakahira, “The performance impact of flexibility in the stanford FLASH multiprocessor,” in Proceedings of the sixth inter...
1994
-
[62]
Cooperative shared memory: Software and hardware support for scalable multiprocesors,
M. D. Hill, J. R. Larus, S. K. Reinhardt, and D. A. Wood, “Cooperative shared memory: Software and hardware support for scalable multiprocesors,” in ASPLOS-V Proceedings - Fifth International Conference on Architectural Support for Programming Languages and Operating Systems, ...
1992
-
[63]
Cooperative shared memory: Software and hardware for scalable multiprocessors,
M. D. Hill, J. R. Larus, S. K. Reinhardt, and D. A. Wood, “Cooperative shared memory: Software and hardware for scalable multiprocessors,” ACM Transactions on Computer Sys- tems, vol. 11, no. 4, pp. 300–318, Nov. 1993, issn: 0734-2071, 1557-7333. doi: 10.1145/ 161541.161544
1993
-
[64]
A 48-core IA-32 message-passing processor with DVFS in 45nm CMOS,
J. Howard, S. Dighe, Y. Hoskote, S. Vangal, D. Finan, G. Ruhl, D. Jenkins, H. Wilson, N. Borkar, G. Schrom, F. Pailet, S. Jain, T. Jacob, S. Yada, S. Marella, P. Salihundam, V. Erraguntla, M. Konow, M. Riepen, G. Droege, J. Lindemann, M. Gries, T. Apel, K. 167 Henriss, T. Lund...
2010
-
[65]
Ieee standard for systemverilog–unified hardware design, specification, and verification lan- guage,
“Ieee standard for systemverilog–unified hardware design, specification, and verification lan- guage,” IEEE Std 1800-2017 (Revision of IEEE Std 1800-2012) , pp. 1–1315, 2018. doi: 10.1109/IEEESTD.2018.8299595
2017
-
[66]
Inclusive cache
“Inclusive cache.” (), [Online]. Available: https://github.com/chipsalliance/rocket- chip-inclusive-cache
-
[67]
IntelSCC
Intel Corporation. “IntelSCC.” (), [Online]. Available: https://www.intel.cn/content/ dam/www/public/us/en/documents/technology- briefs/intel- labs- single- chip- platform-overview-paper.pdf
-
[68]
Distributed-directory scheme: Scalable coherent interface,
D. James, A. Laundrie, S. Gjessing, and G. Sohi, “Distributed-directory scheme: Scalable coherent interface,” Computer, vol. 23, no. 6, pp. 74–77, Jun. 1990, issn: 0018-9162. doi: 10.1109/2.55503
1990 doi
-
[69]
Scalable, programmable and dense: The hammerblade open-source RISC-V manycore,
D. C. Jung, M. Ruttenberg, P. Gao, S. Davidson, D. Petrisko, K. Li, A. K. Kamath, L. Cheng, S. Xie, P. Pan, Z. Zhao, Z. Yue, B. Veluri, S. Muralitharan, A. Sampson, A. Lumsdaine, Z. Zhang, C. Batten, M. Oskin, D. Richmond, and M. B. Taylor, “Scalable, programmable and dense: T...
2024
-
[70]
Implementing a cache consistency protocol,
R. H. Katz, S. J. Eggers, D. A. Wood, C. L. Perkins, and R. G. Sheldon, “Implementing a cache consistency protocol,” ACM SIGARCH Computer Architecture News , vol. 13, no. 3, pp. 276–283, Jun. 1985, issn: 0163-5964. doi: 10.1145/327070.327237
1985
-
[71]
Optimizing co- herence traffic in manycore processors using closed-form caching/home agent mappings,
S. Kommrusch, M. Horro, L.-N. Pouchet, G. Rodriguez, and J. Tourino, “Optimizing co- herence traffic in manycore processors using closed-form caching/home agent mappings,” IEEE Access, vol. 9, pp. 28 930–28 945, 2021, issn: 2169-3536. doi: 10.1109/ACCESS.2021. 3058280
2021 doi
-
[72]
Revisiting the complexity of hardware cache coherence and some implications,
R. Komuravelli, S. V. Adve, and C.-T. Chou, “Revisiting the complexity of hardware cache coherence and some implications,” ACM Transactions on Architecture and Code Optimiza- tion, vol. 11, no. 4, pp. 1–22, Jan. 9, 2015, issn: 1544-3566, 1544-3973. doi: 10 . 1145 / 2663345
2015
-
[73]
Post-silicon microarchitecture,
C. Kumar, A. Chaudhary, S. Bhawalkar, U. Mathur, S. Jain, A. Vastrad, and E. Rotenberg, “Post-silicon microarchitecture,” IEEE Computer Architecture Letters, vol. 19, no. 1, pp. 26– 29, Jan. 1, 2020, issn: 1556-6056, 1556-6064, 2473-2575. doi: 10.1109/LCA.2020.2978841
2020
-
[74]
Post- fabrication microarchitecture,
C. Kumar, A. Seshadri, A. Chaudhary, S. Bhawalkar, R. Singh, and E. Rotenberg, “Post- fabrication microarchitecture,” in MICRO-54: 54th Annual IEEE/ACM International Sym- posium on Microarchitecture, Virtual Event Greece: ACM, Oct. 18, 2021, pp. 1270–1281, isbn: 978-1-4503-855...
2021
-
[75]
Herov2: Full-stack open-source research platform for heterogeneous computing,
A. Kurth, B. Forsberg, and L. Benini, “Herov2: Full-stack open-source research platform for heterogeneous computing,” IEEE Trans. Parallel Distributed Syst., vol. 33, no. 10, pp. 4368– 4382, 2022. doi: 10.1109/TPDS.2022.3189390
2022
-
[76]
HERO: heterogeneous embedded research platform for exploring RISC-V manycore accelerators on FPGA,
A. Kurth, P. Vogel, A. Capotondi, A. Marongiu, and L. Benini, “HERO: heterogeneous embedded research platform for exploring RISC-V manycore accelerators on FPGA,”CoRR, vol. abs/1712.06497, 2017. arXiv: 1712.06497
2017 arXiv
-
[77]
The stan- ford FLASH multiprocessor,
J. Kuskin, D. Ofelt, M. Heinrich, J. Heinlein, R. Simoni, K. Gharachorloo, J. Chapin, D. Nakahira, J. Baxter, M. Horowitz, A. Gupta, M. Rosenblum, and J. Hennessy, “The stan- ford FLASH multiprocessor,” in Proceedings of 21 International Symposium on Computer 168 Architecture,...
1994
-
[78]
THE FLASH MULTIPROCESSOR: DESIGNING a FLEXIBLE AND SCAL- ABLE SYSTEM,
J. Kuskin, “THE FLASH MULTIPROCESSOR: DESIGNING a FLEXIBLE AND SCAL- ABLE SYSTEM,” Ph.D. dissertation, 1997
1997
-
[79]
COMIC: A coherent shared memory interface for cell be,
J. Lee, S. Seo, C. Kim, J. Kim, P. Chun, Z. Sura, J. Kim, and S. Han, “COMIC: A coherent shared memory interface for cell be,” in Proceedings of the 17th international conference on Parallel architectures and compilation techniques , Toronto Ontario Canada: ACM, Oct. 25, 2008,...
2008
-
[80]
An OpenCL framework for homogeneous manycores with no hardware cache coherence,
J. Lee, J. Kim, J. Kim, S. Seo, and J. Lee, “An OpenCL framework for homogeneous manycores with no hardware cache coherence,” in 2011 International Conference on Parallel Architectures and Compilation Techniques, Galveston, TX, USA: IEEE, Oct. 2011, pp. 56–
2011
-
[81]
doi: 10.1109/PACT.2011.12
2011 doi
-
[82]
The stanford dash multiprocessor,
D. Lenoski, J. Laudon, K. Gharachorloo, W.-D. Weber, A. Gupta, J. Hennessy, M. Horowitz, and M. Lam, “The stanford dash multiprocessor,” Computer, vol. 25, no. 3, pp. 63–79, Mar. 1992, issn: 0018-9162. doi: 10.1109/2.121510
1992 doi
-
[83]
Cifer: A cache-coherent 12-nm 16-mm2 soc with four 64-bit risc-v application cores, 18 32-bit risc-v compute cores, and a 1541 lut6/mm2 synthesizable efpga,
A. Li, T.-J. Chang, F. Gao, T. Ta, G. Tziantzioulis, Y. Ou, M. Wang, J. Tu, K. Xu, P. Jackson, A. Ning, G. Chirkov, M. Orenes-Vera, S. Agwa, X. Yan, E. Tang, J. Balkind, C. Batten, and D. Wentzlaff, “Cifer: A cache-coherent 12-nm 16-mm2 soc with four 64-bit risc-v application ...
2023
-
[84]
Sting: A CC-NUMA computer system for the commercial marketplace,
T. Lovett and R. M. Clapp, “Sting: A CC-NUMA computer system for the commercial marketplace,” in Proceedings of the 23rd Annual International Symposium on Computer Architecture, Philadelphia, PA, USA, May 22-24, 1996 , J. Baer, Ed., ACM, 1996, pp. 308–
1996
-
[85]
Token coherence: Decoupling performance and correct- ness,
M. Martin, M. Hill, and D. Wood, “Token coherence: Decoupling performance and correct- ness,” in 30th Annual International Symposium on Computer Architecture, 2003. Proceed- ings., San Diego, CA, USA: IEEE Comput. Soc, 2003, pp. 182–193, isbn: 978-0-7695-1945-6. doi: 10.1109/I...
2003 arXiv
-
[86]
Agile soc development with open ESP : Invited paper,
P. Mantovani, D. Giri, G. D. Guglielmo, L. Piccolboni, J. Zuckerman, E. G. Cota, M. Petracca, C. Pilato, and L. P. Carloni, “Agile soc development with open ESP : Invited paper,” in IEEE/ACM International Conference On Computer Aided Design, ICCAD 2020, San Diego, CA, USA, Nov...
2020 doi
-
[87]
Integrating performance monitoring and com- munication in parallel computers,
M. Martonosi, D. Ofelt, and M. Heinrich, “Integrating performance monitoring and com- munication in parallel computers,” in Proceedings of the 1996 ACM SIGMETRICS inter- national conference on Measurement and modeling of computer systems - SIGMETRICS ’96, Philadelphia, Pennsyl...
1996
-
[88]
Why on-chip cache coherence is here to stay,
M. M. K. Martin, M. D. Hill, and D. J. Sorin, “Why on-chip cache coherence is here to stay,” Communications of the ACM , vol. 55, no. 7, pp. 78–89, Jul. 2012, issn: 0001-0782, 1557-7317. doi: 10.1145/2209249.2209269
2012
-
[89]
Flask coherence: A morphable hybrid coher- ence protocol to balance energy, performance and scalability,
L. G. Menezo, V. Puente, and J.-A. Gregorio, “Flask coherence: A morphable hybrid coher- ence protocol to balance energy, performance and scalability,” in 2015 IEEE 21st Interna- 169 tional Symposium on High Performance Computer Architecture (HPCA) , Burlingame, CA, USA: IEEE,...
2015 doi
-
[90]
Rainbow: A composable coherence proto- col for ¡span style=
L. G. Menezo, V. Puente, and J. A. Gregorio, “Rainbow: A composable coherence proto- col for ¡span style=”font-variant:small-caps;”¿multi-chip¡/span¿ servers,” Concurrency and Computation: Practice and Experience, vol. 32, no. 24, Dec. 25, 2020, issn: 1532-0626, 1532-
2020
-
[91]
Coherence controller architectures for scalable shared-memory multiprocessors,
M. Michael, A. Nanda, and Beng-Hong Lim, “Coherence controller architectures for scalable shared-memory multiprocessors,” IEEE Transactions on Computers, vol. 48, no. 2, pp. 245– 255, Feb. 1999, issn: 00189340. doi: 10.1109/12.752666
1999 doi
-
[92]
Coherence controller architectures for SMP-based CC-NUMA multiprocessors,
M. M. Michael, A. K. Nanda, B.-H. Lim, and M. L. Scott, “Coherence controller architectures for SMP-based CC-NUMA multiprocessors,” ACM SIGARCH Computer Architecture News, vol. 25, no. 2, pp. 219–228, May 1997, issn: 0163-5964. doi: 10.1145/384286.264203
1997
-
[93]
A coherence-capable write-back l1 data cache for ariane,
M. Miceli, “A coherence-capable write-back l1 data cache for ariane,” Politecnico di Torino, 2023
2023
-
[94]
Cramming more components onto integrated circuits,
G. E. Moore, “Cramming more components onto integrated circuits,” Electronics, vol. 38, no. 8, pp. 114–117, 1965
1965
-
[95]
Cramming more components onto integrated circuits,
G. E. Moore, “Cramming more components onto integrated circuits,” Proc. IEEE, vol. 86, no. 1, pp. 82–85, 1998. doi: 10.1109/JPROC.1998.658762
1998
-
[96]
A timestamp-based cache coherence scheme,
S. L. Min and J. Baer, “A timestamp-based cache coherence scheme,” in Proceedings of the International Conference on Parallel Processing, ICPP ’89, The Pennsylvania State University, University Park, PA, USA, August 1989. Volume 1: Architecture , Pennsylvania State University ...
1989
-
[97]
First draft of a report on the EDVAC,
J. von Neumann, “First draft of a report on the EDVAC,” IEEE Ann. Hist. Comput., vol. 15, no. 4, pp. 27–75, 1993. doi: 10.1109/85.238389
1993 doi
-
[98]
The s3.mp scalable shared memory multiprocessor,
A. Nowatzyk, G. Aybay, M. Browne, E. Kelly, D. Lee, and M. Parkin, “The s3.mp scalable shared memory multiprocessor,” in Proceedings of the Twenty-Seventh Hawaii International Conference on System Sciences HICSS-94 , Wailea, HI, USA: IEEE Comput. Soc. Press, 1994, pp. 144–153,...
1994
-
[99]
Nagarajan, D
V. Nagarajan, D. J. Sorin, M. D. Hill, and D. A. Wood, A Primer on Memory Consistency and Cache Coherence (Synthesis Lectures on Computer Architecture). Springer Interna- tional Publishing, 2020. doi: 10.1007/978-3-031-01764-3
2020 doi
-
[100]
Exploiting parallelism in cache coherency protocol engines,
A. Nowatzyk, G. Aybay, M. C. Browne, E. J. Kelly, M. Parkin, B. Radke, and S. Vishin, “Exploiting parallelism in cache coherency protocol engines,” in Euro-Par ’95 Parallel Pro- cessing, First International Euro-Par Conference, Stockholm, Sweden, August 29-31, 1995, Proceeding...
1995 doi
-
[101]
Opensbi
OpenSBI. “Opensbi.” (), [Online]. Available: https : / / github . com / riscv - software - src/opensbi
-
[102]
The s3.mp architecture: A local area multiprocessor,
A. Nowatzyk, M. Monger, M. Parkin, E. Kelly, M. Browne, G. Aybay, and D. Lee, “The s3.mp architecture: A local area multiprocessor,” in Proceedings of the fifth annual ACM symposium on Parallel algorithms and architectures - SPAA ’93 , Velen, Germany: ACM Press, 1993, pp. 140–...
1993
-
[103]
Fast and efficient automatic memory management for GPUs using compiler-assisted runtime coherence scheme,
S. Pai, R. Govindarajan, and M. J. Thazhuthaveetil, “Fast and efficient automatic memory management for GPUs using compiler-assisted runtime coherence scheme,” in Proceedings of the 21st international conference on Parallel architectures and compilation techniques , Minneapoli...
2012
-
[104]
A low-overhead coherence solution for multiprocessors with private cache memories,
M. S. Papamarcos and J. H. Patel, “A low-overhead coherence solution for multiprocessors with private cache memories,” in Proceedings of the 11th annual international symposium 170 on Computer architecture - ISCA ’84 , Not Known: ACM Press, 1984, pp. 348–354, isbn: 978-0-8186-...
1984
-
[105]
OpenSPARC t1 micro architecture specification, 2008
2008
-
[106]
Exploiting transition locality in automatic verification of finite-state concurrent systems,
G. D. Penna, B. Intrigila, I. Melatti, E. Tronci, and M. V. Zilli, “Exploiting transition locality in automatic verification of finite-state concurrent systems,” International Journal on Software Tools for Technology Transfer , vol. 6, no. 4, pp. 320–341, Jul. 2004. doi: 10. 1...
2004
-
[107]
Blackparrot: An agile open-source RISC-V multicore for accelerator socs,
D. Petrisko, F. Gilani, M. Wyse, D. C. Jung, S. Davidson, P. Gao, C. Zhao, Z. Azad, S. Canakci, B. Veluri, T. Guarino, A. Joshi, M. Oskin, and M. B. Taylor, “Blackparrot: An agile open-source RISC-V multicore for accelerator socs,” IEEE Micro, vol. 40, no. 4, pp. 93–102,
-
[108]
Paulin, P
G. Paulin, P. Scheffler, T. Benz, M. A. Cavalcante, T. Fischer, M. Eggimann, Y. Zhang, N. Wistoff, L. Bertaccini, L. Colagrande, G. Ottavi, F. K. G¨ urkaynak, D. Rossi, and L. Benini, “Occamy: A 432-core 28.1 dp-gflop/s/w 83% FPU utilization dual-chiplet, dual- hbm2e risc-v-ba...
2024
-
[109]
Tempest and typhoon: User-level shared memory,
S. Reinhardt, J. Larus, and D. Wood, “Tempest and typhoon: User-level shared memory,” in Proceedings of 21 International Symposium on Computer Architecture, Chicago, IL, USA: IEEE Comput. Soc. Press, 1994, pp. 325–336, isbn: 978-0-8186-5510-4. doi: 10.1109/ISCA. 1994.288138
1994
-
[110]
Splash-3: A properly synchronized benchmark suite for contemporary research,
C. Sakalis, C. Leonardsson, S. Kaxiras, and A. Ros, “Splash-3: A properly synchronized benchmark suite for contemporary research,” in 2016 IEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2016, Uppsala, Sweden, April 17- 19, 2016, IEEE Compu...
2016
-
[111]
SCD: A scalable coherence directory with flexible sharer set encoding,
D. Sanchez and C. Kozyrakis, “SCD: A scalable coherence directory with flexible sharer set encoding,” in IEEE International Symposium on High-Performance Comp Architecture , New Orleans, LA, USA: IEEE, Feb. 2012, pp. 1–12. doi: 10.1109/HPCA.2012.6168950
2012
-
[112]
SWEL: Hardware cache coherence protocols to map shared data onto shared caches,
S. H. Pugsley, J. B. Spjut, D. W. Nellans, and R. Balasubramonian, “SWEL: Hardware cache coherence protocols to map shared data onto shared caches,” in Proceedings of the 19th international conference on Parallel architectures and compilation techniques , Vienna Austria: ACM, ...
2010 doi
-
[113]
T¨ ak¯ o: A polymorphic cache hierarchy for general-purpose optimization of data movement,
B. C. Schwedock, P. Yoovidhya, J. Seibert, and N. Beckmann, “T¨ ak¯ o: A polymorphic cache hierarchy for general-purpose optimization of data movement,” in Proceedings of the 49th Annual International Symposium on Computer Architecture , New York New York: ACM, Jun. 18, 2022, ...
2022
-
[114]
A fully programmable FSM-based process- ing engine for gigabytes/s header parsing,
K. Septinus, P. Pirsch, H. Blume, and U. Mayer, “A fully programmable FSM-based process- ing engine for gigabytes/s header parsing,” in 2010 International Conference on Embedded Computer Systems: Architectures, Modeling and Simulation, Samos, Greece: IEEE, Jul. 2010, pp. 45–54...
2010
-
[115]
SiFive, SiFive TileLink Specification, Version 1.8.1 , 2020. 171
2020
-
[116]
Fine- grain access control for distributed shared memory,
I. Schoinas, B. Falsafi, A. R. Lebeck, S. K. Reinhardt, J. R. Larus, and D. A. Wood, “Fine- grain access control for distributed shared memory,” in Proceedings of the sixth interna- tional conference on Architectural support for programming languages and operating systems - AS...
1994
-
[117]
A class of compatible cache consistency protocols and their support by the IEEE futurebus,
P. Sweazey and A. J. Smith, “A class of compatible cache consistency protocols and their support by the IEEE futurebus,” ACM SIGARCH Computer Architecture News , vol. 14, no. 2, pp. 414–423, May 1986, issn: 0163-5964. doi: 10.1145/17356.17404
1986
-
[118]
Prodigy: Improving the memory latency of data-indirect irregular workloads using hardware-software co-design,
N. Talati, K. May, A. Behroozi, Y. Yang, K. Kaszyk, C. Vasiladiotis, T. Verma, L. Li, B. Nguyen, J. Sun, J. M. Morton, A. Ahmadi, T. M. Austin, M. F. P. O’Boyle, S. A. Mahlke, T. N. Mudge, and R. G. Dreslinski, “Prodigy: Improving the memory latency of data-indirect irregular ...
2021
-
[119]
Cache system design in the tightly coupled multiprocessor system,
C. K. Tang, “Cache system design in the tightly coupled multiprocessor system,” in Pro- ceedings of the June 7-10, 1976, national computer conference and exposition on - AFIPS ’76, New York, New York: ACM Press, 1976, p. 749. doi: 10.1145/1499799.1499901
1976
-
[120]
Specifying and verifying a broadcast and a multicast snooping cache coherence protocol,
D. Sorin, M. Plakal, A. Condon, M. Hill, M. Martin, and D. Wood, “Specifying and verifying a broadcast and a multicast snooping cache coherence protocol,” IEEE Transactions on Parallel and Distributed Systems , vol. 13, no. 6, pp. 556–578, Jun. 2002, issn: 1045-9219. doi: 10.1...
2002 arXiv
-
[121]
Basejump STL: systemverilog needs a standard template library for hardware design,
M. B. Taylor, “Basejump STL: systemverilog needs a standard template library for hardware design,” in Proceedings of the 55th Annual Design Automation Conference, DAC 2018, San Francisco, CA, USA, June 24-29, 2018 , ACM, 2018, 73:1–73:6. doi: 10.1145/3195970. 3199848
2018 doi
-
[122]
Tedeschi, L
R. Tedeschi, L. Valente, G. Ottavi, E. Zelioli, N. Wistoff, M. Giacometti, A. B. Sajjad, L. Benini, and D. Rossi, Culsans: An efficient snoop-based coherency unit for the cva6 open source risc-v application processor, 2024. arXiv: 2407.19895 [eess.SY]
2024 arXiv
-
[123]
Firefly: A multiprocessor workstation,
C. P. Thacker and L. C. Stewart, “Firefly: A multiprocessor workstation,” in Proceedings of the second international conference on Architectual support for programming languages and operating systems, Palo Alto California USA: ACM, Oct. 1987, pp. 164–172. doi: 10.1145/ 36206.36199
1987
-
[124]
ADir/sub p/NB: A cost-effective way to implement full map directory- based cache coherence protocols,
Tao Li and L. John, “ADir/sub p/NB: A cost-effective way to implement full map directory- based cache coherence protocols,” IEEE Transactions on Computers, vol. 50, no. 9, pp. 921– 934, Sep. 2001, issn: 00189340. doi: 10.1109/12.954507
2001 doi
-
[125]
A case for second-level software cache coherency on many-core accelerators,
A. Vianes, F. Petrot, and F. Rousseau, “A case for second-level software cache coherency on many-core accelerators,” in 2022 IEEE International Workshop on Rapid System Proto- typing (RSP), Shanghai, China: IEEE, Oct. 13, 2022, pp. 29–35. doi: 10.1109/RSP57251. 2022.10038999
2022
-
[126]
Efficiently supporting dynamic task parallelism on heterogeneous cache-coherent systems,
M. Wang, T. Ta, L. Cheng, and C. Batten, “Efficiently supporting dynamic task parallelism on heterogeneous cache-coherent systems,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA) , Valencia, Spain: IEEE, May 2020, pp. 173– 186, isbn: 978-1...
2020
-
[127]
A programmable state machine architecture for packet processing,
Wangyang Lai and Chin-Tau Lea, “A programmable state machine architecture for packet processing,” IEEE Micro , vol. 23, no. 4, pp. 32–42, Jul. 2003. doi: 10 . 1109 / MM . 2003 . 1225965
2003
-
[128]
Torvalds
L. Torvalds. “Linux.” (), [Online]. Available: https://github.com/torvalds/linux
-
[129]
The RISC-V instruc- tion set manual, volume i: User-level ISA, version 2.0,
A. Waterman, Y. Lee, R. Avizienis, D. A. Patterson, and K. Asanovi´ c, “The RISC-V instruc- tion set manual, volume i: User-level ISA, version 2.0,” University of California, Berkeley, Tech. Rep. UCB/EECS-2016-118, May 2016. 172
2016
-
[130]
Scalable directories for cache-coherent shared-memory multiprocessors,
W.-D. Weber, “Scalable directories for cache-coherent shared-memory multiprocessors,” Ph.D. dissertation, Stanford University, 1993
1993
-
[131]
The SPLASH-2 programs: Characterization and methodological considerations,
S. C. Woo, M. Ohara, E. Torrie, J. P. Singh, and A. Gupta, “The SPLASH-2 programs: Characterization and methodological considerations,” in Proceedings of the 22nd Annual International Symposium on Computer Architecture, ISCA ’95, Santa Margherita Ligure, Italy, June 22-24, 199...
1995 doi
-
[132]
Design of the RISC-V instruction set architecture,
A. Waterman, “Design of the RISC-V instruction set architecture,” Ph.D. dissertation, Uni- versity of California, Berkeley, USA, 2016
2016
-
[133]
The bedrock cache coherence protocol and system v1.1
M. Wyse. “The bedrock cache coherence protocol and system v1.1.” (2022), [Online]. Avail- able: https://github.com/black-parrot/black-parrot/blob/master/docs/bedrock_ protocol_specification.pdf
2022
-
[134]
M. Wyse, D. Petrisko, F. Gilani, Y.-M. Chueh, P. Gao, D. C. Jung, S. Muralitharan, S. V. Ranga, M. Oskin, and M. Taylor, The blackparrot bedrock cache coherence system , 2022. arXiv: 2211.06390 [cs.AR]
2022 arXiv
-
[135]
Virtual channels and multiple physical networks: Two alternatives to improve noc performance,
Y. Yoon, N. Concer, M. Petracca, and L. P. Carloni, “Virtual channels and multiple physical networks: Two alternatives to improve noc performance,” IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. , vol. 32, no. 12, pp. 1906–1919, 2013. doi: 10 . 1109 / TCAD . 2013 . 2276399
1906
-
[136]
Mechanisms for cooperative shared mem- ory,
D. A. Wood, S. K. Reinhardt, S. Chandra, B. Falsafi, M. D. Hill, J. R. Larus, A. R. Lebeck, J. C. Lewis, S. S. Mukherjee, and S. Palacharla, “Mechanisms for cooperative shared mem- ory,” in Proceedings of the 20th annual international symposium on Computer architecture - ISCA ...
1993
-
[137]
Manticore: A 4096-core RISC-v chiplet architecture for ultraefficient floating-point computing,
F. Zaruba, F. Schuiki, and L. Benini, “Manticore: A 4096-core RISC-v chiplet architecture for ultraefficient floating-point computing,” IEEE Micro, vol. 41, no. 2, pp. 36–42, Mar. 1, 2021, issn: 0272-1732, 1937-4143. doi: 10.1109/MM.2020.3045564
2021
-
[138]
RnR: A software-assisted record-and-replay hard- ware prefetcher,
C. Zhang, Y. Zeng, J. Shalf, and X. Guo, “RnR: A software-assisted record-and-replay hard- ware prefetcher,” in 2020 53rd Annual IEEE/ACM International Symposium on Microarchi- tecture (MICRO), Athens, Greece: IEEE, Oct. 2020, pp. 609–621, isbn: 978-1-72817-383-2. doi: 10.1109...
2020
-
[139]
Fractal coherence: Scalably verifiable cache co- herence,
M. Zhang, A. R. Lebeck, and D. J. Sorin, “Fractal coherence: Scalably verifiable cache co- herence,” in 2010 43rd Annual IEEE/ACM International Symposium on Microarchitecture , Atlanta, GA, USA: IEEE, Dec. 2010, pp. 471–482, isbn: 978-1-4244-9071-4. doi: 10.1109/ MICRO.2010.11
2010
-
[140]
The cost of application-class processing: Energy and performance analysis of a linux-ready 1.7-GHz 64-bit RISC-v core in 22-nm FDSOI technology,
F. Zaruba and L. Benini, “The cost of application-class processing: Energy and performance analysis of a linux-ready 1.7-GHz 64-bit RISC-v core in 22-nm FDSOI technology,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems , vol. 27, no. 11, pp. 2629– 2640, Nov. ...
2019
-
[141]
J. Zhao, B. Korpan, A. Gonzalez, and K. Asanovic, Sonicboom: The 3rd generation berkeley out-of-order machine, May 2020
2020
-
[142]
Software assistance for directory-based caches,
Zhiyuan Li, “Software assistance for directory-based caches,” in Proceedings of 8th Inter- national Parallel Processing Symposium, Cancun, Mexico: IEEE Comput. Soc. Press, 1994, pp. 151–157, isbn: 978-0-8186-5602-6. doi: 10.1109/IPPS.1994.288307
1994
-
[143]
A programmable co-processor for profiling,
C. Zilles and G. Sohi, “A programmable co-processor for profiling,” in Proceedings HPCA Seventh International Symposium on High-Performance Computer Architecture , Monterrey, 173 Mexico: IEEE Comput. Soc, 2001, pp. 241–252, isbn: 978-0-7695-1019-4. doi: 10 . 1109 / HPCA.2001.903267
2001
-
[144]
Constellation: An open-source soc- capable noc generator,
J. Zhao, A. Agrawal, B. Nikolic, and K. Asanovi´ c, “Constellation: An open-source soc- capable noc generator,” in 2022 15th IEEE/ACM International Workshop on Network on Chip Architectures (NoCArc), IEEE, 2022, pp. 1–7
2022
-
[148]
PULP Platform
E. Z¨ urich. “PULP Platform.” (2013), [Online]. Available: https://pulp-platform.org/ index.html. 174
2013
-
[317]
doi: 10.1145/232973.233006
-
[634]
doi: 10.1002/cpe.5947
-
[2020]
doi: 10.1109/MM.2020.2996145
2020
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.