Pith. sign in

Paper Citation Record · LEDGER

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers

As of 20 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2411.19114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19114 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:34:38.587351Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact1
  • verified fuzzy72
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b90477ed-2a68-4d0c-ba1e-76ca3ac8a861 · outbound

This paper cites https://www.graphcore.ai/products/c600.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers https://www.graphcore.ai/products/c600

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.156623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.156623Z digest=sha256:faf0685eb2a4099947d56d4f4c47d88f90ca3eeb717fbfbdcd2c18856856b957

Observation d82c1106-c23d-41e4-8c71-587b09a8c28c · outbound

This paper cites https://github.com/ewan-xu/LibrosaCpp.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers https://github.com/ewan-xu/LibrosaCpp

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.162526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.162526Z digest=sha256:dfb8b9b8c1665e69c578f1ca3abf05ac2db1d3bed8f65b581aa6784a9fa6a527

Observation f1b5dc41-4a46-426c-8537-7c391bd269c8 · outbound

This paper cites https://opencv.org/.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers https://opencv.org/

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.166700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.166700Z digest=sha256:cca7cbaf13991d68aa15a0f4ea220fb258f5d57504dbf0d3c7a2b8149d799074

Observation 947f3b1a-09f7-4dc4-9a1c-5cf1c2adf6b7 · outbound

This paper cites https://rebellions.ai/rebellions-product/atom-2/.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers https://rebellions.ai/rebellions-product/atom-2/

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.171776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.171776Z digest=sha256:c7c4d73ca1c1a0bfd3dae91b5ea2c615f3a5bd23fd33c8a4129636430b8c36be

Observation df85ff40-481d-4136-a44a-070a8911e0fa · outbound

This paper cites an unresolved cited work.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.178223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.178223Z digest=sha256:d913ef32ab7ed59a91ba3c92a3e3a8602eab08c5e9a396b9936a19d28163980a

Observation 368538f5-89a4-4288-a673-ef73c6774dd1 · outbound

This paper cites Albericio, P.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Albericio, P

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.182916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.182916Z digest=sha256:2cba4eb6feaa3d725b8ebf892ba226ef6ff50bca4b27ebabdc9ecccec7b03fa7

Observation 2e256a65-6284-4045-8f32-6fe6c9e75f9e · outbound

This paper cites https://www.amazon.com/NVIDIA-Tesla-A100-Ampere- Graphics/dp/B0BGZJ27SL, 2024.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers https://www.amazon.com/NVIDIA-Tesla-A100-Ampere- Graphics/dp/B0BGZJ27SL, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.188961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.188961Z digest=sha256:dec475857f448fba483c0f2bfe5c5781d2cb67607723ab6369ca2813fc805995

Observation a11e36d1-7753-4423-8ddc-b32f01f93f6c · outbound

This paper cites Alveo U55C High Performance Compute Card.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Alveo U55C High Performance Compute Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.193129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.193129Z digest=sha256:715a39e32019eb03fe0858a87ecc9f7aea2f353a6a5cd334cb0ae4e5b53be2b2

Observation 672c4573-2a89-4d08-a37a-d27400ca1ad8 · outbound

This paper cites AMD Vitis.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers AMD Vitis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.198697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.198697Z digest=sha256:fa8b40f85b4f292d82c33660f2e5e982fa171248c72120c802a5ffabfcf842ad

Observation a71ebe07-1802-401e-9df3-3cbf360f7dc4 · outbound

This paper cites Vitis Libraries.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Vitis Libraries

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.203154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.203154Z digest=sha256:33134f216c90e4bcfef9b349fe160621c92d3dbc882dd06fee55ac70d9a66f32

Observation 2a3346c8-53f2-4514-a7e7-58a0d1211315 · outbound

This paper cites Stanley Williams, Paolo Faraboschi, Wen mei Hwu, John Paul Strachan, Kaushik Roy, and Dejan S Milojicic.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Stanley Williams, Paolo Faraboschi, Wen mei Hwu, John Paul Strachan, Kaushik Roy, and Dejan S Milojicic

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.209239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.209239Z digest=sha256:91d68152916eb04dac546d814d20cf042ebf7b2561e28d3ae5f9f5bd5cb8a4a4

Observation a79bdf04-dd1d-4455-87a2-2eb6c4074f15 · outbound

This paper cites Enabling Programmable Transport Protocols in High-Speed NICs.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Enabling Programmable Transport Protocols in High-Speed NICs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.214168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.214168Z digest=sha256:58ebb919da499fa2eab367b7955d4825f565d98f34ba6bf0bda4c497f91e0d66

Observation b6cce3a4-0d9a-4293-a4c0-460a68386a45 · outbound

This paper cites AWS Nitro System.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers AWS Nitro System

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.224207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.224207Z digest=sha256:4b7c8311f9ec723634829ea0c3ca09d0b610449e3f1671a8548797071e86bd94

Observation d115f016-9194-40aa-812f-a78116f6e820 · outbound

This paper cites AWS inferentia.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers AWS inferentia

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.229356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.229356Z digest=sha256:c8fb68e2a10a92ce5ffc50111b6ec6e7ded222b49b8be6a320ef1aca23d3bfab

Observation ef61dfba-6131-4c19-ac9a-2c29d7576aaa · outbound

This paper cites Microsoft Announces Acquisition of Fungible to Ac- celerate Datacenter Innovation.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Microsoft Announces Acquisition of Fungible to Ac- celerate Datacenter Innovation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.553211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.233414Z digest=sha256:14b05f041382b8a8025b72e7a19c7ba9d5cb2b7d4183748584d9b11b07ca5c4e

Observation 8584f545-c02c-48e1-829c-9c3ab4fe9d59 · outbound

This paper cites F4T: A Fast and Flexible FPGA-based Full- stack TCP Acceleration Framework.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers F4T: A Fast and Flexible FPGA-based Full- stack TCP Acceleration Framework

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.541473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.237357Z digest=sha256:b7ec418f5505a43f4e27e306431edb6a455894785e67a0101f87afbd8f181a6a

Observation 1702821a-82dc-4938-bbd9-35e550b33cd8 · outbound

This paper cites Sheaffer, Sang-Ha Lee, and Kevin Skadron.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Sheaffer, Sang-Ha Lee, and Kevin Skadron

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.530014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.242569Z digest=sha256:30bb5b2ba975836db6b14fa6aa3151ef28289e386f7fbc7e76b7beaa78a9f673

Observation 32591356-8788-4c7c-8d0d-37c218949242 · outbound

This paper cites Prophet: Precise QoS Prediction on Non- Preemptive Accelerators to Improve Utilization in Warehouse-Scale Computers.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Prophet: Precise QoS Prediction on Non- Preemptive Accelerators to Improve Utilization in Warehouse-Scale Computers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.518192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.246224Z digest=sha256:fea2ff3a4414a863a00206e51d56d8f09095110baec3896d8a12e7ac61def90d

Observation 77357582-4c79-43e5-ac8d-2ed79fc1a30f · outbound

This paper cites Baymax: QoS Awareness and Increased Utilization for Non-Preemptive Accelerators in Warehouse Scale Computers.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Baymax: QoS Awareness and Increased Utilization for Non-Preemptive Accelerators in Warehouse Scale Computers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.504878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.250095Z digest=sha256:b7924616dfa75d4c65fbe493e5c484298409da2e1cb4cef2d0c995120ad7c30c

Observation 1a75a5d2-23b6-4a05-a9fc-65f9eeb32449 · outbound

This paper cites an unresolved cited work.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:34:39.493309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.254322Z digest=sha256:101d95aedef56762778751c2e6f677d11fa6fda384aa435f11ab4d81fbeeaf61

Observation 2ac47d41-23e2-4979-a764-805146d33231 · outbound

This paper cites BM-Store: A Transparent and High-performance Lo- cal Storage Architecture for Bare-metal Clouds Enabling Large-scale Deployment.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers BM-Store: A Transparent and High-performance Lo- cal Storage Architecture for Bare-metal Clouds Enabling Large-scale Deployment

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.482848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.258424Z digest=sha256:fb55674cd813e1b1753b16782718ace93d2cc64ca5e3ca784648ce80704154ed

Observation f8fb8185-6de1-4df3-9498-b41c46f0623e · outbound

This paper cites Dlbooster: Boosting End-to- End Deep Learning Workflows with Offloading Data Preprocessing Pipelines.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Dlbooster: Boosting End-to- End Deep Learning Workflows with Offloading Data Preprocessing Pipelines

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.470014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.261755Z digest=sha256:b2a1937c74754926701849056e0ea110e21c0c7da3021dc69e2bf022101bf778

Observation 739c17d8-08da-4a00-a393-bbfa0b18c26e · outbound

This paper cites Lazy Batching: An SLA-aware Batching System for Cloud Machine Learning Inference.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Lazy Batching: An SLA-aware Batching System for Cloud Machine Learning Inference

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.459740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.266916Z digest=sha256:a7b33961b8d2b6e9efb13adf8a1df35cfcd753bab32309910866638d8793009c

Observation 1a82ed02-7bd8-469f-8cc6-5b5f05f7b2a6 · outbound

This paper cites PREMA: A Predictive Multi-task Scheduling Algorithm For Preemptible Neural Processing Units.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers PREMA: A Predictive Multi-task Scheduling Algorithm For Preemptible Neural Processing Units

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.448897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.270814Z digest=sha256:70c47939dd8a2fb667b134dbfbb275b95089b19fdfa9065f842337edeaf96c71

Observation 3826b349-b9a3-423a-901d-23e845b2f3b9 · outbound

This paper cites Clipper: A Low-Latency Online Prediction Serving System.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Clipper: A Low-Latency Online Prediction Serving System

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.438798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.276188Z digest=sha256:569f51672e6e33ebd25173e07188e5cdab0759fa9ed389e788d3e4f0ed53889f

Observation 731d2606-35a2-4c91-9634-e4538e7244aa · outbound

This paper cites Everything you need to know about data center power.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Everything you need to know about data center power

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.427848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.280930Z digest=sha256:a25d0c80f771f18962ff94f8204c81bc30e6689c951de0d29aac2712772121ae

Observation b0d37572-b773-4901-946a-69bd9d71e41d · outbound

This paper cites Neural Cache: Bit-serial In-cache Acceleration of Deep Neural Networks.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Neural Cache: Bit-serial In-cache Acceleration of Deep Neural Networks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.415994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.285003Z digest=sha256:0905a9534c34af3f05bdd17da2489a4f40039e7f77bc735928d4a2f3f301069a

Observation b65bc892-5b1e-421b-a70c-ff4b6ad7efa4 · outbound

This paper cites MAICC : A Lightweight Many-core Architecture with In-Cache Computing for Multi-DNN Parallel Infer- ence.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers MAICC : A Lightweight Many-core Architecture with In-Cache Computing for Multi-DNN Parallel Infer- ence

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.403712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.289893Z digest=sha256:3fb6520f9b7633f7c84ef8fafcd938a365cd68256795e57cb2d044b7a85a0f12

Observation e1c2c093-de63-4593-a9f1-1c718fe0e60c · outbound

This paper cites Low Latency RNN Inference with Cellular Batching.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Low Latency RNN Inference with Cellular Batching

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.392036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.293498Z digest=sha256:92d1681dd38461bd0cafe6c37fa3ba9d7a13c440f372dcb0f7a5fd0cfc377fcf

Observation a31f8d57-7070-4226-916e-b764f620c53c · outbound

This paper cites Low Latency RNN Inference with Cellular Batching.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Low Latency RNN Inference with Cellular Batching

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.377428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.298126Z digest=sha256:ff3d362a8be32c9dff536cfd031b943dc5ff21dffb751d4bda4e590fd55acd93

Observation 8cdfcfe1-a5fe-4430-83b5-007e4564e503 · outbound

This paper cites Cachew: Machine Learning Input Data Processing as a Service.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Cachew: Machine Learning Input Data Processing as a Service

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.364194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.302454Z digest=sha256:d090e84bd21ef0dd76d309f73948187b04130f51fa3c2f822cdb3199da71b902

Observation a4ec59ba-36eb-4c60-a29c-6d4a13c3368f · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.351901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.305978Z digest=sha256:f4ed15434931a379eacd12e6c491e0003ab7fe926ed7ba94b99e957455089a5d

Observation ce46d545-77ae-4572-8c68-c251c500d38b · outbound

This paper cites an unresolved cited work.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:34:39.339793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.310004Z digest=sha256:e342acb5e408bdb088cdcbe328a1fa87fb75a2fbfc1992f19f93686882aef542

Observation ed29906f-fd38-441a-a083-5a8d1eb4845b · outbound

This paper cites Hauswald, Y.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Hauswald, Y

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.328390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.313987Z digest=sha256:3975750d8030d198eb150c0cc8e58f5f6b68b99bae5cb610dd4b2c62677bbb75

Observation cf5c742e-d1f2-4ad5-bf44-4c93cb03ceb8 · outbound

This paper cites Searching for MobileNetV3.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Searching for MobileNetV3

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.318218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.318218Z digest=sha256:be3fc33f68446ecda5f7c85dc695ffd526da1957ebff63d183bc0a6899afa90e

Observation dbeefc74-709a-40cf-9b7a-2bf0c72d087b · outbound

This paper cites SK hynix Develops World’s Best Performing HBM3E, Provides Samples to Customer for Performance Evaluation.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers SK hynix Develops World’s Best Performing HBM3E, Provides Samples to Customer for Performance Evaluation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.316783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.322964Z digest=sha256:a849d200a9a03f32bb257fd4e7f2e36362c24acd2e0e1f683109b99212343ab4

Observation e9d90883-9ca7-4cc6-b82f-7d76822d8ad4 · outbound

This paper cites SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.326898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.326898Z digest=sha256:1439060380b8c6cd74ef5ee4839642dc0c65ae63deabc46ba90b589dfbbf1ced

Observation e8c3500f-08b9-4cdb-9ae1-d42dc4bef9b8 · outbound

This paper cites ImageNet Large Scale Visual Recognition Challenge 2012 (ILSVRC2012).

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers ImageNet Large Scale Visual Recognition Challenge 2012 (ILSVRC2012)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.305630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.331033Z digest=sha256:0a127b71d1117e028f223afad9d3d004e6aef649614259776acdfb41495a4288

Observation fc0708e8-02b4-4e29-97ae-ff3206b4bda7 · outbound

This paper cites Intel Infrastructure Processing Unit (Intel IPU).https://www.intel.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Intel Infrastructure Processing Unit (Intel IPU).https://www.intel

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.294154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.334882Z digest=sha256:7cfb64741e9fb1c9568cb144a9c0774aa01c72c85dcc2d7f54c5a09986bd642a

Observation 2ede64d5-d24c-46ab-b501-9dcd15ecb325 · outbound

This paper cites an unresolved cited work.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:34:39.283057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.338675Z digest=sha256:e101b904830e885dff24a881c8e39841c79751d231fcaca21f15c54046ef500c

Observation cbbbfd37-6d6f-4b72-81f9-dd317f3b7048 · outbound

This paper cites an unresolved cited work.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:34:39.272049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.342379Z digest=sha256:452f36c3971afb7421ddf6263885791842479123d538b32a2295d2777f405e06

Observation 075bd5e1-b316-44a3-ba01-4ea37ae45368 · outbound

This paper cites Rearchitecting the TCP Stack for I/O- Offloaded Content Delivery.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Rearchitecting the TCP Stack for I/O- Offloaded Content Delivery

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.261934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.346313Z digest=sha256:da06a78ea44e194221be348020b06cfe7e3adc91ca3c58c46a150a6fa3b46e7a

Observation 0605111c-8713-4332-bf95-dc44a1f96046 · outbound

This paper cites PARIS and ELSA: An Elastic Scheduling Algorithm for Reconfigurable Multi-GPU Inference Servers.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers PARIS and ELSA: An Elastic Scheduling Algorithm for Reconfigurable Multi-GPU Inference Servers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.250758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.349973Z digest=sha256:f52e8ef56759a074e0016c4901e8ecdb8d1f7040d803d54ba82950a15320f559

Observation 6c10f836-d885-4e91-8053-8466cbde2607 · outbound

This paper cites Dcs-ctrl: A Fast and Flexible Device-Control Mechanism for Device-Centric Server Architecture.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Dcs-ctrl: A Fast and Flexible Device-Control Mechanism for Device-Centric Server Architecture

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.239418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.354723Z digest=sha256:487bb6015962a1d27b5d9f574f0b80f80091a897e674439f1c152bf3445d3ef8

Observation 47a67db7-a1b3-4d76-8f9e-5268677d297b · outbound

This paper cites FVM: FPGA-assisted Virtual Device Emulation for Fast, Scalable, and Flexible Storage Virtualization.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers FVM: FPGA-assisted Virtual Device Emulation for Fast, Scalable, and Flexible Storage Virtualization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.227699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.359899Z digest=sha256:ee0eb2daf8c3d2e35e38e1f16872c979804134ffbfbd1937ba0c8d5408c95daf

Observation 60cadf31-2d4b-4e3b-af78-4274dfba9cad · outbound

This paper cites A Fast and Flexible Hardware-based Virtualization Mechanism for Computational Storage Devices.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers A Fast and Flexible Hardware-based Virtualization Mechanism for Computational Storage Devices

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.216484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.367138Z digest=sha256:0a078c79865145102a458cd92b2288d3727888a5ad614176cb29bda2104d07b8

Observation 69512591-776c-45c6-bce7-f8b536df6196 · outbound

This paper cites Char- acterizing Multi-Instance GPU for Machine Learning Workloads.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Char- acterizing Multi-Instance GPU for Machine Learning Workloads

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.204947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.371821Z digest=sha256:d639547f0bcc502e629e7a5fc470fd3b6549da1979617fc24504590960f96eab

Observation cfc802c0-d64e-469f-8393-c0273ae64d33 · outbound

This paper cites MISO: Exploiting Multi-Instance GPU Capability on Multi-Tenant GPU Clusters.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers MISO: Exploiting Multi-Instance GPU Capability on Multi-Tenant GPU Clusters

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.187682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.377687Z digest=sha256:05629810e893fba868c28954fa7219ec0fc2e3e4f789e8af7482d758a82b54e0

Observation 81a7f121-8198-43ea-8cc9-1a850918d67d · outbound

This paper cites Leapio: Efficient and Portable Virtual NVMe Storage on ARM SOCs.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Leapio: Efficient and Portable Virtual NVMe Storage on ARM SOCs

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.175643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.381917Z digest=sha256:66310d5973647dd1f459f1925bbde4ad8950ccfc0e9bcae21ae9b6facb9736f6

Observation f9f050a1-6ff1-4003-9d8a-7f1af6583a86 · outbound

This paper cites E3: Energy-efficient Microservices on SmartNIC- accelerated Servers.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers E3: Energy-efficient Microservices on SmartNIC- accelerated Servers

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.163812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.386619Z digest=sha256:80fa1b3f872cee79d32e86d55e7b66f515f60efab4e5ec52495fc811f80d7242

Observation b4b12dc2-2edd-4f76-a2a0-08e2baf21b19 · outbound

This paper cites Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.153006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.390620Z digest=sha256:6e95dffd03a7c1e76ba5fe26c98a2074eb2c4de7ccda1e557a4b075fa1d5f560

Observation bfdd653b-1b8b-4674-87a5-f2a67a447350 · outbound

This paper cites Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition, 2021.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition, 2021

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.141964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.395944Z digest=sha256:b4427ae60aacbb98bbd25c1fb72fd94885da1d5b409cae288457e716c3e387c8

Observation ac563915-2e40-4391-a992-f0a22f0db368 · outbound

This paper cites Boost Your Datacenter! https://mangoboost.io/, 2023.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Boost Your Datacenter! https://mangoboost.io/, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.130088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.400054Z digest=sha256:c8d7c933083c497bfc0f2bd6093cef8973a535910346443e565334758d922192

Observation 0ad7be90-137d-47f9-a095-680cdf808718 · outbound

This paper cites Meta MITA v1.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Meta MITA v1

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.118630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.404569Z digest=sha256:05b3f8cce1a551f056993b647458538376e67bd510a7cc967034d66453e51e9f

Observation 8ef49a8a-aba6-470f-b0ff-1f615ae0ecfc · outbound

This paper cites Gimbal: Enabling Multi- Tenant Storage Disaggregation on SmartNIC JBOFs.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Gimbal: Enabling Multi- Tenant Storage Disaggregation on SmartNIC JBOFs

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.105119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.408865Z digest=sha256:d39b3dc98eed005b083a131e9c7278be901ddc37c2446ef7b27e42b1ab714e37

Observation 8d9e749c-6d3b-4056-9abe-ebc78ae46cf7 · outbound

This paper cites Analyzing and Mitigating Data Stalls in DNN Training.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Analyzing and Mitigating Data Stalls in DNN Training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.093021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.413226Z digest=sha256:81cdcc02338f031b18aca329856c20224a0bb182c10c3de127accf9f61af1f86

Observation 29094446-2f0e-4067-8e6e-a5bae921bc45 · outbound

This paper cites AccelTCP: Accelerating Network Applications with Stateful TCP Offloading.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers AccelTCP: Accelerating Network Applications with Stateful TCP Offloading

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.082755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.417003Z digest=sha256:76588558360e98fc75c93c882d605c82640e3e59a2fbde1c4e3ee27665ac8e09

Observation 230cc775-27d7-4dc9-be20-4adc83a21a47 · outbound

This paper cites NVIDIA Triton Inference Server.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers NVIDIA Triton Inference Server

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.072262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.420330Z digest=sha256:b8bbfed70963a30120e527a09b68c04d7116356451656cb6151e15aced0e653e

Observation ebf9d4b2-a68d-4f4a-9949-9e2082919bc9 · outbound

This paper cites NVIDIA TensorRT: Programmable Inference Accelerator.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers NVIDIA TensorRT: Programmable Inference Accelerator

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.062054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.424883Z digest=sha256:a75c48805cf630217b7d25c766561df0aa6b318aa24365a0fdb1313813280fa4

Observation 37694898-1f21-4fa7-a7da-2e310958c5ca · outbound

This paper cites NVIDIA T4.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers NVIDIA T4

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.052078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.429195Z digest=sha256:a57b16e5916e527c75ddd9328c4e5bb317f248fdb06d7f4b9cba220f3cc01acb

Observation 1a0b0e7b-3a66-4831-9904-08523e79708b · outbound

This paper cites NVIDIA A100.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers NVIDIA A100

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.041979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.433167Z digest=sha256:51f9f17d48fd869fc0dce7dbea0cb9952d6e8c49a40ed387d420d0a8e09b1281

Observation 4ea35def-934d-4174-98f4-c5465aa0f2fb · outbound

This paper cites NVIDIA CUDA Programming Guide, 2021.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers NVIDIA CUDA Programming Guide, 2021

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.031803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.437933Z digest=sha256:36fb4ea7f53d933222e513d54c7244e3054bb7979229c82e19466dd53accb334

Observation 32aa194b-d158-45f1-92c4-3d6d87128106 · outbound

This paper cites Multi-Instance GPU.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Multi-Instance GPU

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.021093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.442680Z digest=sha256:808b1c78c11362f9a1d983ce88aa3296641cb014f9232ce4dd7e265f1badb04c

Observation ba7a5b2b-2ed3-4b9b-895a-26814cdba131 · outbound

This paper cites NVIDIA BlueField Data Processing Units.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers NVIDIA BlueField Data Processing Units

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:39.010003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.446521Z digest=sha256:27bd477fac892d4329a1bd335203b858ac718d71bf9f1d152c9fcc3dab3cdbbb

Observation 86b1c471-cdab-43e4-a126-ac9b752cc47c · outbound

This paper cites NVIDIA NeMo.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers NVIDIA NeMo

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.998637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.450518Z digest=sha256:24d557477682b2a3c64096f4e9be1ba2511a7f7e9a1cbf67e38f1bc180da3f4c

Observation fa6b2e4d-492f-4755-a2fd-e2fca17916df · outbound

This paper cites TensorFlow-Serving: Flexible, High-Performance ML Serving.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers TensorFlow-Serving: Flexible, High-Performance ML Serving

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T10:34:38.455548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:34:38.455548Z digest=sha256:feea5dd44eddf2be1964d82414944b08028b2e53f41114a2f800a5aaddfb7e30

Observation 9d3b4759-087a-46ae-9115-77de0426859f · outbound

This paper cites LibriSpeech ASR corpus.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers LibriSpeech ASR corpus

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.986121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.460059Z digest=sha256:26d0144773c6c51d7dfc726294e5f78b58ce6ef6e3bf60c69e13056410aeac08

Observation 6d081e51-fdf0-403a-a9f3-1641b6aeed44 · outbound

This paper cites Parashar, M.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Parashar, M

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.973480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.464443Z digest=sha256:a87da271a9071180f7cdc2dfec1d0d90b30bb4a0e8c17020d4eca6b33332fb42

Observation 50003c18-731b-4de5-9459-d40cb480b3a9 · outbound

This paper cites TrainBox: An Extreme-Scale Neural Network Training Server Architecture by Sys- tematically Balancing Operations.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers TrainBox: An Extreme-Scale Neural Network Training Server Architecture by Sys- tematically Balancing Operations

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.960473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.470075Z digest=sha256:4b183afa9aeca35e2d08c1bf365b057cd6481faf721e705f6bfd2d5d3c2df7c8

Observation 88c4e2be-6983-41d9-a772-a8745c792b6f · outbound

This paper cites Single-Root Input/Output Virtualization.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Single-Root Input/Output Virtualization

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.948323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.475804Z digest=sha256:6e60c8cae34929a92b1a0e1fe532897e4f2fabf9f6a1c2e8a5575d2ff53fe609

Observation a0d28059-fd07-4130-9587-4f0a99fb847f · outbound

This paper cites Torchaudio Transforms.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Torchaudio Transforms

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.937202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.480250Z digest=sha256:279f914b5bb24c43ede346064dce759baa8130f4c7a60337ee7e859ee93e219f

Observation b98e3363-1adf-4287-bc8d-fddb5cf5900c · outbound

This paper cites Transforming and Augmenting Images.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Transforming and Augmenting Images

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.926170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.485978Z digest=sha256:cea6df96c4aa34ae6d973e96a0a4935b8dd5e4223599b6abd681a3dd44998a5b

Observation 1aa15da8-656c-4ba4-a670-acf6c222edc5 · outbound

This paper cites PyTorch Hub.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers PyTorch Hub

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.914300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.490272Z digest=sha256:a3eb8af033490c877d4313fa1364d93ae3aef8a4fe1adb0a8bac2d6b4c57864c

Observation fac3e955-133f-4561-8b55-a997476d8ac5 · outbound

This paper cites Scott Gardner, Itay Hubara, Sachin Idgunji, Thomas B.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Scott Gardner, Itay Hubara, Sachin Idgunji, Thomas B

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.903823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.494647Z digest=sha256:f7ee87cd73337fbb749211ee14f2e198fd7945ca7870d35cd466ae9c90722320

Observation 307c311e-298d-49b3-b08c-e858e69cec4a · outbound

This paper cites An Analysis of Collocation on GPUs for Deep Learning Training.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers An Analysis of Collocation on GPUs for Deep Learning Training

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:34:38.631590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.498606Z digest=sha256:c8459b9c24eb9f9538c7169bf942c7033a859f8fee367aa5cb4bb8b0cdd56171

Observation 51d51f1a-b225-4511-8290-5b529f91c2e0 · outbound

This paper cites Samsung Develops Industry’s First GDDR7 DRAM To Un- lock the Next Generation of Graphics Performance.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Samsung Develops Industry’s First GDDR7 DRAM To Un- lock the Next Generation of Graphics Performance

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.893368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.507544Z digest=sha256:2516ea863875c3f08a4d424df613c182743c0b65234e2d60227288b611179f56

Observation f94ae1fa-158f-454b-89f0-147983990e66 · outbound

This paper cites Tell, Yanqing Zhang, William J.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Tell, Yanqing Zhang, William J

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.880489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.513530Z digest=sha256:dc2f55d5f544f9603ea2bedcc5fdfe1e5a338e35dcf547426dac08242074229f

Observation 72a17066-2bb1-434e-90ea-b2b6941735a9 · outbound

This paper cites Laconic Deep Learning Inference Acceleration.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Laconic Deep Learning Inference Acceleration

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.869518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.518104Z digest=sha256:169f51a3a0df2580b6fc91a87a14683385c4af1568c96003bf7cbb164346d053

Observation 8a812d19-29e9-4d44-8b0a-07dfb5aab145 · outbound

This paper cites Bit Fu- sion: Bit-Level Dynamically Composable Architecture for Accelerating Deep Neural Network.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Bit Fu- sion: Bit-Level Dynamically Composable Architecture for Accelerating Deep Neural Network

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.857136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.522338Z digest=sha256:c28a3b1ef5a37413b2c2d148c4e826319e4ae53a5b6ad6b7f0e86acaf9045466

Observation cc97a1a7-0cf6-43d6-b238-24d2dfcf411c · outbound

This paper cites FlexTOE: Flexible TCP Offload with Fine-Grained Parallelism.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers FlexTOE: Flexible TCP Offload with Fine-Grained Parallelism

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.845985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.527085Z digest=sha256:0eac9d5ca43a207a0410b29d7425f862f446d422f59f3abfd9ffb1490e4e7328

Observation 3bf04da7-a059-4668-83d4-e3b9dc783821 · outbound

This paper cites Microsoft Maia v1.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Microsoft Maia v1

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.834515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.531687Z digest=sha256:7a1775d9dd2b6a3fdb81547a481cfd84bd62d7302ab8fc702c1d53b8af671ae9

Observation 8e2b71e1-ff23-4c46-a190-abf344b7dd05 · outbound

This paper cites https://store.supermicro.com/us_en/mainstream-amd- 2u-as-2024s-tr.html , 2024.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers https://store.supermicro.com/us_en/mainstream-amd- 2u-as-2024s-tr.html , 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.822832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.537507Z digest=sha256:3ce4e878a0f180e457888884aa087841750df8aeb1215155e43021e27685a0bd

Observation 4258ab49-cc20-4166-b261-9185d28a3387 · outbound

This paper cites Lynx: A SmartNIC- Driven Accelerator-Centric Architecture for Network Servers.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Lynx: A SmartNIC- Driven Accelerator-Centric Architecture for Network Servers

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.812152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.541987Z digest=sha256:0de3f61a9b301e5f2d807a310afa8f3158b3cafd39eafb4ed9a7ceac1918fcfa

Observation 92292c9b-1935-4906-a1a8-17ec111a052c · outbound

This paper cites FastFlow: Accelerating Deep Learn- ing Model Training with Smart Offloading of Input Data Pipeline.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers FastFlow: Accelerating Deep Learn- ing Model Training with Smart Offloading of Input Data Pipeline

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.801010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.547356Z digest=sha256:cd0b691222a6006c68142494852d4e6e00c4269e23a7e85eeddbbb63fbcdcd5b

Observation 6452af10-447a-42f2-8132-05dff8e11a02 · outbound

This paper cites Real-time Meets Approximate Computing: An Elastic CNN Inference Accelerator with Adaptive Trade-off Between QoS and QoR.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Real-time Meets Approximate Computing: An Elastic CNN Inference Accelerator with Adaptive Trade-off Between QoS and QoR

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.789779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.551240Z digest=sha256:d278ed99e0db879ad94c3d01ea4119a05701358f004b4e4c537c10671aca01bb

Observation 616a5dd7-cffc-42c8-898f-10819e1288ec · outbound

This paper cites A None- Sparse Inference Accelerator that Distills and Reuses the Computation Redundancy in CNNs.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers A None- Sparse Inference Accelerator that Distills and Reuses the Computation Redundancy in CNNs

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.778962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.555406Z digest=sha256:1cc179786e7f53803d0689d127326d687784820c2cfcee159836c90da99bcfb2

Observation 4aba7ce0-7905-4094-bb76-180883163ca9 · outbound

This paper cites FpgaNIC: An FPGA-based Versatile 100Gb SmartNIC for GPUs.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers FpgaNIC: An FPGA-based Versatile 100Gb SmartNIC for GPUs

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.768076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.559296Z digest=sha256:1f792b459565aaddaf93aefad874e8f69d281b6793bc5e09f28b0bc2f215ca27

Observation 7be053fa-4b55-48a0-a761-e4abf03e2e85 · outbound

This paper cites Xilinx OpenCL Extension.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Xilinx OpenCL Extension

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.756232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.562983Z digest=sha256:4f7bb44f6380f0a284526de32286f0321e46d7655c0bb889ffac349ee6624c94

Observation 343db716-947b-429c-a570-db8f2b3895ef · outbound

This paper cites Vitis High-level Synthesis User Guide.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Vitis High-level Synthesis User Guide

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.744636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.567086Z digest=sha256:89dd86dd01328508d253e5aac3b4ba517e69d61b725716de0175240ce63d4778

Observation b52a22d2-5abb-4f34-9cde-e356361ec547 · outbound

This paper cites https://www.xilinx.com/products/boards-and-kits/alveo/u55c.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers https://www.xilinx.com/products/boards-and-kits/alveo/u55c

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.730853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.571345Z digest=sha256:a1705e1de180d9c630a7d6c606f8cda1ff842418a938aecc5200051daf185bb0

Observation 7bc2c7ec-8ae4-43e7-a186-b3e5de010601 · outbound

This paper cites Sinclair, Bradford M.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Sinclair, Bradford M

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.717003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.575360Z digest=sha256:021b5c8850829d7cc42294dfbfc68eb0dd31b57ccd78931bca3095ae6d40c845

Observation 59896571-7772-473c-aada-a84e5b34bf33 · outbound

This paper cites Orca: A Distributed Serving System for Transformer-Based Generative Models.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Orca: A Distributed Serving System for Transformer-Based Generative Models

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.703916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.579709Z digest=sha256:dfb358623d129ab5528fccdb26797349b64dfc6372b0e9d5da8e7c0a3e299c83

Observation a0149199-d212-4d45-b3ae-9edf222e5098 · outbound

This paper cites MArk: Ex- ploiting Cloud Services for Cost-Effective,SLO-Aware Machine Learn- ing Inference Serving.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers MArk: Ex- ploiting Cloud Services for Cost-Effective,SLO-Aware Machine Learn- ing Inference Serving

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.691935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.583231Z digest=sha256:5814003c1b412e6697bf875ceff1fdf9261ef81abb2dfdc31bb996e18df0ca12

Observation 7c445e07-0e05-4986-a570-9095071f01e6 · outbound

This paper cites Understand- ing Data Storage and Ingestion for Large-Scale Deep Recommendation Model Training: Industrial Product.

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Understand- ing Data Storage and Ingestion for Large-Scale Deep Recommendation Model Training: Industrial Product

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:34:38.679897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:34:38.587351Z digest=sha256:1c8e6e67295da41b8bbb1cfee5ca700ff9108508770334d35fc9b5b7d22cae53

Pith citing papers

No inbound Pith citation observations are available.