Pith. sign in

Paper Citation Record · LEDGER

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization

As of 8 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2506.23714.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23714 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:38:45.547295Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c8d0ab6-eabf-45ec-ba85-c2c154aa2a93 · outbound

This paper cites Apostolidis, E.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Apostolidis, E

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:51.946959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:41.726899Z digest=sha256:d386e36f6655fce8416c93d3159c9a5f526b992eaa48aed253816c80b2b158ee

Observation 2eb97bf5-b907-4c40-9252-8890bcc32079 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:51.675229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:41.843334Z digest=sha256:aafe5baba6366cda2c83a68f3499027b72a5ca30d6464d2f7b7cdf57fceaf20d

Observation a9d4485a-7722-4d00-85f7-da7d4868509a · outbound

This paper cites Tiwari, C.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Tiwari, C

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:51.479100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:41.927525Z digest=sha256:7d95f13546989d07d9659ef7c1baf16ad7c06ad7a9c9cc1a6080af0a35c9df0a

Observation 381d26d5-de0e-4068-917a-b53ec001a0dc · outbound

This paper cites Otani, Y.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Otani, Y

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:51.283393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:42.019442Z digest=sha256:f7cb62404efeb8d5bce1ef9f0abdad05499bc707e58c1a49e798b3fa2eba072f

Observation cbc56d9b-5ee8-48fc-8314-4b61e76cd7bd · outbound

This paper cites Rochan, L.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Rochan, L

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:51.086223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:42.083343Z digest=sha256:451c00afbc872bcb3b9776c614fa5a36e3c236302e2fd87f93718ef8846e591d

Observation 961b9393-b0aa-4f4b-9f6b-1dc8cdf10302 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:50.835567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:42.160765Z digest=sha256:0477565199d383b45b4e00ae1dc795b1645f7898d846fabf01b7bba50ef612dd

Observation 0c249360-1a60-4e4d-8517-c8cd63ef52e5 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:50.613929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:42.271990Z digest=sha256:2f5b9782787444e392bdd12caa622b7159f2b1438eb7c1be3abfe34121ad0db7

Observation f5e1c958-ed12-4208-b7d2-94f2b619d2b5 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:42.386654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:42.386654Z digest=sha256:fd1d33526a306b925ac0e975a3528838f7144e6ab0e0744a82ea4cf84d400890

Observation 2c3d545d-d505-4b1e-9333-520bc826e5bb · outbound

This paper cites Get To The Point: Summarization with Pointer-Generator Networks.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Get To The Point: Summarization with Pointer-Generator Networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:42.501169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:42.501169Z digest=sha256:51abd44b9e5ae2bd334809dca3cc6668f70cd1e0894c3eebb3c28ceca411169e

Observation 0d9b036c-f9b5-483c-886a-d230ec05a78b · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:50.445452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:42.593261Z digest=sha256:669a1b3e83bb7890ef5a9eeb761b6dad5856d60e683a9a0c22e707e55913c984

Observation a207b5fe-3468-416a-982f-cee44c9e27ea · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:50.194173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:42.792890Z digest=sha256:e7913382b8da820ea18cac4a0d831aaccb9d5d81e2fda1bce28cb46cec541175

Observation a72f83fe-2201-4393-bfcd-80fed780a55d · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:50.036229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:42.831226Z digest=sha256:52aeff27ce10214d30e956f03308cbd578e8900550c9f57ec146f945b5977833

Observation c798f581-c284-4204-adee-12e3d0c0733a · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:49.782366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:42.910882Z digest=sha256:b904be8b843bc3fb522fad43dfe59223ecfcdbe841a91efe12f2852b062af979

Observation 5a0af3e5-a26c-4430-9d5a-89f6a6ffd8f0 · outbound

This paper cites Zhang, W.-L.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Zhang, W.-L

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:49.528283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:43.006112Z digest=sha256:a191cb7cdc91652f364a1ee5fdc7f974e3f73d0c78e382f3787813377a645746

Observation 62b27317-ba1b-4d54-8481-cb8ebca2c983 · outbound

This paper cites Saini, K.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Saini, K

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:49.287084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:43.118186Z digest=sha256:618d083916d5eb52fc109a09b08039e38e24b588f5abd5886c423aefc854cf66

Observation 886df382-a896-4c47-9f4c-2b180d455ba8 · outbound

This paper cites Evangelopoulos, A.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Evangelopoulos, A

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:49.078689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:43.213503Z digest=sha256:d06f863db8d519b516ca04c7c7993c33821d9cb03314fcd4285def86eb722e4b

Observation fb2656d3-fa37-472f-b2af-f9f51d9aa4df · outbound

This paper cites Multimodal Frame-Scoring Transformer for Video Summarization.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Multimodal Frame-Scoring Transformer for Video Summarization

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:38:45.965102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:43.276662Z digest=sha256:76c4aea5e92dd2b72f8e8eaddeb778db4576d4b8b5ca39d807cf3dc34a2ec593

Observation d39ed602-c3d9-42ff-bcc5-f0bda509258b · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:48.856571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:43.444350Z digest=sha256:f544b493f7308a09eda17054696ecd385707a06b028c723a98d66bd3edb2ed82

Observation b1503f76-54af-49d9-b4d3-70b0802bf9d5 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:48.607542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:43.569978Z digest=sha256:f8b776c0f3bfd1cf86449bd0d5fa3bf2a0c4e698b45610ed63628a13544e9520

Observation a9f653a7-cd91-4f50-9d39-55e416539141 · outbound

This paper cites Psallidas, P.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Psallidas, P

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:48.323566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:43.730639Z digest=sha256:f3d4deaea14239d3d5848d392b1f27cb5391b7c77936b78894ff97b7e2a01138

Observation cbc3dd90-8b34-4a14-ac97-60f9b3d3be1a · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:48.174758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:43.802627Z digest=sha256:de6ce6d19ffdffa25d7ca52772bc77a810e189c0c9be86a325505bae926bca41

Observation 144dfa5c-f4b8-4c54-a3bf-1e44247daee2 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:48.054461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:43.934293Z digest=sha256:358cd0541d6865b2b8606ebef8fe0d6e0a7ea87ff039eaeb3ffe8a4cb108123f

Observation 8c7936bb-aefd-439f-97de-1c0c8bb55e04 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:47.918043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:44.117750Z digest=sha256:7da70c23d34299ca90f0398eeb555dc8b8ca13458f127f09d55b458b4473d7e8

Observation e6bc727e-9258-4621-9e86-eb581face9c4 · outbound

This paper cites Ponce-López, B.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Ponce-López, B

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:47.765723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:44.274131Z digest=sha256:5e25dfc0f1ea54b44bdd340b843246128b90e3d4bb82e66879e988bd2e0228a6

Observation 454caf69-4c5d-4748-8eef-257f5696f746 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Robust Speech Recognition via Large-Scale Weak Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:44.378445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:44.378445Z digest=sha256:da476e99771b8fdf029d70cabd16b1be3d2b9d47e0a9859ad92fefd92648577d

Observation 40129d1c-c0d4-4d33-a40f-6d5796b06ee4 · outbound

This paper cites Honnibal, I.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Honnibal, I

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:47.609541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:44.467892Z digest=sha256:b1f654e57057e6cc1c80c5b2d9c2d205b581dbfa6f63fc00131f54da910303b2

Observation 15ad8d3d-7527-4caa-b676-3655989e9ea0 · outbound

This paper cites McAuliffe, M.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization McAuliffe, M

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:47.442804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:44.581457Z digest=sha256:2031e7312a38f206c437291697c155692a003fbef922ae6eb4899dfed0427252

Observation 2b019d12-fe56-4340-aa0d-f2d290cdfe2a · outbound

This paper cites Bradski, The opencv library., Dr.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Bradski, The opencv library., Dr

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:47.182495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:44.656403Z digest=sha256:2cabea661eb7185a483ca49449d17db443a94f282bcbced7591255763f304554

Observation bdce2f8c-fb21-465d-a7e6-e7799bf975fc · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization MediaPipe: A Framework for Building Perception Pipelines

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:44.780495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:44.780495Z digest=sha256:52ec2755dd1fe24acf7ba10e3e1bca2b7abfe32106196adc55fd3556463e1dfb

Observation 5c5b0eb6-fc0a-4835-9f06-09b9bbe92be0 · outbound

This paper cites Serengil, A.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Serengil, A

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:44.868440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:44.868440Z digest=sha256:9873a9ccdbbf3929cc528320cae7473121798096e6a3cc1d8f6e1856cce8b1e2

Observation 5bc15467-43d0-45ed-b8e9-6b5c2583ff94 · outbound

This paper cites Kasi, Yet another algorithm for pitch tracking (yaapt), 2002.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Kasi, Yet another algorithm for pitch tracking (yaapt), 2002

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:46.996042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:45.019358Z digest=sha256:acd9b9474f281e817589ffc73f8dafba98702eda2ede5ef428d7d3b8dcbe33aa

Observation 65596910-48c2-42f9-aad3-7f074369f7c3 · outbound

This paper cites Eyben, M.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Eyben, M

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:46.771570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:45.121470Z digest=sha256:22f1c4733fc4bdcd9e26859ae930ccba40d5a0b3a2b05aab84bfc8b936589953

Observation a4398015-e2c8-4dc7-b105-5bb6393f2704 · outbound

This paper cites Sparck Jones, A statistical interpretation of term specificity and its application in retrieval, Journal of documentation 28 (1972) 11–21.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Sparck Jones, A statistical interpretation of term specificity and its application in retrieval, Journal of documentation 28 (1972) 11–21

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:46.554911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:45.264167Z digest=sha256:ee5d5a3d748008a3580b6e4997fc56a21cc1d472826dc37edb1f1eadfb3eaee3

Observation 8296f4a2-12d4-42c6-89b9-ca931e52cdeb · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:46.397336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:45.362663Z digest=sha256:9bf2da4d38d69d3641d842840f49fa62f96680c6d71386a20a0a57c43b7171a8

Observation a69251c5-26de-4f6c-b9ca-fa59d31aa6b9 · outbound

This paper cites Scaling Up Video Summarization Pretraining with Large Language Models.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Scaling Up Video Summarization Pretraining with Large Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:38:45.783605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:45.460292Z digest=sha256:3e3807949b2a3ff17361d0e08ef211c245c7c36f1af5f7f0c586e357ab1f7bca

Observation d4c6b550-d87d-4990-bc63-c22d7e5f54d0 · outbound

This paper cites Narasimhan, A.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Narasimhan, A

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:46.223123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:38:45.547295Z digest=sha256:fe267f5df440ba6795cd1546c92327afc34b97e5d6bb165357e3ca2b82ee02dd

Pith citing papers

No inbound Pith citation observations are available.