{"paper":{"title":"Can LLMs Reason About Attention? Towards Zero-Shot Analysis of Multimodal Classroom Behavior","license":"http://creativecommons.org/licenses/by/4.0/","headline":"LLMs can perform zero-shot analysis of student attention using only pose and gaze coordinates from classroom videos.","cross_cats":["cs.AI","cs.CV"],"primary_cat":"cs.HC","authors_text":"Alp Tural, Andrew Katz, Elif Tural, Nada Basit, Nolan Platt, Saad Nizamani, Sehrish Nizamani, Yoonje Lee","submitted_at":"2026-04-03T19:04:31Z","abstract_excerpt":"Understanding student engagement usually requires time-consuming manual observation or invasive recording that raises privacy concerns. We present a privacy-preserving pipeline that analyzes classroom videos to extract insights about student attention, without storing any identifiable footage. Our system runs on a single GPU, using OpenPose for skeletal extraction and Gaze-LLE for visual attention estimation. Original video frames are deleted immediately after pose extraction, thus only geometric coordinates (stored as JSON) are retained, ensuring compliance with FERPA. The extracted pose and "},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Our preliminary findings suggest that LLMs may show promise for multimodal behavior understanding, although they still struggle with spatial reasoning about classroom layouts.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That geometric pose and gaze coordinates extracted by pre-trained models are sufficient for an LLM to perform accurate zero-shot inference of student attention levels without domain-specific fine-tuning or spatial context.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"A pipeline uses OpenPose and Gaze-LLE to extract pose and gaze data from classroom videos, deletes the raw footage, and applies an LLM for zero-shot behavioral analysis of student attention.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"LLMs can perform zero-shot analysis of student attention using only pose and gaze coordinates from classroom videos.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"a70198af33a525b49ba474c2c05c0430462dc3f9144f1bd46c5dd1b591fe7046"},"source":{"id":"2604.03401","kind":"arxiv","version":4},"verdict":{"id":"b6770b65-783e-43ff-9de4-247ab4620d90","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-13T18:16:41.233773Z","strongest_claim":"Our preliminary findings suggest that LLMs may show promise for multimodal behavior understanding, although they still struggle with spatial reasoning about classroom layouts.","one_line_summary":"A pipeline uses OpenPose and Gaze-LLE to extract pose and gaze data from classroom videos, deletes the raw footage, and applies an LLM for zero-shot behavioral analysis of student attention.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That geometric pose and gaze coordinates extracted by pre-trained models are sufficient for an LLM to perform accurate zero-shot inference of student attention levels without domain-specific fine-tuning or spatial context.","pith_extraction_headline":"LLMs can perform zero-shot analysis of student attention using only pose and gaze coordinates from classroom videos."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2604.03401/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}