REVIEW 3 major objections 5 minor 72 references
Report of 2017 NSF Workshop on Multimedia Challenges, Opportunities and Research Roadmaps
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A consensus workshop report sets out the multimedia research agenda for the next 15 years, centered on seven cross-cutting areas and four application domains.
desk verdict A coherent 2017 NSF workshop roadmap for multimedia research, useful as a snapshot and teaching resource, but its 'most important topics' claim rests on undocumented consensus among 23 mostly US invitees. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Structurally, the argument is carried by the consensus workshop process: 23 invited researchers reviewed candidate topics, selected priorities through discussion, and refined them in breakout groups before the whole group. The report's internal engine is a five-part taxonomy of multimodal learning—representation, alignment, fusion, translation, and co-learning—which organizes the foundational chapter and recurs as the frame for challenges. Each topic's roadmap table then pairs current state of the art with 5-, 10-, and 15-year milestones, making the priorities concrete and checkable.
What would settle it
A concrete test would be a retrospective comparison, at the 10-year mark, of progress and funding yield in the report's seven listed cross-cutting areas versus areas it deferred (privacy, social-network multimedia, IoT, public safety). If unlisted areas produced more transformative advances per dollar, the consensus priority ordering would be falsified.
Extended reading notes
Core claim
The report's central claim is that the future of multimedia research lies in deep integration of multiple modalities—learning, aligning, fusing, translating, and co-learning across heterogeneous data—rather than in unimodal advances. It asserts that the field is now mature enough to move from single-modality successes to joint representations spanning vision, language, and acoustics; to ground language in physical scenes and actions; to infer psychological characteristics from multimodal behavior; to generate personalized and verifiable content; and to build systems whose quality of experience is tied to semantics and context. For each area the report provides a stated timeline, with milestones such as trimodal representations and audio-visual temporal pattern discovery in 5 years, interpretable video models and limited-data learning in 10, and multimodal machine translation and human-behavior generation in 15. The paper's evidence is a collective assessment of state of the art plus a roadmap, not a new experiment.
Load-bearing premise
The roadmap's authority rests on the assumption that a consensus among 23 invited researchers, reached by discussion rather than by data or bibliometric analysis, is a valid way to predict the most important topics over a 10- to 15-year horizon.
Editorial extensions
If this is right
- If the roadmap is followed, the next five years should produce demonstrable trimodal representations (language, vision, acoustics) and temporal pattern discovery in audio-visual streams.
- Within ten years, grounding should reach unconstrained physical environments, enabling high-accuracy visual question answering on video and automatic knowledge construction from loosely coupled multimodal sources.
- Within fifteen years, the report expects multimodal machine translation across languages and cultures, natural multimodal communication with robots, and continuous knowledge discovery from large multilingual streaming sources.
- For the application areas, the report implies that multimodal analytics can yield ultra-reliable predictions of human state—intention, emotion, cognition, health—and systems-level theories in education and medicine.
- On content generation, the milestones imply automatic modality recoding, provenance verification, and personalized generation that preserves a common basis for shared experience.
Reading between the lines
- I read the report's list of 'areas for future discussion'—privacy, social networks, IoT, public safety—as a hedge: these were acknowledged but not roadmapped, and an updated version might well have put privacy at the center.
- The report implicitly predicts that multimodal integration, not further unimodal scaling, will be the main driver of AI progress in perceptual tasks; this is testable by comparing benchmark gains in visual question answering and captioning against advances in single-modality recognition.
- A likely blind spot of a consensus of mostly academic participants is industry deployment; the emphasis on provenance and content authenticity suggests the authors already saw this gap, but the roadmap itself contains no measures of industrial adoption.
- If the roadmap is right, then multimodal training data with natural supervision (for example, captioned video and descriptive audio) should become a first-class research infrastructure priority, since several milestones hinge on learning with limited labeled data.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports the outcomes of a two-day NSF-sponsored workshop held in March 2017 in Washington, DC, with the goal of identifying research areas most important to multimedia and multimodal (MM) research over the next 10–15 years. The report describes the workshop's consensus process, lists 23 invited participants, and identifies seven cross-cutting research areas (foundational methods, knowledge discovery, grounding, person-centered interaction, content generation, systems, and data) and four application areas (education, healthcare, smart infrastructure, and social good). For each area, a chapter summarizes the state of the art, key challenges, and proposed 5-, 10-, and 15-year roadmaps. A final section lists topics identified but not assigned breakout groups, such as privacy, multimodal IoT, and public safety. The document is explicitly a consensus-based roadmap rather than a technical research contribution.
Significance. If its roadmap is accepted by the community, this report could influence research priorities, funding decisions, and collaboration patterns in multimedia and multimodal research. Its strengths include a clearly organized structure, a transparent participant list, and a broad survey of the state of the art as of 2017, with each chapter offering a useful taxonomy of challenges and milestones. The report also explicitly acknowledges its own limitations by listing areas for future discussion. However, its significance is constrained by the lack of documented methodology for topic selection and by several unsupported quantitative claims, which limit its authority as a roadmap. The report contains no equations or derivations, so circularity concerns are minimal; its claims are expert consensus statements rather than proven results.
major comments (3)
- [Section 1 (Introduction)] The central claim that the listed research areas are the 'most important' for the next 10–15 years rests entirely on undocumented consensus among 23 invited participants. The Introduction states that topics were 'selected through discussion and consensus' but provides no selection criteria, voting record, discussion protocol, or participant diversity information. This is load-bearing because the report explicitly aims to guide community effort and funding. Moreover, the report itself lists 'Areas for Future Discussion' (privacy, multimodal social networks/HCI, multimodal IoT, and public safety) that were identified but not assigned breakout groups. Without a documented prioritization method, the completeness and representativeness of the selected set cannot be assessed. I recommend adding a methodology section describing how topics were proposed, debated, and selected, and how the listed future-discussion areas were determined to be lower priority.
- [Multimodal Knowledge Discovery, Section 3.2 (Empirical-Statistical Techniques)] The claim that 'Analysis of combined multimodal data can yield substantially higher reliabilities than unimodal information sources, in some cases achieving 98-100% accuracies' is unsupported. No specific studies, datasets, tasks, or evaluation conditions are cited for this quantitative range. This claim is used to justify the roadmap's milestones for 'ultra-reliable multimodal-multisensor systems' and 'ultra-reliable models & predictors,' so it is not a peripheral statement. Either provide references to the studies achieving 98–100% accuracy or qualify the number as an aspirational goal, and define what 'ultra-reliable accuracy' means in the summary table.
- [Cross-Cutting Research chapters (e.g., Person-Centered Multimodal Interaction, Section 6)] The roadmap milestones are stated as declarative predictions ('will be achieved') without measurable indicators or evaluation criteria. For example, the 5-year milestone 'Detection of 50 basic human multimodal activities' does not specify a benchmark, dataset, or definition of 'basic activity,' and the 10-15 year milestones are similarly unfalsifiable. For a roadmap intended to shape a decade of research, I recommend adding concrete, testable criteria for at least the most prominent milestones (e.g., a named benchmark or a defined accuracy target). This would also make it possible to evaluate the roadmap's own success retrospectively.
minor comments (5)
- [Full document] The report contains numerous typographical and formatting errors, including inconsistent spacing around hyphens, garbled figure captions (e.g., the multi-level analysis figure in the Knowledge Discovery chapter), and a misspelled word in the Fundamental Methods chapter ('starting pointing' instead of 'starting point'). A careful proofreading pass is needed.
- [References (various chapters)] Several references are cited as 'in press' to The Handbook of Multimodal-Multisensor Interfaces without volume, page, or year details. Please complete these citations or note that they are forthcoming.
- [Multimedia and Multimodal Systems chapter] The chapter editors (Wu-Chi Feng, Ketan Mayer-Patel, Balakrishnan Prabhakaran) are not listed in the workshop participant list on the first page. Please clarify their roles (e.g., as invited participants not listed, or as contributing authors).
- [Title and citation] The title on the first page uses 'Research Roadmaps' while the report citation on the same page uses 'Research Directions.' Standardize the title across the document and the citation.
- [Table of contents] The table of contents lists Chapter 8, 'Data and Challenges,' but the provided text excerpt does not include it. Please confirm that the complete document contains all 14 chapters listed.
Circularity Check
No circularity: this workshop report is an expert-consensus roadmap with no fitted predictions, derived equations, or load-bearing self-citations.
full rationale
The report's central claim is that the listed cross-cutting areas and application areas are important for multimedia research over the next 10-15 years, and the stated basis is explicit: 'Important topics were selected through discussion and consensus, and then discussed in depth in breakout groups' (Section 1, Introduction). This is a method of expert deliberation, not a derivation from assumed premises, so the roadmap does not reduce to its own inputs in the sense relevant to circularity analysis. There are no equations, fitted parameters, or predicted quantities, and no claim is made that the topic list is forced by a theorem or by prior published results. The chapters cite prior surveys and technical work, including some works by workshop participants, but those citations support state-of-the-art descriptions and milestone proposals rather than the selection of the roadmap topics; none is invoked to forbid alternative topic choices. The report itself identifies a limitation on the consensus method: 'Due to the limited time of the workshop, a few areas were identified but were not assigned to breakout discussion groups during the workshop,' listing Privacy/Personalized MM, Multimodal-multimedia social network and HCI, Multimodal Internet of things, and Public safety utilizing MM data. That is an admitted coverage limitation, not a circular step. The underlying assumption that 23 invited participants can predict field importance is an evidential weakness, not a self-definitional equivalence, and per the hard rules it belongs under correctness risk rather than circularity. Accordingly, no circular steps are identified and the score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The taxonomy of multimodal learning (representation, alignment, fusion, translation, co-learning) is an adequate organizing framework for the field.
- domain assumption Expert consensus of 23 invited participants is a reliable method for identifying the most important research areas.
- domain assumption Current trends in computing hardware, data availability, and machine learning will continue for 10-15 years.
Cite this review
Pith. "Pith review of Report of 2017 NSF Workshop on Multimedia Challenges, Opportunities and Research Roadmaps." pith.science (2026). https://pith.science/paper/IGDHWWQX
@misc{pith2026190802308,
author = {Pith},
title = {Pith review of: Report of 2017 NSF Workshop on Multimedia Challenges, Opportunities and Research Roadmaps},
year = {2026},
howpublished = {\url{https://pith.science/paper/IGDHWWQX}},
note = {Machine review of arXiv:1908.02308}
}
read the original abstract
With the transformative technologies and the rapidly changing global R&D landscape, the multimedia and multimodal community is now faced with many new opportunities and uncertainties. With the open source dissemination platform and pervasive computing resources, new research results are being discovered at an unprecedented pace. In addition, the rapid exchange and influence of ideas across traditional discipline boundaries have made the emphasis on multimedia multimodal research even more important than before. To seize these opportunities and respond to the challenges, we have organized a workshop to specifically address and brainstorm the challenges, opportunities, and research roadmaps for MM research. The two-day workshop, held on March 30 and 31, 2017 in Washington DC, was sponsored by the Information and Intelligent Systems Division of the National Science Foundation of the United States. Twenty-three (23) invited participants were asked to review and identify research areas in the MM field that are most important over the next 10-15 year timeframe. Important topics were selected through discussion and consensus, and then discussed in depth in breakout groups. Breakout groups reported initial discussion results to the whole group, who continued with further extensive deliberation. For each identified topic, a summary was produced after the workshop to describe the main findings, including the state of the art, challenges, and research roadmaps planned for the next 5, 10, and 15 years in the identified area.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Aethon, http://www.aethon.com
-
[2]
Intelligent multimedia surveillance system for safer cities
Takeshi Arikuma and Yasunori Mochizuki, “Intelligent multimedia surveillance system for safer cities” APSIPA Transactions on Signal and Information Processing 5 (2016): 1–8
work page 2016
-
[4]
Big Op‐Ed: Shifting Opinions On Surveillance Cameras,
“Big Op‐Ed: Shifting Opinions On Surveillance Cameras,”, Talk of the Nation, NPR, April 22, 2013, accessed August 1, 2016, http://www.npr.org/2013/04/22/178436355/big‐op‐ed‐ shifting‐opinions‐on‐surveillance‐cameras
work page 2013
-
[5]
Carnegie Learning, “Resources and Support,” https://www. carnegielearning.com/resources‐support/
- [6]
- [7]
-
[9]
Effects of Changing Reliability on Trust of Robot Systems,
Munjal Desai, Mikhail Medvedev, Marynel Vázquez, Sean McSheehy, Sofia Gadea‐ Omelchenko, Christian Bruggeman, Aaron Steinfeld, Holly Yanco, “Effects of Changing Reliability on Trust of Robot Systems,” HRI 2012: Proceedings of the 7th ACM/IEEE Int’l Conference on Human Robot Interaction, 2012
work page 2012
-
[10]
John M. Evans and Bala Krishnamurthy, “HelpMate®, the trackless robotic courier: A perspective on the development of a commercial autonomous mobile robot,” Lecture Notes in Control and Information Sciences 236, June 18, 2005 (Springer‐Verlag London Limited, 1998), 182–210, accessed August 1, 2016, http://link.springer.com/chapter/10.1007%2FBFb0030806
work page 2005
Show all 72 references
-
[14]
Intuitive Surgical, accessed August 1, 2016, http://www.intuitivesurgical.com
2016
-
[15]
Predicting Salient Updates for Disaster Summarization,
Chris Kedzie, Kathleen McKeown and Fernando Diaz, “Predicting Salient Updates for Disaster Summarization,” Proceedings of the Association for Computational Linguistics, Beijing, China, July 2015. 136
2015
-
[16]
Internet image archaeology: Automatically tracing the manipulation history of photographs on the web,
Lyndon Kennedy, and Shih‐Fu Chang. "Internet image archaeology: Automatically tracing the manipulation history of photographs on the web," ACM conference on Multimedia, 2008
2008
-
[17]
LAW‐TRAIN, http://www.law‐train.eu/
-
[18]
Combining Acoustic‐Prosodic, Lexical, and Phonotactic Features for Automatic Deception Detection,
Sarah Ita Levitan, Guozhen An, Min Ma, Rivka Levitan, Andrew Rosenberg and Julia Hirschberg “Combining Acoustic‐Prosodic, Lexical, and Phonotactic Features for Automatic Deception Detection,” Interspeech 2016, September. San Francisco, CA
2016
-
[19]
SHERLOCK: A Coached Practice Environment for an Electronics Troubleshooting Job,
Alan Lesgold, Suzanne Lajoie, Marilyn Bunzo, and Gary Eggan, “SHERLOCK: A Coached Practice Environment for an Electronics Troubleshooting Job,” in J. H. Larkin and R. W. Chabay, eds., Computer‐Assisted Instruction and Intelligent Tutoring Systems: Shared Goals and Complementar...
1988
-
[20]
Medical Post 23:1985
"Medical Post 23:1985" (PDF))
1985
-
[21]
Meet Dash,
“Meet Dash,” Wonder Workshop, https://www.makewonder.com/ dash
-
[22]
TSA’s 2017 Budget—A Commitment to Security (Part I),
Peter Neffenger, “TSA’s 2017 Budget—A Commitment to Security (Part I),” Department of Homeland Security, March 1, 2016, accessed August 1, 2016, https://www.tsa.gov/news/ testimony/2016/03/01/hearing‐fy17‐budget‐request‐transportation‐security‐ administration
2017
-
[24]
Passive‐blind image forensics,
Tian‐Tsong Ng, Shih‐Fu Chang, Ching‐Yung Lin, and Qibin Sun. "Passive‐blind image forensics,” Multimedia Security Technologies for Digital Rights 15 (2006): 383‐412
2006
-
[25]
Physics‐motivated features for distinguishing photographic images and computer graphics,
Tian‐Tsong Ng, Shih‐Fu Chang, Jessie Hsu, Lexing Xie, and Mao‐Pei Tsui. "Physics‐motivated features for distinguishing photographic images and computer graphics," ACM conference on Multimedia, 2005
2005
-
[26]
Ozobot, http://ozobot.com/
-
[27]
The Role of Crime Forecasting in Law Enforcement Operations,
Walter L. Perry, Brian McInnis, Carter C. Price, Susan Smith, and John S. Hollywood, “The Role of Crime Forecasting in Law Enforcement Operations,” Rand Corporation Report 233 (2013)
2013
-
[28]
Pleo rb,
“Pleo rb,” Innvo Labs, http://www.pleoworld.com/pleo_rb/eng/ lifeform.php
-
[29]
ROBODOC, http://www.robodoc.com/professionals.html
-
[30]
Milind Tambe, Security and Game Theory: Algorithms, Deployed Systems, Lessons Learned (New York: Cambridge University Press, 2011)
2011
-
[31]
Intuitive Surgical Maintains Its Growth Momentum With Strong Growth In Procedure Volumes,
Trefis Team, “Intuitive Surgical Maintains Its Growth Momentum With Strong Growth In Procedure Volumes,” Forbes, January 22, 2016, accessed August 1, 2016, http://www.forbes.com/ sites/greatspeculations/2016/01/22/intuitive‐surgical‐maintains‐ its‐growth‐momentum‐with‐strong‐g...
2016
-
[32]
The Architecture of Why2‐Atlas: A Coach for Qualitative Physics Essay Writing,
Kurt VanLehn, Pamela W. Jordan, Carolyn P. Rosé, Dumisizwe Bhembe, Michael Böttner, Andy Gaydos, Maxim Makatchev, Umarani Pappuswamy, Michael Ringenberg, Antonio 137 Roque, Stephanie Siler, and Ramesh Srivastava, “The Architecture of Why2‐Atlas: A Coach for Qualitative Physics...
2002
-
[33]
Virtual Humans, http://ict.usc.edu/groups/virtual‐humans/
-
[34]
Learning Discriminative and Transformation Covariant Local Feature Detectors,
Xu Zhang, Felix Yu, Svebor Karaman, and Shih‐Fu Chang. "Learning Discriminative and Transformation Covariant Local Feature Detectors," IEEE Computer Vision and Pattern Recognition (CVPR), 2017
2017
-
[35]
M., Lindsay, J
DePaulo, B. M., Lindsay, J. J., Malone, B. E., Muhlenbruck, L., Charlton, K., & Cooper, H. (2003). Cues to deception. Psychological bulletin, 129(1), 74
2003
-
[36]
Verification and implementation of language‐based deception indicators in civil and criminal,
J. Bachenko, E. Fitzpatrick, and M. Schonwetter. 2008. “Verification and implementation of language‐based deception indicators in civil and criminal,” International Conference on Computational Linguistics (Manchester), 1, 41‐48
2008
-
[37]
Human behavior and deception detection,
M. Frank, M. O’Sullivan, and M. Menasco. 2008. “Human behavior and deception detection,” in J. G. Voeller, ed., Wiley Handbook of Science and Technology for Homeland Security, New York: John Wiley & Sons
2008
-
[38]
On lying and being lied to: A linguistic analysis of deception
J. Hancock, L. Curry, S. Goorha, and M. Woodworth. 2008. “On lying and being lied to: A linguistic analysis of deception.” Discourse Processes, 45, 1‐23,
2008
-
[39]
Distinguishing deceptive from non‐deceptive speech
J. Hirschberg et al. 2005. “Distinguishing deceptive from non‐deceptive speech.” Interspeech 2005 (Lisbon)
2005
-
[40]
Lying words: Predicting deception from linguistic style
M. Newman, J. Pennebaker, D. Berry, and J. Richards. 2003. “Lying words: Predicting deception from linguistic style.” Personality and Social Psychology Bulletin, 29. 665‐675
2003
-
[41]
Combining Acoustic‐Prosodic, Lexical, and Phonotactic Features for Automatic Deception Detection
“Combining Acoustic‐Prosodic, Lexical, and Phonotactic Features for Automatic Deception Detection” Sarah Ita Levitan, Guozhen An, Min Ma, Rivka Levitan, Andrew Rosenberg and Julia Hirschberg. Interspeech 2016, September. San Francisco, CA
2016
-
[42]
Passive‐blind image forensics,
Tian‐Tsong Ng, Shih‐Fu Chang, Ching‐Yung Lin, and Qibin Sun "Passive‐blind image forensics,". Multimedia Security Technologies for Digital Rights 15 (2006): 383‐412
2006
-
[43]
Internet image archaeology: Automatically tracing the manipulation history of photographs on the web,
Lyndon Kennedy, and Shih‐Fu Chang "Internet image archaeology: Automatically tracing the manipulation history of photographs on the web,". ACM conference on Multimedia, 2008
2008
-
[44]
Learning Discriminative and Transformation Covariant Local Feature Detectors,
Xu Zhang, Felix Yu, Svebor Karaman, and Shih‐Fu Chang."Learning Discriminative and Transformation Covariant Local Feature Detectors," IEEE Computer Vision and Pattern Recognition (CVPR), 2017
2017
-
[45]
Physics‐motivated features for distinguishing photographic images and computer graphics,
Tian‐Tsong Ng, Shih‐Fu Chang, Jessie Hsu, Lexing Xie, and Mao‐Pei Tsui."Physics‐motivated features for distinguishing photographic images and computer graphics," ACM conference on Multimedia, 2005
2005
-
[46]
Predicting Salient Updates for Disaster Summarization,
Chris Kedzie, Kathleen McKeown and Fernando Diaz, “Predicting Salient Updates for Disaster Summarization,” Proceedings of the Association for Computational Linguistics, Beijing, China, July 2015. 138
2015
-
[47]
F Dzogang, T Lansdall‐Welfare, FMPN Team, N Cristianini, Discovering Periodic Patterns in Historical News, PloS one 11 (11), e0165736, 2016
2016
-
[48]
1275‐1277, 2011; 10.1145/1989323.1989474
FLAOUNAS Ilias; ALI Omar; TURCHI MARCO; SNOWSIL Tristan; NICART Florent; DEBIE Tijl; CRISTIANINI Nello; NOAM: News Outlets Analysis and Monitoring System, Proceedings of the 2011 ACM SIGMOD International Conference on Management of data p. 1275‐1277, 2011; 10.1145/1989323.1989474
2011
-
[49]
S Sudhahar, G De Fazio, R Franzosi, N Cristianini, ; Network analysis of narrative content in large corporaNatural Language Engineering 21 (1), 81‐112; 2015
2015
-
[50]
Horta, N., Luís, R., & Simões, D. (2004). Enhancing the SCORM Modelling Scope. ICALT
2004
-
[51]
R.Udall, How the Pioneers of the MOOC Got it Wrong; IEEE Spectrum, 16 Jan 2017; https://spectrum.ieee.org/tech‐talk/at‐work/education/how‐the‐pioneers‐of‐the‐mooc‐ got‐it‐wrong
2017
-
[52]
Video Synchronization and Sound Search for Human Rights Documentation and Conflict Monitoring
Junwei Liang, Susanne Burger, Alex Hauptmann and Jay D. Aronson, “Video Synchronization and Sound Search for Human Rights Documentation and Conflict Monitoring” CMU CHRS Technical Report(June 2016)
2016
-
[53]
Summary Report: Workshop on Artificial Intelligence, Video Analysis, and Human Rights, June 6‐7, 2017
Enrique Piracés, "Summary Report: Workshop on Artificial Intelligence, Video Analysis, and Human Rights, June 6‐7, 2017”, CMU‐CHRS, (August 2017)
2017
-
[54]
Gibney, the Scientist who spots fake videos, Nature, October 6 2017, doi:10.1038/nature.2017.22784
E. Gibney, the Scientist who spots fake videos, Nature, October 6 2017, doi:10.1038/nature.2017.22784
2017
-
[55]
DARPA MEDIA FORENSICS, https://www.darpa.mil/program/media‐forensics, Accessed 12/18/2017
2017
-
[56]
Smart Parking
Faiz Shaikh, Nikhilkumar, Omkar Kulkarni, Pratik Jadhav, Saideep Bandarkar A Survey on “Smart Parking” Systems,; International Journal of Innovative Research in Science, Engineering and Technology, Vol. 4, Issue 10, October 2015, DOI:10.15680/IJIRSET.2015.04
2015 doi
-
[57]
& Fornalski, P., A smart camera for the surveillance of vehicles in intelligent transportation systems; Multimed Tools Appl (2016) 75: 10471
Baran, R., Rusc, T. & Fornalski, P., A smart camera for the surveillance of vehicles in intelligent transportation systems; Multimed Tools Appl (2016) 75: 10471. https://doi.org/10.1007/s11042‐015‐3151‐y
2016 doi
-
[58]
Michael Bommes, Adrian Fazekas, Tobias Volkenhoff, Markus Oeser, Video Based Intelligent Transportation Systems – State of the Art and Future Development, In Transportation Research Procedia, Volume 14, 2016, Pages 4495‐4504, ISSN 2352‐1465, https://doi.org/10.1016/j.trpro.2...
2016 doi
-
[59]
MEDRESPOND, 2017 https://www.medrespond.com/case‐studies/ Accessed 12/20/2017
2017
-
[60]
In: CMU‐LTI‐14‐002 Technical Report (2014)
Wang, Y., Hauptmann, A.: An assistive system for monitoring asthma inhaler usage. In: CMU‐LTI‐14‐002 Technical Report (2014)
2014
-
[61]
(2015) Monitoring and Coaching the Use of Home Medical Devices
Cai Y., Yang Y., Hauptmann A., Wactlar H. (2015) Monitoring and Coaching the Use of Home Medical Devices. In: Briassouli A., Benois‐Pineau J., Hauptmann A. (eds) Health Monitoring and Personalized Feedback using Multimedia Data. Springer, Cham
2015
-
[62]
HALYARD, 2017 https://www.halyardhealth.com/solutions/infection‐ prevention/compliance‐monitoring.aspx, Accessed 12/20/2017 139
2017
-
[63]
CENTRAK, 2017, https://www.centrak.com/handhygiene‐compliance/ Accessed 12/20/2017
2017
-
[64]
The Internet of Things for Health Care: A Comprehensive Survey,
S. M. R. Islam, D. Kwak, M. H. Kabir, M. Hossain and K. S. Kwak, "The Internet of Things for Health Care: A Comprehensive Survey," in IEEE Access, vol. 3, pp. 678‐708, 2015. doi: 10.1109/ACCESS.2015.2437951
2015
-
[65]
IROBOT, 2017; http://www.irobot.com/?_ga=2.150362998.1174985209.1499047579‐ 1611701288.1499047579 Accessed 12/20/2017
2017
-
[66]
CYBERNETICZOO, 2017; http://cyberneticzoo.com/tag/electrolux‐trilobite‐robotic‐vacuum‐ cleaner/ Accessed 12/20/2017 140 Training, Infrastructure and Funding 141 Training the Next Generation Chapter Editors: Louis‐Philippe Morency, Carnegie Mellon University Sharon Oviatt, Mona...
2017
-
[67]
Introduction Training the next generation of workers is a fundamental challe nge for our modern society. This training enables not only the young generation to learn skills and knowledge for their future career, but it also enables current workers to transition to new jobs and...
-
[68]
The focus of these pedagogical resources has focused on building systems and technologies to index and retrieve specific instances from a large collection of videos
State of the art There have been a considerable number of courses and textbooks developed to train students in the specialized field of multimedia indexing and retrieval. The focus of these pedagogical resources has focused on building systems and technologies to index and ret...
2015
-
[69]
Milestones This topic requires multidisciplinary training, which could begin in high school. The topic requires groups with complementary expertise, so Institute‐level organizations would be ideal, and they could be supported internationally and include international scientist...
-
[70]
International dimension One of the essential ingredients of multimodal and multimedia is multi‐cultural. Where mono‐ disciplinary research often has clear problems set and tasks to solve, a multi‐modal, multi‐media and multi‐cultural setting is often the playground for emergin...
-
[71]
Spence, & B.E
References Calvert, G., C. Spence, & B.E. Stein, eds. (2004) The Handbook of Multisensory Processing. MIT Press: Cambridge, MA. Chang, S. F. (2018). Frontiers of Multimedia Research. Morgan & Claypool. Oviatt, S.L. & Cohen, P.R. The Paradigm Shift to Multimodality in Contempor...
2004
-
[72]
Many conferences, like ACM Multimedia, since inception have been held with regular location rotation over different geographical areas
Introduction Multimedia and multimodal communities have enjoyed broad participation of researchers and practitioners from many regions in the world. Many conferences, like ACM Multimedia, since inception have been held with regular location rotation over different geographical...
-
[73]
Man‐machine interaction is important, but it usually focuses on a controlled environment for a restricted set of purposes
Examples of Current International Multimedia Initiatives Under the EU Cordis program, the header of man‐machine interaction, multimodal interaction is one of the themes. Man‐machine interaction is important, but it usually focuses on a controlled environment for a restricted s...
2011
-
[74]
At the moment, most of that work is still unimodal (e.g., analyzing speech signal features to detect Parkinson’s)
New Challenges and Problems Calling for International Collaboration ● Multimodal analytics for diagnosing and monitoring medical and health‐related intervention progress are critical areas that are beginning to emerge. At the moment, most of that work is still unimodal (e.g., ...
-
[75]
As argued above, these are the times to install these programs between national science foundations across continents
Mechanisms for Stimulating International Collaboration ● If anything would be appropriate for an internationally funded program of research it would be cross‐cultural multimodal research. As argued above, these are the times to install these programs between national science f...
-
[76]
EU CORDIS program, (http://cordis.europa.eu/home_en.html
-
[77]
TREC Video Retrieval Evaluation, http://trecvid.nist.gov/
-
[78]
opt‐in for research
Robot Operation System, http://www.ros.org/ 148 Data and Computing Infrastructure Chapter Editor: Alex Hauptmann, Carnegie Mellon University This section represents a summary of informal replies by workshop participants to questions on future data requirements, access to compu...
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.