Pith. sign in

REVIEW 3 major objections 5 minor 72 references

Report of 2017 NSF Workshop on Multimedia Challenges, Opportunities and Research Roadmaps

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A consensus workshop report sets out the multimedia research agenda for the next 15 years, centered on seven cross-cutting areas and four application domains.

desk verdict A coherent 2017 NSF workshop roadmap for multimedia research, useful as a snapshot and teaching resource, but its 'most important topics' claim rests on undocumented consensus among 23 mostly US invitees. read the letter →

arxiv 1908.02308 v1 pith:IGDHWWQX submitted 2019-08-06 cs.MM

classification cs.MM
keywords multimediaresearchroadmapmultimodalmachinelearninggroundingknowledgediscoveryinteractionsystemscontentgenerationagenda
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This report tries to establish which research questions will matter most to the multimedia and multimodal field over the next 10 to 15 years, and what progress should be achievable at 5, 10, and 15 years. It argues that seven cross-cutting areas—foundational multimodal methods, knowledge discovery, grounding language to perception, person-centered interaction, content generation, systems, and data and evaluation—deserve concentrated effort, alongside four application areas: education, healthcare, smart infrastructure, and social good. A sympathetic reader would take the report as a community-negotiated priority list, built to guide funding and research planning.

What carries the argument

Structurally, the argument is carried by the consensus workshop process: 23 invited researchers reviewed candidate topics, selected priorities through discussion, and refined them in breakout groups before the whole group. The report's internal engine is a five-part taxonomy of multimodal learning—representation, alignment, fusion, translation, and co-learning—which organizes the foundational chapter and recurs as the frame for challenges. Each topic's roadmap table then pairs current state of the art with 5-, 10-, and 15-year milestones, making the priorities concrete and checkable.

What would settle it

A concrete test would be a retrospective comparison, at the 10-year mark, of progress and funding yield in the report's seven listed cross-cutting areas versus areas it deferred (privacy, social-network multimedia, IoT, public safety). If unlisted areas produced more transformative advances per dollar, the consensus priority ordering would be falsified.

Watch

Extended reading notes

Core claim

The report's central claim is that the future of multimedia research lies in deep integration of multiple modalities—learning, aligning, fusing, translating, and co-learning across heterogeneous data—rather than in unimodal advances. It asserts that the field is now mature enough to move from single-modality successes to joint representations spanning vision, language, and acoustics; to ground language in physical scenes and actions; to infer psychological characteristics from multimodal behavior; to generate personalized and verifiable content; and to build systems whose quality of experience is tied to semantics and context. For each area the report provides a stated timeline, with milestones such as trimodal representations and audio-visual temporal pattern discovery in 5 years, interpretable video models and limited-data learning in 10, and multimodal machine translation and human-behavior generation in 15. The paper's evidence is a collective assessment of state of the art plus a roadmap, not a new experiment.

Load-bearing premise

The roadmap's authority rests on the assumption that a consensus among 23 invited researchers, reached by discussion rather than by data or bibliometric analysis, is a valid way to predict the most important topics over a 10- to 15-year horizon.

Editorial extensions

If this is right

  • If the roadmap is followed, the next five years should produce demonstrable trimodal representations (language, vision, acoustics) and temporal pattern discovery in audio-visual streams.
  • Within ten years, grounding should reach unconstrained physical environments, enabling high-accuracy visual question answering on video and automatic knowledge construction from loosely coupled multimodal sources.
  • Within fifteen years, the report expects multimodal machine translation across languages and cultures, natural multimodal communication with robots, and continuous knowledge discovery from large multilingual streaming sources.
  • For the application areas, the report implies that multimodal analytics can yield ultra-reliable predictions of human state—intention, emotion, cognition, health—and systems-level theories in education and medicine.
  • On content generation, the milestones imply automatic modality recoding, provenance verification, and personalized generation that preserves a common basis for shared experience.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • I read the report's list of 'areas for future discussion'—privacy, social networks, IoT, public safety—as a hedge: these were acknowledged but not roadmapped, and an updated version might well have put privacy at the center.
  • The report implicitly predicts that multimodal integration, not further unimodal scaling, will be the main driver of AI progress in perceptual tasks; this is testable by comparing benchmark gains in visual question answering and captioning against advances in single-modality recognition.
  • A likely blind spot of a consensus of mostly academic participants is industry deployment; the emphasis on provenance and content authenticity suggests the authors already saw this gap, but the roadmap itself contains no measures of industrial adoption.
  • If the roadmap is right, then multimodal training data with natural supervision (for example, captioned video and descriptive audio) should become a first-class research infrastructure priority, since several milestones hinge on learning with limited labeled data.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This manuscript reports the outcomes of a two-day NSF-sponsored workshop held in March 2017 in Washington, DC, with the goal of identifying research areas most important to multimedia and multimodal (MM) research over the next 10–15 years. The report describes the workshop's consensus process, lists 23 invited participants, and identifies seven cross-cutting research areas (foundational methods, knowledge discovery, grounding, person-centered interaction, content generation, systems, and data) and four application areas (education, healthcare, smart infrastructure, and social good). For each area, a chapter summarizes the state of the art, key challenges, and proposed 5-, 10-, and 15-year roadmaps. A final section lists topics identified but not assigned breakout groups, such as privacy, multimodal IoT, and public safety. The document is explicitly a consensus-based roadmap rather than a technical research contribution.

Significance. If its roadmap is accepted by the community, this report could influence research priorities, funding decisions, and collaboration patterns in multimedia and multimodal research. Its strengths include a clearly organized structure, a transparent participant list, and a broad survey of the state of the art as of 2017, with each chapter offering a useful taxonomy of challenges and milestones. The report also explicitly acknowledges its own limitations by listing areas for future discussion. However, its significance is constrained by the lack of documented methodology for topic selection and by several unsupported quantitative claims, which limit its authority as a roadmap. The report contains no equations or derivations, so circularity concerns are minimal; its claims are expert consensus statements rather than proven results.

major comments (3)
  1. [Section 1 (Introduction)] The central claim that the listed research areas are the 'most important' for the next 10–15 years rests entirely on undocumented consensus among 23 invited participants. The Introduction states that topics were 'selected through discussion and consensus' but provides no selection criteria, voting record, discussion protocol, or participant diversity information. This is load-bearing because the report explicitly aims to guide community effort and funding. Moreover, the report itself lists 'Areas for Future Discussion' (privacy, multimodal social networks/HCI, multimodal IoT, and public safety) that were identified but not assigned breakout groups. Without a documented prioritization method, the completeness and representativeness of the selected set cannot be assessed. I recommend adding a methodology section describing how topics were proposed, debated, and selected, and how the listed future-discussion areas were determined to be lower priority.
  2. [Multimodal Knowledge Discovery, Section 3.2 (Empirical-Statistical Techniques)] The claim that 'Analysis of combined multimodal data can yield substantially higher reliabilities than unimodal information sources, in some cases achieving 98-100% accuracies' is unsupported. No specific studies, datasets, tasks, or evaluation conditions are cited for this quantitative range. This claim is used to justify the roadmap's milestones for 'ultra-reliable multimodal-multisensor systems' and 'ultra-reliable models & predictors,' so it is not a peripheral statement. Either provide references to the studies achieving 98–100% accuracy or qualify the number as an aspirational goal, and define what 'ultra-reliable accuracy' means in the summary table.
  3. [Cross-Cutting Research chapters (e.g., Person-Centered Multimodal Interaction, Section 6)] The roadmap milestones are stated as declarative predictions ('will be achieved') without measurable indicators or evaluation criteria. For example, the 5-year milestone 'Detection of 50 basic human multimodal activities' does not specify a benchmark, dataset, or definition of 'basic activity,' and the 10-15 year milestones are similarly unfalsifiable. For a roadmap intended to shape a decade of research, I recommend adding concrete, testable criteria for at least the most prominent milestones (e.g., a named benchmark or a defined accuracy target). This would also make it possible to evaluate the roadmap's own success retrospectively.
minor comments (5)
  1. [Full document] The report contains numerous typographical and formatting errors, including inconsistent spacing around hyphens, garbled figure captions (e.g., the multi-level analysis figure in the Knowledge Discovery chapter), and a misspelled word in the Fundamental Methods chapter ('starting pointing' instead of 'starting point'). A careful proofreading pass is needed.
  2. [References (various chapters)] Several references are cited as 'in press' to The Handbook of Multimodal-Multisensor Interfaces without volume, page, or year details. Please complete these citations or note that they are forthcoming.
  3. [Multimedia and Multimodal Systems chapter] The chapter editors (Wu-Chi Feng, Ketan Mayer-Patel, Balakrishnan Prabhakaran) are not listed in the workshop participant list on the first page. Please clarify their roles (e.g., as invited participants not listed, or as contributing authors).
  4. [Title and citation] The title on the first page uses 'Research Roadmaps' while the report citation on the same page uses 'Research Directions.' Standardize the title across the document and the citation.
  5. [Table of contents] The table of contents lists Chapter 8, 'Data and Challenges,' but the provided text excerpt does not include it. Please confirm that the complete document contains all 14 chapters listed.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: this workshop report is an expert-consensus roadmap with no fitted predictions, derived equations, or load-bearing self-citations.

full rationale

The report's central claim is that the listed cross-cutting areas and application areas are important for multimedia research over the next 10-15 years, and the stated basis is explicit: 'Important topics were selected through discussion and consensus, and then discussed in depth in breakout groups' (Section 1, Introduction). This is a method of expert deliberation, not a derivation from assumed premises, so the roadmap does not reduce to its own inputs in the sense relevant to circularity analysis. There are no equations, fitted parameters, or predicted quantities, and no claim is made that the topic list is forced by a theorem or by prior published results. The chapters cite prior surveys and technical work, including some works by workshop participants, but those citations support state-of-the-art descriptions and milestone proposals rather than the selection of the roadmap topics; none is invoked to forbid alternative topic choices. The report itself identifies a limitation on the consensus method: 'Due to the limited time of the workshop, a few areas were identified but were not assigned to breakout discussion groups during the workshop,' listing Privacy/Personalized MM, Multimodal-multimedia social network and HCI, Multimodal Internet of things, and Public safety utilizing MM data. That is an admitted coverage limitation, not a circular step. The underlying assumption that 23 invited participants can predict field importance is an evidential weakness, not a self-definitional equivalence, and per the hard rules it belongs under correctness risk rather than circularity. Accordingly, no circular steps are identified and the score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities appear because the document contains no derivations or proposed mechanisms. The axioms listed are the background assumptions that give the roadmap its authority.

assumptions (3)
  • domain assumption The taxonomy of multimodal learning (representation, alignment, fusion, translation, co-learning) is an adequate organizing framework for the field.
    Adopted from Baltrušaitis et al. (2017) in the Fundamental Methods chapter, Section 1; the entire chapter's discussion is structured by this taxonomy.
  • domain assumption Expert consensus of 23 invited participants is a reliable method for identifying the most important research areas.
    The Introduction states topics were 'selected through discussion and consensus' among invited participants; the roadmap's authority rests on this premise.
  • domain assumption Current trends in computing hardware, data availability, and machine learning will continue for 10-15 years.
    The roadmaps in every chapter project 5, 10, and 15 year milestones based on extrapolation of current capabilities; the Introduction describes the rapidly changing R&D landscape.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Report of 2017 NSF Workshop on Multimedia Challenges, Opportunities and Research Roadmaps." pith.science (2026). https://pith.science/paper/IGDHWWQX

@misc{pith2026190802308,
  author       = {Pith},
  title        = {Pith review of: Report of 2017 NSF Workshop on Multimedia Challenges, Opportunities and Research Roadmaps},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IGDHWWQX}},
  note         = {Machine review of arXiv:1908.02308}
}
read the original abstract

With the transformative technologies and the rapidly changing global R&D landscape, the multimedia and multimodal community is now faced with many new opportunities and uncertainties. With the open source dissemination platform and pervasive computing resources, new research results are being discovered at an unprecedented pace. In addition, the rapid exchange and influence of ideas across traditional discipline boundaries have made the emphasis on multimedia multimodal research even more important than before. To seize these opportunities and respond to the challenges, we have organized a workshop to specifically address and brainstorm the challenges, opportunities, and research roadmaps for MM research. The two-day workshop, held on March 30 and 31, 2017 in Washington DC, was sponsored by the Information and Intelligent Systems Division of the National Science Foundation of the United States. Twenty-three (23) invited participants were asked to review and identify research areas in the MM field that are most important over the next 10-15 year timeframe. Important topics were selected through discussion and consensus, and then discussed in depth in breakout groups. Breakout groups reported initial discussion results to the whole group, who continued with further extensive deliberation. For each identified topic, a summary was produced after the workshop to describe the main findings, including the state of the art, challenges, and research roadmaps planned for the next 5, 10, and 15 years in the identified area.

Figures

Figures reproduced from arXiv: 1908.02308 by the authors.

Figure 1
Figure 1. Taxonomy for multimodal learning. Problems on multimodal processing may involve combination of these categories. Importantly, crosscutting research in these areas will open opportunities to better understand and interpret multimodal data across domains, serving as instrumental tool for the community. These tools can be generic, working across problems. They can also be specific to determined problems or modalities. … view at source ↗
Figure 1
Figure 1. Synchronized multimodal data involving images, writin [PITH_FULL_IMAGE:figures/full_fig_p021_1.png] view at source ↗
Figure 1
Figure 1. An example of grounding verb frames from a sentence to the video Recent years have seen an increasing amount of work on grounding referring expressions to the perceived environment (Liu and Chai 2015) and grounding sentences (e.g., language instructions) to perceived activities in images or videos or action execution in the environment (Chen and Mooney 2011, Tellex et al., 2011; Artzi and Zottlemoyer 2013; Krishnamu… view at source ↗
Figures from the paper (8 more)
Figure 2
Figure 2. Figure 2: Sample Image Caption Generation from Vinyals et al. (2015) One of the most well investigated problems in language grounding is the automatic generation of natural language descriptions (i.e. captions) of images and videos. There has been significant progress on this ta…
Figure 3
Figure 3. Figure 3: Sample VQA Images and Questions/Answers 2.3 Visual Question Answering Another important task that clearly requires connecting linguistic symbols to continuous perceptual data is answering questions about images and video, i.e. visual question answering (VQA) as shown i…
Figure 4
Figure 4. Figure 4: Visual Relation Detection and Applications in Visual [PITH_FULL_IMAGE:figures/full_fig_p039_4.png]
Figure 1
Figure 1. Figure 1: Generic models for generating content summaries (Mone [PITH_FULL_IMAGE:figures/full_fig_p064_1.png]
Figure 2
Figure 2. Figure 2: Generic multi‐view projections based on narrative str [PITH_FULL_IMAGE:figures/full_fig_p065_2.png]
Figure 4
Figure 4. Figure 4: Generating a collaborative version of Là ci darem la mano from Don Giovanni, by two artists who have never met, using Smule. 3. Challenges 3.1 Labeling multimedia content that can serve as the basis for generation The proliferation of low‐cost, high‐quality media captu…
Figure 1
Figure 1. Figure 1: The Window of Scarcity Framework [Jeffay 1986] One way to reason about multimedia systems research is within a framework often expressed as “The Window of Scarcity” illustrated in the figure above [Jeffay1986]. This conceptual framework describes three main development…
Figure 1
Figure 1. Figure 1: Eight‐year‐old conversing with animated characters, f [PITH_FULL_IMAGE:figures/full_fig_p101_1.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

72 extracted references · 70 canonical work pages

  1. [1]

    Aethon, http://www.aethon.com

  2. [2]

    Intelligent multimedia surveillance system for safer cities

    Takeshi Arikuma and Yasunori Mochizuki, “Intelligent multimedia surveillance system for safer cities” APSIPA Transactions on Signal and Information Processing 5 (2016): 1–8

  3. [4]

    Big Op‐Ed: Shifting Opinions On Surveillance Cameras,

    “Big Op‐Ed: Shifting Opinions On Surveillance Cameras,”, Talk of the Nation, NPR, April 22, 2013, accessed August 1, 2016, http://www.npr.org/2013/04/22/178436355/big‐op‐ed‐ shifting‐opinions‐on‐surveillance‐cameras

  4. [5]

    Resources and Support,

    Carnegie Learning, “Resources and Support,” https://www. carnegielearning.com/resources‐support/

  5. [6]

    CompStat,

    “CompStat,” Wikipedia, last modified July 28, 2016, accessed August 1, 2016, https:// en.wikipedia.org/wiki/CompStat

  6. [7]

    Cubelets,

    “Cubelets,” Modular Robotics, accessed August 1, 2016, http://www.modrobotics.com/ cubelets

  7. [9]

    Effects of Changing Reliability on Trust of Robot Systems,

    Munjal Desai, Mikhail Medvedev, Marynel Vázquez, Sean McSheehy, Sofia Gadea‐ Omelchenko, Christian Bruggeman, Aaron Steinfeld, Holly Yanco, “Effects of Changing Reliability on Trust of Robot Systems,” HRI 2012: Proceedings of the 7th ACM/IEEE Int’l Conference on Human Robot Interaction, 2012

  8. [10]

    HelpMate®, the trackless robotic courier: A perspective on the development of a commercial autonomous mobile robot,

    John M. Evans and Bala Krishnamurthy, “HelpMate®, the trackless robotic courier: A perspective on the development of a commercial autonomous mobile robot,” Lecture Notes in Control and Information Sciences 236, June 18, 2005 (Springer‐Verlag London Limited, 1998), 182–210, accessed August 1, 2016, http://link.springer.com/chapter/10.1007%2FBFb0030806

Show all 72 references
  1. [14]

    Intuitive Surgical, accessed August 1, 2016, http://www.intuitivesurgical.com

  2. [15]

    Predicting Salient Updates for Disaster Summarization,

    Chris Kedzie, Kathleen McKeown and Fernando Diaz, “Predicting Salient Updates for Disaster Summarization,” Proceedings of the Association for Computational Linguistics, Beijing, China, July 2015. 136

  3. [16]

    Internet image archaeology: Automatically tracing the manipulation history of photographs on the web,

    Lyndon Kennedy, and Shih‐Fu Chang. "Internet image archaeology: Automatically tracing the manipulation history of photographs on the web," ACM conference on Multimedia, 2008

  4. [17]

    LAW‐TRAIN, http://www.law‐train.eu/

  5. [18]

    Combining Acoustic‐Prosodic, Lexical, and Phonotactic Features for Automatic Deception Detection,

    Sarah Ita Levitan, Guozhen An, Min Ma, Rivka Levitan, Andrew Rosenberg and Julia Hirschberg “Combining Acoustic‐Prosodic, Lexical, and Phonotactic Features for Automatic Deception Detection,” Interspeech 2016, September. San Francisco, CA

  6. [19]

    SHERLOCK: A Coached Practice Environment for an Electronics Troubleshooting Job,

    Alan Lesgold, Suzanne Lajoie, Marilyn Bunzo, and Gary Eggan, “SHERLOCK: A Coached Practice Environment for an Electronics Troubleshooting Job,” in J. H. Larkin and R. W. Chabay, eds., Computer‐Assisted Instruction and Intelligent Tutoring Systems: Shared Goals and Complementar...

  7. [20]

    Medical Post 23:1985

    "Medical Post 23:1985" (PDF))

  8. [21]

    Meet Dash,

    “Meet Dash,” Wonder Workshop, https://www.makewonder.com/ dash

  9. [22]

    TSA’s 2017 Budget—A Commitment to Security (Part I),

    Peter Neffenger, “TSA’s 2017 Budget—A Commitment to Security (Part I),” Department of Homeland Security, March 1, 2016, accessed August 1, 2016, https://www.tsa.gov/news/ testimony/2016/03/01/hearing‐fy17‐budget‐request‐transportation‐security‐ administration

  10. [24]

    Passive‐blind image forensics,

    Tian‐Tsong Ng, Shih‐Fu Chang, Ching‐Yung Lin, and Qibin Sun. "Passive‐blind image forensics,” Multimedia Security Technologies for Digital Rights 15 (2006): 383‐412

  11. [25]

    Physics‐motivated features for distinguishing photographic images and computer graphics,

    Tian‐Tsong Ng, Shih‐Fu Chang, Jessie Hsu, Lexing Xie, and Mao‐Pei Tsui. "Physics‐motivated features for distinguishing photographic images and computer graphics," ACM conference on Multimedia, 2005

  12. [26]

    Ozobot, http://ozobot.com/

  13. [27]

    The Role of Crime Forecasting in Law Enforcement Operations,

    Walter L. Perry, Brian McInnis, Carter C. Price, Susan Smith, and John S. Hollywood, “The Role of Crime Forecasting in Law Enforcement Operations,” Rand Corporation Report 233 (2013)

  14. [28]

    Pleo rb,

    “Pleo rb,” Innvo Labs, http://www.pleoworld.com/pleo_rb/eng/ lifeform.php

  15. [29]

    ROBODOC, http://www.robodoc.com/professionals.html

  16. [30]

    Milind Tambe, Security and Game Theory: Algorithms, Deployed Systems, Lessons Learned (New York: Cambridge University Press, 2011)

  17. [31]

    Intuitive Surgical Maintains Its Growth Momentum With Strong Growth In Procedure Volumes,

    Trefis Team, “Intuitive Surgical Maintains Its Growth Momentum With Strong Growth In Procedure Volumes,” Forbes, January 22, 2016, accessed August 1, 2016, http://www.forbes.com/ sites/greatspeculations/2016/01/22/intuitive‐surgical‐maintains‐ its‐growth‐momentum‐with‐strong‐g...

  18. [32]

    The Architecture of Why2‐Atlas: A Coach for Qualitative Physics Essay Writing,

    Kurt VanLehn, Pamela W. Jordan, Carolyn P. Rosé, Dumisizwe Bhembe, Michael Böttner, Andy Gaydos, Maxim Makatchev, Umarani Pappuswamy, Michael Ringenberg, Antonio 137 Roque, Stephanie Siler, and Ramesh Srivastava, “The Architecture of Why2‐Atlas: A Coach for Qualitative Physics...

  19. [33]

    Virtual Humans, http://ict.usc.edu/groups/virtual‐humans/

  20. [34]

    Learning Discriminative and Transformation Covariant Local Feature Detectors,

    Xu Zhang, Felix Yu, Svebor Karaman, and Shih‐Fu Chang. "Learning Discriminative and Transformation Covariant Local Feature Detectors," IEEE Computer Vision and Pattern Recognition (CVPR), 2017

  21. [35]

    M., Lindsay, J

    DePaulo, B. M., Lindsay, J. J., Malone, B. E., Muhlenbruck, L., Charlton, K., & Cooper, H. (2003). Cues to deception. Psychological bulletin, 129(1), 74

  22. [36]

    Verification and implementation of language‐based deception indicators in civil and criminal,

    J. Bachenko, E. Fitzpatrick, and M. Schonwetter. 2008. “Verification and implementation of language‐based deception indicators in civil and criminal,” International Conference on Computational Linguistics (Manchester), 1, 41‐48

  23. [37]

    Human behavior and deception detection,

    M. Frank, M. O’Sullivan, and M. Menasco. 2008. “Human behavior and deception detection,” in J. G. Voeller, ed., Wiley Handbook of Science and Technology for Homeland Security, New York: John Wiley & Sons

  24. [38]

    On lying and being lied to: A linguistic analysis of deception

    J. Hancock, L. Curry, S. Goorha, and M. Woodworth. 2008. “On lying and being lied to: A linguistic analysis of deception.” Discourse Processes, 45, 1‐23,

  25. [39]

    Distinguishing deceptive from non‐deceptive speech

    J. Hirschberg et al. 2005. “Distinguishing deceptive from non‐deceptive speech.” Interspeech 2005 (Lisbon)

  26. [40]

    Lying words: Predicting deception from linguistic style

    M. Newman, J. Pennebaker, D. Berry, and J. Richards. 2003. “Lying words: Predicting deception from linguistic style.” Personality and Social Psychology Bulletin, 29. 665‐675

  27. [41]

    Combining Acoustic‐Prosodic, Lexical, and Phonotactic Features for Automatic Deception Detection

    “Combining Acoustic‐Prosodic, Lexical, and Phonotactic Features for Automatic Deception Detection” Sarah Ita Levitan, Guozhen An, Min Ma, Rivka Levitan, Andrew Rosenberg and Julia Hirschberg. Interspeech 2016, September. San Francisco, CA

  28. [42]

    Passive‐blind image forensics,

    Tian‐Tsong Ng, Shih‐Fu Chang, Ching‐Yung Lin, and Qibin Sun "Passive‐blind image forensics,". Multimedia Security Technologies for Digital Rights 15 (2006): 383‐412

  29. [43]

    Internet image archaeology: Automatically tracing the manipulation history of photographs on the web,

    Lyndon Kennedy, and Shih‐Fu Chang "Internet image archaeology: Automatically tracing the manipulation history of photographs on the web,". ACM conference on Multimedia, 2008

  30. [44]

    Learning Discriminative and Transformation Covariant Local Feature Detectors,

    Xu Zhang, Felix Yu, Svebor Karaman, and Shih‐Fu Chang."Learning Discriminative and Transformation Covariant Local Feature Detectors," IEEE Computer Vision and Pattern Recognition (CVPR), 2017

  31. [45]

    Physics‐motivated features for distinguishing photographic images and computer graphics,

    Tian‐Tsong Ng, Shih‐Fu Chang, Jessie Hsu, Lexing Xie, and Mao‐Pei Tsui."Physics‐motivated features for distinguishing photographic images and computer graphics," ACM conference on Multimedia, 2005

  32. [46]

    Predicting Salient Updates for Disaster Summarization,

    Chris Kedzie, Kathleen McKeown and Fernando Diaz, “Predicting Salient Updates for Disaster Summarization,” Proceedings of the Association for Computational Linguistics, Beijing, China, July 2015. 138

  33. [47]

    F Dzogang, T Lansdall‐Welfare, FMPN Team, N Cristianini, Discovering Periodic Patterns in Historical News, PloS one 11 (11), e0165736, 2016

  34. [48]

    1275‐1277, 2011; 10.1145/1989323.1989474

    FLAOUNAS Ilias; ALI Omar; TURCHI MARCO; SNOWSIL Tristan; NICART Florent; DEBIE Tijl; CRISTIANINI Nello; NOAM: News Outlets Analysis and Monitoring System, Proceedings of the 2011 ACM SIGMOD International Conference on Management of data p. 1275‐1277, 2011; 10.1145/1989323.1989474

  35. [49]

    S Sudhahar, G De Fazio, R Franzosi, N Cristianini, ; Network analysis of narrative content in large corporaNatural Language Engineering 21 (1), 81‐112; 2015

  36. [50]

    Horta, N., Luís, R., & Simões, D. (2004). Enhancing the SCORM Modelling Scope. ICALT

  37. [51]

    R.Udall, How the Pioneers of the MOOC Got it Wrong; IEEE Spectrum, 16 Jan 2017; https://spectrum.ieee.org/tech‐talk/at‐work/education/how‐the‐pioneers‐of‐the‐mooc‐ got‐it‐wrong

  38. [52]

    Video Synchronization and Sound Search for Human Rights Documentation and Conflict Monitoring

    Junwei Liang, Susanne Burger, Alex Hauptmann and Jay D. Aronson, “Video Synchronization and Sound Search for Human Rights Documentation and Conflict Monitoring” CMU CHRS Technical Report(June 2016)

  39. [53]

    Summary Report: Workshop on Artificial Intelligence, Video Analysis, and Human Rights, June 6‐7, 2017

    Enrique Piracés, "Summary Report: Workshop on Artificial Intelligence, Video Analysis, and Human Rights, June 6‐7, 2017”, CMU‐CHRS, (August 2017)

  40. [54]

    Gibney, the Scientist who spots fake videos, Nature, October 6 2017, doi:10.1038/nature.2017.22784

    E. Gibney, the Scientist who spots fake videos, Nature, October 6 2017, doi:10.1038/nature.2017.22784

  41. [55]

    DARPA MEDIA FORENSICS, https://www.darpa.mil/program/media‐forensics, Accessed 12/18/2017

  42. [56]

    Smart Parking

    Faiz Shaikh, Nikhilkumar, Omkar Kulkarni, Pratik Jadhav, Saideep Bandarkar A Survey on “Smart Parking” Systems,; International Journal of Innovative Research in Science, Engineering and Technology, Vol. 4, Issue 10, October 2015, DOI:10.15680/IJIRSET.2015.04

  43. [57]

    & Fornalski, P., A smart camera for the surveillance of vehicles in intelligent transportation systems; Multimed Tools Appl (2016) 75: 10471

    Baran, R., Rusc, T. & Fornalski, P., A smart camera for the surveillance of vehicles in intelligent transportation systems; Multimed Tools Appl (2016) 75: 10471. https://doi.org/10.1007/s11042‐015‐3151‐y

  44. [58]

    Michael Bommes, Adrian Fazekas, Tobias Volkenhoff, Markus Oeser, Video Based Intelligent Transportation Systems – State of the Art and Future Development, In Transportation Research Procedia, Volume 14, 2016, Pages 4495‐4504, ISSN 2352‐1465, https://doi.org/10.1016/j.trpro.2...

  45. [59]

    MEDRESPOND, 2017 https://www.medrespond.com/case‐studies/ Accessed 12/20/2017

  46. [60]

    In: CMU‐LTI‐14‐002 Technical Report (2014)

    Wang, Y., Hauptmann, A.: An assistive system for monitoring asthma inhaler usage. In: CMU‐LTI‐14‐002 Technical Report (2014)

  47. [61]

    (2015) Monitoring and Coaching the Use of Home Medical Devices

    Cai Y., Yang Y., Hauptmann A., Wactlar H. (2015) Monitoring and Coaching the Use of Home Medical Devices. In: Briassouli A., Benois‐Pineau J., Hauptmann A. (eds) Health Monitoring and Personalized Feedback using Multimedia Data. Springer, Cham

  48. [62]

    HALYARD, 2017 https://www.halyardhealth.com/solutions/infection‐ prevention/compliance‐monitoring.aspx, Accessed 12/20/2017 139

  49. [63]

    CENTRAK, 2017, https://www.centrak.com/handhygiene‐compliance/ Accessed 12/20/2017

  50. [64]

    The Internet of Things for Health Care: A Comprehensive Survey,

    S. M. R. Islam, D. Kwak, M. H. Kabir, M. Hossain and K. S. Kwak, "The Internet of Things for Health Care: A Comprehensive Survey," in IEEE Access, vol. 3, pp. 678‐708, 2015. doi: 10.1109/ACCESS.2015.2437951

  51. [65]

    IROBOT, 2017; http://www.irobot.com/?_ga=2.150362998.1174985209.1499047579‐ 1611701288.1499047579 Accessed 12/20/2017

  52. [66]

    CYBERNETICZOO, 2017; http://cyberneticzoo.com/tag/electrolux‐trilobite‐robotic‐vacuum‐ cleaner/ Accessed 12/20/2017 140 Training, Infrastructure and Funding 141 Training the Next Generation Chapter Editors: Louis‐Philippe Morency, Carnegie Mellon University Sharon Oviatt, Mona...

  53. [67]

    Introduction Training the next generation of workers is a fundamental challe nge for our modern society. This training enables not only the young generation to learn skills and knowledge for their future career, but it also enables current workers to transition to new jobs and...

  54. [68]

    The focus of these pedagogical resources has focused on building systems and technologies to index and retrieve specific instances from a large collection of videos

    State of the art There have been a considerable number of courses and textbooks developed to train students in the specialized field of multimedia indexing and retrieval. The focus of these pedagogical resources has focused on building systems and technologies to index and ret...

  55. [69]

    Milestones This topic requires multidisciplinary training, which could begin in high school. The topic requires groups with complementary expertise, so Institute‐level organizations would be ideal, and they could be supported internationally and include international scientist...

  56. [70]

    International dimension One of the essential ingredients of multimodal and multimedia is multi‐cultural. Where mono‐ disciplinary research often has clear problems set and tasks to solve, a multi‐modal, multi‐media and multi‐cultural setting is often the playground for emergin...

  57. [71]

    Spence, & B.E

    References Calvert, G., C. Spence, & B.E. Stein, eds. (2004) The Handbook of Multisensory Processing. MIT Press: Cambridge, MA. Chang, S. F. (2018). Frontiers of Multimedia Research. Morgan & Claypool. Oviatt, S.L. & Cohen, P.R. The Paradigm Shift to Multimodality in Contempor...

  58. [72]

    Many conferences, like ACM Multimedia, since inception have been held with regular location rotation over different geographical areas

    Introduction Multimedia and multimodal communities have enjoyed broad participation of researchers and practitioners from many regions in the world. Many conferences, like ACM Multimedia, since inception have been held with regular location rotation over different geographical...

  59. [73]

    Man‐machine interaction is important, but it usually focuses on a controlled environment for a restricted set of purposes

    Examples of Current International Multimedia Initiatives Under the EU Cordis program, the header of man‐machine interaction, multimodal interaction is one of the themes. Man‐machine interaction is important, but it usually focuses on a controlled environment for a restricted s...

  60. [74]

    At the moment, most of that work is still unimodal (e.g., analyzing speech signal features to detect Parkinson’s)

    New Challenges and Problems Calling for International Collaboration ● Multimodal analytics for diagnosing and monitoring medical and health‐related intervention progress are critical areas that are beginning to emerge. At the moment, most of that work is still unimodal (e.g., ...

  61. [75]

    As argued above, these are the times to install these programs between national science foundations across continents

    Mechanisms for Stimulating International Collaboration ● If anything would be appropriate for an internationally funded program of research it would be cross‐cultural multimodal research. As argued above, these are the times to install these programs between national science f...

  62. [76]

    EU CORDIS program, (http://cordis.europa.eu/home_en.html

  63. [77]

    TREC Video Retrieval Evaluation, http://trecvid.nist.gov/

  64. [78]

    opt‐in for research

    Robot Operation System, http://www.ros.org/ 148 Data and Computing Infrastructure Chapter Editor: Alex Hauptmann, Carnegie Mellon University This section represents a summary of informal replies by workshop participants to questions on future data requirements, access to compu...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.