Hausarbeiten logo
Shop
Shop
Tutorials
De En
Shop
Tutorials
  • How to find your topic
  • How to research effectively
  • How to structure an academic paper
  • How to cite correctly
  • How to format in Word
Trends
FAQ
Go to shop › Computer Sciences - Artificial Intelligence

Multimodal Framework for Video Summarization Using Transformers and Deep Learning

Integrating Visual, Audio, and Textual Data

Title: Multimodal Framework for Video Summarization Using Transformers and Deep Learning

Research Paper (postgraduate) , 2026 , 58 Pages

Autor:in: Suryakanthi Tangirala (Author), Jinaga Neeraja (Author)

Computer Sciences - Artificial Intelligence

Excerpt & Details   Look inside the ebook
Summary Details

This text aims to develop a multimodal deep learning framework for automatically generating concise and meaningful summaries of lengthy video content. The proposed approach combines visual, audio, and textual information to identify relevant video segments while preserving essential context and reducing redundancy.

The framework integrates pretrained models such as Vision Transformers (ViT) for visual feature extraction, Whisper for speech recognition, and BART or DistilBERT for textual and contextual processing. Transformer-based architectures combine these representations to assess segment relevance and generate coherent summaries. The approach can support applications in education, entertainment, surveillance, news analysis, sports, and multimedia information retrieval.

Details

Title
Multimodal Framework for Video Summarization Using Transformers and Deep Learning
Subtitle
Integrating Visual, Audio, and Textual Data
Authors
Suryakanthi Tangirala (Author), Jinaga Neeraja (Author)
Publication Year
2026
Pages
58
Catalog Number
V1759421
ISBN (eBook)
9783389206256
ISBN (Book)
9783389206263
Language
English
Tags
Multimodal Video Summarization Deep Learning Transformers Vision Transformer Automatic Speech Recognition
Product Safety
GRIN Publishing GmbH
Quote paper
Suryakanthi Tangirala (Author), Jinaga Neeraja (Author), 2026, Multimodal Framework for Video Summarization Using Transformers and Deep Learning, Munich, GRIN Verlag, https://www.hausarbeiten.de/document/1759421
Look inside the ebook
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
Excerpt from  58  pages
Hausarbeiten logo
  • Facebook
  • Instagram
  • TikTok
  • Shop
  • Tutorials
  • FAQ
  • Payment & Shipping
  • About us
  • Contact
  • Privacy
  • Terms
  • Imprint
  • Withdraw Contract