Machine Learning for Audio, Image and Video Analysis

CHF 85.05
Auf Lager
SKU
T9AVUBH6E71
Stock 1 Verfügbar
Geliefert zwischen Fr., 23.01.2026 und Mo., 26.01.2026

Details

This second edition focuses on audio, image and video data, the three main types of input that machines deal with when interacting with the real world. A set of appendices provides the reader with self-contained introductions to the mathematical background necessary to read the book.
Divided into three main parts, From Perception to Computation introduces methodologies aimed at representing the data in forms suitable for computer processing, especially when it comes to audio and images. Whilst the second part, Machine Learning includes an extensive overview of statistical techniques aimed at addressing three main problems, namely classification (automatically assigning a data sample to one of the classes belonging to a predefined set), clustering (automatically grouping data samples according to the similarity of their properties) and sequence analysis (automatically mapping a sequence of observations into a sequence of human-understandable symbols). The third part Applications shows how the abstract problems defined in the second part underlie technologies capable to perform complex tasks such as the recognition of hand gestures or the transcription of handwritten data.

Machine Learning for Audio, Image and Video Analysis is suitable for students to acquire a solid background in machine learning as well as for practitioners to deepen their knowledge of the state-of-the-art. All application chapters are based on publicly available data and free software packages, thus allowing readers to replicate the experiments.


Presents techniques for extracting features from audio recordings, images and videos Provides the mathematical background required to use the techniques described Covers the most important machine learning techniques for classification, clustering and sequence analysis Includes supplementary material: sn.pub/extras

Klappentext

This second edition focuses on audio, image and video data, the three main types of input that machines deal with when interacting with the real world. A set of appendices provides the reader with self-contained introductions to the mathematical background necessary to read the book.

Divided into three main parts, From Perception to Computation introduces methodologies aimed at representing the data in forms suitable for computer processing, especially when it comes to audio and images. Whilst the second part, Machine Learning includes an extensive overview of statistical techniques aimed at addressing three main problems, namely classification (automatically assigning a data sample to one of the classes belonging to a predefined set), clustering (automatically grouping data samples according to the similarity of their properties) and sequence analysis (automatically mapping a sequence of observations into a sequence of human-understandable symbols). The third part Applications shows how the abstract problems defined in the second part underlie technologies capable to perform complex tasks such as the recognition of hand gestures or the transcription of handwritten data.

Machine Learning for Audio, Image and Video Analysis is suitable for students to acquire a solid background in machine learning as well as for practitioners to deepen their knowledge of the state-of-the-art. All application chapters are based on publicly available data and free software packages, thus allowing readers to replicate the experiments.



Inhalt
Introduction.- Part I: From Perception to Computation.- Audio Acquisition, Representation and Storage.- Image and Video Acquisition, Representation and Storage.- Part II: Machine Learning.- Machine Learning.- Bayesian Theory of Decision.- Clustering Methods.- Foundations of Statistical Learning and Model Selection.- Supervised Neural Networks and Ensemble Methods.- Kernel Methods.- Markovian Models for Sequential Data.- Feature Extraction Methods and Manifold Learning Methods.- Part III: Applications.- Speech and Handwriting Recognition.- Speech and Handwriting Recognition.- Video Segmentation and Keyframe Extraction.- Real-Time Hand Pose Recognition.- Automatic Personality Perception.- Part IV: Appendices.- Appendix A: Statistics.- Appendix B: Signal Processing.- Appendix C: Matrix Algebra.- Appendix D: Mathematical Foundations of Kernel Methods.- Index.

Weitere Informationen

  • Allgemeine Informationen
    • GTIN 09781447168409
    • Genre Information Technology
    • Auflage 2. Aufl.
    • Lesemotiv Verstehen
    • Anzahl Seiten 561
    • Größe H33mm x B153mm x T234mm
    • Jahr 2016
    • EAN 9781447168409
    • Format Kartonierter Einband
    • ISBN 978-1-4471-6840-9
    • Titel Machine Learning for Audio, Image and Video Analysis
    • Autor Francesco Camastra , Alessandro Vinciarelli
    • Untertitel Theory and Applications
    • Gewicht 878g
    • Herausgeber Springer
    • Sprache Englisch

Bewertungen

Schreiben Sie eine Bewertung
Nur registrierte Benutzer können Bewertungen schreiben. Bitte loggen Sie sich ein oder erstellen Sie ein Konto.
Made with ♥ in Switzerland | ©2025 Avento by Gametime AG
Gametime AG | Hohlstrasse 216 | 8004 Zürich | Schweiz | UID: CHE-112.967.470