Full-Text Article
The complete peer-reviewed manuscript is published as a PDF — this is the official version of record for this article.
Devle et al. Res. Trends Int. J. Technol. Innov., July - September 2026, 1 (3) : 11-17
Department of Entc Engineering, Pune, Maharashtra, India
Article History
Accepted : 06 Aug 2026
Published : 08 Aug 2026
Publication Issue
Volume 1, Issue 3
July - September 2026
Page Number11–17
Paper NumberRTIJTI-2026-000046
Recommending music that matches or regulates a listener's emotional state requires knowing that state first, and most commercial recommenders skip this step entirely, relying instead on listening history and collaborative filtering that say nothing about how the listener feels right now. This paper presents a multimodal emotion-driven recommendation system that infers a listener's affective state from three signals captured during a session: facial expression from the device camera, sentiment extracted from any text the listener enters (search queries, chat, or captions), and, where available, vocal tone from a short voice sample. Each modality is scored independently by a lightweight classifier and the three scores are combined through a confidence-weighted late fusion step, so a modality that is unavailable or low-confidence in a given session, such as a blank camera feed, does not silently distort the result. The fused output is a position on a two-dimensional valence-arousal plane rather than a single discrete label, which is then matched against a track catalogue annotated with the same valence-arousal coordinates derived from audio features such as tempo, energy, and mode. We describe the fusion architecture, the mapping from emotional state to track features, and a feedback loop that adjusts future recommendations when a listener skips a suggested track quickly. We evaluate the system on a constructed listening-session dataset for classification accuracy per modality, fusion accuracy against self-reported mood, and recommendation acceptance rate compared to a collaborative-filtering baseline. Index Terms - Emotion Recognition, Multimodal Fusion, Music Recommendation, Valence-Arousal Model, Affective Computing, Session-Based Feedback
Keywords - Emotion Recognition, Multimodal Fusion, Music Recommendation, Valence-Arousal Model, Affective Computing, Session-Based Feedback
The complete peer-reviewed manuscript is published as a PDF — this is the official version of record for this article.
© 2026 The Author(s). Published by RTIJTI Editorial Office. This is an open access article under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
Chetan Devle , Prof. Shelke S. A (2026). Emotion-Driven Music Recommendation System Using Multimodal Data. Research Trends International Journal of Technology and Innovation, 1(3), 11-17. https://doi.org/10.5555/rtijti.2026.admin.7b5701ab