Publications
.
2024. Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training. Proc. of IEEE Tencon, Singapore.
.
2024. Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder. Proc. of IEEE Tencon, Singapore.
.
2020. Acoustic prediction of flowrate: varying liquid jet stream onto a free surface. IEEE International Conference on Signal Processing and Communications (SPCOM).
preprint flow.pdf (1.01 MB)
.
2021. aiSTROM - A roadmap for developing a successful AI strategy. IEEE Access.
.
2026. Aligning Generative Music AI with Human Preferences: Methods and Challenges. Proceedings of AAAI, senior member track.
2511.15038v1.pdf (417.24 KB)
.
2025. Analysis and Synthesis of Audio with AI: from Neurological Disease to Accented Speech and Music.
thesis_Jan.pdf (26.4 MB)
.
2026. APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music. arXiv:2605.03395.
2605.03395v1 (1).pdf (292.45 KB)
.
2025. Are we there yet? A brief survey of Music Emotion Prediction Datasets, Models and Outstanding Challenges IEEE Transactions on Affective Computing.
2406.08809v1.pdf (156.19 KB)
.
2020. Asthmatic versus healthy child classification based on cough and vocalised /a:/ sounds. The Journal of the Acoustical Society of America (JASA). 148, EL253
.
2021. AttendAffectNet – Emotion Prediction of Movie Viewers Using Multimodal Fusion with Self-attention. Sensors. Special issue on Intelligent Sensors: Sensor Based Multi-Modal Emotion Recognition.
sensors-21-08356.pdf (1.03 MB)
.
2021. AttendAffectNet: Self-Attention based Networks for Predicting Affective Responses from Movies. Proceedings of the International Conference on Pattern Recognition (ICPR2020).
2010.11188.pdf (7.07 MB)
.
2025. BandCondiNet: Parallel Transformers-based Conditional Popular Music Generation with Multi-View Features. Expert Systems with Applications. 130059
2407.10462v2.pdf (2.6 MB)
.
2018. Blacklisted speaker identification using triplet neural networks. MCE2018 competition.
SUTD_description.pdf (133.08 KB)
.
2015. Classification and generation of composer-specific music using global feature models and variable neighborhood search. Computer Music Journal. 39(3):91.
papercmj-dh_preprint.pdf (637.63 KB)
.
2025. Coarse-to-Fine Text-to-Music Latent Diffusion. Proceedings of ICASSP.
.
2024. Coarse-to-Fine Text-to-Music Latent Diffusion. Audio Imagination: NeurIPS 2024 Workshop.
.
2015. Compose ≡ compute. 4OR. 13:335–336.
.
2014. Compose=Compute - Computer Generation And Classification Of Music Through Operations Research Methods. PhD Thesis, University of Antwerp. :250.
.
2015. Composer Classification Models for Music-Theory Building. Computational Music Analysis.
Chapter_HerremansEtAl_preprint.pdf (475.26 KB)
.
2012. Composing counterpoint musical scores with variable neighborhood search. Annual Conference of the Belgian Operation Research Society (ORBEL26).
orbel26abs_vnsforcp.pdf (116.85 KB)
.
2013. Composing Fifth Species Counterpoint Music With A Variable Neighborhood Search Algorithm. Expert Systems with Applications. 40
paper_preprint_cp5.pdf (405.75 KB)
.
2012. Composing Fifth Species Counterpoint Music With Variable Neighborhood Search.
wp_cp5.pdf (508 KB)
.
2012. Composing first species counterpoint musical scores with a variable neighbourhood search algorithm. Journal of Mathematics and the Arts. 6:169-189.
.
2022. Computationally Efficient Physics Approximating Neural Networks for Highly Nonlinear Maps. 2022 International Conference on Research in Adaptive and Convergent Systems.
.
2022. Conditional Drums Generation using Compound Word Representations. EvoMUSART (EVO*) - Lecture Notes in Computer Science.
2202.04464.pdf (525.36 KB)
.
2023. Constructing Time-Series Momentum Portfolios with Deep Multi-Task Learning. Expert Systems with Applications. 230(120587)
2306.13661.pdf (707.95 KB)
.
2014. Dance hit song prediction. Journal of New music Research. 43:302.
wp_hit.pdf (689.07 KB)
.
2013. Dance Hit Song Science. International Workshop on Music and Machine Learning.
abstract_preprint_MML2013_DH.pdf (194.82 KB)
.
2024. DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech. Audio Imagination: NeurIPS 2024 Workshop.
.
2020. Data-driven 3D Scene Understanding. PhD
.
2020. A dataset and classification model for Malay, Hindi, Tamil and Chinese music. 13th Workshop on music and machine learning (MML) as part of ECML/PKDD.
2009.04459.pdf (234.8 KB)
.
2021. Deep Neural Network Based Respiratory Pathology Classification Using Cough Sounds. Sensors. 21(16):5555.
2106.12174.pdf (6.52 MB)
.
2024. DeepUnifiedMom: Unified Time-series Momentum Portfolio Construction via Multi-Task Learning with Multi-Gate Mixture of Experts. arXiv:2406.08742.
2406.08742v1.pdf (1.06 MB)
.
2025. Demystifying deep search: a holistic evaluation with hint-free multi-hop questions and factorised metrics. arXiv:2510.05137.
.
2026. Development of Interpretable Deep Learning-based Segmentation Algorithm for Automated Assessment of Oral Diadochokinesis in Progressive Neurological Diseases. Journal of Speech, Language, and Hearing Research.
.
2019. Development of Machine Learning for asthmatic and healthy voluntary cough - a proof of concept study. Applied Sciences. 9(14)
applsci-09-02833.pdf (2.06 MB)
.
2023. DiffRoll: Diffusion-based Generative Music Transcription with Unsupervised Pretraining Capability. ICASSP.
diffroll.pdf (2.2 MB)
.
2024. DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage. Proc. of IEEE Tencon, Singapore.
.
2023. A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music Modeling. Proceedings of the 37th AAAI Conference on Artificial Intelligence.
2212.00973.pdf (1.74 MB)
.
2019. Doppler Invariant Demodulation for Shallow Water Acoustic Communications Using Deep Belief Networks. 16th IEEE Asia Pacific Wireless Communications Symposium (APWCS).
1909.02850.pdf (790.54 KB)
.
2022. Downscaling using Deep Convolutional Autoencoders, a case study for South East Asia. Egusphere preprint.
egusphere-2022-234.pdf (8.99 MB)
.
2015. The effect of repetitive structure on enjoyment and altered states in uplifting trance music. 2nd International Conference on Music and Consciousness (MUSCON 2), Brighton.
AgresEtAl_muscon.pdf (12.47 KB)
.
2016. The Effect of Repetitive Structure on Enjoyment in Uplifting Trance Music. 14th International Conference for Music Perception and Cognition (ICMPC). :280-282.
preprint_trance.pdf (139.27 KB)
.
2021. The Effect of Spectrogram Reconstructions on Automatic Music Transcription:An Alternative Approach to Improve Transcription Accuracy. Proceedings of the International Conference on Pattern Recognition (ICPR2020).
2010.09969.pdf (3.46 MB)
.
2019. The emergence of deep learning: new opportunities for music and audio technologies. Neural Computing and Applications.
main_preprint.pdf (102.16 KB)
.
2026. Emerging AI Technologies for Music: Towards Controllable, Collaborative, and Creative Systems. Proceedings of Machine Learning Research, PMLR 303:1-5, 2026.
bhandari26a.pdf (161.47 KB)
.
2022. EmoMV: Affective Music-Video Correspondence Learning Datasets for Classification and Retrieval. Information Fusion.
SSRN-id4189323.pdf (2.01 MB)
.
2025. End-to-End Text-to-SQL with Dataset Selection: Leveraging LLMs for Adaptive Query Generation. Proceedings of IJCNN, Rome, Italy.
.
2021. Evaluating the Effectiveness of an Augmented Reality Game Promoting Environmental Action. Sustainability. 13(24):13912.
sustainability-13-13912.pdf (16.23 MB)
.
2025. An exploration of controllability in symbolic music infilling. IEEE Access.
.
2013. First species counterpoint generation with VNS and vertical viewpoints. Digital Music Research Network (DMNR+8).
dnmr8_dh_dc.pdf (147.73 KB)
.
2014. First species counterpoint generation with VNS and vertical viewpoints. Annual Conference of the Belgian Operation Research Society (ORBEL28).
orbel28_dh.pdf (216.63 KB)
.
2025. Forecasting Bitcoin Volatility Spikes from Whale Transactions and Cryptoquant Data Using Synthesizer Transformer Models. IEEE Access. 13:117788-117807.
SSRN-id4247684.pdf (5.05 MB)
.
2026. Formalizing Semi-Structured Interviews for Design Requirement Discovery: A Multi-Agent Framework for Clarification and Empathic Interaction. Available at SSRN 7199726.
.
2018. From Context to Concept: Exploring Semantic Relationships in Music with Word2Vec. Neural Computing and Applications.
paper.pdf (1.64 MB)
.
2017. A Functional Taxonomy of Music Generation Systems. ACM Computing Surveys. 50(5):30.
music_generation_survey_dh_preprint.pdf (349.15 KB)
.
2013. FuX, an Android app that generates counterpoint. IEEE Symposium on Computational Intelligence for Creativity and Affective Computing (CICAC). :48-55.
wp_fux.pdf (486.27 KB)
.
2024. Gamification and skills tree. Trends and Foresight Report on Cyber-Physical Learning.
.
2022. A Gaussian mixture classifier model to differentiate respiratory symptoms using phonated /ɑː/ sounds. The 18th Australasian International Conference on Speech Science and Technology (SST).
ahsounds.pdf (1018.01 KB)
.
2015. Generating Fingerings for Polyphonic Piano Music with a Tabu Search Algorithm. Mathematics and Computation in Music. 9110:149-160.
paper_mcm_preprint.pdf (405.73 KB)
.
2017. Generating guitar solos by integer programming. Journal of the Operational Research Society. :971-985.
preprint_guitar_solo_generation_dh.pdf (772.59 KB)
.
2021. Generating Lead Sheets with Affect: A Novel Conditional seq2seq Framework. Proceedings of the International Joint Conference on Neural Networks (IJCNN).
2104.13056.pdf (857.78 KB)
.
2015. Generating music with an optimization algorithm using a Markov based objective function. ORBEL29, Belgian Conference on Operations Research.
orbel29abs.pdf (138.67 KB)
.
2015. Generating structured music for bagana using quality metrics based on Markov models. Expert Systems With Applications. 42 (21)(21):424–7435.
paper-bagana.pdf (1.73 MB)
.
2014. Generating structured music using quality metrics based on Markov models.
wp_bagana.pdf (1.7 MB)
.
2026. Generative AI in Education for SDG 4: Insights from Indonesia and Kazakhstan. Proceedings of the Pacific Asia Conference on Information Systems (PACIS)..
.
2020. Generative Modelling for Controllable Audio Synthesis of Expressive Piano Performance. Workshop on Machine Learning for Music Discover (ML4MD) as part of ICML.
2006.09833.pdf (2.81 MB)
.
2017. Harmonic Structure Predicts the Enjoyment of Uplifting Trance Music. Frontiers in Psychology, Cognitive Science. 7(1999)
agres16ut.pdf (1.15 MB)
.
2022. HEAR 2021: Holistic Evaluation of Audio Representations. Proceedings of Machine Learning Research (PMLR): NeurIPS 2021 Competition Track.
2203.03022.pdf (406.58 KB)
.
2025. HHNAS-AM: Hierarchical Hybrid Neural Architecture Search using Adaptive Mutation Policies. arXiv:2508.14946.
.
2021. Hierarchical Recurrent Neural Networks for Conditional Melody Generation with Long-term Structure. Proceedings of the International Joint Conference on Neural Networks (IJCNN).
2102.09794.pdf (1015.73 KB)
.
2017. Hit Song Prediction Based on Early Adopter Data and Audio Features. The 18th International Society for Music Information Retrieval Conference (ISMIR) - Late Breaking Demo.
paper_preprint_hit.pdf (221.73 KB)
.
2019. A Hybrid Fuzzy Logic-Neural Network Approach For Multi-path Separation Of Underwater Acoustic Signals. 89th IEEE Vehicular Technology Conference.
fuzzy logic.pdf (1.66 MB)
.
2017. IMMA-Emo: A Multimodal Interface for Visualising Score- and Audio-synchronised Emotion Annotations. Audio Mostly.
IMMA-emo_preprint.pdf (1.4 MB)
.
2020. The impact of Audio input representations on neural network based music transcription. Proceedings of the International Joint Conference on Neural Networks (IJCNN).
2001.09989.pdf (1.87 MB)
.
2019. The impact of musical structure on enjoyment and absorptive listening states in trance music. Music and Consciousness 2 - Worlds, Practices, Modalities.
.
2025. ImprovNet: Generating Controllable Musical Improvisations with Iterative Corruption Refinement. Proceedings of IJCNN.
.
2026. JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment. Empirical Methods in Natural Language Processing (EMNLP).
.
2025. JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata. Proceedings of IJCNN, Rome, Italy.
.
2026. KARMA-MV: A Benchmark for Causal Question Answering on Music Videos. arXiv:2605.08175.
2605.08175v1.pdf (3.32 MB)
.
2019. Latent space representation for multi-target speaker detection and identification with a sparse dataset using Triplet neural networks. IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2019).
1910.01463.pdf (934.76 KB)
.
2023. Learning accent representation with multi-level VAE towards controllable speech synthesis. IEEE Spoken Language Technology (SLT) Workshop.
.
2019. Learning Disentangled Representations of Timbre and Pitch for Musical Instrument Sounds Using Gaussian Mixture Variational Autoencoders. ISMIR.
jyun-ismir.pdf (5.62 MB)
.
2026. Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction. Conference on AI Music Creativity (AIMC).
.
2022. A Machine Learning Approach for MIDI to Guitar Tablature Conversion. Sound and Music Computing Conference (SMC).
25.pdf (528.42 KB)
.
2019. Machine Learning Research that Matters for Music Creation: A Case Study. Journal of New Music Research. 48(1):36-55.
concert_paper_preprint.pdf (1.6 MB)
.
2014. Markov Based Quality Metrics For Generating Structured Music With Optimization Techniques. Digital Music Research Network (DMNR+9).
dmrn9_dh.pdf (133.29 KB)
.
2026. Measuring and Mitigating Rapport Bias of Large Language Models under Multi-Agent Social Interactions. Proceedings of ICLR.
.
2026. MelodySim: Measuring Melody-aware Music Similarity for Plagiarism Detection. IEEE Tencon.
.
2023. MERP: A Music Dataset with Emotion Ratings and Raters’ Profile Information. Sensors - Intelligent Sensors. 23(1)
sensors-23-00382 (2).pdf (1.21 MB)
.
2019. Midi Miner – A Python library for tonal tension and track classification. ISMIR - Late Breaking Demo.
midi_miner.pdf (83.7 KB)
.
2024. MidiCaps — A large-scale MIDI dataset with text captions. ISMIR.
2406.02255v1.pdf (699.83 KB)
.
2018. Minimally Simple Binaural Room Modelling Using a Single Feedback Delay Network. Journal of the Audio Engineering Society. 66(10):791-807.
angus_jaes_preprint.pdf (6.39 MB)
.
2024. MIRFLEX: Music Information Retrieval Feature Library for Extraction. ISMIR, Late Breaking Demos.
2411.00469v1.pdf (89.86 KB)
.
2017. Modeling Musical Context with Word2vec. First International Workshop On Deep Learning and Music. 1:11-18.
herremans2017work2vec.pdf (745.8 KB)
]