Music signal processing · Interspeech 2026

BeatGain: rhythm-aware music enhancement for cochlear-implant listening

BeatGain extends stem-based music preprocessing by using metrical information to emphasize rhythm-relevant percussive events and reduce weaker off-grid events.

This website accompanies the Interspeech 2026 paper on BeatGain. Its main purpose is to make the evaluated processing conditions directly audible through the listening examples below.

It also provides a concise, interactive overview of the motivation, the rhythm-aware algorithm, the comparison methods, and the central study results. The page is intended as an accessible companion to the paper rather than a replacement for the full publication.

Cochlear implants can restore access to sound and often support speech understanding, but music is frequently perceived as less natural or enjoyable. Broad stimulation regions, place–frequency mismatches, and compression reduce spectral resolution. Pitch and timbre are particularly affected, whereas temporal-envelope information is comparatively better preserved [1] [2] [3] [4].

Consequently, CI listeners tend to value a stable beat, prominent vocals, and reduced instrumental complexity [1] [5] [6] [7]. Existing music enhancement work has mainly followed two paths: source separation and remixing, or spectral simplification through methods such as principal-component and subspace approaches. HPSS and stem separation can independently rebalance harmonic and percussive material [8] [9]; spectral methods reduce competing structure. These methods do not explicitly target the temporal organization of rhythm.

BeatGain addresses that complement. In a common 4/4 meter, quarter-note beats are strongest, followed by regular eighth- and sixteenth-note subdivisions. Events on weaker metrical positions can create syncopation and increase perceived rhythmic complexity [10] [11] [12]. The hypothesis is therefore deliberately narrow: amplify percussive events aligned with salient positions and attenuate weaker off-grid events, while retaining the balance of a stem-based remix.

The input mixture x(n) is separated into vocals v(n), bass b(n), drums d(n), and other o(n) stems with Spleeter [13]. Each stem is then divided into harmonic and percussive components using HPSS [9]. The full source-faithful processing structure is shown below in an interactive figure.

Figure 1 / Signal flow diagram

Hover to inspect · click to pin

vh(n) vp(n)
bh(n) bp(n)
dh(n) dp(n)
oh(n) op(n)

Beat grid

BeatThis! estimates quarter-note beat times and their metrical positions [14]. Linear interpolation subdivides these times into sixteenth-note positions t(k), with discrete sample locations nk = t(k)fs and position labels w(k).

Windowed gains

A Hann window hk(n) is centered at each nk. Its length adapts to the local sixteenth-note duration: Nk = R(nk+1 − nk−1). The parameter R ∈ [0,1] controls the window size and overlap.

Gain table and scaling

The gain table G ∈ ℝ16×1 assigns a factor to every metrical sixteenth position. Values below one attenuate and values above one amplify. The resulting γ(n) is scaled for each component as γu(n) = (γ(n)−1)gu+1, where gu ∈ [0,1].

hk(n) is a centered Hann window with support around nk; Nk = R(nk+1 − nk−1).

γ(n) = ∏k=1K((G(w(k))−1)hk(n)+1), then γu(n) = (γ(n)−1)gu+1.

Figure 2 / Beat-adaptive gain signals

Metric eighth- and quarter-note pattern Illustrative alternative pattern
Gain table
w(k)=z12345678910111213141516
Gmetric(z)2020202020202020
Galternative(z)20.510.510.510.51.50.510.51.250.510.5
One 4/4 measure at 60 beats/min. The blue signal strengthens metrically salient quarter- and eighth-note positions. The orange dashed comparison uses distinct quarter-note gains while retaining a gain of 1 for eighths and 0.5 for off-grid sixteenths.

The evaluated pattern uses a factor of two on the quarter-note pulse and its eighth-note subdivisions, while intermediate sixteenth-note positions are attenuated to zero. This structured pattern is combined with stem-specific controls rather than applied to all components indiscriminately.

Evaluated parameterization

Vocals retain both harmonic and percussive components with βvh = βvp = 1. Harmonic bass, drums, and other components are suppressed. Their percussive components share βbp = βdp = βop = α. BeatGain applies metrical modulation to these percussive parts; the reference V+αP uses the same stem factors but disables modulation. V+P and V+2P denote α = 1 and α = 2. BeatGain denotes α = 1. The implementation uses the paper’s R = 0.9 setting, 22.05 kHz sampling rate, and HPSS configuration.

Experimental design

The objective measures were computed over α from 0 to 5 on the IKA CI Pop Music Dataset [18]. Estimated instrumental complexity combines tempo, roughness, pulse clarity, and crest factor on a 0–100 scale [15], an approach motivated in part by the relationship between complexity ratings of CI users and normal-hearing listeners [16]. SI-SDR measures structural deviation from the original mix while compensating for global level differences [17].

The listening experiment compared Unproc, V+P, V+2P, and BeatGain in a two-alternative forced-choice design. 17 normal-hearing participants heard ten short pop/rock excerpts through an 8-band noise-excited vocoder [19], with bands spaced on the Greenwood frequency-place scale [20]. Signals were loudness-normalized using the ITU-R BS.1770 recommendation [21].

Compared conditions

Unproc

The original, loudness-normalized mixture without stem remixing or rhythm-dependent processing.

V+P

Vocals are retained while the percussive components of bass, drums, and other stems are mixed in with α = 1. No beat-adaptive gain pattern is applied. This setting corresponds to the best-performing remix reported by Gauer et al. [8].

V+2P

The same static stem remix as V+P, but with α = 2, so the non-vocal percussive components receive twice the reference weight.

BeatGain

Uses the α = 1 stem balance of V+P and additionally applies time-varying gains: salient quarter- and eighth-note positions are emphasized while weaker sixteenth-note positions are reduced.

Original paper composite crop of four spectrogram panels with axes, panel labels, and colorbar: (a) Input Signal, (b) V+P, (c) V+2P, and (d) BeatGain.
Figure 3. Spectrograms of the original music signal excerpt and processed versions of “Invisible Familiars – Disturbing Wildlife” (sample 6 from the IKA CI Pop Music Dataset [18]): (a) original Input Signal, (b) V+P, (c) V+2P, and (d) BeatGain.

Figure 4 / Instrumental measures

V+αP BeatGainα
Instrumental measures for BeatGainα and the pure stem-mixing parameterization V+αP. α amplifies the percussive drum, bass, and other components. Larger SI-SDR values indicate less deviation from the unprocessed mixture; lower complexity values indicate simpler rhythm.

Instrumental measures

BeatGainα produced lower estimated complexity than the V+αP reference across the evaluated α range, and both processing families were below the unprocessed condition. Lower complexity came with a deviation trade-off: stronger changes can reduce complexity while moving further from the original mix.

Figure 5 / Listening-experiment preference scores

Preference for the first method in each comparison. The vertical reference marks equal preference at 50%.

Preference scores averaged across all listeners and excerpts. Significance: * p<0.05, ** p<0.01, *** p<0.001 (Bonferroni–Holm corrected).

Subjective preferences

Processed conditions were preferred over Unproc for both overall impression and rhythmic clarity. Compared with V+2P, BeatGain was significantly preferred for rhythmic clarity, while the overall-impression difference was not significant.

BeatGain extends remixing by making metrical position explicit. The results support rhythm-focused processing as a possible complement to spectral and stem-based strategies, but the perceptual benefit is multidimensional: clearer rhythm does not automatically imply a significant improvement in overall impression.

The method is deliberately restricted to a 4/4 meter and short pop/rock excerpts. Extreme α values can alter musical character through the complexity–deviation trade-off. Future work should examine other meters, signal- and listener-specific parameters, and direct studies with CI users.

Choose an excerpt and compare the four variants. The central player is shared by every button. The static header preserves the preference-ranking presentation order; the dynamic matrix is populated by app.js. Switch to CI simulation to hear the vocoded mode used in the listening experiment.

Now listening
Example 1 - Unproc

BeatGain · Lentz et al.

BeatGain — A Rhythmic Pattern Enhancement Algorithm for Music Listening with Cochlear Implants

Accepted at Interspeech 2026; publication details and DOI forthcoming.

Dataset attribution

Examples use the IKA CI Pop Music Dataset: Johannes Gauer, Anil Nagathil, Benjamin Lentz, and Rainer Martin (2022), derived from MedleyDB by Rachel M. Bittner et al. The 12-second excerpts and remixes were modified from the original dataset. The dataset is licensed CC BY-NC-SA 4.0; use and sharing are non-commercial and share-alike. License details · Dataset DOI.

References

  1. Gfeller et al. Musical backgrounds, listening habits, and aesthetic enjoyment of adult cochlear implant recipients. Journal of the American Academy of Audiology, 11(07), 390–406 (2000). Local BibTeX key: gfeller2000musical; no external DOI link is provided here.
  2. Alexander et al. From fragments to the whole: a comparison between cochlear implant users and normal-hearing listeners in music perception and enjoyment. Journal of Otolaryngology-Head & Neck Surgery, 40(1), 1–7 (2011). Local BibTeX key: alexander2011fragments; no external DOI link is provided here.
  3. Limb and Roy. Technological, biological, and acoustical constraints to music perception in cochlear implant users. Hearing Research, 308, 13–26 (2014). DOI
  4. Nogueira et al. Making Music More Accessible for Cochlear Implant Listeners: Recent Developments. IEEE Signal Processing Magazine, 36(1), 115–127 (2018). DOI
  5. Lassaletta et al. Changes in listening habits and quality of musical sound after cochlear implantation. Otolaryngology—Head and Neck Surgery, 138(3), 363–367 (2008). Local BibTeX key: lassaletta2008changes; no external DOI link is provided here.
  6. Kohlberg et al. Music engineering as a novel strategy for enhancing music enjoyment in the cochlear implant recipient. Behavioural Neurology, 2015, 829680 (2015). Local BibTeX key: kohlberg2015music; no external DOI link is provided here.
  7. Buyens et al. Music mixing preferences of cochlear implant recipients: A pilot study. International Journal of Audiology, 53(5), 294–301 (2014). Local BibTeX key: buyens2014music; no external DOI link is provided here.
  8. Gauer et al. A versatile deep-neural-network-based music preprocessing and remixing scheme for cochlear implant listeners. The Journal of the Acoustical Society of America, 151(5), 2975–2986 (2022). DOI
  9. Driedger et al. Extending Harmonic-Percussive Separation of Audio Signals. Proc. ISMIR, 611–616 (2014). DOI
  10. Toussaint. The Geometry of Musical Rhythm: What Makes a “Good” Rhythm Good? CRC Press (2019). DOI
  11. de Fleurian et al. Information-theoretic measures predict the human judgment of rhythm complexity. Cognitive Science, 41(3), 800–813 (2017). Local BibTeX key: de2017information; no external DOI link is provided here.
  12. Thul and Toussaint. Rhythm Complexity Measures: A Comparison of Mathematical Models of Human Perception and Performance. Proc. ISMIR, 663–668 (2008). Local BibTeX key: thul2008rhythm; no external DOI link is provided here.
  13. Hennequin et al. Spleeter: a fast and efficient music source separation tool with pre-trained models. Journal of Open Source Software, 5(50), 2154 (2020). DOI
  14. Foscarin et al. Beat This! Accurate Beat Tracking without DBN Postprocessing. Proc. ISMIR (2024). DOI
  15. Buyens et al. A model for music complexity applied to music preprocessing for cochlear implants. Proc. EUSIPCO, 971–975 (2017). Local BibTeX key: buyens2017model; no external DOI link is provided here.
  16. Gfeller et al. The effects of familiarity and complexity on appraisal of complex songs by cochlear implant recipients and normal hearing adults. Journal of Music Therapy, 40(2), 78–112 (2003). Local BibTeX key: gfeller2003effects; no external DOI link is provided here.
  17. Le Roux et al. SDR—half-baked or well done? Proc. ICASSP, 626–630 (2019). DOI
  18. Gauer et al. IKA CI Pop Music Dataset. Zenodo (2022). Dataset DOI
  19. Bingabr, Espinoza-Varas, and Loizou. Simulating the effect of spread of excitation in cochlear implants. Hearing Research, 241(1), 73–79 (2008). DOI
  20. Greenwood. A cochlear frequency-position function for several species—29 years later. The Journal of the Acoustical Society of America, 87(6), 2592–2605 (1990). Local BibTeX key: greenwood1990cochlear; no external DOI link is provided here.
  21. ITU-R BS.1770-4. Algorithms to measure audio programme loudness and true-peak audio level. (2015). Recommendation