Unproc
The original, loudness-normalized mixture without stem remixing or rhythm-dependent processing.
Music signal processing · Interspeech 2026
BeatGain extends stem-based music preprocessing by using metrical information to emphasize rhythm-relevant percussive events and reduce weaker off-grid events.
This website accompanies the Interspeech 2026 paper on BeatGain. Its main purpose is to make the evaluated processing conditions directly audible through the listening examples below.
It also provides a concise, interactive overview of the motivation, the rhythm-aware algorithm, the comparison methods, and the central study results. The page is intended as an accessible companion to the paper rather than a replacement for the full publication.
Cochlear implants can restore access to sound and often support speech understanding, but music is frequently perceived as less natural or enjoyable. Broad stimulation regions, place–frequency mismatches, and compression reduce spectral resolution. Pitch and timbre are particularly affected, whereas temporal-envelope information is comparatively better preserved [1] [2] [3] [4].
Consequently, CI listeners tend to value a stable beat, prominent vocals, and reduced instrumental complexity [1] [5] [6] [7]. Existing music enhancement work has mainly followed two paths: source separation and remixing, or spectral simplification through methods such as principal-component and subspace approaches. HPSS and stem separation can independently rebalance harmonic and percussive material [8] [9]; spectral methods reduce competing structure. These methods do not explicitly target the temporal organization of rhythm.
BeatGain addresses that complement. In a common 4/4 meter, quarter-note beats are strongest, followed by regular eighth- and sixteenth-note subdivisions. Events on weaker metrical positions can create syncopation and increase perceived rhythmic complexity [10] [11] [12]. The hypothesis is therefore deliberately narrow: amplify percussive events aligned with salient positions and attenuate weaker off-grid events, while retaining the balance of a stem-based remix.
The input mixture x(n) is separated into vocals v(n), bass b(n), drums d(n), and other o(n) stems with Spleeter [13]. Each stem is then divided into harmonic and percussive components using HPSS [9]. The full source-faithful processing structure is shown below in an interactive figure.
Figure 1 / Signal flow diagram
Hover to inspect · click to pin
BeatThis! estimates quarter-note beat times and their metrical positions [14]. Linear interpolation subdivides these times into sixteenth-note positions t(k), with discrete sample locations nk = t(k)fs and position labels w(k).
A Hann window hk(n) is centered at each nk. Its length adapts to the local sixteenth-note duration: Nk = R(nk+1 − nk−1). The parameter R ∈ [0,1] controls the window size and overlap.
The gain table G ∈ ℝ16×1 assigns a factor to every metrical sixteenth position. Values below one attenuate and values above one amplify. The resulting γ(n) is scaled for each component as γu(n) = (γ(n)−1)gu+1, where gu ∈ [0,1].
hk(n) is a centered Hann window with support around nk; Nk = R(nk+1 − nk−1).
γ(n) = ∏k=1K((G(w(k))−1)hk(n)+1), then γu(n) = (γ(n)−1)gu+1.
Figure 2 / Beat-adaptive gain signals
| w(k)=z | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gmetric(z) | 2 | 0 | 2 | 0 | 2 | 0 | 2 | 0 | 2 | 0 | 2 | 0 | 2 | 0 | 2 | 0 |
| Galternative(z) | 2 | 0.5 | 1 | 0.5 | 1 | 0.5 | 1 | 0.5 | 1.5 | 0.5 | 1 | 0.5 | 1.25 | 0.5 | 1 | 0.5 |
The evaluated pattern uses a factor of two on the quarter-note pulse and its eighth-note subdivisions, while intermediate sixteenth-note positions are attenuated to zero. This structured pattern is combined with stem-specific controls rather than applied to all components indiscriminately.
Vocals retain both harmonic and percussive components with βvh = βvp = 1. Harmonic bass, drums, and other components are suppressed. Their percussive components share βbp = βdp = βop = α. BeatGain applies metrical modulation to these percussive parts; the reference V+αP uses the same stem factors but disables modulation. V+P and V+2P denote α = 1 and α = 2. BeatGain denotes α = 1. The implementation uses the paper’s R = 0.9 setting, 22.05 kHz sampling rate, and HPSS configuration.
The objective measures were computed over α from 0 to 5 on the IKA CI Pop Music Dataset [18]. Estimated instrumental complexity combines tempo, roughness, pulse clarity, and crest factor on a 0–100 scale [15], an approach motivated in part by the relationship between complexity ratings of CI users and normal-hearing listeners [16]. SI-SDR measures structural deviation from the original mix while compensating for global level differences [17].
The listening experiment compared Unproc, V+P, V+2P, and BeatGain in a two-alternative forced-choice design. 17 normal-hearing participants heard ten short pop/rock excerpts through an 8-band noise-excited vocoder [19], with bands spaced on the Greenwood frequency-place scale [20]. Signals were loudness-normalized using the ITU-R BS.1770 recommendation [21].
The original, loudness-normalized mixture without stem remixing or rhythm-dependent processing.
Vocals are retained while the percussive components of bass, drums, and other stems are mixed in with α = 1. No beat-adaptive gain pattern is applied. This setting corresponds to the best-performing remix reported by Gauer et al. [8].
The same static stem remix as V+P, but with α = 2, so the non-vocal percussive components receive twice the reference weight.
Uses the α = 1 stem balance of V+P and additionally applies time-varying gains: salient quarter- and eighth-note positions are emphasized while weaker sixteenth-note positions are reduced.
Figure 4 / Instrumental measures
BeatGainα produced lower estimated complexity than the V+αP reference across the evaluated α range, and both processing families were below the unprocessed condition. Lower complexity came with a deviation trade-off: stronger changes can reduce complexity while moving further from the original mix.
Figure 5 / Listening-experiment preference scores
Preference for the first method in each comparison. The vertical reference marks equal preference at 50%.
Processed conditions were preferred over Unproc for both overall impression and rhythmic clarity. Compared with V+2P, BeatGain was significantly preferred for rhythmic clarity, while the overall-impression difference was not significant.
BeatGain extends remixing by making metrical position explicit. The results support rhythm-focused processing as a possible complement to spectral and stem-based strategies, but the perceptual benefit is multidimensional: clearer rhythm does not automatically imply a significant improvement in overall impression.
The method is deliberately restricted to a 4/4 meter and short pop/rock excerpts. Extreme α values can alter musical character through the complexity–deviation trade-off. Future work should examine other meters, signal- and listener-specific parameters, and direct studies with CI users.
Choose an excerpt and compare the four variants. The central player is shared by every button. The static header preserves the preference-ranking presentation order; the dynamic matrix is populated by app.js. Switch to CI simulation to hear the vocoded mode used in the listening experiment.
BeatGain · Lentz et al.
Accepted at Interspeech 2026; publication details and DOI forthcoming.
Examples use the IKA CI Pop Music Dataset: Johannes Gauer, Anil Nagathil, Benjamin Lentz, and Rainer Martin (2022), derived from MedleyDB by Rachel M. Bittner et al. The 12-second excerpts and remixes were modified from the original dataset. The dataset is licensed CC BY-NC-SA 4.0; use and sharing are non-commercial and share-alike. License details · Dataset DOI.
gfeller2000musical; no external DOI link is provided here.alexander2011fragments; no external DOI link is provided here.lassaletta2008changes; no external DOI link is provided here.kohlberg2015music; no external DOI link is provided here.buyens2014music; no external DOI link is provided here.de2017information; no external DOI link is provided here.thul2008rhythm; no external DOI link is provided here.buyens2017model; no external DOI link is provided here.gfeller2003effects; no external DOI link is provided here.greenwood1990cochlear; no external DOI link is provided here.