one slider, two acts: F2 sweeps while F1 holds — then F1 sweeps while F2 holds
AbstractPhonetic experiments need to change a single acoustic cue precisely while keeping the speech natural. Signal-processing tools control transparently but lose quality; controllable neural synthesizers sound natural but must learn their controls from large corpora. We present ARIS, a glass-box neural source–filter synthesizer: an encoder estimates frame-level parameters, and a deterministic synthesizer renders them through a glottal source, a noise source and a cascade of explicit formant resonators. Because every control is a synthesizer coefficient, the control space need not be covered by training data, and each model uses at most 1 h of one speaker. On five corpora in three languages, ARIS reproduces sentences and Mandarin syllables with a lower log-spectral distance than WORLD (7.38 vs. 7.62 dB on CSMSC) and higher UTMOS and SQUIM. When single formants are scaled, the median F2/F3 errors on CSMSC are 37.8/110.1 Hz, against 72.3/346.9 Hz for Praat KlattGrid, and editing one formant moves the others by at most 0.04%. HiFi-Glot, pre-trained on 1664 h and fine-tuned on the same 1 h, attains higher predicted MOS but reproduces recordings less faithfully and edits formants with more than twice the error of ARIS.
What is ARIS?
A neural encoder estimates frame-level controls; a deterministic source–filter synthesizer renders them. Every control is a synthesizer coefficient (F0, Rd, one resonator per formant), so an edit is applied exactly, with a model trained on at most 1 h of one speaker.
★ editable · dashed: learned, not exposed
Objective results
Error of the edited cue when F0–F3 are scaled by 0.7–1.3 (60 sentences per corpus).
Copy synthesis
Unedited resynthesis by each system. All files at −26 dBFS.
You don't know the middle classes as well as I do.
HiFi-TTS · 16 kHz · original
Scale
ARIS
KlattGrid
HiFi-Glot (FT)
×0.8
×1 unedited
×1.2
Scale
ARIS
KlattGrid
HiFi-Glot (FT)
×0.8
×1 unedited
×1.2
Scale
ARIS
KlattGrid
HiFi-Glot (FT)
×0.8
×1 unedited
×1.2
Scale
ARIS
KlattGrid
HiFi-Glot (FT)
×0.8
×1 unedited
×1.2
They dislike finality in this country.
HiFi-TTS · 16 kHz · original
Scale
ARIS
KlattGrid
HiFi-Glot (FT)
×0.8
×1 unedited
×1.2
Scale
ARIS
KlattGrid
HiFi-Glot (FT)
×0.8
×1 unedited
×1.2
Scale
ARIS
KlattGrid
HiFi-Glot (FT)
×0.8
×1 unedited
×1.2
Scale
ARIS
KlattGrid
HiFi-Glot (FT)
×0.8
×1 unedited
×1.2
Vowel & formant continua
One natural token, one control moved in equal steps. Values under each step are measured.
Mandarin vowel space ǎ → yǐ (Bark-interpolated F1–F2 trajectory)
ǎ → yǐ啊 (interjection) → 椅 ‘chair’
F024 · manipulated: F1,F2,F3 (Hz) · base token F024_a3
F1 and F2 move from the natural ǎ targets towards those of yǐ from the same speaker, with targets equally spaced in Bark.
details
F3 is also interpolated in Bark to avoid F2–F3 overlap near yǐ; F0 and source controls are held fixed. These are formant edits, not a complete reproduction of the natural endpoint vowel.
MALD · manipulated: F1,F2,F3 (ratio) · base token LEMON
F1, F2 and F3 of ‘lemon’ are multiplied by a shared factor from 0.85 to 1.15.
details
F0 and source controls stay fixed, and the eight residual pole pairs are not scaled. This is a partial spectral manipulation, not a complete vocal-tract length transformation; a change in perceived speaker size is not established by this demo.
F0 is interpolated in semitones between natural dá and dǎ contours from the same speaker.
details
The contours are time-normalised to the base token's voiced span, with period-doubled frames in the Tone 3 dip bridged. The base token's spectral-envelope, voicing, duration and amplitude controls are held fixed; this illustrates an F0 edit, without a claim about which tone pair is most confusable.
Only the final voiced stretch of ‘now’ (at most 400 ms) receives an F0 edit.
details
Its contour is interpolated in semitones from the natural contour towards a curved rise ending 9 semitones above an anchor taken from the region's first 50 ms. F0 outside the region and the spectral-envelope, timing and voicing controls are held fixed; the demo does not measure statement–question judgements.
The model's frame-wise R_d control for sustained ā is scaled by factors from 0.6 to 1.6 through a shift on its logarithmic lookup table.
details
Values are limited to the table's range. F0, formant, noise-branch and duration controls are held fixed; H1–H2 and HNR below are measured outputs, not guaranteed monotonic changes or perceptual labels.
F024 · manipulated: noise gain (dB) · base token F024_ta1 (train split)
The gain of ARIS's spectrally shaped noise branch is shifted from −12 to +12 dB for tā.
details
Harmonic-branch, F0, formant and timing controls are held fixed. HNR is measured on the vowel, alongside pre-voicing energy relative to vowel energy; the manipulation changes noise level rather than aspiration duration.