ARIS: Low-Resource Glass-Box Neural Source–Filter Synthesis for Phonetic Stimulus Manipulation

Yiran Ding, Wenwei Xu

Leiden University Centre for Linguistics (LUCL), Leiden University, The Netherlands · ICASSP 2027 submission

Paper (PDF) · soonCode · soonAudio samples

F2 F1 F0 — control ·F2F1×0.7×1.0×1.3
one slider, two acts: F2 sweeps while F1 holds — then F1 sweeps while F2 holds
AbstractPhonetic experiments need to change a single acoustic cue precisely while keeping the speech natural. Signal-processing tools control transparently but lose quality; controllable neural synthesizers sound natural but must learn their controls from large corpora. We present ARIS, a glass-box neural source–filter synthesizer: an encoder estimates frame-level parameters, and a deterministic synthesizer renders them through a glottal source, a noise source and a cascade of explicit formant resonators. Because every control is a synthesizer coefficient, the control space need not be covered by training data, and each model uses at most 1 h of one speaker. On five corpora in three languages, ARIS reproduces sentences and Mandarin syllables with a lower log-spectral distance than WORLD (7.38 vs. 7.62 dB on CSMSC) and higher UTMOS and SQUIM. When single formants are scaled, the median F2/F3 errors on CSMSC are 37.8/110.1 Hz, against 72.3/346.9 Hz for Praat KlattGrid, and editing one formant moves the others by at most 0.04%. HiFi-Glot, pre-trained on 1664 h and fine-tuned on the same 1 h, attains higher predicted MOS but reproduces recordings less faithfully and edits formants with more than twice the error of ARIS.

What is ARIS?

A neural encoder estimates frame-level controls; a deterministic source–filter synthesizer renders them. Every control is a synthesizer coefficient (F0, Rd, one resonator per formant), so an edit is applied exactly, with a model trained on at most 1 h of one speaker.

ARIS architecture: Conformer encoder with source, tract and noise observations feeding per-component heads; decoder with LF glottal wavetable, noise FIR and a cascade of formant resonators, residual poles and gated zeros.

★ editable · dashed: learned, not exposed

Objective results

Error of the edited cue when F0–F3 are scaled by 0.7–1.3 (60 sentences per corpus).

Control precision (RMSE95 in Hz) of ARIS, Praat KlattGrid and HiFi-Glot (fine-tuned) for F0-F3 scaling on CSMSC and HiFi-TTS; ARIS has the lowest error for F1-F3.

Copy synthesis

Unedited resynthesis by each system. All files at −26 dBFS.

CSMSC Mandarin sentences · 24 kHz

ItemNaturalARISWORLDKlattGridHiFi-Glot (FT)
大量土方涌进犬舍。dà liàng tǔ fāng yǒng jìn quǎn shè000255referenceLSD 6.66LSD 6.86LSD 46.48LSD 9.47
欢迎各位朋友常来逛逛聊聊。huān yíng gè wèi péng you cháng lái guàng guang liáo liao000195referenceLSD 7.10LSD 7.29LSD 47.87LSD 8.95
主要有鲤鱼、银鲫、鲶鱼、草鱼、泥鳅等。zhǔ yào yóu lǐ yú yín qí nián yú cǎo yú ní qiu děng000930referenceLSD 7.15LSD 7.33LSD 47.38LSD 9.08

LSD in dB (lower = closer to the recording) · training items with the lowest ARIS LSD

HiFi-TTS English sentences · 16 kHz

ItemNaturalARISWORLDKlattGridHiFi-Glot (FT)
You understand what I mean, Mister Verloc?secretagent_02_conrad_0699referenceLSD 6.93LSD 7.56LSD 33.05LSD 8.81
What with one sort of attachment and another you are doing away with your usefulness.secretagent_02_conrad_0893referenceLSD 6.94LSD 7.54LSD 36.47LSD 8.93
It isn't very wise to call me up like this.secretagent_02_conrad_0616referenceLSD 6.98LSD 7.69LSD 36.82LSD 8.72

LSD in dB (lower = closer to the recording) · training items with the lowest ARIS LSD

F024 Mandarin monosyllables · 16 kHz

ItemNaturalARISWORLDKlattGrid
pǐF024_pi3referenceLSD 6.71LSD 7.57LSD 33.13
mǎiF024_mai3referenceLSD 6.75LSD 7.58LSD 34.09
miǎoF024_miao3referenceLSD 6.80LSD 7.68LSD 34.73

LSD in dB (lower = closer to the recording) · training items with the lowest ARIS LSD

MALD English isolated words · 16 kHz

ItemNaturalARISWORLDKlattGrid
planetaryPLANETARYreferenceLSD 7.04LSD 7.10LSD 36.36
vileVILEreferenceLSD 7.10LSD 7.01LSD 32.42
unnaturalUNNATURALreferenceLSD 7.11LSD 7.44LSD 36.53

LSD in dB (lower = closer to the recording) · training items with the lowest ARIS LSD

BALDEY Dutch isolated words · 16 kHz

ItemNaturalARISWORLDKlattGrid
beitsbeitsreferenceLSD 7.35LSD 7.59LSD 36.33
gekeurdgekeurdreferenceLSD 7.41LSD 7.71LSD 40.07
tijtijreferenceLSD 7.42LSD 7.57LSD 36.36

LSD in dB (lower = closer to the recording) · training items with the lowest ARIS LSD

Spectrogram comparison

The same sentence rendered by each system.

CSMSC

“大量土方涌进犬舍。”

024681012kHz00.511.522.5s
Natural
024681012kHz00.511.522.5s
ARIS
024681012kHz00.511.522.5s
HiFi-Glot (FT)

HiFi-TTS

“You understand what I mean, Mister Verloc?”

02468kHz00.511.52s
Natural
02468kHz00.511.52s
ARIS
02468kHz00.511.52s
HiFi-Glot (FT)

Single-cue editing

One cue scaled at a time; each button shows the measured change.

Button = measured change of the edited cue: ±2 pts 2–5 pts > 5 pts from the requested ±20%.

写后放在座位的右边,用以自警。

xiě hòu fàng zài zuò wèi de yòu biān yòng yǐ zì jǐng

CSMSC · 24 kHz · original

ScaleARISKlattGridHiFi-Glot (FT)
×0.8
×1 unedited
×1.2

非要说是扮演贺龙的杜源呢?

fēi yào shuō shì bàn yǎn hè lóng de dù yuán ne

CSMSC · 24 kHz · original

ScaleARISKlattGridHiFi-Glot (FT)
×0.8
×1 unedited
×1.2

You don't know the middle classes as well as I do.

HiFi-TTS · 16 kHz · original

ScaleARISKlattGridHiFi-Glot (FT)
×0.8
×1 unedited
×1.2

They dislike finality in this country.

HiFi-TTS · 16 kHz · original

ScaleARISKlattGridHiFi-Glot (FT)
×0.8
×1 unedited
×1.2

Vowel & formant continua

One natural token, one control moved in equal steps. Values under each step are measured.

Mandarin vowel space ǎ → yǐ (Bark-interpolated F1–F2 trajectory)

ǎ → yǐ 啊 (interjection) → 椅 ‘chair’

F024 · manipulated: F1,F2,F3 (Hz) · base token F024_a3

F1 and F2 move from the natural ǎ targets towards those of yǐ from the same speaker, with targets equally spaced in Bark.

details

F3 is also interpolated in Bark to avoid F2–F3 overlap near yǐ; F0 and source controls are held fixed. These are formant edits, not a complete reproduction of the natural endpoint vowel.

Reference background

Source of the Bark conversion used for target spacing; not a validation of this Mandarin vowel path. Traunmüller, H. (1990). Analytical expressions for the tonotopic sensory scale. JASA, 88(1), 97–100.

1500200025005001000F2 (Hz) ←F1 (Hz) ↓step 1: F1 1122 Hz, F2 1593 Hzstep 2: F1 980 Hz, F2 1649 Hzstep 3: F1 908 Hz, F2 1778 Hzstep 4: F1 815 Hz, F2 1881 Hzstep 5: F1 699 Hz, F2 2008 Hzstep 6: F1 666 Hz, F2 2138 Hzstep 7: F1 540 Hz, F2 2284 Hzstep 8: F1 463 Hz, F2 2442 Hzstep 9: F1 427 Hz, F2 2613 Hznatural ǎ: F1 1111, F2 1577 Hzǎnatural yǐ: F1 295, F2 2663 Hzyǐ

ǎ → yǐ

02468kHz00.20.40.60.81sstep 1
step 1 · ǎ
02468kHz00.20.40.60.81sstep 9
step 9 · yǐ
ǎ
F1 1122
F2 1593
F3 3163
F1 980
F2 1649
F3 3210
F1 908
F2 1778
F3 3260
F1 815
F2 1881
F3 3331
F1 699
F2 2008
F3 3390
F1 666
F2 2138
F3 3434
F1 540
F2 2284
F3 3493
F1 463
F2 2442
F3 3577
F1 427
F2 2613
F3 3627
yǐ

Mandarin vowel space ǎ → wǔ (Bark-interpolated F1–F2 trajectory)

ǎ → wǔ 啊 (interjection) → 五 ‘five’

F024 · manipulated: F1,F2 (Hz) · base token F024_a3

F1 and F2 move from the natural ǎ targets towards those of wǔ from the same speaker, with targets equally spaced in Bark.

details

F3, F0 and source controls are held fixed, so the endpoint is an F1–F2 approximation rather than a full reproduction of wǔ.

Reference background

Source of the Bark conversion used for target spacing; not a validation of this Mandarin vowel path. Traunmüller, H. (1990). Analytical expressions for the tonotopic sensory scale. JASA, 88(1), 97–100.

500100015005007501000F2 (Hz) ←F1 (Hz) ↓step 1: F1 1122 Hz, F2 1593 Hzstep 2: F1 951 Hz, F2 1394 Hzstep 3: F1 902 Hz, F2 1242 Hzstep 4: F1 761 Hz, F2 1119 Hzstep 5: F1 684 Hz, F2 966 Hzstep 6: F1 654 Hz, F2 833 Hzstep 7: F1 493 Hz, F2 712 Hzstep 8: F1 455 Hz, F2 706 Hzstep 9: F1 433 Hz, F2 694 Hznatural ǎ: F1 1111, F2 1577 Hzǎnatural wǔ: F1 429, F2 638 Hzwǔ

ǎ → wǔ

02468kHz00.20.40.60.81sstep 1
step 1 · ǎ
02468kHz00.20.40.60.81sstep 9
step 9 · wǔ
ǎ
F1 1122
F2 1593
F1 951
F2 1394
F1 902
F2 1242
F1 761
F2 1119
F1 684
F2 966
F1 654
F2 833
F1 493
F2 712
F1 455
F2 706
F1 433
F2 694
wǔ

Joint F1–F3 scaling

lemon [ˈlɛmən] ‘lemon’

MALD · manipulated: F1,F2,F3 (ratio) · base token LEMON

F1, F2 and F3 of ‘lemon’ are multiplied by a shared factor from 0.85 to 1.15.

details

F0 and source controls stay fixed, and the eight residual pole pairs are not scaled. This is a partial spectral manipulation, not a complete vocal-tract length transformation; a change in perceived speaker size is not established by this demo.

Reference background

Background on spectral scaling and perceived speaker size. This demo scales only F1–F3 and does not reproduce that experiment. Smith, D. R. R., Patterson, R. D., Turner, R., Kawahara, H., & Irino, T. (2005). The processing and perception of size information in speech sounds. JASA, 117(1), 305–318.

1000125015001750300400500600F2 (Hz) ←F1 (Hz) ↓step 1: F1 419 Hz, F2 1168 Hzstep 2: F1 452 Hz, F2 1242 Hzstep 3: F1 476 Hz, F2 1311 Hzstep 4: F1 506 Hz, F2 1370 Hzstep 5: F1 514 Hz, F2 1458 Hzstep 6: F1 539 Hz, F2 1524 Hzstep 7: F1 565 Hz, F2 1601 Hznatural ×0.85 · lower formants: F1 473, F2 1360 Hz×0.85 · lower formants

×0.85 · lower formants → ×1.15 · higher formants

02468kHz00.2sstep 1
step 1 · ×0.85 · lower formants
02468kHz00.2sstep 7
step 7 · ×1.15 · higher formants
×0.85 · lower formants
F1 419
F2 1168
F3 2166
F1 452
F2 1242
F3 2270
F1 476
F2 1311
F3 2408
F1 506
F2 1370
F3 2574
F1 514
F2 1458
F3 2722
F1 539
F2 1524
F3 2882
F1 565
F2 1601
F3 2993

Tone & intonation continua

F0 contours interpolated in semitones; spectral envelope, timing and voicing unchanged.

Mandarin Tone 2 – Tone 3

dá → dǎ ‘answer’ (答 dá) vs ‘hit’ (打 dǎ)

F024 · manipulated: F0 contour (st) · base token F024_da2 (validation split)

F0 is interpolated in semitones between natural dá and dǎ contours from the same speaker.

details

The contours are time-normalised to the base token's voiced span, with period-doubled frames in the Tone 3 dip bridged. The base token's spectral-envelope, voicing, duration and amplitude controls are held fixed; this illustrates an F0 edit, without a claim about which tone pair is most confusable.

00.250.50.75105101520normalized time (voiced span)F0 (st re 100 Hz)step 1step 2step 3step 4step 5step 6step 7step 8step 9natural dánatural dǎ

dá → dǎ

02468kHz00.20.40.60.81sstep 1
step 1 · dá
02468kHz00.20.40.60.81sstep 9
step 9 · dǎ
dá
261→351
min 235
257→340
min 232
256→331
min 224
252→320
min 214
252→309
min 204
249→302
min 195
247→292
min 186
241→282
min 177
242→275
min 169
dǎ

Mandarin Tone 1 – Tone 2

yī → yí ‘clothing’ (衣 yī) vs ‘aunt’ (姨 yí)

F024 · manipulated: F0 contour (st) · base token F024_yi1 (train split)

F0 is interpolated in semitones from the natural yī contour towards the natural yí contour of the same speaker.

details

The target contour is time-normalised to the base voiced span; spectral-envelope, duration and amplitude controls are held fixed.

Reference background

Background on level–rising pitch perception and language experience. The study used linear pitch ramps; this demo interpolates natural contours and does not establish categorical perception. Xu, Y., Gandour, J. T., & Francis, A. L. (2006). Effects of language experience and stimulus complexity on the categorical perception of pitch direction. JASA, 120(2), 1063–1074.

00.250.50.7511517202225normalized time (voiced span)F0 (st re 100 Hz)step 1step 2step 3step 4step 5step 6step 7step 8step 9natural yīnatural yí

yī → yí

02468kHz00.20.40.60.81sstep 1
step 1 · yī
02468kHz00.20.40.60.81sstep 9
step 9 · yí
yī
286→332
min 281
280→339
min 276
280→347
min 279
273→356
min 271
268→362
min 268
264→372
min 261
258→381
min 253
255→387
min 244
251→397
min 234
yí

Final-word F0 rise

But he was resigned now. [bʌt hi wəz ɹɪˈzaɪnd ˈnaʊ] natural contour → final rise

HiFi-TTS · manipulated: F0 contour (final word) (st) · base token secretagent_03_conrad_0024 (validation split)

Only the final voiced stretch of ‘now’ (at most 400 ms) receives an F0 edit.

details

Its contour is interpolated in semitones from the natural contour towards a curved rise ending 9 semitones above an anchor taken from the region's first 50 ms. F0 outside the region and the spectral-envelope, timing and voicing controls are held fixed; the demo does not measure statement–question judgements.

00.250.50.751101520normalized time (voiced span)F0 (st re 100 Hz)step 1step 1step 1step 2step 2step 2step 3step 3step 3step 4step 4step 4step 5step 5step 5step 6step 6step 6step 7step 7step 7natural natural F0 contournatural natural F0 contournatural natural F0 contour

natural F0 contour → rising F0 contour

02468kHz00.511.5sstep 1
step 1 · natural F0 contour
02468kHz00.511.5sstep 7
step 7 · rising F0 contour
natural F0 contour
end 179 Hz
max 244
end 190 Hz
max 244
end 200 Hz
max 244
end 212 Hz
max 244
end 229 Hz
max 256
end 245 Hz
max 272
end 260 Hz
max 303

Voice-source continua

Glottal pulse shape (Rd) and noise gain; formants and F0 unchanged.

Glottal pulse shape (R_d)

ā interjection (啊 ā)

F024 · manipulated: Rd (ratio) · base token F024_a1 (train split)

The model's frame-wise R_d control for sustained ā is scaled by factors from 0.6 to 1.6 through a shift on its logarithmic lookup table.

details

Values are limited to the table's range. F0, formant, noise-branch and duration controls are held fixed; H1–H2 and HNR below are measured outputs, not guaranteed monotonic changes or perceptual labels.

Reference background

Background on the LF source model. Linked through the author's KTH bibliography. Fant, G. (1995). The LF-model revisited. Transformations and frequency domain analysis. STL-QPSR, 36(2–3), 119–156.

1234567−505stepH1–H2 (dB)step 1: -3.8 dBstep 2: -2.3 dBstep 3: -0.1 dBstep 4: 1.8 dBstep 5: 3.9 dBstep 6: 5.5 dBstep 7: 6.8 dB

R_d ×0.6 → R_d ×1.6

02468kHz00.511.5sstep 1
step 1 · R_d ×0.6
02468kHz00.511.5sstep 7
step 7 · R_d ×1.6
H1–H2 −3.8
HNR 27.8
H1–H2 −2.3
HNR 27.3
H1–H2 −0.1
HNR 26.2
H1–H2 1.8
HNR 25.3
H1–H2 3.9
HNR 24.3
H1–H2 5.5
HNR 23.4
H1–H2 6.8
HNR 22.6

Aspiration / breath noise gain

tā ‘he’ (他 tā)

F024 · manipulated: noise gain (dB) · base token F024_ta1 (train split)

The gain of ARIS's spectrally shaped noise branch is shifted from −12 to +12 dB for tā.

details

Harmonic-branch, F0, formant and timing controls are held fixed. HNR is measured on the vowel, alongside pre-voicing energy relative to vowel energy; the manipulation changes noise level rather than aspiration duration.

Reference background

Evidence that aspiration noise contributes to perceived breathiness. This gain edit does not manipulate voice onset time or establish a consonant category boundary. Klatt, D. H., & Klatt, L. C. (1990). Analysis, synthesis, and perception of voice quality variations among female and male talkers. JASA, 87(2), 820–857.

123456789−40−30−20stepunvoiced / voiced energy (dB)step 1: -38.3 dBstep 2: -35.3 dBstep 3: -32.3 dBstep 4: -29.3 dBstep 5: -26.3 dBstep 6: -23.3 dBstep 7: -20.3 dBstep 8: -17.4 dBstep 9: -14.5 dB

noise −12 dB → noise +12 dB

02468kHz00.20.40.60.81sstep 1
step 1 · noise −12 dB
02468kHz00.20.40.60.81sstep 9
step 9 · noise +12 dB
uv −38.3 dB
HNR 23.1
uv −35.3 dB
HNR 22.4
uv −32.3 dB
HNR 21.4
uv −29.3 dB
HNR 20.0
uv −26.3 dB
HNR 21.4
uv −23.3 dB
HNR 19.4
uv −20.3 dB
HNR 16.6
uv −17.4 dB
HNR 13.7
uv −14.5 dB
HNR 11.4