Naomi at LabPhon20
June 19, 2026
Modeling the flexible perception of speech sounds.
June 26-28, Naomi is an invited speaker at LabPhon20, held this year at the Coeur des Sciences at UQAM in downtown Montréal, and jointly organized by McGill, Carleton and Ottawa. Her talk, "Beyond long-term statistics: A mismatch in feature weighting between models and humans," is abstracted below.
Statistical learning models that are trained on speech specialize for the 'native' language they're trained on. However, these models do not generally match human listeners' reliance on different speech features, such as duration, formants, aspiration, and so on. This difference between models and humans poses a challenge to theories that hypothesize that listeners' reliance on speech features emerges from the long-term statistical properties of their input. We introduce a model, formalized using rate distortion theory, that can learn feature weights that deviate from the long-term statistics of the training data, formalizing a way in which selective attention could mediate listeners' reliance on speech features. We then show how this strategy could be beneficial to listeners by enabling them to train a flexible perceptual system that can easily adapt to new listening conditions. Our work highlights a possible computational strategy for increasing the robustness of a perceptual system and has implications for how closely we should expect statistical learning models' perceptual systems to match those of human listeners.