The Research Foundation

Sixty years of speech science, in full — not a highlight reel.

Goeckoh is built on published, peer-reviewed research spanning motor control theory, auditory neuroscience, acoustic phonetics, digital signal processing, and applied behavior analysis. This page is that research, presented completely: the mechanism, the specific studies, what they actually found, and how each piece connects to what the product does. Nothing here is simplified for marketing. If you want the short version, the homepage has it. If you want to actually evaluate the science — as a parent, a clinician, or anyone deciding whether to trust a product that touches your child's voice — this is that page.

How to read this page: each section below separates two different things — the established science (what's been published, replicated, and is not in serious scientific dispute) from Goeckoh's engineering application of it (how we built a consumer product around these principles, which is our work, not a peer-reviewed claim in itself). We think that distinction matters. A product that touches something as personal as a child's voice should tell you plainly which parts are borrowed from decades of university research and which parts are ours to stand behind.
01 · Motor Control Theory

Corollary discharge: the brain's internal copy of its own voice

Why the brain needs to predict what it's about to hear before it hears it.

Every voluntary movement your body makes is accompanied by a second, invisible signal. When your motor cortex sends a command to your muscles — reach for a cup, blink, speak a word — it simultaneously routes a copy of that same command to the sensory areas of your brain that will process the movement's consequences. This copy is called an efference copy, and the prediction it generates about incoming sensory input is the corollary discharge. The concept predates modern neuroscience: it was proposed independently by Erich von Holst and Horst Mittelstaedt in 1950, and by Roger Sperry the same year, to explain a basic puzzle — how animals tell the difference between the world moving and themselves moving through it.

In speech, this mechanism has a very specific job. Before you hear a single sound of your own voice, your brain has already generated a prediction of exactly what that sound should be — its pitch, its loudness, its formant structure — based on the motor command it just sent to your lips, tongue, jaw, and larynx. When the actual sound arrives at your ears a few milliseconds later, your auditory cortex doesn't process it from scratch. It compares what arrived against what was predicted. If the two match, the auditory response is dampened — your brain already "knew" what it was about to hear, so there's less new information to process. If the two don't match, that mismatch is itself a signal: something about the intended movement didn't come out right, and the motor system has a chance to correct it before the next attempt.

What the research specifically shows

Katherine Niziolek, Srikantan Nagarajan, and John Houde published a key piece of this picture in 2013 in the Journal of Neuroscience. Using a paradigm that perturbed speakers' formant feedback in real time, they showed that the brain's efference copy isn't a coarse, generic "I'm about to talk" signal — it carries fine acoustic detail specific to the exact vowel being produced. Speakers whose own productions were more variable showed correspondingly more variable neural responses to feedback perturbation, meaning the prediction is precise enough to track trial-by-trial differences in a person's own articulation. This matters because it rules out a simpler explanation (that the brain just suppresses all self-generated sound equally) and supports a model where the comparison between prediction and reality is acoustically detailed enough to drive genuine motor learning.

John Houde and Edward Chang's 2015 review in Current Biology, "The cortical computations underlying feedback control in vocal production," lays out how this predictive loop is implemented across the speech motor and auditory cortical networks — not as a single switch, but as a distributed system that continuously updates its predictions as a person learns, adapts, or compensates for a change in their own vocal tract or hearing.

Why this matters for Goeckoh: in several neurodivergent and neurological presentations — including profiles associated with autism, and speech changes following stroke or traumatic brain injury — this predictive comparison loop is theorized to be attenuated or less reliably engaged, which would make it harder for the brain to detect and correct its own speech errors from feedback alone. Goeckoh's entire design rationale is to give that comparison loop better raw material to work with: an auditory signal that has already been corrected toward the intended target, delivered back to the ear fast enough to still participate in the same predictive window the brain is already using.
References
  • Niziolek, C.A., Nagarajan, S.S., & Houde, J.F. (2013). What does motor efference copy represent? Evidence from speech production. Journal of Neuroscience, 33(41), 16110–16116.
  • Houde, J.F., & Chang, E.F. (2015). The cortical computations underlying feedback control in vocal production. Current Biology, 25(21), R1024–R1035.
  • von Holst, E., & Mittelstaedt, H. (1950). Das Reafferenzprinzip. Naturwissenschaften, 37, 464–476.
  • The von Holst & Mittelstaedt and Sperry papers are the foundational, decades-old origin of the corollary discharge concept in general motor control — cited here for completeness, not as speech-specific evidence.
02 · Auditory Neuroscience

N1 suppression: the measurable signature of self-recognition

How corollary discharge shows up as an actual, recordable brain signal.

Corollary discharge is a theoretical mechanism. N1 suppression is the evidence that it actually happens, measured directly from the brain. The N1 (sometimes called N100) is a negative-going deflection in the electroencephalogram (EEG) or magnetoencephalogram (MEG) that occurs roughly 100 milliseconds after any auditory event — a click, a tone, a syllable. Its amplitude reflects how much the auditory cortex "reacts" to that sound.

The consistent finding across decades of research is this: the N1 response to a sound is smaller when that sound is self-generated than when an identical sound is played back to a passive listener. Speak a vowel, and your auditory cortex's N1 response to hearing your own voice is measurably suppressed compared to the N1 response if that exact same recording were played to you a moment later while you sat silently. This is the auditory system doing exactly what the corollary discharge model predicts — using an internal prediction to partially "explain away" expected sensory input, freeing up neural resources to flag anything unexpected instead.

What the research specifically shows

John Houde, Srikantan Nagarajan, Kensuke Sekihara, and Michael Merzenich demonstrated this using MEG in a foundational 2002 study, showing the magnetic equivalent of N1 suppression (the M100 component) during speaking compared to passive listening to playback of one's own voice. Ramesh Behroozmand and colleagues have extended this line of work extensively using ERP methods, including studies examining how the N1/P2 complex changes when a speaker's own pitch or formant feedback is artificially perturbed in real time — directly testing how the suppression response depends on whether the returned sound matches what was predicted.

A separate but closely related and highly influential study is John Houde and Michael Jordan's 1998 paper in Science, "Sensorimotor adaptation in speech production." Rather than measuring N1 directly, this study demonstrated something that depends on the same underlying comparison mechanism: when speakers hear their own vowel formants shifted in real time through headphones, they unconsciously adjust their articulation to compensate — moving their formants in the opposite direction of the shift, over repeated trials, without being told to do anything differently. This is direct behavioral proof that the brain uses corrected auditory feedback to update its motor commands for speech, which is the exact mechanism Goeckoh's correction loop is designed to engage deliberately, rather than by accident of a lab perturbation experiment.

Why this matters for Goeckoh: the Houde & Jordan finding is arguably the single most direct piece of evidence that real-time formant-shifted feedback changes what a speaker does next. It's also the reason correction speed matters so much: the adaptation effect depends on the corrected signal arriving close enough to natural speech timing that the brain treats it as feedback about its own voice, not as a separate, delayed sound. That's the entire reason Goeckoh is built around a sub-50 millisecond correction latency rather than a system that corrects speech after the fact.
References
  • Houde, J.F., Nagarajan, S.S., Sekihara, K., & Merzenich, M.M. (2002). Modulation of the auditory cortex during speech: an MEG study. Journal of Cognitive Neuroscience, 14(8), 1125–1138.
  • Behroozmand, R., Karvelis, L., Liu, H., & Larson, C.R. (2009). Vocalization-induced enhancement of the auditory cortex responsiveness during voice F0 feedback perturbation. Clinical Neurophysiology, 120(7), 1303–1312.
  • Houde, J.F., & Jordan, M.I. (1998). Sensorimotor adaptation in speech production. Science, 279(5354), 1213–1216.
03 · Computational Neuroscience

The DIVA model: a unified account of how speech is learned

Turning the prediction-and-comparison loop into a working, testable model of the whole speech system.

Corollary discharge and N1 suppression describe a mechanism. The Directions Into Velocities of Articulators (DIVA) model, developed by Frank Guenther and colleagues at Boston University over several decades, is an attempt to describe the entire system that mechanism operates within — a neural network model, grounded in known brain anatomy, of how a person acquires and continuously fine-tunes the motor skill of speaking.

DIVA proposes that speech is controlled by two parallel systems working together: a feedforward system, which stores learned motor commands for producing familiar sounds efficiently and quickly, and a feedback system, which monitors auditory and somatosensory (touch and proprioceptive) consequences of speech in real time and corrects the feedforward commands when they drift off target. Early in life — or whenever adapting to a change, such as a growing vocal tract, a new set of teeth, or a neurological change following injury — the feedback system does most of the work, actively listening and correcting. Over time, successful corrections get folded back into the feedforward system, making the skill increasingly automatic. This is the same basic architecture used to explain motor learning in reaching and other skilled movements, applied specifically to the vocal tract.

What the research specifically shows

Guenther's 2006 paper in the Journal of Communication Disorders, "Cortical interactions underlying the production of speech sounds," lays out the model's architecture and maps its components onto specific brain regions — including Broca's area, premotor cortex, and superior temporal auditory regions — rather than treating it as a purely abstract computational diagram. A companion body of work, including fMRI studies published in Cerebral Cortex, has tested DIVA's predictions against real brain-imaging data collected while participants spoke, finding activation patterns broadly consistent with the model's proposed feedback and feedforward pathways.

DIVA has since been used to model and explain a range of speech disorders by simulating what happens when specific components of the model are disrupted — including stuttering, apraxia of speech, and hearing-impairment-related speech changes — by degrading or removing specific feedback pathways in simulation and observing whether the resulting model output resembles the disordered speech pattern seen clinically.

Why this matters for Goeckoh: DIVA is the reason Goeckoh is designed as a feedback-correction tool rather than a feedforward-replacement tool. The model predicts that if you improve the accuracy and timeliness of the auditory feedback someone receives about their own speech, you strengthen the same natural learning loop that gradually shapes feedforward motor programs — rather than trying to override or replace a person's speech motor system from the outside. That's a meaningfully different goal than devices that generate or synthesize speech on someone's behalf.
References
  • Guenther, F.H. (2006). Cortical interactions underlying the production of speech sounds. Journal of Communication Disorders, 39(5), 350–365.
  • Guenther, F.H., Ghosh, S.S., & Tourville, J.A. (2006). Neural modeling and imaging of the cortical interactions underlying syllable production. Brain and Language, 96(3), 280–301.
  • Golfinopoulos, E., Tourville, J.A., & Guenther, F.H. (2010). The integration of large-scale neural network modeling and functional brain imaging in speech motor control. NeuroImage, 52(3), 862–874.
04 · Acoustic Phonetics

Formant acoustics: the physical substrate of a vowel's identity

What Goeckoh is actually measuring and correcting, at the level of sound waves.

Everything above describes a neural loop. This section describes the acoustic signal that loop is actually processing. When air is pushed from the lungs through the vocal folds, it produces a buzzing source sound rich in harmonics. The vocal tract — throat, mouth, tongue position, lip shape — acts as a resonant filter on that source sound, amplifying certain frequency bands and damping others. Those amplified frequency bands are called formants, and this description of speech production as a source (the vocal folds) filtered by a resonator (the vocal tract) is known as source-filter theory, formalized by Gunnar Fant in his 1960 book Acoustic Theory of Speech Production — still the foundational text of modern acoustic phonetics.

The first two formants, F1 and F2, are what your brain uses to identify which vowel it's hearing, almost independent of who is speaking. F1 corresponds inversely to tongue height (a low F1 corresponds to a high tongue position, as in "ee"; a high F1 corresponds to a low tongue position, as in "ah"). F2 corresponds to tongue frontness (a high F2 corresponds to a front vowel like "ee"; a low F2 corresponds to a back vowel like "oo"). Plot F1 against F2 for every vowel a speaker produces, and you get that speaker's personal vowel space — a map of exactly where their vocal tract places each vowel acoustically.

What the research specifically shows

James Hillenbrand, Laura Getty, Michael Clark, and Kimberlee Wheeler published the modern normative reference for American English vowels in 1995 in the Journal of the Acoustical Society of America: "Acoustic characteristics of American English vowels." Using recordings from 139 speakers (men, women, and children) producing a standard set of vowels, they established the statistical distribution of formant frequencies for each vowel category — in effect, a map of where a given vowel "should" sit acoustically, and how much natural variation exists around that target across speakers of different ages, sexes, and vocal tract sizes. This dataset (and its extensions) remains the standard reference point for what counts as a clearly identifiable vowel versus one that has drifted toward a neighboring vowel's acoustic territory.

This is directly clinically relevant: many of the speech differences associated with dysarthria, apraxia, and some ASD-related speech patterns show up acoustically as vowels that have drifted from their expected F1/F2 territory — a centralized or compressed vowel space, where distinct vowels acoustically overlap and become harder for a listener to tell apart, even though the intended word was correct.

Why this matters for Goeckoh: "correcting a voice" is a vague phrase until it's grounded in formant space. What Goeckoh's engine actually does, moment to moment, is estimate a speaker's current formant positions and reshape them toward the target vowel region for the sound the person is producing — without altering pitch, speaking rate, or the broader spectral qualities that make a voice recognizably that person's own. Formant correction, not voice replacement, is the specific acoustic operation happening in real time.
References
  • Fant, G. (1960). Acoustic Theory of Speech Production. Mouton, The Hague.
  • Hillenbrand, J., Getty, L.A., Clark, M.J., & Wheeler, K. (1995). Acoustic characteristics of American English vowels. Journal of the Acoustical Society of America, 97(5), 3099–3111.
  • Kent, R.D., & Kim, Y.J. (2003). Toward an acoustic typology of motor speech disorders. Clinical Linguistics & Phonetics, 17(6), 427–445.
05 · Digital Signal Processing

Linear predictive coding: extracting and reshaping formants fast enough to matter

The engineering method that turns formant theory into something that can run in real time on a phone.

Knowing that formants exist and knowing where they should sit doesn't, by itself, give you a way to measure and change them in the roughly 20-millisecond windows available during live conversation. That's a signal-processing problem, and the dominant solution — used in telephony, speech coding, and speech analysis since the 1970s — is linear predictive coding (LPC).

LPC works by modeling the vocal tract as an "all-pole" filter: a mathematical system where each sample of the speech waveform can be predicted as a weighted sum of the samples immediately before it. Solving for the weights that best predict a short window of real speech gives you a set of filter coefficients that describe the shape of the vocal tract's resonances at that instant — which is mathematically equivalent to recovering the formant frequencies directly from the waveform, without needing a separate acoustic model of the throat and mouth. Because the underlying math (solving a linear system) is computationally cheap, LPC analysis can run on a short sliding window of audio fast enough for real-time use, which is precisely why it became the backbone of both mid-20th-century telephone compression and modern real-time speech analysis.

What the research specifically shows

Bishnu Atal's foundational work at Bell Labs in the late 1960s and early 1970s established both the analysis method (extracting LPC coefficients from a speech signal) and the synthesis method (using those coefficients, plus a simplified source signal, to regenerate intelligible speech) — the same two operations, analysis and resynthesis, that sit at the center of any real-time formant-correction pipeline. John Makhoul's widely cited 1975 tutorial in the Proceedings of the IEEE, "Linear Prediction: A Tutorial Review," remains the standard reference explaining why this particular mathematical approach became the dominant method in the field, and how its coefficients map back onto genuine acoustic properties of the vocal tract rather than being an arbitrary compression trick.

Because LPC coefficients correspond to actual resonance frequencies, they can be manipulated directly — shifted toward a target formant position. One classical approach, dating back to Atal's original work, is to then resynthesize the speech signal from those adjusted coefficients plus a modeled source signal. This is what makes LPC-based methods suited to formant correction specifically, as opposed to other voice modification techniques that operate on pitch or timing instead — though as the next section explains, Goeckoh itself uses a simpler, more direct variant of this idea rather than full resynthesis.

Why this matters for Goeckoh: this is the actual engineering underneath the product. Goeckoh's correction pipeline analyzes short windows of live microphone audio using LPC-family techniques to estimate current formant positions, computes the adjustment needed to move them toward the target vowel region established by acoustic-phonetics research (see Formant Acoustics, above), and applies that adjustment as a direct real-time filter on the original signal — rather than regenerating it from scratch — so pitch, voicing, and timing are preserved automatically. Audio pass-through adds only a few milliseconds; locking onto a newly changed target formant takes longer, typically well under 150ms (see the Science page's latency budget for the full breakdown), still comfortably inside the biological feedback window described in the corollary discharge and N1 suppression sections below. Every other section on this page describes why that correction should help; this section is how it's technically possible to do it live.
References
  • Atal, B.S., & Hanauer, S.L. (1971). Speech analysis and synthesis by linear prediction of the speech wave. Journal of the Acoustical Society of America, 50(2B), 637–655.
  • Makhoul, J. (1975). Linear prediction: A tutorial review. Proceedings of the IEEE, 63(4), 561–580.
06 · Clinical Measurement

Applied Behavior Analysis: measuring progress instead of guessing at it

Why session data is structured the way it is, and what it's actually tracking.

Everything above explains why real-time formant correction should help. It doesn't, by itself, tell a family or a clinician whether it's actually working for a specific person over time. That's a measurement problem, and Goeckoh's answer to it is structured around Applied Behavior Analysis (ABA) — not as a therapy method Goeckoh delivers, but as the measurement discipline behind how session data is recorded.

ABA's core methodological contribution, independent of any specific intervention technique, is insisting that target behaviors be defined operationally (specific, observable, and countable) and tracked systematically over time, so that whether a skill is actually improving is a question answered by data rather than by subjective impression. Cooper, Heron, and Heward's textbook Applied Behavior Analysis (now in its third edition, 2020) is the standard graduate-level reference for this methodology and is used broadly across special education and clinical behavior analysis training.

What this looks like in practice

Applied to Goeckoh specifically: every session logs concrete, countable events rather than a vague sense of "it seemed to go well" — which specific phonological targets (vowel or consonant productions) were attempted and how close each landed to its acoustic target, discrete fluency events (hesitations, self-corrections, successful productions), and co-regulation markers relevant to emotional state during a session. This turns each session into structured, comparable data points rather than an isolated anecdote, so a caregiver or clinician can look at a trend across weeks instead of relying on memory of how last Tuesday went.

Why this matters for Goeckoh: this is the difference between a product that claims to help and a product that shows its work. Because every session's acoustic targets, attempts, and outcomes are logged as structured data on-device, a parent or a speech-language pathologist can review actual progress over time — whether a specific vowel contrast is stabilizing, whether fluency events are decreasing, whether a particular correction profile is or isn't producing change — the same way any accountable clinical intervention should be reviewable, instead of asking a family to simply trust that something is working.
References
  • Cooper, J.O., Heron, T.E., & Heward, W.L. (2020). Applied Behavior Analysis (3rd ed.). Pearson.

That's the full foundation. Now hear it work.

Every mechanism on this page is what the live demo is actually running — not a simulation of it. Try it with your own voice, free, no account required.

Try the Demo → See Pricing