Goeckoh is built on published, peer-reviewed research spanning motor control theory, auditory neuroscience, acoustic phonetics, digital signal processing, and applied behavior analysis. This page is that research, presented completely: the mechanism, the specific studies, what they actually found, and how each piece connects to what the product does. Nothing here is simplified for marketing. If you want the short version, the homepage has it. If you want to actually evaluate the science — as a parent, a clinician, or anyone deciding whether to trust a product that touches your child's voice — this is that page.
Why the brain needs to predict what it's about to hear before it hears it.
Every voluntary movement your body makes is accompanied by a second, invisible signal. When your motor cortex sends a command to your muscles — reach for a cup, blink, speak a word — it simultaneously routes a copy of that same command to the sensory areas of your brain that will process the movement's consequences. This copy is called an efference copy, and the prediction it generates about incoming sensory input is the corollary discharge. The concept predates modern neuroscience: it was proposed independently by Erich von Holst and Horst Mittelstaedt in 1950, and by Roger Sperry the same year, to explain a basic puzzle — how animals tell the difference between the world moving and themselves moving through it.
In speech, this mechanism has a very specific job. Before you hear a single sound of your own voice, your brain has already generated a prediction of exactly what that sound should be — its pitch, its loudness, its formant structure — based on the motor command it just sent to your lips, tongue, jaw, and larynx. When the actual sound arrives at your ears a few milliseconds later, your auditory cortex doesn't process it from scratch. It compares what arrived against what was predicted. If the two match, the auditory response is dampened — your brain already "knew" what it was about to hear, so there's less new information to process. If the two don't match, that mismatch is itself a signal: something about the intended movement didn't come out right, and the motor system has a chance to correct it before the next attempt.
Katherine Niziolek, Srikantan Nagarajan, and John Houde published a key piece of this picture in 2013 in the Journal of Neuroscience. Using a paradigm that perturbed speakers' formant feedback in real time, they showed that the brain's efference copy isn't a coarse, generic "I'm about to talk" signal — it carries fine acoustic detail specific to the exact vowel being produced. Speakers whose own productions were more variable showed correspondingly more variable neural responses to feedback perturbation, meaning the prediction is precise enough to track trial-by-trial differences in a person's own articulation. This matters because it rules out a simpler explanation (that the brain just suppresses all self-generated sound equally) and supports a model where the comparison between prediction and reality is acoustically detailed enough to drive genuine motor learning.
John Houde and Edward Chang's 2015 review in Current Biology, "The cortical computations underlying feedback control in vocal production," lays out how this predictive loop is implemented across the speech motor and auditory cortical networks — not as a single switch, but as a distributed system that continuously updates its predictions as a person learns, adapts, or compensates for a change in their own vocal tract or hearing.
How corollary discharge shows up as an actual, recordable brain signal.
Corollary discharge is a theoretical mechanism. N1 suppression is the evidence that it actually happens, measured directly from the brain. The N1 (sometimes called N100) is a negative-going deflection in the electroencephalogram (EEG) or magnetoencephalogram (MEG) that occurs roughly 100 milliseconds after any auditory event — a click, a tone, a syllable. Its amplitude reflects how much the auditory cortex "reacts" to that sound.
The consistent finding across decades of research is this: the N1 response to a sound is smaller when that sound is self-generated than when an identical sound is played back to a passive listener. Speak a vowel, and your auditory cortex's N1 response to hearing your own voice is measurably suppressed compared to the N1 response if that exact same recording were played to you a moment later while you sat silently. This is the auditory system doing exactly what the corollary discharge model predicts — using an internal prediction to partially "explain away" expected sensory input, freeing up neural resources to flag anything unexpected instead.
John Houde, Srikantan Nagarajan, Kensuke Sekihara, and Michael Merzenich demonstrated this using MEG in a foundational 2002 study, showing the magnetic equivalent of N1 suppression (the M100 component) during speaking compared to passive listening to playback of one's own voice. Ramesh Behroozmand and colleagues have extended this line of work extensively using ERP methods, including studies examining how the N1/P2 complex changes when a speaker's own pitch or formant feedback is artificially perturbed in real time — directly testing how the suppression response depends on whether the returned sound matches what was predicted.
A separate but closely related and highly influential study is John Houde and Michael Jordan's 1998 paper in Science, "Sensorimotor adaptation in speech production." Rather than measuring N1 directly, this study demonstrated something that depends on the same underlying comparison mechanism: when speakers hear their own vowel formants shifted in real time through headphones, they unconsciously adjust their articulation to compensate — moving their formants in the opposite direction of the shift, over repeated trials, without being told to do anything differently. This is direct behavioral proof that the brain uses corrected auditory feedback to update its motor commands for speech, which is the exact mechanism Goeckoh's correction loop is designed to engage deliberately, rather than by accident of a lab perturbation experiment.
Turning the prediction-and-comparison loop into a working, testable model of the whole speech system.
Corollary discharge and N1 suppression describe a mechanism. The Directions Into Velocities of Articulators (DIVA) model, developed by Frank Guenther and colleagues at Boston University over several decades, is an attempt to describe the entire system that mechanism operates within — a neural network model, grounded in known brain anatomy, of how a person acquires and continuously fine-tunes the motor skill of speaking.
DIVA proposes that speech is controlled by two parallel systems working together: a feedforward system, which stores learned motor commands for producing familiar sounds efficiently and quickly, and a feedback system, which monitors auditory and somatosensory (touch and proprioceptive) consequences of speech in real time and corrects the feedforward commands when they drift off target. Early in life — or whenever adapting to a change, such as a growing vocal tract, a new set of teeth, or a neurological change following injury — the feedback system does most of the work, actively listening and correcting. Over time, successful corrections get folded back into the feedforward system, making the skill increasingly automatic. This is the same basic architecture used to explain motor learning in reaching and other skilled movements, applied specifically to the vocal tract.
Guenther's 2006 paper in the Journal of Communication Disorders, "Cortical interactions underlying the production of speech sounds," lays out the model's architecture and maps its components onto specific brain regions — including Broca's area, premotor cortex, and superior temporal auditory regions — rather than treating it as a purely abstract computational diagram. A companion body of work, including fMRI studies published in Cerebral Cortex, has tested DIVA's predictions against real brain-imaging data collected while participants spoke, finding activation patterns broadly consistent with the model's proposed feedback and feedforward pathways.
DIVA has since been used to model and explain a range of speech disorders by simulating what happens when specific components of the model are disrupted — including stuttering, apraxia of speech, and hearing-impairment-related speech changes — by degrading or removing specific feedback pathways in simulation and observing whether the resulting model output resembles the disordered speech pattern seen clinically.
What Goeckoh is actually measuring and correcting, at the level of sound waves.
Everything above describes a neural loop. This section describes the acoustic signal that loop is actually processing. When air is pushed from the lungs through the vocal folds, it produces a buzzing source sound rich in harmonics. The vocal tract — throat, mouth, tongue position, lip shape — acts as a resonant filter on that source sound, amplifying certain frequency bands and damping others. Those amplified frequency bands are called formants, and this description of speech production as a source (the vocal folds) filtered by a resonator (the vocal tract) is known as source-filter theory, formalized by Gunnar Fant in his 1960 book Acoustic Theory of Speech Production — still the foundational text of modern acoustic phonetics.
The first two formants, F1 and F2, are what your brain uses to identify which vowel it's hearing, almost independent of who is speaking. F1 corresponds inversely to tongue height (a low F1 corresponds to a high tongue position, as in "ee"; a high F1 corresponds to a low tongue position, as in "ah"). F2 corresponds to tongue frontness (a high F2 corresponds to a front vowel like "ee"; a low F2 corresponds to a back vowel like "oo"). Plot F1 against F2 for every vowel a speaker produces, and you get that speaker's personal vowel space — a map of exactly where their vocal tract places each vowel acoustically.
James Hillenbrand, Laura Getty, Michael Clark, and Kimberlee Wheeler published the modern normative reference for American English vowels in 1995 in the Journal of the Acoustical Society of America: "Acoustic characteristics of American English vowels." Using recordings from 139 speakers (men, women, and children) producing a standard set of vowels, they established the statistical distribution of formant frequencies for each vowel category — in effect, a map of where a given vowel "should" sit acoustically, and how much natural variation exists around that target across speakers of different ages, sexes, and vocal tract sizes. This dataset (and its extensions) remains the standard reference point for what counts as a clearly identifiable vowel versus one that has drifted toward a neighboring vowel's acoustic territory.
This is directly clinically relevant: many of the speech differences associated with dysarthria, apraxia, and some ASD-related speech patterns show up acoustically as vowels that have drifted from their expected F1/F2 territory — a centralized or compressed vowel space, where distinct vowels acoustically overlap and become harder for a listener to tell apart, even though the intended word was correct.
The engineering method that turns formant theory into something that can run in real time on a phone.
Knowing that formants exist and knowing where they should sit doesn't, by itself, give you a way to measure and change them in the roughly 20-millisecond windows available during live conversation. That's a signal-processing problem, and the dominant solution — used in telephony, speech coding, and speech analysis since the 1970s — is linear predictive coding (LPC).
LPC works by modeling the vocal tract as an "all-pole" filter: a mathematical system where each sample of the speech waveform can be predicted as a weighted sum of the samples immediately before it. Solving for the weights that best predict a short window of real speech gives you a set of filter coefficients that describe the shape of the vocal tract's resonances at that instant — which is mathematically equivalent to recovering the formant frequencies directly from the waveform, without needing a separate acoustic model of the throat and mouth. Because the underlying math (solving a linear system) is computationally cheap, LPC analysis can run on a short sliding window of audio fast enough for real-time use, which is precisely why it became the backbone of both mid-20th-century telephone compression and modern real-time speech analysis.
Bishnu Atal's foundational work at Bell Labs in the late 1960s and early 1970s established both the analysis method (extracting LPC coefficients from a speech signal) and the synthesis method (using those coefficients, plus a simplified source signal, to regenerate intelligible speech) — the same two operations, analysis and resynthesis, that sit at the center of any real-time formant-correction pipeline. John Makhoul's widely cited 1975 tutorial in the Proceedings of the IEEE, "Linear Prediction: A Tutorial Review," remains the standard reference explaining why this particular mathematical approach became the dominant method in the field, and how its coefficients map back onto genuine acoustic properties of the vocal tract rather than being an arbitrary compression trick.
Because LPC coefficients correspond to actual resonance frequencies, they can be manipulated directly — shifted toward a target formant position. One classical approach, dating back to Atal's original work, is to then resynthesize the speech signal from those adjusted coefficients plus a modeled source signal. This is what makes LPC-based methods suited to formant correction specifically, as opposed to other voice modification techniques that operate on pitch or timing instead — though as the next section explains, Goeckoh itself uses a simpler, more direct variant of this idea rather than full resynthesis.
Why session data is structured the way it is, and what it's actually tracking.
Everything above explains why real-time formant correction should help. It doesn't, by itself, tell a family or a clinician whether it's actually working for a specific person over time. That's a measurement problem, and Goeckoh's answer to it is structured around Applied Behavior Analysis (ABA) — not as a therapy method Goeckoh delivers, but as the measurement discipline behind how session data is recorded.
ABA's core methodological contribution, independent of any specific intervention technique, is insisting that target behaviors be defined operationally (specific, observable, and countable) and tracked systematically over time, so that whether a skill is actually improving is a question answered by data rather than by subjective impression. Cooper, Heron, and Heward's textbook Applied Behavior Analysis (now in its third edition, 2020) is the standard graduate-level reference for this methodology and is used broadly across special education and clinical behavior analysis training.
Applied to Goeckoh specifically: every session logs concrete, countable events rather than a vague sense of "it seemed to go well" — which specific phonological targets (vowel or consonant productions) were attempted and how close each landed to its acoustic target, discrete fluency events (hesitations, self-corrections, successful productions), and co-regulation markers relevant to emotional state during a session. This turns each session into structured, comparable data points rather than an isolated anecdote, so a caregiver or clinician can look at a trend across weeks instead of relying on memory of how last Tuesday went.
Every mechanism on this page is what the live demo is actually running — not a simulation of it. Try it with your own voice, free, no account required.