AI-powered speech therapy games Download free

SpeechLP data report · Version 1.1

State of Kids' Speech Practice 2026

How often, how long, when and on which sounds children practice speech at home with SpeechLP, and what model-scored accuracy can and cannot tell us about progress.

This report describes about 61,000 practice words from about 950 children on SpeechLP family accounts, January to October 2026. The median child practiced 16 words per active week, the median session lasted under 2.5 minutes, and practice peaked in the evening. Of children with enough follow-up, 36% practiced again after week 1. Model-scored accuracy stayed about the same across sessions.

Key findings

  • 16 words Children practicing speech sounds at home with SpeechLP in 2026 practiced a median of 16 words in each week that they practiced. (n = 950 children, approximately)
  • 36% Only 36% of children who started practicing speech sounds with SpeechLP practiced in the app again after their first week. Among the 686 children followed long enough, 19% practiced again after their seventh week. (n = 855 children with enough follow-up time)
  • 8 words A typical home practice session on SpeechLP covered a median of 8 practice words in under 2.5 minutes. (n = 950 children, approximately)
  • 35% 35% of home practice sessions on SpeechLP started in the evening, between 5:00 and 9:59 pm local time. (n = 950 children, approximately)
  • 10% The /r/ sound made up 10% of practice words on SpeechLP family accounts in this report, more than any other sound. (n = 950 children, approximately)
  • 72% Model-scored accuracy on a sound stayed about the same across a child's SpeechLP sessions: about 72% in each of the first four sessions, and 2.9 points higher in sessions 4 and later than in the first (95% CI roughly -1 to +7), a difference not distinguishable from zero. Accuracy here is scored by SpeechLP's on-device model, not by a clinician. (n = 190 children, approximately)
  • 82% In a small 2025 internal comparison on 49 recorded child word attempts from very few speakers, scored by one speech-language pathologist, the phoneme model SpeechLP selected agreed with the clinician on whether the target sound was produced 82% of the time (Cohen's kappa 0.59), compared with 57% (kappa 0.04) for a general-purpose speech-to-text system. The model tested was an earlier version than the one in the app today. (n = 49 recorded word attempts)

Scope: aggregate data from SpeechLP practice sessions on family (parent) accounts, January 22 to October 7, 2026. Internal, staff, test, demo and guest accounts are excluded, and accounts set up by speech-language pathologists are excluded from per-child figures. Every published figure from practice data covers at least 5 children. Accuracy is scored by SpeechLP's on-device phoneme model, not by a clinician, and is not first-try accuracy. These are observational data from families who chose to use the app, so they cannot show cause and effect.

Methods

Data source. SpeechLP is a speech practice app in which children say target words aloud and SpeechLP's on-device phoneme model judges whether the target sound was produced. Each practice card (one word) is logged as an analytics event. We analyzed every practice card logged between January 22 and October 7, 2026 (Toronto time). SpeechLP moved scoring to its on-device phoneme model in January 2026, and the window starts after that change. Practice plans, assigned homework and other data stored outside the analytics system were not available for this report.

Exclusions. We removed internal, staff, test, demo and guest accounts using the same rules as SpeechLP's internal reporting, and cards with no app version recorded. We excluded screening and onboarding cards (they are assessments, not practice) and cards that target more than one sound at once. Per-child figures use family (parent) accounts only; accounts set up by speech-language pathologists were excluded from them. After exclusions, the per-child analysis covers about 950 children and about 61,000 practice words.

Definitions. A child is one child profile within one account. Age is the age entered for the child at their first card in the window, grouped as 2 to 5, 6 to 8 and 9 to 12 years; children with no age, or an age outside 2 to 12, are included in overall figures only. A session is a run of practice within one app session, split wherever more than 30 minutes pass between cards. Weeks are counted from each child's first card (week 1 is days 0 to 6). An active week is a week with at least one session. Session length is the time from the start of the first card to the end of the last card. Sounds are written with familiar symbols: /r/ stands for the English r sound (IPA /ɹ/), and "th", "sh" and "ch" for the sounds in think or this, ship and chip.

Model-scored accuracy. A practice word counts as correct when SpeechLP's on-device phoneme model accepted it within that card's attempts. The number of attempts per card is not logged, so model-scored accuracy is not first-try accuracy, and it is not a clinician's judgment. The section How accurate is the scoring model? describes what is known about how the model compares with a speech-language pathologist.

Statistics. We report medians or means across children, with 95% confidence intervals from a bootstrap that resamples whole children (2,000 resamples), so that children who practice a lot do not dominate. Changes in model-scored accuracy compare each child with themselves on the same sound, and are reported in percentage points. Three separate analyses were run on the same frozen data extract with written conventions, and headline figures were re-derived by a separate verification analysis from its own data pull before publication. All of this work was done by or for SpeechLP, using AI-assisted analysis tools; every figure comes from saved, re-runnable code, and the published wording was checked against the data before release. Where these figures differed slightly because of small differences in definitions, we publish the figure at a precision on which both agree (for example, "about 950 children"). Two follow-up analyses on practice amount and frequency were specified in writing before they were computed and are reported regardless of their direction.

Privacy. Only aggregate statistics are published. Every published figure from practice data covers at least 5 children, with secondary suppression so that no smaller group can be worked out by subtraction. No audio, names, words as said, individual timelines or exact dates tied to small groups are published.

Results

The results cover about 950 children and about 61,000 practice words on family accounts between January 22 and October 7, 2026. The first sections describe how much, how consistently, how long and when children practiced, and which sounds they worked on. The last section looks at whether model-scored accuracy changed across a child's sessions on a sound.

How much children practice

In weeks when they practiced, children practiced a median of 16 words (about 950 children). The mean was higher, 27.9 words (95% CI 25.4 to 30.6), because a minority of children practiced far more than most.

In a typical active week, children had one practice session (about 950 children; median 1.0 sessions per active week, 95% CI 1.0 to 1.3; mean 1.88, 95% CI 1.74 to 2.07). Measured by time, the median child practiced 4.0 minutes in each active week (95% CI 3.6 to 4.5).

How consistently children practice

Practice was uneven over time. Among children who had at least 28 days since their first card, children practiced in an average of 34% of their first four weeks (about 780 children). Most of these children practiced in only one of those four weeks (74.4%, 95% CI 71.3% to 77.4%), and 2.7% practiced in all four (95% CI 1.7% to 3.9%). Over the first eight weeks, children practiced in an average of 20.2% of weeks (95% CI 19.0% to 21.3%; n = 686).

Many children practiced in a short burst and then stopped using the app, at least for a while. Among children with enough follow-up time, 36% practiced again after their first week (95% CI 33% to 40%; n = 855), 27% after their third week (95% CI 24% to 30%; n = 780), 19% after their seventh week (95% CI 16% to 22%; n = 686) and 14% after their eleventh week (95% CI 12% to 17%; n = 634). These figures count practice in SpeechLP only; families may also practice in other ways that the app cannot see, so they are a floor, not a measure of all home practice.

Figure 1. How many of their first four weeks children practiced Practice done away from the app is not recorded, so these are minimum counts. 95% confidence intervals from a bootstrap over children.

Horizontal bar chart of 780 children who had at least 28 days of follow-up after their first practice word. It shows the share of children by how many of their first four weeks (7-day blocks counted from the first practice word) included at least one practice session: one week 74% (580 children), two weeks 17% (133), three weeks 6% (46), all four weeks 3% (21). On average, 34% of a child's first four weeks included practice. In weeks with any practice, the median child practiced 16 words.

Show the data for Figure 1
Children by number of their first four weeks with practice (n = 780 children)
Weeks with practiceShare of children95% CIChildren
1 of 4 weeks74%71% to 77%580
2 of 4 weeks17%14% to 20%133
3 of 4 weeks6%4% to 8%46
All 4 weeks3%2% to 4%21
Figure 2. How many children kept practicing after their first weeks Weeks are 7-day blocks from each child's first practice word. 95% confidence intervals from a bootstrap over children.

Horizontal bar chart of the share of children with any SpeechLP practice in a later week, counting weeks as 7-day blocks from each child's first practice word. After week 1 (any practice in week 2 or later): 36% of 855 children (95% CI 33% to 40%). After week 3: 27% of 780 children (24% to 30%). After week 7: 19% of 686 children (16% to 22%). Each row includes only children observed long enough to reach that week, so the number of children differs by row.

Show the data for Figure 2
Children with any practice after week 1, 3 and 7
Practiced againShare of children95% CIChildren observed
After week 1 (week 2 or later)36%33% to 40%855
After week 3 (week 4 or later)27%24% to 30%780
After week 7 (week 8 or later)19%16% to 22%686

How long a practice session lasts

The median practice session covered 8 practice words and lasted under 2.5 minutes (about 950 children). Sessions varied widely, from a handful of words to several minutes of practice.

These times run from the first card to the last card in a session, so they leave out time spent on menus, rewards and instructions. Measured the way SpeechLP's internal weekly reports measure it (from first to last game event, sessions with a finished game only), the median session was 4.3 minutes (95% CI 3.9 to 4.7; n = 864 children).

When families practice

Evenings were the busiest time for each hour of the day. 35% of sessions started between 5:00 and 9:59 pm on the device's local clock (about 950 children).

Weekends were slightly quieter than weekdays: 25% of sessions fell on a Saturday or Sunday, a little below the 29% (2 days out of 7) expected if practice were spread evenly across the week.

Figure 3. When children practice: share of sessions by time of day Session = a child's practice words within one app session, split on gaps over 30 minutes. Times are device-local. 95% confidence intervals from a bootstrap over children.

Histogram of about 4,000 practice sessions from about 950 children, by the start time on the child's device. Because the time blocks have different lengths, bar height shows the share of sessions per hour and the label shows the block's total share. 6 to 9 am: 10%. 9 am to 3 pm: 35%. 3 to 5 pm: 14%. 5 to 8 pm: 25%, the busiest hours. 8 to 10 pm: 10%. 10 pm to 6 am: 6%. Evening sessions (5 to 10 pm) make up 35% of all sessions. Weekend sessions make up 25%, a little below the 29% expected if practice were spread evenly over the week.

Show the data for Figure 3
Practice sessions by start time on the child's device
Start timeShare of sessions95% CIShare per hour
6 to 9 am10%8% to 14%3.5%
9 am to 3 pm35%31% to 38%5.8%
3 to 5 pm14%13% to 16%7.1%
5 to 8 pm25%22% to 28%8.3%
8 to 10 pm10%9% to 11%5%
10 pm to 6 am6%5% to 8%0.8%

Which sounds children practice

The /r/ sound was the most practiced sound by number of words: 10% of all practice words targeted /r/ (about 950 children), ahead of /s/ (7.1%, 95% CI 6.5% to 7.9%) and /l/ (6.9%, 95% CI 6.2% to 7.6%). By number of children, practice was spread widely: 55% of children practiced /r/, and a similar share practiced /k/ (56.5%, 95% CI 53.1% to 59.7%).

The focus shifts with age. Among children aged 2 to 5, early sounds such as /p/, /t/, /k/ and /n/ led; 46% practiced /r/, and /r/ made up 4.2% of their practice words (95% CI 3.3% to 5.6%; n = 437). Among children aged 6 to 8, 64% practiced /r/, and it made up 12.7% of practice words (95% CI 9.0% to 16.7%; n = 265). Among children aged 9 to 12, about 7 in 10 practiced /r/, and it made up 16.0% of practice words (95% CI 9.9% to 24.3%; n = 146).

Most children worked on several sounds. Only 12.7% of children with at least 10 practice words spent half or more of their words on a single sound (95% CI 10.3% to 15.4%; n = 667); among those focused children, /r/ was the most common focus sound.

Pooled across children, model-scored accuracy on /r/ looked lower for vocalic /r/, the r-colored vowels in words like car and bird (59.2%, 95% CI 54.5% to 64.6%; 411 children), than for /r/ before a vowel, as in red (65.1%, 95% CI 59.7% to 69.7%; 352 children), though the confidence intervals overlap. These contexts were assigned by an automatic rule that has not been reviewed by an SLP, the scores are model-scored, not clinician-scored, and this comparison was not re-derived in the separate verification analysis.

For context, in the U.S. studies reviewed by Crowe and McLeod (2020), the age by which 90% of typically developing children produced /r/ correctly averaged 5;6 (years;months), one of the latest consonants to develop. In this data, /r/ was the most practiced sound among children aged 6 and older.

Figure 4. Most-practiced sounds by age: the r sound dominates after age 6 Single-sound practice cards only; age at the child's first practice word. Children with unknown age are excluded from this figure. 95% confidence intervals from a bootstrap over children.

Three small bar charts, one per age band, each showing the five sounds with the largest share of that age group's practice words, plus /ɹ/ (the English r sound) where it is not in the top five. Ages 2 to 5 (437 children): /p/ 8%, /s/ 7%, /n/ 7%, /k/ 7%, /t/ 6%, /ɹ/ 4%. Ages 6 to 8 (265 children): /ɹ/ 13%, /l/ 8%, /s/ 7%, /t/ 6%, /b/ 5%. Ages 9 to 12 (146 children): /ɹ/ 16%, /l/ 8%, /t/ 8%, /s/ 7%, /p/ 6%. At ages 2 to 5, practice is spread across early sounds such as /p/, /s/ and /n/. From age 6, /ɹ/ is the most-practiced sound, well ahead of the next sound. The share of children who practiced /ɹ/ at least once rises from 46% at ages 2 to 5 to 64% at 6 to 8 and about 7 in 10 at 9 to 12.

Show the data for Figure 4
Share of practice words by sound and age band (top five sounds plus /ɹ/)
Age bandSoundShare of practice words95% CIChildren practicing the sound
Ages 2 to 5 (437 children)/p/8.1%6.6% to 10%273
Ages 2 to 5 (437 children)/s/7.5%6.4% to 8.7%246
Ages 2 to 5 (437 children)/n/7.3%6.6% to 8.1%263
Ages 2 to 5 (437 children)/k/6.8%6% to 7.7%263
Ages 2 to 5 (437 children)/t/6.3%5.7% to 7%266
Ages 2 to 5 (437 children)/ɹ/4.2%3.3% to 5.6%46% of children
Ages 6 to 8 (265 children)/ɹ/12.7%9% to 16.7%169
Ages 6 to 8 (265 children)/l/7.6%6.5% to 8.9%167
Ages 6 to 8 (265 children)/s/6.9%5.9% to 8%160
Ages 6 to 8 (265 children)/t/5.6%4.5% to 6.7%151
Ages 6 to 8 (265 children)/b/5.5%4.6% to 6.5%151
Ages 9 to 12 (146 children)/ɹ/16%9.9% to 24.3%104
Ages 9 to 12 (146 children)/l/7.6%5.1% to 11.1%79
Ages 9 to 12 (146 children)/t/7.5%4.8% to 10%77
Ages 9 to 12 (146 children)/s/7.3%5% to 10.7%73
Ages 9 to 12 (146 children)/p/5.6%3.6% to 7.8%68

Finishing games, by account type

We measured how often a game that was started was finished: the share of game starts whose next game event was a finished game (first-round completion). Account type is the type recorded on the account today, which may differ from what it was when the game was played, and the unit is accounts, not children.

Parent accounts, whether or not they were linked to an SLP, finished more than half of the games they started (56% to 60%, depending on the group; about 890 parent accounts). Accounts set up by SLPs finished fewer than half (about 130 accounts). SLP accounts may often start a game to preview it rather than to practice; the data cannot tell these uses apart.

Model-scored accuracy over time

Across all accounts, about 74% of single-target practice words were scored correct by SpeechLP's on-device phoneme model within the card's attempts (about 1,060 children; 95% CI 68% to 78%). This is not first-try accuracy and not a clinician's score.

Model-scored accuracy stayed about the same across a child's sessions on a sound. Comparing each child with themselves on the same sound, accuracy in sessions 4 and later was 2.9 points higher than in the first session (95% CI roughly -1 to +7, which includes zero; about 190 children). Average accuracy was about 72% in each of a child's first four sessions on a sound (see the table below). Comparing each child's first and last session on a sound gave a difference of 0.1 points (95% CI -2.3 to 2.6; 491 children), and statistical models using every session found no meaningful trend.

Later points on the curve include only the children who got that far, so they describe a smaller and more selected group: children who went on to four or more sessions started 6.6 points lower in their first session than children who stopped sooner (95% CI 2.8 to 10.1; n = 629). When the second session is used as the baseline instead of the first, the difference to sessions 4 and later is -0.3 points (95% CI -4.6 to 4.0; 190 children), so whatever difference exists sits between the first and second session, possibly while children are still getting used to a sound and the app.

Where later sessions did score higher, it was on words the child had already practiced in the first session (10.5 points, 95% CI 3.9 to 16.9; 110 children), not on new words (2.3 points, 95% CI -1.5 to 6.3; 189 children). That pattern is consistent with familiarity with specific words rather than carryover to new ones.

Few children practiced long enough to apply a common mastery convention (80% model-scored accuracy in each of 3 consecutive sessions on the same sound and position, with at least 5 words per session): fewer than 4% of children had enough sessions to test it (35 of 958; 95% CI 2.5% to 4.9%).

Sensitivity to analysis choices. Results depend on how many words a session must contain on the sound to count. Requiring at least 5 words in each compared period gives a larger difference, about 9 points (59 children), but across minimums from 1 to 10 words the difference ranged from 1.3 to 9.1 points, and its confidence interval excluded zero only for minimums of 4 to 6 words. That threshold was not set in advance, so we make no improvement claim. For context only, a rough calculation we derived from published age-of-acquisition norms (Crowe & McLeod, 2020) suggests about 1 point of change from development alone over spans like these; those norms are cross-sectional and do not measure individual growth, other reasonable calculations give larger values, and neither is a measured comparison group.

No dose-response. After accounting for starting accuracy, age band and app version, the number of sessions on a sound was not associated with larger change: each doubling of sessions went with 0.2 points more change (95% CI -2.3 to 2.8; n = 123). Two follow-up analyses specified before they were run did not change this picture. The first followed the children who practiced a sound the most: 8 or more sessions of at least 5 words on the sound, spanning at least 21 days (10 children, 30 child and sound pairs, a median of about 4 months between the first and last sessions compared). Their model-scored accuracy rose from about 62% in the first session to about 76% in sessions 8 and later, and the last 2 sessions scored 17 points higher than the first 2 (95% CI 5 to 28). That change did not survive the checks set in advance. These children started 17 points below other children in their first session (95% CI 4 to 30 points lower), leaving more room to rise. Comparing the same words over the same span, the change was 9 points before adjustment (95% CI -2 to 26), -3 points after adjusting for app version and word difficulty (95% CI -16 to 19), and +1 point after also adjusting for calendar month (95% CI -27 to 26). Over about 4 months, typical development alone could account for up to 18 points on these sounds (observed minus that upper bound: -1 point, 95% CI -11 to 10). With 10 children and no comparison group, this result cannot show that practice caused the change. A second comparison of frequent (weekly or more) with infrequent (monthly or less) practice over the same span of time (n = 33) found, if anything, smaller change with frequent practice, but that result did not hold in a split-half check or over a longer span.

Bottom line. In this data, model-scored accuracy on a sound did not measurably change over the sessions children completed. That is not evidence that practice does not work: most children practiced only briefly, the scoring and content of the app changed during 2026, the model's score passes a word within the card's attempts rather than on the first try, and observational data could not show cause and effect in either direction. A blinded, controlled study with untrained probe words scored by speech-language pathologists is planned to answer the question properly.

Figure 5. Model-scored accuracy by session number on a sound Model-scored accuracy is the share of practice words SpeechLP's on-device phoneme model scored as correct within the card's attempts; it is not a clinician's rating and not first-try accuracy. Session 10+ pools all sessions from the 10th on. 95% confidence intervals from a bootstrap over children.

Line chart of mean per-child model-scored accuracy at each session on a given sound, with a shaded 95% confidence band. The number of children at each point falls as children stop practicing a sound: session 1, 72% (954 children); session 2, 72% (491 children); session 3, 71% (287 children); session 4, 72% (190 children); session 5, 75% (132 children); session 6, 69% (97 children); session 7, 74% (80 children); session 8, 76% (65 children); session 9, 79% (55 children); session 10+, 71% (45 children). The line is roughly flat, between about 69% and 79%. Because different children make up each point, the line mixes change within children with changes in who is still practicing. Following only the 45 children who reached 10 or more sessions on a sound, accuracy changed by +2.8 points from their first session to their 10th and later sessions (95% CI −6.8 to 11.6).

Show the data for Figure 5
Mean model-scored accuracy by session number on a sound
Session on the soundModel-scored accuracy95% CIChildrenPractice words
172%70% to 74%95422268
272%70% to 74%4919905
371%68% to 74%2876139
472%68% to 75%1904100
575%70% to 79%1323027
669%64% to 74%972254
774%68% to 79%801740
876%71% to 81%651332
979%74% to 85%55993
10+71%63% to 78%459153

How accurate is the scoring model?

Every accuracy figure in this report comes from SpeechLP's on-device phoneme model, which listens to each attempt and decides whether the target sound was produced. A practice word counts as correct if the model accepted it within the card's attempts. It is not first-try accuracy and it is not a clinician's score, so it should be read as a practice signal, not a clinical measure.

The only published paired comparison. In a small 2025 internal comparison, a speech-language pathologist transcribed 49 recorded child word attempts from very few speakers, each containing a target sound such as /s/, /z/, /f/, /v/, "th", "sh" or "ch". We then checked whether each system's output agreed with the clinician on whether the target sound was produced. The phoneme model SpeechLP selected in that comparison, an earlier version than the one in the app today, agreed with the clinician on 82% of the attempts (Cohen's kappa 0.59, 95% CI about 0.34 to 0.81). A general-purpose speech-to-text system, of the kind that scored practice in early versions of the app, agreed 57% of the time (kappa 0.04, no better than chance).

How does that compare? A systematic review of automated speech analysis tools for children applied an 80% agreement threshold, based on commonly accepted agreement between two human raters of 75 to 85% for perceptual speech judgments (McKechnie et al., 2018, citing Charter, 2003, and Cucchiarini, 1996). The point estimate clears that bar, but the sample is small, the lower end of its uncertainty range falls below 80%, and a kappa of 0.59 means agreement well above chance but far from perfect.

This comparison has real weaknesses. It used a small set of recordings from very few speakers, scored by one speech-language pathologist, and when three SLPs transcribed two of the same recordings, they disagreed on some sounds. The scoring rule checked whether the sound appeared anywhere in the word rather than in its target position, and the model tested was an earlier version than the one in the app today. It should not be summarized as "SpeechLP is 82% accurate."

Internal testing. In internal testing, SpeechLP's team, including its speech-language pathologists, checked the model's detection on every target sound across hundreds of trials with about 10 children aged 2 to 7. That testing found detection performing as intended on all target sounds except /s/ produced with a lateral lisp. Detection of lateral /s/ errors is a known weakness under active development. Item-level results from this testing were not recorded in a form we can publish or independently analyze, so no agreement statistic is reported from it. This testing was done by SpeechLP and was not independent.

What comes next. A formal, blinded study of agreement between the model and speech-language pathologists, with a larger and more varied group of children, is planned. Until then, model-scored accuracy is best read as a signal of how practice is going, not as a measure of clinical progress.

Figure 6. 2025 model comparison: agreement with a speech-language pathologist Clip-level comparison from 2025 with one SLP rater; it is not a measure of the model in the app today.

Dot plot of raw agreement between two systems and a speech-language pathologist's transcription on whether the target sound was produced, across 49 recorded child word productions from SpeechLP's 2025 internal model comparison. A vertical line marks the 80% agreement benchmark. The model SpeechLP selected in that comparison agreed with the clinician on 82% of clips (Cohen's kappa 0.59). A general-purpose speech-to-text system, scored on whether it recognized the target word, agreed on 57% (kappa 0.04, close to chance). The clip set is small and the comparison predates the model used in the app today.

Show the data for Figure 6
Agreement with an SLP on target-sound presence, 2025 internal comparison
SystemAgreement with the SLPCohen's kappaClips
SpeechLP's selected model82%0.5949
General-purpose speech-to-text57%0.0449

What this means for parents

Speech sound difficulties are common: the National Institute on Deafness and Other Communication Disorders estimates that 8 to 9% of young children have a speech sound disorder (NIDCD, 2025). If you are practicing at home, here is what the data suggest.

  • Short sessions are common. A typical session in this data lasted a couple of minutes. This report cannot say how much practice is enough, so ask your speech-language pathologist how often to practice and how many words to aim for.
  • Keep going past the first week. Most children in this data practiced in a short burst and then stopped. The first week or two is when routines are easiest to lose, so plan the next few sessions ahead.
  • Frequency mattered in clinical research. In one clinical trial with preschoolers, children seen three times a week did better than children who had the same number of sessions once a week (Allen, 2013). This report's app data could not confirm a frequency effect, so discuss a schedule with your speech-language pathologist.
  • Evenings work for many families. Evening was the most common practice time in this data. Pick a time that fits your routine, such as after dinner or before bath, and keep it predictable.
  • /r/ takes time. The r sound is one of the last sounds children learn: in U.S. studies, the age by which 90% of children said it correctly averaged about 5;6 (Crowe & McLeod, 2020). In this data, most children aged 6 and older practiced /r/. If your child is past that age and still working on /r/, a speech-language pathologist can advise you.
  • Treat the app's score as a practice signal. SpeechLP's score shows how practice is going; it is not a diagnosis and not a clinician's assessment. Scores move around from session to session, and even a trend over weeks is a model's score, not a clinical measure of progress.
  • Keep your speech-language pathologist in the loop. Share which sounds and words your child is practicing and how often. If you have concerns about your child's speech and do not have an SLP, a licensed speech-language pathologist can evaluate your child.

What this means for SLPs

Dose at home is small. Published phonological intervention research typically delivered 2 to 3 sessions a week of about 100 production trials each (Sugden et al., 2018). In this data, the median home session covered 8 practice words and the median child practiced 16 words in an active week. A practice word can include more than one attempt, and attempts are not logged, so production trials are undercounted; even so, the gap between research dose and home dose is large. Reviews of a small evidence base suggest that greater treatment intensity tends to go with better outcomes in speech sound disorders (Kaipa & Peterson, 2016; Williams, 2012), and in one trial with preschoolers, more frequent sessions did better at equal total sessions (Allen, 2013).

Frequency is the weak point. Children practiced in about a third (34%) of their first four weeks (about 780 children), and most practiced in only one of those weeks. Home practice is rarely described in detail in research: of 176 phonological intervention papers, 61 reported parent involvement or home tasks, mostly with limited detail (Sugden et al., 2016), and surveyed SLPs report that following up on homework completion is inconsistent, with school-based SLPs less likely to follow up (Tambyraja, 2020). App practice logs give you a frequency record to review with families at each visit.

What model-scored accuracy can tell you: whether a child is practicing, on which sounds and words, and the direction of their in-app scores on a sound over weeks. What it cannot tell you: first-try accuracy (attempts are not logged), accuracy as you would score it, accuracy in a target position (agreement was checked on presence of the sound in the word), carryover to connected speech, or generalization outside the app. Detection of lateral /s/ errors is a known weakness. A mastery criterion such as 80% across three consecutive sessions should not be applied to model-scored accuracy as if it were clinician-scored.

On progress, the honest reading is narrow. Model-scored accuracy on a sound stayed about the same across sessions, a small difference that does appear is tied to repeated words rather than new ones, and we found no dose-response. Fewer than 4% of children practiced long enough to test a three-session mastery criterion at all. These data neither show that home practice in SpeechLP changes how a child speaks nor show that it does not; with home doses this small and this short, a null in observational data is expected either way. A blinded, controlled study with SLP-scored untrained probe words is planned.

Keeping families practicing is a practical lever. Only about a third of children practiced in the app again after their first week (n = 855). Building the next two or three sessions into the plan, checking in early, and choosing words with the child's carryover goals in mind are practical steps that do not depend on any app.

Target selection. /r/ dominated practice for children aged 6 and older, at or past the average age (5;6) by which 90% of children in U.S. norm studies produce it (Crowe & McLeod, 2020), while younger children practiced a broad mix of early sounds. Word selection is itself an active ingredient in phonological learning (Storkel, 2018), so the words assigned for home practice deserve the same attention as the sound.

Limitations

This is a descriptive report of product data. Its main limitations:

  • Observational data with no control group. Nothing in this report shows that practice in SpeechLP causes change in a child's speech. A blinded, controlled study with SLP-scored untrained probe words is planned.
  • No IRB review and no controlled trial. The data were collected in the normal course of using the app and analyzed in aggregate.
  • Internal analysis and review. The report was produced and funded by SpeechLP, the analysis was done by SpeechLP, and the clinical reviewers are SpeechLP co-founders and co-CEOs, so the review was not independent. No external peer review has been done.
  • Model-scored, not clinician-scored. Accuracy comes from SpeechLP's on-device phoneme model. Its agreement with clinicians has been checked only in a small 2025 comparison of an earlier model version and in internal testing without recorded item-level results; no paired comparison of the model in the app today has been published. Detection of lateral /s/ errors is a known weakness under active development.
  • Attempts are not logged. A word counts as correct if it passed within the card's attempts, so model-scored accuracy is not first-try accuracy, and more retries or prompting would raise it.
  • Family accounts only. Per-child figures exclude children on accounts set up by speech-language pathologists, so they may not describe children who practice under an SLP's account.
  • Start of the window. Children who began using the app before January 22, 2026 are counted from their first card in the window, so for some children the first week in this data is not their true first week.
  • Account type is current-state. Whether an account is a parent account, linked to an SLP or an SLP account reflects its setting today, not necessarily at the time of practice.
  • Analysis choices. The within-child comparison of model-scored accuracy is sensitive to the minimum number of words a session must contain, a threshold that was not set in advance (see Results). No improvement claim is made.
  • The app changed over the window. Scoring, content, games and experiments changed during 2026, which affects comparisons over time.
  • Data stored outside the analytics system, such as practice plans and assigned homework, were not available, so we cannot compare practice with what was assigned.
  • Self-selected users. Families who choose and keep using a practice app differ from families who do not, and children who kept practicing differ from those who stopped.
  • No diagnosis or clinical data. We do not know children's diagnoses, severity, outside therapy, or who helped them practice.
  • Practice outside the app is invisible, so weeks without app practice are not necessarily weeks without practice.
  • Ages are as entered by families, and age bands with too few children are not reported.

How these numbers relate to earlier SpeechLP posts

This report is the first time SpeechLP has published statistics on how children practice in the app. Figures that appear in SpeechLP's app store listing, website, ads or social posts, such as word and game counts, describe the product rather than practice, cover different periods and accounts, and were not produced with this report's exclusions and methods, so they should not be compared with the numbers here. When citing SpeechLP practice data, cite this report, with its window, its exclusions and the number of children behind each figure.

Data availability

State of Kids' Speech Practice 2026: aggregate tables

Aggregate statistics on at-home speech sound practice in the SpeechLP app from January 22 to October 7, 2026, mainly covering about 950 children on family accounts: practice volume, consistency, session length and timing, sounds practiced by age, and within-child change in model-scored accuracy, plus the 2025 model comparison. Each row is one published statistic with its estimate, 95% confidence interval and number of children; game completion by account type appears only as ranges in the notes. Every cell based on practice data covers at least 5 children. Released under CC BY 4.0.

The tables behind every chart and headline figure are available as a CSV file, and the full report as a PDF, using the download buttons on this page. Both are released under the Creative Commons Attribution 4.0 International license (CC BY 4.0): you may share and adapt them with credit to SpeechLP.

To protect children's privacy, row-level data are not released, and every published cell based on practice data covers at least 5 children (or 5 accounts). Questions about the methods can be sent to support@speechlp.com.

Period covered
January 22, 2026 to October 7, 2026
Cost
Free to access

Variables in the dataset

Practice words per active week (words per week)
Practice words (one card each) per child in weeks with at least one session; median and mean across children.
Share of first four weeks with practice (percent)
Share of a child's first four weeks with at least one practice session, among children with at least 28 days of follow-up.
Children still practicing (percent)
Share of children with any SpeechLP practice in week 2, 4 or 8 after their first card, or later.
Session length (minutes)
Minutes from the start of the first card to the end of the last card in a session, and practice words per session.
Session start time (percent)
Share of practice sessions starting in each block of the device's local time, including evening (17:00 to 21:59) and weekend shares.
Sounds practiced (percent)
Share of practice words and share of children for each target sound, overall and by age band.
First-round game completion (percent)
Share of game starts whose next game event was a finished game, by account type recorded today. Given only as ranges (more than half, fewer than half) in the notes, not as estimates.
Model-scored accuracy (percent)
Share of practice words that SpeechLP's on-device phoneme model accepted within the card's attempts. Not first-try accuracy and not clinician-scored.
Within-child change in model-scored accuracy (percentage points)
Difference between a child's later sessions and first session on the same sound, with sensitivity analyses.

As heard on

References

  1. Allen, M. M. (2013). Intervention efficacy and intensity for children with speech sound disorder. Journal of Speech, Language, and Hearing Research, 56(3), 865-877. https://doi.org/10.1044/1092-4388(2012/11-0076)
  2. Charter, R. A. (2003). A breakdown of reliability coefficients by test type and reliability method, and the clinical implications of low reliability. The Journal of General Psychology, 130(3), 290-304. https://doi.org/10.1080/00221300309601160
  3. Crowe, K., & McLeod, S. (2020). Children's English consonant acquisition in the United States: A review. American Journal of Speech-Language Pathology, 29(4), 2155-2169. https://doi.org/10.1044/2020_AJSLP-19-00168
  4. Cucchiarini, C. (1996). Assessing transcription agreement: Methodological aspects. Clinical Linguistics & Phonetics, 10(2), 131-155. https://doi.org/10.3109/02699209608985167
  5. Kaipa, R., & Peterson, A. M. (2016). A systematic review of treatment intensity in speech disorders. International Journal of Speech-Language Pathology, 18(6), 507-520. https://doi.org/10.3109/17549507.2015.1126640
  6. McKechnie, J., Ahmed, B., Gutierrez-Osuna, R., Monroe, P., McCabe, P., & Ballard, K. J. (2018). Automated speech analysis tools for children's speech production: A systematic literature review. International Journal of Speech-Language Pathology, 20(6), 583-598. https://doi.org/10.1080/17549507.2018.1477991
  7. National Institute on Deafness and Other Communication Disorders. (2025, July 8). Quick statistics about voice, speech, language. https://www.nidcd.nih.gov/health/statistics/quick-statistics-voice-speech-language
  8. Storkel, H. L. (2018). Implementing evidence-based practice: Selecting treatment words to boost phonological learning. Language, Speech, and Hearing Services in Schools, 49(3), 482-496. https://doi.org/10.1044/2017_LSHSS-17-0080
  9. Sugden, E., Baker, E., Munro, N., & Williams, A. L. (2016). Involvement of parents in intervention for childhood speech sound disorders: A review of the evidence. International Journal of Language & Communication Disorders, 51(6), 597-625. https://doi.org/10.1111/1460-6984.12247
  10. Sugden, E., Baker, E., Munro, N., Williams, A. L., & Trivette, C. M. (2018). Service delivery and intervention intensity for phonology-based speech sound disorders. International Journal of Language & Communication Disorders, 53(4), 718-734. https://doi.org/10.1111/1460-6984.12399
  11. Tambyraja, S. R. (2020). Facilitating parental involvement in speech therapy for children with speech sound disorders: A survey of speech-language pathologists' practices, perspectives, and strategies. American Journal of Speech-Language Pathology, 29(4), 1987-1996. https://doi.org/10.1044/2020_AJSLP-19-00071
  12. Williams, A. L. (2012). Intensity in phonological intervention: Is there a prescribed amount? International Journal of Speech-Language Pathology, 14(5), 456-461. https://doi.org/10.3109/17549507.2012.688866

How to cite

SpeechLP Research. (2026). State of Kids' Speech Practice 2026 (Version 1.1). SpeechLP. https://speechlp.com/research/state-of-kids-speech-practice-2026/

Please cite the canonical page URL and the version number shown above.

Cite the version you used. If the report is revised, the version number and update date on this page will change.

Free iOS app

Speech games kids love, designed with SLPs

Download SpeechLP and start practicing at home today.