← Back to Blog Research

Does Practice Improve Speaking? What 12,660 Sessions Mean for L&D

· 6 min read
William Burden
William Burden Founder @ Elqo

“Practice makes perfect” is the kind of claim learning teams inherit and rarely get to test. We tested it for speaking, on scored sessions run in Elqo.

We pulled every scored speaking session and looked at how scores changed as people completed more of them. 12,660 scored attempts. 1,033 people with at least three sessions. Every session graded on the same 0–100 scale by the same model. The question a program owner actually has to answer: do scores go up with reps, and is a block of deliberate practice worth putting on a cohort calendar?

Short answer: yes, they do, and the effect is larger and more consistent than we expected.

The headline numbers

Before the methodology, here is what 1,033 practice histories actually say.

MetricValue
Users with 3+ scored sessions1,033
Users with an improving trend666 (64%)
Users with a declining trend354 (34%)
Average points gained per additional session+1.9
Mean score at session 143.3
Mean score at session 1055.9
Mean score at session 2060.1
Mean score at session 3062.6
Absolute lift, session 1 → 30+19.3 points

Two out of three people improve. Early on, the average gain is roughly two points per additional session. In aggregate, scores climb from the low 40s to the low 60s over the first thirty sessions, a 45% relative improvement.

That last number looks decisive. It is also the one a skeptical L&D lead should pressure first.

What survivor bias does to the average

The aggregate average has a built-in problem: by session 30, only 46 of the original 1,033 people are still in the dataset. The other 987 stopped somewhere along the way. It is possible that the people who keep practising are the people who were already strong, or already improving, and that the score lift is self-selection.

So we re-ran the analysis with the same people tracked from start to finish. For each cohort below, we took only the people who reached at least K sessions, and compared their average score at session 1 with their average score at session K. The same people at both ends.

CohortUsersSession 1 avgSession K avgLift
Reached 5+ sessions44545.9652.04 (sess 5)+6.1
Reached 10+ sessions17148.6455.91 (sess 10)+7.3
Reached 20+ sessions7349.4860.05 (sess 20)+10.6

The 73 people who reached at least 20 sessions improved by 10.6 points from their first session to their twentieth. The same group, measured at both ends of their own practice arc. The improvement holds when the cohort is held constant.

The session-by-session path for that 20+ cohort:

Session #Avg scoreSession #Avg score
149.481158.77
250.991259.66
353.321357.99
455.401461.71
558.791562.05
657.961657.95
756.621760.75
859.211860.03
959.191960.23
1059.742060.05

Two things to notice. The trajectory is bumpy and unmistakably upward. And most of the lift comes in the first five sessions, from 49.5 to 58.8. After that, the gains are smaller and slower, layered on top.

Design the first five sessions as a sprint

The steepest gains in these 12,660 sessions show up early. Elqo scores each rep on the same 0–100 scale, so a cohort can see the curve as it forms.

Book a demo

Improvement holds at every engagement level

The next worry was that improvement might be a short-stay effect: people start low, pick up a few quick wins, then flatten. So we grouped people by total session count and compared the mean of the first half of each person's sessions with the mean of their second half.

Total sessionsUsersFirst half avgSecond half avgΔ% improved
3–458841.6545.75+4.1160%
5–927446.8350.52+3.6962%
10–199849.5153.16+3.6463%
20–495256.8860.52+3.6469%
50+2161.1664.13+2.9771%

Two clean findings:

  1. Every engagement bucket improves. The second half is higher than the first in every row. There is no ceiling group in this set. People with 50+ sessions are still gaining about 3 points between their early and late attempts.
  2. The share of people who improve rises with engagement. 60% of people with 3–4 sessions improve. At 50+ sessions, it is 71%. More reps line up with a higher chance that a given person is on an upward trajectory.

Heavier users also start higher. Someone who eventually completes 50+ sessions begins around 61, versus about 42 for someone who only does 3–4. People who are already stronger tend to stay longer. The improvement still happens inside each group, including for people who started below average.

Diminishing returns, and what they mean for program length

The “+1.9 points per session” headline hides a steep drop in the per-session slope.

Total sessions bucketAvg slope (points per session)
3–4+2.65
5–9+1.26
10–19+0.47
20–49+0.24
50++0.10

Most of the gain happens in the first 10–15 sessions. After about session 20, scores settle in the low-to-mid 60s and creep up from there. That is a normal skill-acquisition curve: rapid early learning, then a long grind for the marginal points.

For program design, that is a length decision. About 10–15 scored sessions covers the steep part of the curve. Two or three short sessions a week puts a cohort on that part of the curve within a month. By session 30, only 46 of the original 1,033 people are still in the set, which is a reason to put the effort into the early reps, where the scores actually move, and to make those reps easy to finish.

What an L&D team should do with this

The first 10 sessions matter more than the next 90. The largest, fastest gains in this data come from moving a cohort from “we have never recorded this” to “we have recorded it ten times.” If a leadership team is unsure whether deliberate practice belongs in the calendar, ten scored reps will answer it for that cohort.

Three design choices follow directly from the curve:

  1. Treat the first five sessions as a sprint. In the cohort that reached 20 sessions, the average moved from 49.5 to 58.8 between session 1 and session 5, about 9 points, which is where most of the early lift sits. Volume beats perfectionism in that window. Do not ask the cohort to agonise over each take.
  2. Judge the program after about 10 sessions. 64% of people with at least three scored sessions show an improving trend overall, and the early sessions are noisy. Three sessions is too few to know whether a cohort is moving. Ten usually is.
  3. After 20 sessions, change the goal. The slope in the 20–49 bucket is +0.24 points per session. Chasing the headline score will feel slow. Move people onto targeted drills for two or three specific weak spots. The companion analysis of pace, filler words, and visual delivery is the shortlist those drills should come from.

A coaching institution can run the same arc with a client cohort. A research-training program can run it as the speaking practicum. The session count, the early sprint, and the later shift from volume to targeted drills are the design.

A note on the method

Every “session” here is one row in lesson_attempts: one full scored recording at a lesson. Scores come from the same 0–100 model applied uniformly across the period. We did not use the progress table, which only stores the latest score per lesson, or lesson_completion_events, which has no score column, because neither lets you reconstruct a per-session trajectory. The 1,033-person set is everyone with at least three scored attempts, the minimum for a meaningful per-person trend line.

The confound we cannot fully separate is engagement and ability. People who complete 50+ sessions also start about 20 points higher than people who only do 3–4. “More practice, more improvement” is partly tangled with “stronger speakers tend to practise more.” The within-cohort improvement is still real at every level, including for people who started below average. That within-group lift is the number a program owner can plan around.

The decision for the program

The curve flattens well short of a perfect score. What the data supports is specific: deliberate practice produces measurable, repeated improvement for two out of three people who stay with it.

An L&D team should put that on the calendar as a finite program. Ten to fifteen scored sessions, two or three short reps a week, the first five treated as a sprint, and a review that waits until the trend is readable. Sixty-four percent of people with at least three sessions were on an improving trend. That is the base rate to plan for. The way to see it in a live cohort is to give them the reps and read the same curve.

Put a deliberate-practice curve on the cohort calendar

See how an L&D team, a coaching institution, or a research-training program runs scored speaking practice and reads the trajectory session by session.

Book a demo