← Back to Blog Research

What 18,219 Speaking Sessions Mean for a Cohort Practice Program

· 7 min read
William Burden
William Burden Founder @ Elqo

After enough speaking practice has been recorded, a program owner can stop guessing what a cohort will struggle with. The recordings answer it.

We analysed recordings completed on Elqo from June 2025 through April 2026. 18,219 recordings. 3,221 people. Anonymised, aggregated, and read against the same objective signals, and against the coaching log of recurring challenges per person.

Three bottlenecks stood out. They differ from the goals people name when they join a speaking program. Here is the evidence, and what a cohort practice program should do with it.

Three habits dominate

A curriculum can spend its time on nerves, freezing, accent, or vocabulary. Across 18,219 sessions and 3,221 people, the patterns that actually dominate are simpler and more stubborn:

  1. Pacing — almost half of all sessions land outside the natural speaking band, and “too slow” outweighs “too fast” nearly four to one.
  2. Filler words — the cleanest signal in the set, confirmed both by per-session counts and by the coaching log of recurring challenges.
  3. Visual delivery — gestures, eye contact, and posture cluster as a tight band of recurring challenges, even though most people never list them as the reason they are in the program.

Pronunciation, accuracy, vocabulary, and fluency sit well behind these three. In practice, “speaking better” almost always means working on one of these three first. That is the scope a cohort program should declare up front.

Finding 1: Pacing is the most under-recognised issue

Pacing is tracked on every session as words per minute, and it has been measured since launch, so coverage is deep: 10,805 recordings with a valid WPM reading. The distribution:

Pacing bandSessionsAverage WPM
Ideal (100–160 WPM)5,635124.5
Too slow (<100 WPM)3,96072.8
Too fast (>160 WPM)1,052305.9

That is 47% of sessions outside the natural conversational band. Too-slow sessions outnumber too-fast ones by nearly 4:1. The average too-slow session sits at 72.8 WPM, below the 100 WPM floor where listeners start to drift.

The sharper finding is the gap between how often pacing is off and how often it is flagged as a recurring, person-level issue. Only 13.5% of people had pacing logged as a recurring challenge, while 47% of sessions were actually off-target.

Pacing is the most prevalent, most under-recognised issue in the set. Most people do not hear it in themselves. A program that waits for participants to name pace as their goal will miss the bottleneck that shows up most often in the recordings.

Finding 2: Filler words are the cleanest signal

Pacing is the issue people do not know they have. Filler words are the issue everyone recognises. Two different views of the data agree:

  • Per-session view: 34% of sessions contain 3 or more filler words. The overall average is 2.49 fillers per session.
  • Per-person view, from the coaching log: 38.6% of people with any tagged challenges have “filler words” listed as a recurring issue, the most common real challenge in that log.

Two independent measurements within about 4 percentage points of each other. Filler words are the most defensible “most common speaking issue” claim in this data. They sit in a useful range for a program: frequent enough that listeners hear them, and specific enough to drill.

Distribution of fillers per session:

  • 0 fillers: 3,318 sessions
  • 1–2 fillers: 3,790 sessions
  • 3–5 fillers: 2,599 sessions
  • 6–10 fillers: 973 sessions
  • 11+ fillers: 125 sessions

Most sessions are not extremely filler-heavy. About a third cross the line where listeners start to tag the speaker as hesitant or unprepared, often without saying so.

Finding 3: Visual delivery is larger than intake forms suggest

The third pattern is how often the coaching log tags visual challenges — gestures, eye contact, posture — as recurring issues. Almost nobody joins a speaking program to “fix body language.”

Among people with any tagged recurring challenges:

Recurring challenge% of tagged users
Filler words38.6%
Visual presence36.1%
Gestures26.4%
Eye contact25.0%
Posture22.1%
Speaking pace13.5%

Three of the top six recurring challenges are visual. Stacked together, visual delivery is as common a bottleneck as anything verbal. It almost never appears in the goal people write at intake. They arrive wanting to sound more confident. The review often shows that their hands have not moved.

Visual signals carry disproportionate weight in how listeners judge competence. Neutral words with strong visual presence read as confidence. The same words with weak visual delivery read as nerves. A cohort review that includes the video will see the cluster the coaching log already ranks next to filler words.

Train the three bottlenecks as a cohort

Pace, filler words, and visual delivery can be reviewed on every rep. Elqo gives a program the session view so a cohort can practise them together.

Book a demo

What the cohort program should make visible

The useful response to 18,219 sessions is a curriculum aimed at the three bottlenecks, with the signals on the table after every rep. Three standards are enough:

1. A visible WPM band, and a pacing flag when the cohort trends slow

Pacing is the most prevalent issue in the set and the most under-flagged in the recurring-challenge log: 13.5% of people tagged, 47% of sessions off-target. The program should show the 100–160 WPM band on the session, and call pacing out when takes trend slow. If “too slow” is the pattern, the cohort should see it after a session, and the program lead should see it across the group.

2. Filler words as a headline metric

Both views of the data agree this is the cleanest, most widespread speaking issue. Treat it that way in the review: count, density, and the specific words a person over-uses. Point the drills at the techniques in How to Reduce Filler Words, and make that track part of the cohort, especially once the early volume of reps is done.

3. A video pass for eye contact, gestures, and posture

Visual challenges show up throughout the person-level coaching log. Record the cohort on video, and review eye contact, gestures, and posture in the same conversation as pace and filler words. People will not put body language on the intake form. The program has to put it on the score anyway, because the log says it is already one of the recurring issues.

What to run with the cohort

These are assignments for a cohort, drawn from the patterns in the 18,219 sessions. Same prompt for everyone. Same review. The program lead looks at the group, and each person looks at their own take.

  1. Ninety seconds on a real piece of work. The default prompt is “tell me about a project you're proud of,” or a prompt taken from the company's own material. Play it back at normal speed. If the speaker can finish the sentence in their head before the recording does, the take is in the too-slow band. The average too-slow session in this set is 72.8 WPM, against an ideal band of 100–160. Run the same prompt again, deliberately faster, and compare the two takes. Too-slow sessions outnumber too-fast ones nearly four to one, so this is the pacing drill to run first.
  2. Replace one filler with a pause. One word, for the whole cohort, for the week. In this data, “um” and “like” dominate. The instruction is to close the mouth instead of saying the word. The first few pauses feel long to the speaker. To listeners they are ordinary. That swap is how the habit fades, and 34% of sessions are at 3 or more fillers, which is the threshold worth posting on the cohort review.
  3. Record on video. Video is what puts the visual cluster on the review: visual presence at 36.1% of tagged people, gestures at 26.4%, eye contact at 25.0%, posture at 22.1%. Watch the body with the cohort, so visual delivery is a shared standard alongside pace and fillers.

Run the three as a program, across repeated reps. Practising with feedback is what closes the gap between knowing the standard and hitting it on the next take. Pair the drills with a finite session count: the 12,660-session practice curve puts the steep gains in the first 10–15 reps, so these bottlenecks belong in that early block, while the cohort is still together.

A coaching institution can run this as the spine of a speaking engagement. A research office or graduate school can run it against talks, defences, and lab meetings. An L&D team can run it against the conversations managers and specialists actually have. The prompts change. The three bottlenecks do not.

Where a cohort should start

The common speaking issues are the stubborn ones. Across these recordings, the picture is consistent: pace, fillers, and visual delivery. Energy, fluency, and accuracy show up in narrower, more situational ways. These three appear nearly everywhere.

If a cohort can train one thing first, train pacing. It is the most prevalent miss, and the easiest to ignore until a WPM number is on the review. If there is room for two, add filler words. If there is room for three, record on video and watch what the body is doing.

That is what 18,219 sessions say a cohort practice program should train, and the order it should train them.

Build the cohort practice program

Use the pace, filler-word, and visual-delivery patterns from 18,219 sessions to decide what your next cohort practises, and score the reps against those signals.

Book a demo