Every articulation session produces data whether you write it down or not. The real question is whether that data survives the trip from the therapy table to the IEP progress report — or whether it evaporates into “did well today” and has to be reconstructed from memory three months later at a review meeting.
This guide walks the full speech therapy data collection pipeline an SLP actually needs: score a trial (correct or incorrect, at a named cue level, in a specific word position), let those trials accumulate into a session history, turn that history into an auto-generated SOAP note, and roll it forward into an IEP goal progress update — one connected flow rather than four disconnected chores. Along the way it covers the mechanics most quick guides gloss over: the accuracy formula and where averaging misleads, the two-dimensional cue-level-by-position view, defensible baselines grounded in developmental norms, and the criterion critique that questions the reflexive “80% accuracy” goal.
It is written for SLPs — school-based and clinic alike — so the jargon stays (probe, cue hierarchy, PCC, PLOP). Where a number is app-specific, it is Sound Safari’s; where it is clinical, it is cited to ASHA, Crowe & McLeod, or Shriberg & Kwiatkowski.
Why Speech Therapy Data Collection Matters
Data collection is not busywork bolted onto therapy — it is the thing that makes therapy legally and clinically defensible. Three concrete reasons.
IEP progress monitoring is a legal requirement. Under IDEA, every IEP must state how progress toward each measurable annual goal will be measured and when periodic reports go to parents. The IRIS Center at Vanderbilt lays this out plainly: a goal is only as strong as the data behind it. No data, no defensible goal.
Eligibility and placement rest on the baseline. A weak or vague present level of academic achievement and functional performance (PLAAFP) undermines everything downstream — a goal built on a fuzzy baseline cannot be shown to be appropriate, and placement decisions built on that goal inherit the weakness.
Defensibility is the whole point. Data has to survive scrutiny — a review meeting, a skeptical parent, a due-process hearing. “The student is doing better” loses that room; “the student moved from 30% to 78% accuracy on initial /k/ across the last three sessions given moderate cues” wins it.
Two distinctions worth previewing. Objective data (a tallied accuracy percentage) outranks subjective data (a clinical impression) for defensibility, though the subjective still explains the objective when numbers wobble.
And validity (are you measuring the skill the goal names?) is different from reliability (would another clinician, or you next week, score it the same way?). Reliability hinges on something concrete — identical cue-level definitions across sessions, settings, and clinicians — a thread picked up in the cue-level section below.
One disclaimer, stated once and meant: Sound Safari is a clinical tool, not a medical device. It does not provide diagnoses or treatment recommendations. Clinical decisions should always be made by qualified professionals.
Trial-by-Trial vs Probe Data: A Decision Guide
Two tracks of data answer two different questions, and confusing them is one of the most common measurement errors on an articulation caseload.
Trial-by-trial data is every response scored during teaching. It is sensitive to session-to-session change — you see movement quickly — but it is collected while you are actively cueing, so it can be inflated by the very support you are providing. A great trial-by-trial day at maximum cue is not evidence of independent skill.
Probe data is a periodic, standardized, usually uncued sample of the target skill. It is a cleaner read on true, generalized ability precisely because you take your hands off the wheel. Baselines are probes; progress checks are probes.
| Trial-by-trial | Probe (baseline / progress) | |
|---|---|---|
| What it is | Every response scored during teaching | Periodic, standardized sample of the target skill |
| When to use | Daily/session teaching and shaping | Baseline before instruction; scheduled progress + generalization checks |
| How it graphs | Dense, session-by-session line | Sparse points at set intervals |
| Effort | Higher — score every trial | Lower per event, but scheduled |
| What it measures | Responsiveness to teaching (cue-influenced) | True / generalized ability (usually uncued) |
The workflow most SLPs actually run is a hybrid: a cold, uncued probe at the start of a session (before any teaching contaminates the read), cued trials during teaching, and a short probe at the end. A common rule of thumb keeps the teaching trials in the student’s zone of proximal development — roughly 60-80% accuracy, hard enough to be work but not so hard it collapses. And “less is more”: a focused set of well-scored trials beats a giant tally you cannot trust.
Sound Safari tracks both tracks separately. Ordinary practice builds the trial-by-trial session history. Dedicated Probe sessions form the probe track, with six probe types — Initial Baseline, Re-Baseline, Progress Probe, Criterion Probe, Maintenance Probe, and Generalization (two of those, the baselines, anchor where a goal starts). Probes run on a standard 10-word stimulus set, apply an 80% success threshold on Criterion and Maintenance probes, and code errors with SODA — Substitution, Omission, Distortion, Addition — so an error pattern, not just a wrong count, comes out the other side.
Tally and Frequency Methods
Before percentages come the raw marks. The common marking systems are simple by design: check/x, plus/minus (+/−), and frequency tallies (a running count of correct productions). The practical tricks that make real-time data survive a busy group: take data on one goal at a time, mark in the moment rather than reconstructing after, and batch-calculate the percentages at the end rather than doing mental math mid-session.
Sound Safari’s per-trial scale is Correct / Incorrect, with Approximate available as an opt-in third value (the default cycle is just Correct/Incorrect). Scoring is tap-to-cycle: tapping a response advances it through the enabled values, so a single thumb keeps pace with a fast drill.
Why offer a three-point scale at all? Approximate gives partial credit for a shaping attempt — an emerging production that is not fully correct — without inflating the “correct” count and, with it, the accuracy percentage. That protects the integrity of the number.
And the reliability rule applies here with force: whatever your score definitions are, keep them identical across sessions, settings, and clinicians. An “approximate” that means one thing on Monday and another on Thursday is not data — it is noise wearing a number’s clothes.
Tracking by Cue Level and Word Position
Now the two dimensions that change what an accuracy number means.
Cue level. Sound Safari uses a four-step, hierarchy-ordered prompt scale that fades support from most to least:
- Maximum — Picture + Word + Model (the student sees the picture, gets the word, and has a clinician model to imitate)
- Moderate — Picture + Word (no model)
- Minimal — Word only
- Independent — Picture only (the student names it with no word and no model)
Word position. The app tracks accuracy in all three positions — Initial, Medial, and Final. Most articulation goals target the initial position first, then medial and final.
Here is what a single ”% correct” column cannot capture: 80% at Maximum cue and 80% Independent are not the same clinical fact, and 80% in the initial position is not the same as 80% in the final position. Cue level and word position each change the meaning of the number. Collapse them and you lose the very information that tells you whether a student is ready to advance.
The two-dimensional view is a simple grid — cue level down, word position across — and one worked cell shows how to read it:
| Cue level | Initial | Medial | Final |
|---|---|---|---|
| Maximum | — | — | — |
| Moderate | — | — | — |
| Minimal | — | /k/ 12/16 = 75% | — |
| Independent | — | — | — |
That filled cell reads: /k/, medial position, Minimal cue — 12 of 16 correct = 75%. Nothing about it means the same as 75% at Maximum cue, or 75% in the initial position.
This is usually the first thing to fall off a paper data sheet, because filling a 4×3 grid by hand for every target is more paperwork than a session allows. Sound Safari captures the prompt level and the word position on every practice trial automatically, so the grid is a byproduct of scoring rather than a second job.
Calculating Accuracy Percentages
The core formula is the one every data sheet assumes and few explain:
accuracy = (correct ÷ total) × 100
16 correct out of 20 trials → 16 ÷ 20 = 0.80 → 80%
Sound Safari stores this as an integer percentage (correct trials ÷ total trials × 100), so a goal’s entries read as clean whole numbers.
Now the pitfall that quietly corrupts progress data: averaging session percentages instead of pooling trials. Suppose a goal has three recent entries — 3/3 (100%), 0/1 (0%), and 8/10 (80%). Average the three percentages and you get (100 + 0 + 80) ÷ 3 = 60%. Pool the trials and you get (3 + 0 + 8) ÷ (3 + 1 + 10) = 11/14 = 78%.
Those are wildly different stories, and the averaged one is wrong — it lets a single 0-of-1 trial count as much as a full 10-trial session. Sound Safari derives a goal’s current accuracy by pooling the last three progress entries (sum of correct over sum of total), not by averaging the three percentages, which is why a stray low-N session cannot hijack the number.
When a whole-sample measure is better: PCC
For children with many errors, single-phoneme accuracy is not enough — you want a whole-sample measure. Percentage of Consonants Correct (PCC) applies the same correct-over-attempted logic to a connected speech sample: count every consonant the child attempts, count how many are correct, divide, and multiply by 100. Introduced by Shriberg & Kwiatkowski (1982), PCC comes with severity bands:
| PCC | Severity |
|---|---|
| > 85% | Mild |
| 65-85% | Mild-moderate |
| 50-65% | Moderate-severe |
| < 50% | Severe |
Every error type counts against the score — substitutions, omissions, distortions, and additions all reduce PCC.
When to use which: reach for whole-sample PCC when a child has multiple errors or a phonological profile; reach for single-phoneme accuracy when you have a targeted articulation goal on one sound.
Setting Defensible Baselines and Criteria
A baseline is an uncued probe taken before instruction, dated, and tied to the goal’s exact skill — the specific sound, position, and linguistic level the goal names. A missing, undated, or vague baseline is one of the most common defensibility failures, because every later progress claim is measured against it.
What number should a baseline be? That depends on the target and on developmental norms. Anchor expectations in Crowe & McLeod (2020), the U.S. consonant-acquisition review the app uses for its 90%-mastery ages (the 2018 McLeod & Crowe cross-linguistic review is the companion, not the source of the U.S. ages). A /p/ goal for a five-year-old starts from a very different place than an /r/ goal — /r/ and voiced “th” are not at 90% mastery until 6;0, and voiceless “th” not until 7;0. Cross-reference the on-site speech sound milestones by age when you set the expectation.
Sound Safari attaches suggested baseline ranges to its goal templates — a concrete starting point. Two are worth quoting as verified app defaults: the “Sound in Word Position” template suggests a 0-40% baseline, and the phonology “Fronting” and “Stopping” templates suggest 0-30%. The app attaches other ranges too — the “Cluster Reduction” and “Gliding” templates suggest 0-25% — so treat only the per-goal example baselines you see elsewhere (say, 20-40% or 30-50% in an IEP goals guide) as illustrative rather than fixed app values.
The mastery criterion is where the received wisdom deserves a hard look. Sound Safari’s default is 80% accuracy held across three consecutive sessions (many clinicians instead write the goal as 80% across three of four consecutive sessions). But Rebecca Moore, writing in the ASHA Leader (2018), makes the case that 80% is essentially arbitrary and not strongly evidence-based. The defensible move is to pick a functional, individualized criterion scaled to context: you might demand higher accuracy at isolation and tolerate lower accuracy in conversation, where the linguistic load is heavier. Pairing how you measure with what criterion is defensible is the part that’s easy to overlook.
If you do not have a baseline yet, the built-in screener is the natural on-ramp: it samples 24 sounds across initial, medial, and final word positions, giving you a dated, uncued starting point tied to real productions. The articulation screening checklist walks the workflow.
Turning Raw Data into Progress Reports and IEP Updates
Here is the step that ties it together: converting raw tallies into a defensible progress narrative.
Start with the trend. Sound Safari auto-computes a goal’s direction — Improving, Stable, or Declining (plus Insufficient when there is too little data). It needs at least three data points; the classifier splits the entries into a first half and a second half and applies a ±5-percentage-point threshold: a gain greater than +5 reads as Improving, a drop worse than −5 as Declining, and anything in between as Stable. That threshold is deliberate — it stops normal session-to-session noise from being mislabeled “progress.”
The trend then becomes a sentence. A present-level (PLOP) or progress-update line writes itself from the same data:
“Given moderate cues, [Student] produced /k/ in the initial position of words at 78% accuracy across the last three sessions — an improving trend from a 30% baseline.”
That is a defensible statement: it names the cue level, the sound, the position, the pooled accuracy, the trend, and the baseline it is measured against.
The same session data auto-generates a SOAP note. The note has four sections — Subjective, Objective, Assessment, Plan — and ships in two built-in templates (Standard SOAP and Brief SOAP). About 20 merge fields auto-fill from the session — {accuracy}, {correctWords}, {totalWords}, {promptingLevelName}, {progressStatement}, {scoreBreakdown}, and more — so the numbers you scored flow into sentences instead of being retyped. A four-point engagement rating (Excellent / Good / Fair / Poor) feeds the Subjective and Assessment sections. For the full section-by-section anatomy, see the SOAP notes guide.
State the spine plainly, because it is the whole argument of this article: a practice trial (scored with a cue level and a position) becomes session history, which becomes an auto-generated SOAP note, which becomes an IEP goal progress update — captured during practice and carried forward automatically, not rebuilt from memory at reporting time. The IEP goals for articulation guide covers the goal-writing end of that pipeline, and the clinical feature set shows the capture in action.
Common Data-Collection Pitfalls (Defensibility Checklist)
A bookmarkable checklist of the failure modes that actually sink articulation data — each with its one-line fix.
- Measuring a different skill than the goal states. Fix: score the exact sound, position, and level the goal names — nothing adjacent.
- Inconsistent cue-level definitions across settings or clinicians. Fix: write the cue definitions down and apply them identically everywhere (the reliability rule).
- Recency bias with no dated baseline. Fix: take a dated, uncued baseline before instruction and measure every later gain against it.
- Taking data only on trained items. Fix: always probe untrained items too — generalization is the point, and trained-only data hides its absence.
- Averaging small-N session percentages instead of pooling trials. Fix: pool correct over total across sessions (11/14, not the mean of 100%, 0%, and 80%).
- Calling noise “progress.” Fix: do not declare a trend on fewer than three data points, and hold a real threshold (±5 points) before naming a direction.
One compliance beat belongs on every SLP’s checklist: guard the identifiers. Sound Safari is not HIPAA compliant and no Business Associate Agreement is available, so use student initials rather than full names in any data record you keep or export. That is not a knock on the data — it is how you handle who the data is about.
Frequently Asked Questions
What is the best data collection method for speech therapy?
There is no single best method — the right one depends on the question you are asking. Use trial-by-trial data during teaching to track responsiveness to instruction, and probe data (an uncued, standardized sample) to measure true, generalized skill. Most SLPs run a hybrid: a cold probe at session start, cued trials during teaching, and a short probe at the end. Whatever you choose, keep your score definitions consistent so the data stays reliable across sessions and clinicians.
How do you calculate accuracy percentage in speech therapy?
Divide the number of correct productions by the total number of trials and multiply by 100. Sixteen correct out of twenty trials is 16 ÷ 20 × 100 = 80%. When you combine several sessions, pool the trials (sum of correct over sum of total) rather than averaging the session percentages — averaging lets a tiny 0-of-1 session distort the result.
What is the difference between trial-by-trial data and probe data?
Trial-by-trial data is every response scored while you are actively teaching and cueing; it is sensitive to change but can be inflated by the support you are giving. Probe data is a periodic, usually uncued sample of the target skill, so it is a cleaner read on independent, generalized ability. Baselines and progress checks are probes; day-to-day teaching data is trial-by-trial.
Is 80% accuracy an evidence-based goal criterion?
Not especially. Rebecca Moore, writing in the ASHA Leader (2018), argues that the reflexive 80% accuracy criterion is largely arbitrary and not strongly evidence-based. The better practice is to set a functional, individualized criterion scaled to context — often higher at isolation and lower in conversation, where the linguistic demand is greater.
What is Percentage of Consonants Correct (PCC) and how is it calculated?
PCC is a whole-sample measure: count the consonants a child produces correctly in a connected speech sample, divide by the consonants attempted, and multiply by 100. Introduced by Shriberg & Kwiatkowski (1982), it uses severity bands — mild above 85%, mild-moderate 65-85%, moderate-severe 50-65%, and severe below 50% — with substitutions, omissions, distortions, and additions all counting as errors. Use PCC for children with multiple errors or a phonological profile; use single-phoneme accuracy for a targeted articulation goal.
How do you turn session data into an IEP progress update?
Pool the accuracy across recent sessions, compare it to the goal’s dated baseline, and classify the trend (Improving, Stable, or Declining) using a threshold so noise is not mistaken for progress. Then write it as one sentence naming the cue level, sound, position, accuracy, and trend. In Sound Safari, auto-captured session data flows into an auto-generated SOAP note and an IEP goal progress update, so the reporting sentence is built from the numbers you already scored.
How often should SLPs collect baseline and progress data?
Take a baseline once, before instruction begins on a goal, and date it. Collect progress data frequently enough to show a trend — because a defensible trend needs at least three data points, space your probes so you accumulate three or more before a reporting period. Trial-by-trial data can be taken every session; formal probes are usually scheduled at intervals.
What should be included on a speech therapy data collection sheet?
At minimum: the target sound, the word position (initial, medial, or final), the linguistic level (isolation through conversation), the cue or prompting level, and the correct-over-total tally that yields an accuracy percentage. Add the date and a baseline reference so progress can be measured against a fixed point. Capturing cue level and position — not just a bare percentage — is what lets the same data support both a SOAP note and an IEP update.