Baseline data is the least glamorous number in an articulation IEP — and the one everything else leans on. Collect it cleanly and the rest of the year has a defensible reference point. Collect it sloppily, or skip it, and every progress claim you make is arguable at the review table.
This guide follows one running case — a student producing /s/ in the initial position of words — from a clean probe, through the accuracy calculation, to a norm-checked target, into a finished measurable goal, and out the other end as progress-monitoring data. Along the way it draws the distinctions most baseline guides skip: probe data versus first-session data, per-position versus global accuracy, and how to tell when a baseline has gone stale.
Every template, default, and suggested range in this guide is a starting point — not a substitute for clinical judgment, your district’s IEP conventions, or the student’s own data.
A clean baseline is the single number that makes an articulation IEP goal defensible — it sets the floor the entire year of progress is measured against.
What a Baseline Is — and Why It Anchors the Goal
A baseline is performance measured before intervention, under probe conditions — no teaching, no cueing, no feedback. The idea traces to single-subject and multiple-baseline research design, where repeated pre-treatment observations establish a stable reference before the treatment phase begins. In speech therapy, that same logic underpins data-based decision-making — using a documented starting point and repeated measures to guide treatment choices (ASHA, American Journal of Speech-Language Pathology). Strip out the teaching and what you are left with is the student’s unassisted floor.
That floor is not optional paperwork. Under IDEA — whose measurable-goals requirement was added in the 1997 reauthorization — an IEP goal is built from the Present Levels of Academic Achievement and Functional Performance (PLAAFP), and the PLAAFP is precisely where the baseline lives.
The IRIS Center at Vanderbilt frames a measurable goal as four elements: the target behavior, the conditions, the criterion for acceptable performance, and the timeframe. Each of those elements only means something relative to a starting number.
A goal that reads “80% accuracy” is unmeasurable if you never recorded where the student began. In practice, a defensible articulation goal cannot exist without a baseline.
The baseline does more than justify the goal — it becomes the math underneath it. When you later plot progress, the baseline is the bottom of the ruler: progress is computed as (current accuracy − baseline) ÷ (target − baseline).
Move the baseline and you move the denominator that every progress percentage is divided by. The full worked calculation comes later in this guide; for now, just hold the idea that the baseline anchors the progress bar directly.
None of this changes where the goal itself gets written — that is a separate craft, and our articulation IEP goals guide covers the goal language sound by sound. The baseline is the number you drop into it.
Probe Data vs. First-Session Data: What Counts as a Clean Baseline
Here is the distinction most baseline guides blur, and it is the one that decides whether your number holds up. A probe baseline is collected with no feedback, cueing, or teaching. First-session data is usually contaminated — warm-up, a couple of models, a prompt or two — because a first session is often part evaluation, part therapy. If you scored the student after you started shaping the sound, you measured a moving target, not a floor.
The nuance worth stating carefully: you can establish a baseline on the very first session, as long as you run a no-teaching probe before any instruction. The rule is method, not calendar day. Probe first, teach second, and the first session yields a clean baseline. Teach first and the numbers describe initial performance, not a baseline.
Sound Safari enforces this distinction in its wording. When it drafts a SOAP note from a treatment session, it labels that session’s numbers “initial performance,” not “baseline,” because treatment data has prompts baked in — the “treatment is not testing” principle. It also types probes by purpose: of the six probe purposes, exactly two establish a baseline — Initial Baseline (“first assessment before intervention begins”) and Re-Baseline (“new baseline after a break or goal revision”). The other four — Progress Probe, Criterion Probe, Maintenance Probe, and Generalization — measure movement away from the baseline, not the baseline itself.
| Probe baseline | First-session (treatment) data | |
|---|---|---|
| Teaching | None before scoring | Instruction interleaved with scoring |
| Cueing | None (or a fixed, recorded level) | Warm-up prompts and models baked in |
| Feedback | None during the probe | Corrective feedback throughout |
| Valid for | The measurable “before” floor | Clinical impression, session planning |
| Labeled | ”Baseline" | "Initial performance,” not “baseline” |
Bookmark that table. The single most common way a baseline gets challenged at a review is a parent advocate noticing the “before” number was collected with help.
How to Collect a Clean Baseline, Step by Step
A repeatable method matters more than any single number. Here is the sequence.
-
Pick the sound and position(s). Baseline each word position separately — initial, medial, and final — because accuracy routinely differs by position and a single global number hides that. This matches Sound Safari’s data model and its built-in screener, which covers 24 sounds across word-initial, medial, and final positions. If you are not sure which positions are affected, screen first, then baseline the ones that break down.
-
Present the same stimuli you will reuse for progress monitoring. Whatever pictures, words, or sentence frames you use for the baseline become the stimuli you re-present at every progress check. Change the stimuli mid-year and the delta is uninterpretable — you are no longer comparing like to like (Old School Speech on clean, apples-to-apples data). Consistency is the whole game.
-
Score under probe conditions, and record the cue level. No teaching, no feedback. Then write down the prompting level you used — map it to the four levels Sound Safari uses: Maximum, Moderate, Minimal, or Independent. A baseline of 30% at the independent level is a completely different floor than 30% with a verbal-plus-visual cue, and a goal written off the wrong one will look either impossible or already met. State the cue level explicitly, every time.
-
Capture the error type. For each miss, note the error using SODA — Substitution, Omission, Distortion, or Addition (distortions carry subtypes: lateralized, dentalized, palatalized, nasalized, weak, or other). The baseline percentage tells you how much; the error pattern tells you what to target first.
-
Keep the trial count consistent — not fixed. No authority mandates a trial count. The ten-trial convention is popular only because it converts cleanly to a percentage, and the frequently quoted “100 trials per session” figure is explicitly something you do not need to fully record. Treat trial counts and monitoring cadence as common practice, not a rule. Pick a count you can reuse every time, and reuse it.
If you want a ready-made stimulus set to baseline from, the articulation screening checklist doubles as a position-by-position probe source, and our data-collection guide covers monitoring cadence once the baseline is set.
Calculating Baseline Accuracy %
The formula is simple and worth stating exactly:
Baseline accuracy % = (correct productions ÷ total opportunities) × 100. Example: 3 correct out of 10 independent probe trials = 3/10 = 30%.
That is the same calculation Sound Safari runs under the hood (correct trials × 100 ÷ total trials), and it is the convention SLP resources converge on precisely because a ten-item probe reads straight off as a percentage.
Here is the running case this guide carries the rest of the way. On a given date, a 6-year-old student produces /s/ in the initial position of words on 3 of 10 independent probe trials — a 30% baseline, collected with no cueing or feedback. That single number, tied to that date and that cue level, is now the floor.
Report accuracy per position and, when it differs, per cue level. The same student might sit at 30% independent but 70% with a verbal cue — two different floors that imply two different goals and two different fade plans. This is exactly why the cue level from Step 3 is not optional: without it, “30%” is ambiguous.
In Sound Safari, this percentage is computed for you as you score each trial, so the baseline number is calculated rather than tallied by hand — and it is stored with the date and cue level attached. Sound Safari is a clinical tool, not a medical device; it does not provide diagnoses or treatment recommendations, and the number it computes is only as clean as the probe conditions you set.
Setting the Target From the Baseline
With a floor in hand, you set a ceiling. Sound Safari’s built-in IEP templates default to an 80% accuracy target and goals default to 3 consecutive sessions as the mastery criterion — the familiar “[accuracy]% across [sessions] consecutive sessions” pattern. Those are sensible defaults; they are also just defaults.
Each template also suggests a starting baseline range, and these are examples rather than a fixed menu. For instance, the flagship “Sound in Word Position” articulation template suggests a 0–40% starting range, while phonology-process templates such as “Fronting” and “Stopping” suggest 0–30%. The app carries other suggested ranges for other templates too — treat these as illustrative anchors, not the complete set.
The step most goal banks skip: reality-check the target against the sound’s age of acquisition. A target and timeframe are only “achievable” if the sound is developmentally due. Using U.S. 90% mastery ages from Crowe & McLeod (2020) — the U.S. companion review to the broader McLeod & Crowe (2018) 27-language cross-linguistic scoping review — the later sounds land at /s/ 5;0, /l/ 5;0, /r/ 6;0, voiced th 6;0, and voiceless th 7;0 (the last English consonant to stabilize). So a 5-year-old scoring 20% on /r/ is age-appropriate-emerging — the 90% age is not until 6;0 — not necessarily delayed, and the goal should reflect a realistic runway rather than a steep one-year climb. Our speech sound milestones by age lays out the full norm table. (Do not attach the older Sander- or Smit-era ages, which place /r/ and /s/ years later, to Crowe & McLeod — those are a different, outdated source.)
Back to the running case: our 6-year-old is past the /s/ 90% mastery age of 5;0, so targeting /s/ is appropriate. We take the 30% baseline to an 80% target across 3 consecutive sessions, using the “Sound in Word Position” template, which carries an estimated 12–36 week window — a realistic default range you narrow to the student.
Because Sound Safari’s built-in goal bank pre-fills the target, criterion, and timeframe, you spend your judgment on the number rather than rebuilding the scaffold each time. The scaffold is a default; the clinical call is yours.
Writing the Baseline Into a Defensible Goal
Now assemble the running case into one finished, measurable goal that carries all four IDEA elements plus an explicit baseline statement. Using the app’s actual fill-in-the-blank structure:
Goal: [Student] will produce the /s/ sound in the initial position of words with 80% accuracy across 3 consecutive sessions.
Baseline statement: Given 10 independent probe trials on [date], [Student] produced /s/ in the initial position of words with 30% accuracy (3/10), with no cueing or feedback.
That baseline statement is the part goal banks leave out. Model it every time: conditions (10 independent probe trials), cue level (independent, no feedback), the fraction (3/10), and the date. Use student initials rather than full names in anything you export.
What makes this durable is where the number lives. Sound Safari stores the baseline as first-class, linked data — the baseline percentage, the date it was collected, the notes, and a link to the exact probe that established it — not as free text buried in a comment field.
Copying a probe’s result into a goal carries that percentage straight through, and the goal stays tied to the probe that set it. That traceability is what lets you answer “where did 30% come from?” at a review with a specific probe on a specific date, rather than a shrug.
For context on how much scaffolding sits behind that, the IEP goal bank spans 8 categories and ships a full library of built-in templates covering articulation, phonology, and beyond — but the baseline mechanics above are identical whichever template you start from. The template gives you structure; the baseline gives it meaning. For finished goal language across every commonly targeted sound, see the articulation IEP goals guide.
Common Baseline Mistakes (and How to Avoid Them)
Most weak baselines fail in one of a handful of predictable ways. Use this as a pre-flight checklist.
- Collecting the baseline with cues or feedback. This inflates the floor and shrinks the measurable progress window — you leave yourself less room to show growth, and the number is challengeable. Probe conditions, always.
- Changing stimuli or method between baseline and progress monitoring. The delta becomes uninterpretable. Whatever you used to baseline, reuse to monitor.
- Believing a “valid” baseline needs a mandated trial count. It does not. Ten trials is convention, not law; you are not obligated to record 100 trials. Consistency beats volume.
- Recording one global accuracy % instead of per-position baselines. A single number masks a position-specific need — a student can be fine in the initial position and break down in the final. Baseline initial, medial, and final separately.
- Not checking the baseline against developmental norms. Targeting a still-emerging sound as if it were a delay sets an unrealistic timeframe. Tie the target back to Crowe & McLeod (2020) before you commit.
- Trusting a stale baseline. After summer regression or a mid-year transfer, the old floor may no longer describe the student. That is a re-baseline situation, covered next.
- Skipping the baseline because “the goal is obvious.” Collecting a baseline before the IEP meeting sometimes reveals the goal is already met or mis-targeted. The quick decision cue: if you would be surprised by the number, you needed the number.
Sound Safari can score and store these baselines position by position as you go, which makes the per-position and cue-level habits above the path of least resistance rather than extra bookkeeping.
How the Baseline Travels Through the IEP Year
Here is the pipeline Sound Safari connects end to end: baseline → SOAP documentation → progress trend → the clinician’s continue/adjust/exit call. The baseline is the fixed point every downstream number is measured against.
The progress math. Progress against the goal is computed as (current accuracy − baseline) ÷ (target − baseline) × 100, clamped between 0 and 100%. With the running case — baseline 30, target 80 — a probe at 55% lands at (55 − 30) ÷ (80 − 30) = 25 ÷ 50 = 50% of the way to mastery. The baseline literally is the bottom of that fraction; get it wrong and every progress percentage for the year is wrong with it.
The SOAP note. When Sound Safari drafts a session note, it pulls the baseline into the Assessment using a defined order of preference: first the active IEP goal’s recorded baseline for that sound and position; if there isn’t one, the most recent probe; and if there’s no probe either, the first prior scored session — which the note labels “initial performance,” not “baseline,” because treatment data has prompts baked in. The note then states the change in percentage points from that reference to the current session. It is deliberately not a rolling average — it anchors to a real starting point so the comparison stays honest. How that note gets written out is covered in the SOAP notes guide.
The trend. Across probes, the trajectory from baseline to current is classified into five states: Improving, Stable, Declining, Variable, and Insufficient Data. A simple first-to-last read calls it Improving above +5 points and Declining below −5; a more careful half-versus-half analysis uses a ±10-point band and flags Variable when the spread is wide (a standard deviation over 25).
The app also surfaces a higher-confidence “meaningful progress” flag — a practical heuristic, not an inferential significance test — when all three hold: at least a 15-point gain, at least 3 probes, and a standard deviation under 20. Mastery is 80% sustained across 3 consecutive probes at or above criterion.
Those states and gates inform the clinician’s continue/adjust/exit call — they don’t make it. Sound Safari surfaces the math; clinical decisions should always be made by qualified professionals.
When to re-baseline. After a break — summer regression is the classic case — or a goal revision, run a Re-Baseline probe rather than trusting the old number. A stale baseline quietly distorts every progress percentage that follows it. Pair this with a sensible monitoring cadence from the data-collection guide.
Because session data flows into the goal and the drafted SOAP note automatically, the probe data behind all of this stays current without re-entry — the baseline set in September is still the anchor in May.
Frequently Asked Questions
What is baseline data in articulation therapy?
Baseline data is a student’s accuracy on a target sound measured before intervention, under probe conditions — no teaching, cueing, or feedback. It mirrors single-subject research design, where repeated pre-treatment observations establish a stable reference. In an IEP, the baseline is the “before” number the whole year’s progress is measured against.
How do you calculate baseline percentage in speech therapy?
Divide correct productions by total opportunities and multiply by 100. Three correct out of ten independent probe trials is 3/10, or 30% accuracy. Report the percentage per word position and note the cue level, because 30% independent is a different floor than 30% with a verbal-plus-visual cue.
What is the difference between baseline data and probe data?
A probe is any structured, no-feedback sample of a target behavior. A baseline is the specific probe (or probes) collected before intervention begins — the reference point. All baselines are probes; not all probes are baselines. Later progress and criterion probes measure movement away from that baseline.
Can you collect baseline data during the first therapy session?
Yes, if you run a no-teaching probe before any instruction. A clean baseline is defined by method, not the calendar. What contaminates first-session data is warm-up, modeling, and prompts baked into treatment — so once you have taught, that session’s numbers reflect initial performance, not a baseline.
How many trials do you need to take a baseline?
No authority mandates a trial count. The ten-trial convention is popular only because it converts cleanly to a percentage. You do not need to record 100 trials per session. Treat trial counts and monitoring cadence as common practice, not a rule, and keep them consistent between baseline and progress checks.
How do you set a target accuracy from baseline data?
Start from a common default like 80% accuracy across three consecutive sessions, then reality-check it against the sound’s age of acquisition. Per Crowe & McLeod (2020), /s/ reaches 90% mastery at 5;0 and /r/ at 6;0 — a target on a still-emerging sound needs a longer timeframe, not a lower bar.
Does every IEP goal need baseline data?
Yes. Under IDEA, a measurable goal is built from the Present Levels statement, which defines the “before” state progress is measured against. Without a baseline number there is nothing to measure movement from, and the goal is not defensible at a review. Collecting it first sometimes reveals the goal is already met or mis-targeted.
How is baseline used to monitor progress over the IEP year?
Progress is computed as (current accuracy − baseline) ÷ (target − baseline) × 100. A student at 55% with a 30% baseline and an 80% target is 50% of the way to mastery. The baseline also anchors each SOAP note’s comparison statement and the improving/stable/declining trend that informs the clinician’s continue, adjust, or exit decision.
Closing
A clean baseline is cheap to collect and expensive to fake later — it is the number the whole IEP year quietly leans on. The pipeline is short and worth memorizing: probe baseline → accuracy % → norm-checked target → defensible goal → auto-drafted SOAP note and trend → continue, adjust, or exit.
Get the floor right once, under real probe conditions, and every downstream number inherits that credibility. Sound Safari’s IEP goal bank and auto-generated SOAP notes keep that baseline linked to the exact probe that set it, so the number you defend in May is the number you collected in September. The templates and defaults are starting points — the clinical judgment about what to target, and how fast, stays yours.