The distinction: a software feature is not a learning result
“Adaptive” usually describes a decision made by software. The app records something a child does, applies a rule, and changes a later item, hint, level, sequence or pace. That can be a precise description of the product. It is not, by itself, a finding about learning.
A learning-outcome claim is a different kind of statement. It says that children who used the app acquired, retained or could apply knowledge or a skill better than they otherwise would have done. To support that statement, someone must define the outcome, measure it, and compare app use with a meaningful alternative. The alternative matters: an app can look effective against doing nothing, while showing no advantage over an adult reading with a child, ordinary classroom teaching, or a non-adaptive version of the same material.
The practical error is to let a plausible mechanism become a conclusion. If an app gives easier items after errors and harder items after success, it may reduce repeated failure or prevent a child from spending time on items already answered correctly. Those are possible routes to a benefit. They do not establish that the route worked, for which children, or on which outcome.
A small child’s attention is organised for coverage rather than completion. When adults expect persistence, an adaptive sequence that keeps changing may appear engaging without showing that the child has learned or remembered what changed.
Read the wording closely. “Adjusts to your child”, “personalised pathway” and “responds to performance” are usually feature claims. “Improves reading”, “accelerates progress” or “builds lasting number knowledge” are outcome claims. The second category requires substantially stronger evidence than the first.
What an adaptive feature can actually mean
Adaptive systems vary greatly. A simple version may move a child up after a fixed number of correct answers and down after errors. Another may select questions from tagged sets, alter the spacing of review, provide extra prompts, or infer a level from several recent responses. Some systems adapt only the order of content. Others change the task itself. A broad label conceals these differences.
Before asking whether adaptation improves learning, identify what the software observes. Is it merely right and wrong answers? Does it record response time? Can a child tap repeatedly until an answer is found? Does an adult choose the starting point? Each detail affects what the system can infer. A correct answer may reflect knowledge, a lucky guess, reading assistance, remembered screen position, or a cue in the item. Software cannot automatically distinguish these possibilities.
Then identify what changes. Adaptation may concern difficulty, repetition, feedback, pacing or presentation. These are not interchangeable. More repetition after an error is not necessarily a difficulty adjustment. A faster sequence is not necessarily a better-matched sequence. A badge awarded after a run of responses may be motivational design rather than adaptation.
Pibblewick includes Learning Lane, a zone of adaptive flashcards.
That sentence identifies a feature, not its educational effect. A reader can ask what signals choose each flashcard, whether the system revisits errors, and whether the next card differs in content or only in order. Those questions make the claim inspectable without assuming that personalisation produces a learning gain.
What evidence of improved learning must show
Evidence for an outcome starts with a defined outcome. “Learning” is too broad to test until it becomes something observable: recognising taught letter-sound correspondences, recalling number facts after a delay, understanding a word in a new sentence, or solving an unfamiliar problem. A score inside the app is useful for operating the app, but it is often a weak test of independent learning because children may become familiar with its formats, feedback and answer locations.
Stronger studies use an assessment that is not simply the app’s next screen. They measure children before and after use, and preferably assess retention later. They also ask whether learning transfers. Transfer means using the knowledge with unfamiliar examples, materials or wording. An app that teaches the correct response to its own prompts may still leave a child unable to use the idea elsewhere.
A comparison group is central. The question is not only whether children’s scores rose. Children change over time, receive teaching elsewhere and may improve from taking a similar test twice. A useful comparison could receive ordinary teaching, a similar amount of non-adaptive digital practice, or the same content in a different format. The right comparison depends on the claim being made.
Allocation matters too. If families or teachers choose who uses an app, users may differ beforehand in access, confidence, prior attainment or adult support. Random allocation can reduce some of those differences, though it does not cure every problem. Studies also need enough participants and clear reporting of missing data. A result from a narrowly selected group should not quietly become a claim about all children.
A screenshot rule for reading claims and studies
Use the table as two separate tests. The first asks whether the statement accurately describes the app. The second asks whether research supports a learning outcome. Passing the first test does not move a claim automatically into the second column.
| Question to ask | If the answer is clear | What it supports | What it does not support |
|---|---|---|---|
| What child response does the system use? | The rule names correct answers, errors, timing or another defined signal. | A stated adaptive mechanism. | That the signal accurately represents understanding. |
| What changes after that response? | The content, level, spacing, hint or sequence is specified. | A checkable description of the feature. | That the change is educationally beneficial. |
| What exact outcome was measured? | The skill, test and timing are named. | A bounded outcome claim. | A general claim that the app improves learning. |
| What was the comparison? | Children receiving the app are compared with a defined alternative. | Evidence of difference from that alternative. | Superiority to every other approach. |
| Was the assessment independent of the app? | The measure uses new items or another format. | Evidence beyond simple familiarity with the app. | Proof of broad transfer unless transfer was directly measured. |
| Who took part and for how long? | Age, setting, support and duration are reported. | A limit on where the finding may apply. | That it applies to children outside that group. |
One further check concerns the role of adults. If a study includes regular adult conversation, prompts or feedback alongside app use, its result concerns that whole arrangement. It cannot establish that software adaptation alone caused the outcome. Similarly, a study of a supervised setting does not automatically answer what happens when a child uses the app independently at home.
How to interpret the common forms of evidence
Product descriptions are useful for identifying intended design, but they are not independent outcome evidence. Demonstrations can show that a system changes material after an answer. They cannot show whether children retain the material. Usage figures can show returns, sessions or completion of activities. They may describe engagement, although even engagement needs careful definition. They do not measure learning unless linked to a suitable assessment.
Testimonials describe individual experience. They can suggest questions for research, such as whether children find a feedback style confusing or whether adults can see what has been practised. They cannot separate the app’s role from family routines, prior interest, adult help or normal development. Before-and-after scores have a similar limitation when there is no comparison group.
A study can be useful while remaining narrow. For example, a controlled study may support a claim about performance on a particular early skill after a stated period of supported use. It may not tell a reader about motivation over a year, effects on wider language, outcomes for children with different needs, or the consequences of replacing other activities with screen time. The conclusion should remain the size of the study’s question.
Look also for selective reporting. If many outcomes were measured but only the most favourable result is presented, the reported finding may give an incomplete picture. If researchers, teachers or families know which group is using the app, expectations can affect how some outcomes are recorded. These are not reasons to dismiss every study. They are reasons to ask what the design can and cannot distinguish.
Limits: what this framework does not decide
This framework does not decide whether any particular app suits a particular child, household, class or support plan. It does not assess interfaces, safeguarding arrangements, accessibility, data practices, content quality, or the balance of a child’s day. Those are separate questions requiring their own information. It also does not claim that a non-adaptive activity cannot support learning. A book, conversation, game, teacher explanation or repeated hands-on task may use feedback and adjustment without software.
It is not a rule that adaptation is ineffective. Some adaptive choices may prove useful for specific outcomes under specific conditions. The point is narrower: the presence of an adaptive feature cannot settle that question. Evidence has to test the feature, or the intervention containing it, against a relevant alternative.
The framework applies least well when a claim is undefined. Words such as “confidence”, “potential”, “readiness” and “love of learning” can refer to important experiences, but they need a stated meaning and a way of being assessed before they become research outcomes. It also cannot turn a short study into evidence of long-term effects.
For children with additional needs, bilingual experience, sensory differences or substantial variation in prior knowledge, a general study may be especially limited if it did not include or report relevant participants. The appropriate conclusion is not that an effect will be the same or different. It is that the available evidence may not answer the question for that child.