What Happens When AI Learns From AI?
Model collapse can occur in recursive training, but synthetic data is not automatically harmful. Three research papers explain the crucial differences.

Global / Technology
The internet is filling with generated text and images. Whether that damages future models depends on what gets collected, retained, checked and learned.
AI model collapse is a failure in which repeated training on model-generated data causes later models to lose important features of the original data distribution. Research has demonstrated this under particular training conditions. It does not establish that every use of synthetic data damages a model or that all AI systems inevitably deteriorate.
The distinction matters because “AI learns from AI” describes many different processes. Replacing an original dataset with unchecked model output is one. Adding carefully selected synthetic examples while retaining original data is another. Using generated exercises to teach a narrow skill is different again.
The material may share an origin while the training process produces very different results.
Why the idea is so unsettling
A generated article can become a web page. A generated image can become part of a public collection. If future datasets include that material, a model may learn from outputs produced by earlier models.
That possibility creates an intuitive worry: what happens when a system increasingly encounters versions of the world that other systems have already approximated?
The concern is more precise than the familiar complaint that some online content is poor. A dataset can lose variety even when its examples look individually plausible. If the uncommon cases disappear, the next model may have less evidence from which to represent them.
For a publication, the analogy is a newsroom that gradually replaces direct reporting with summaries of previous summaries. The problem is not that every summary is useless. It is that the chain can become detached from the original observations and the details those observations contained.
That analogy helps explain the concern, but it is not a mathematical proof. The research has to establish which processes fail and under what conditions.
What the original model-collapse research found
Ilia Shumailov and colleagues investigated recursive training in research published in Nature in 2024 as AI models collapse when trained on recursively generated data. The associated research preprint examines how successive generations can lose information, including less common parts of the original distribution, when learning from generated data.
The result concerns a training process across generations. It should not be read as a diagnosis that a chatbot has “collapsed” whenever it produces a bad answer. Nor does it mean the mere presence of one generated paragraph in a dataset guarantees failure.
A useful way to report the study is to name the mechanism: repeated reliance on generated samples can erode representation of the original data under the conditions studied. The conditions are part of the finding, not an optional qualification to be removed from the headline.
That leaves a substantial concern while avoiding an unsupported prediction about every future AI system.
A small example makes the loss easier to see
Imagine a fictional collection of 100 travel accounts. Ninety describe common journeys and ten describe unusual ones. A summarising system produces ten new examples that happen to cover only the common journeys.
If the original collection is discarded and the next system receives only those ten examples, it has no direct evidence of the unusual journeys. It may become very fluent about the common routes while representing the range of experiences less well.
This is an illustrative story, not the experimental setup or a quantitative result from the paper. The numbers are chosen to make one idea visible: an apparently sensible sample can omit rare information.
Now imagine keeping the original 100 accounts and adding the ten generated examples. The information available to the next system is different. The unusual accounts have not automatically disappeared simply because synthetic examples were added.
That does not guarantee the new training run succeeds. It does show why retaining and weighting original data are substantive design decisions. “Contains AI text” is too broad a description to determine the outcome by itself.
Another study tested a different arrangement
In Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data, Matthias Gerstgrasser and colleagues examined the difference between replacing datasets and accumulating data while retaining original examples. Their experiments found that accumulation could avoid collapse in the settings they studied.
This is an important counterpoint to an inevitable-decline narrative. It is not proof that any mixture, at any scale, with any training method will be safe. It identifies a consequential difference between experimental arrangements.
The two research directions therefore need not be presented as a simple contradiction. One demonstrates a failure mode. Another investigates conditions under which that failure can be avoided.
For readers, the practical lesson is to ask what happened to the original data. Was it retained? Was it replaced? How much synthetic material was added, and how was it selected? Those questions move the discussion from a slogan to a process that can actually be examined.
Synthetic data can be deliberately useful
Generated examples are not always an accidental contaminant. Researchers can create them for a defined purpose, select them and use them in a controlled training procedure.
Self-Instruct: Aligning Language Models with Self-Generated Instructions, by Yizhong Wang and colleagues, describes a process that generates instructions and examples, filters unsuitable or overly similar material and uses the resulting data to improve instruction following. The reported gains concern that method and evaluation, rather than a universal endorsement of generated training data.
The difference from indiscriminate collection is the designed procedure. The researchers are not merely assuming that everything a model writes should become training material of equal value.
A useful comparison is educational publishing. An invented maths exercise can be valuable even though it is not a record of a real shopping trip. Its value depends on whether it correctly teaches the intended skill. An invented eyewitness account presented as a real event creates a different problem.
Origin matters, but so do purpose, correctness and the way an example is used.
Three situations that should not share one headline
| Situation | The question to ask |
|---|---|
| A dataset is repeatedly replaced by generated samples | What information is lost across generations? |
| Synthetic examples are added while original data remains | How are the sources balanced and evaluated? |
| Generated examples teach a defined task | Are the examples correct, varied and relevant to that task? |
These are simplified categories, but they make a useful editorial test. If an article cites evidence from one situation and draws a conclusion about all three, it has probably skipped an important step.
The same care is needed with model types. A result from a particular experimental system is evidence about that system and mechanism. Extending it to every deployed model requires additional reasoning and evidence.
This is not an argument for ignoring research until it reproduces every commercial product. Controlled experiments are valuable precisely because they isolate mechanisms. The responsibility is to report their scope accurately.
Why average performance may not tell the whole story
Return to the fictional travel collection. A test made mostly of common journeys could make a system look capable even after it loses the unusual ones. The evaluation would be asking questions concentrated in the part it still represents well.
This is a consequence of the example, not a claim that a particular commercial benchmark has concealed collapse. It illustrates why the test set matters alongside the training set.
If you care about rare situations, uncommon language varieties or unusual combinations of requirements, you need evaluations that actually include them. A high overall score can coexist with weaknesses in a smaller category.
The issue is familiar outside AI. A transport service can have a strong average punctuality figure while performing badly on one route. The average is informative; it does not replace the route-level question.
For AI, a credible account of performance should therefore explain what was measured and which parts of the intended use remain uncertain.
A bad answer is not enough to diagnose collapse
When an AI system gives an incorrect answer, several explanations are possible. The prompt, available context, training, retrieval and application design may all be relevant. Observing the output alone does not establish a history of recursive training failure.
Model collapse is a claim about a mechanism and its consequences. To support it in a specific system, you would need evidence about the training process and changes in the system's representation or performance.
A screenshot of an absurd response may demonstrate that the response is absurd. It does not reveal what was in the training data or prove which process produced the error.
This distinction matters in public debate because vivid examples travel quickly. They can motivate investigation, but they cannot substitute for the investigation. If the diagnosis is more specific than the evidence, the language should become more cautious.
What publishers can control
A publisher cannot determine every future use of material on the internet. It can control how its own work is produced and how readers can inspect its claims.
For a researched article, preserve the links to original evidence. Distinguish direct findings from interpretation. Label hypothetical examples as hypothetical. Avoid manufacturing a quotation, interview or first-hand experience to make an article feel more authoritative.
Those are editorial practices rather than a technical guarantee against model collapse. Their value is immediate: a reader can follow the evidence and see where the publication's own reasoning begins.
The same principles shape our coverage of why AI uses electricity. A past-year estimate, a future scenario and an illustrative calculation are useful in different ways. Treating them as interchangeable makes the article easier to write and harder to trust.
Original reporting also remains distinct from rewriting available material. A fresh observation adds something that a chain of summaries cannot create merely by becoming longer.
Questions worth asking an AI provider
A provider may not disclose its full dataset, but the questions can still be clear. Does it distinguish original and generated material? What quality checks are applied? How are duplicates and highly similar examples handled? What evidence shows that the resulting model improves on the tasks it is intended to perform?
Ask about evaluations as well as training descriptions. A statement that data was carefully curated is less informative than an explanation of the checks and the results they produced.
Also ask what the provider does not know. Uncertainty about provenance can be a real limitation. Acknowledging that limitation is more useful than replacing it with an absolute claim that all material is pristine or that synthetic content is harmless in every circumstance.
For users comparing tools, these questions may not produce a complete audit. They can still distinguish a substantive technical account from a reassuring phrase.
The future is conditional, not predetermined
The research gives us a failure mode, evidence about ways to avoid it in particular settings, and examples of synthetic data being used productively. Taken together, that is a more useful picture than either inevitable collapse or effortless self-improvement.
The quality of future training material depends on choices: what is retained, what is generated, what is checked and what counts as success. Those choices have costs, including the computational costs discussed in our energy guide, and they deserve scrutiny.
The question is not answered by spotting an AI-generated sentence online. It is answered by examining what happens when that sentence enters a dataset and what the next system is asked to learn from it.
Sources checked 20 September 2026. The research findings are attributed to their experimental settings; the travel-account example is an original illustration, not a reported experiment.

