Model Collapse

Model Collapse

Model collapse refers to the gradual quality decline of AI systems that are trained predominantly on text or images that themselves originate from AI. Over multiple generations, such systems lose rare details and ultimately produce only monotonous, average-like outputs.

Programs like ChatGPT learn by analyzing huge amounts of text and images from the internet. Until recently, almost all of this came from humans. Meanwhile, a growing share of the internet is itself machine-generated. When a new program learns mainly from such machine-generated content, its quality deteriorates from generation to generation. This exact effect is called model collapse. At the end of this chain stand systems that sound fluent but are content-wise flat and monotonous.

Why the internet as training material is becoming scarce

The performance leaps of recent years were largely based on sheer quantity. More text, more images, more compute yielded better systems. However, high-quality, human-written text is a finite resource. Some researchers estimate that the usable supply on the web will be largely exhausted still within this decade.

At the same time, the web is filling up with automatically generated articles, product descriptions, and images. Anyone collecting data from the web today can barely distinguish cleanly what originates from humans and what from machines. This threatens not only to make the supply scarce, but also contaminated.

For companies, this is a tangible economic risk. A model whose training costs hundreds of millions of dollars must not be worse than its predecessor. That’s why companies are now paying dearly for rights to newspaper archives, forums, and book collections. Human-generated data is becoming a raw material with a market price.

How the rare cases disappear

An AI model doesn’t store texts, but probabilities. It learns which phrasings occur frequently and which occur rarely. When it generates something itself, it mostly chooses the probable one. Unusual words, rare opinions, and edge cases therefore appear less often in its outputs than they occurred in the original.

If a new model is now trained on these outputs, it knows the edge cases even more weakly. Its own texts then become even more uniform. After several rounds, a large part of the diversity has simply vanished. Experts compare this to a photocopy of a photocopy: each pass seems harmless on its own, but after ten passes the text is barely readable anymore.

A common misconception is that the models suddenly start producing nonsense. The decay is rather inconspicuous. The answers remain grammatically correct and sound confident. What gets lost is what the average doesn’t capture: the unusual illness, the dialect, the minority position. A model affected by model collapse appears competent and has nevertheless become poorer.

Remedies and debates in the industry

The term became well-known in 2023 through a widely cited study by British and Canadian researchers and has since regularly appeared in business news. It usually concerns data licenses: when an AI company signs a contract with a publisher or a news agency, model collapse is part of the justification. Human-generated content is the insurance against the creeping loss of quality.

In research, the situation is less dramatic than early headlines suggested. The decay occurs mainly when new data completely replaces the old. If machine-generated data is mixed with human originals, quality remains largely stable. Moreover, companies deliberately use generated training data with success, for instance in mathematics, where every solution can be automatically verified.

In practice, this means: what matters is not the origin of the data alone, but the control over it. Whoever checks and filters can use machine-generated data without risk. Whoever collects everything unfiltered from the web takes on a risk. This is why some platforms now mark AI content with invisible tags, so-called watermarks, so that it can later be filtered out again.

Related Products

Latest News

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.