Why AI Models Are Starting to Sound More Alike

AI models from rival companies are beginning to produce more similar answers, even when asked to handle open-ended creative tasks. New research from Duke University found a statistically significant decline in the diversity of model outputs across both everyday prompts and a standard psychology test.
The study, published on August 21, 2026, examined 69 models from 12 provider families. The models came from Anthropic, Cohere, DeepSeek, Google, Meta, MiniMax, Mistral AI, Moonshot AI, OpenAI, Qwen, xAI, and Z.ai, with releases spanning March 2023 to July 2026.
A Three-Year Look at Model Creativity
Researchers tested 122 runs of models distributed across that three-year period. Some providers released models on increasingly frequent schedules, giving the study a way to compare creative answers across different points in the development cycle.
The team used two kinds of prompts. One set came from real-world, open-ended questions that asked models to generate responses without a single correct answer. The other came from a standard creativity test focused on unusual uses for ordinary objects. In both cases, the researchers measured how far apart the answers were in meaning.
That distance shrank over time. In plain terms, responses from different models began to resemble one another more closely, even though the models came from competing companies and different provider families.
The finding does not say that every answer from every model is identical. It shows that the range of answers is becoming narrower across the models tested. LLM responses to open-ended prompts have become more similar over time, and the same pattern appeared in the creativity test.
Why Similar Answers Matter
The authors connect this convergence to a decline in creativity when LLMs handle tasks that require open-ended responses. Their concern reaches beyond whether one answer sounds original. If several systems produce the same kinds of ideas, users may see fewer possibilities when they use those systems for writing, brainstorming, or other creative work.
“Our findings, though preliminary, raise concerns about the long-term usefulness of LLMs as creative partners,” the authors of the new paper wrote.
That concern also affects the people using the models. Converging outputs could limit the range of possibilities users encounter and narrow the breadth of their own thinking. An LLM that offers a smaller set of recurring ideas may still produce polished work, but it can expose users to less variety.
The research adds a broader question to the debate about AI tools: does better performance come with less creative difference between systems? The study does not answer why the outputs are converging. It does show that the pattern appeared across real-world prompts and a psychology-based test, rather than in only one narrow task.
Earlier Signs of the Same Pattern
The new study follows isolated incidents that pointed toward similar behavior. A Cornell University study in May 20026 examined similarities in LLM responses to prompts involving “lighthouse keepers” and names such as “Mara” and “Elias.” Those examples drew attention to the possibility that different systems may return overlapping creative ideas.
Duke University’s research tested that possibility across a much larger set of models and a longer period. Instead of examining only a few memorable responses, it compared three years of model outputs and measured how their meanings moved closer together over time.
That approach matters because a single shared answer can be an isolated event. A repeated decline in output diversity across 69 models offers a broader view of how LLM behavior is changing. The results raise concerns about the long-term role of these systems as creative partners, especially when users depend on them to expand the range of ideas.
The Bigger Question for AI Users
AI models now come from a wide group of providers, including companies behind Claude Mythos, GPT-5.6 Sol, and ChatGPT. The Duke University research suggests that provider choice may not always produce as much creative variety as users expect, because answers from different systems are moving toward similar meanings.
For people using LLMs to explore open-ended questions, the value of a response may depend on more than its clarity or polish. Variety matters too. A creative partner is useful partly because it can offer an angle a person did not consider, but converging outputs may reduce the number of angles placed in front of them.
The authors describe their findings as preliminary, but the study presents a clear trend across the models and tests examined. As providers continue releasing models, the question is no longer only what an LLM can produce. It is also whether different LLMs will keep giving users different ways to think.
Based on
- AI Models’ ‘Creative’ Output is Becoming Similar Across Providers — unite.ai
- AI firms are watermarking generated text – here’s why it won’t work | New Scientist — newscientist.com
- AI lab’s safety systems are falling behind | Fortune — fortune.com
- Agentic AI and cybersecurity, the story so far | Nature Machine Intelligence — nature.com




