Problem Patterns in AI: Beyond Hallucinations

By now, you probably know AI slop when you see it. Bloated articles about nothing much at all. Uncanny images of too-glossy “people” smiling into the middle distance. LinkedIn comments (90% of them). Slop has become an entire genre of multimodal machine-generated text, and most people can spot it at twenty paces even if they’d struggle to explain exactly what gives it away.

A sloppy AI article might contain a fabricated statistic or a citation to a paper that doesn’t exist, and I’ve written about why those “hallucinations” happen and why the problem is a feature of Large Language Models (LLMs). But slop isn’t just about hallucinations: strip out all the factual errors and the feeling of slop remains: the genericity, the fluent-seeming but empty prose, the structures and rhetorical techniques. Whatever is wrong with that writing, “the AI got its facts wrong” doesn’t cover it.

For several years we have focused our attention on hallucinations. Entire risk frameworks have been written around mitigating or attempting to stop hallucinations. But these “undesirable outcomes” are just one member of a much larger family of recurring, recognisable problems. Most of those problems have no agreed names, which makes them hard to talk about, hard to teach, and hard to manage. So, in this article, I’ll identify some of the broader risks posed by careless or problematic use of GenAI.

Hallucination is too narrow

If you’re a regular GenAI user, or even a reluctant AI reader, think about the outputs that cause irritation in day-to-day AI use. Some are hallucinations, certainly. But most of the time, the problematic features of GenAI are stylistic, aesthetic, or can’t-quite-put-my-finger-on-it uncanny.

Maybe the model has:

  • Editorialised, adding its own interpretation you never asked for
  • Expanded three bullet points into three paragraphs of padding
  • Overstated the certainty of a conclusion in a summary
  • Smoothed away a nuance you specifically wanted preserved
  • Flattened your voice into the house style of nowhere-in-particular
  • Buried the lede as a side effect of presenting a “fair” argument
  • Produced reasoning that is fluent, persuasive, and just ever-so-slightly wrong

I have a recurring example of my own: use Claude to build a slide deck and, left to its own devices, it will add interpretive subtitles to the section slides. A slide that should just say “Assessment” becomes “Assessment: rethinking what we value and why it matters”. I absolutely didn’t ask for – or want – the editorial gloss. I like a minimalist slide, with no superfluous text and certainly no “why it matters” AI bullshittery. This problem of slide ad-libbing is so persistent that even if you include explicit instructions to the contrary, build a detailed skill, or yell at the model in all caps, it will eventually drift back.

TIL that the superscript title-above-the-title Claude loves using is called an “eyebrow”…

The Wharton School’s Ethan Mollick reckons Claude and ChatGPT have gotten good at slides, but the Odysseus-themed deck offered as evidence commits every one of the sins listed above, and more.

None of the items on that list is a factual error, and calling any of them a hallucination stretches the word past usefulness. These are what I’m calling “problem patterns”, in the literal sense that they result from the pattern-matching, probability-conforming, middle of the road beigeness that LLMs are famous for.

And I think they present more of a reputational risk than hallucinations alone.

Problem patterns

Leon’s Dictionary of AI Words That Mean Things defines problem patterns, as of this morning, as follows:

Problem patterns are recurring, recognisable forms in AI-generated content that may cause problems or create risk depending on context. Problem patterns arise from training that optimises for statistical plausibility and approval.

Context is everything. The same hallucinated citation is trivial in a private brainstorm and catastrophic in filed court documents. The same generic rule-of-three, “it’s not x it’s y” style is irrelevant in meeting notes, but costly when spotted out in the wild on your company’s blog. A problem pattern is a potential problem; how much of a problem it becomes in practice depends on where the output goes and who is reading it.

There are many types of problem pattern, so I’ve attempted to categorise them under seven headings: truth, faithfulness, reasoning, judgement, knowledge, communication, and aesthetic.

Truth problems. Whether claims correspond to facts and evidence. This is the hallucination problem: fabrication, factual error, invented citations, unsupported specificity. The reason hallucinations happen is explained further in my earlier article, but in a nutshell it’s an unavoidable technical limitation of predictive text models that contain no “ground truth”.

Faithfulness problems. Whether the output stays faithful to the source material, the instructional prompt, and the context. This includes instruction drift, editorialisation, omission, distortion, unrequested expansion, steam-rolling over nuance. This category also includes drifting away from contextual materials, whether uploaded into a chat or included as documents in a process like Retrieval Augmented Generation (RAG).

Reasoning problems. How claims and conclusions are connected. These are plausible but unsupported inferences, weak causal logic, internally coherent arguments with subtle (or unsubtle) flaws, shallow arguments, and other logical issues. More capable models “reason” by passing their own output back to themselves in a loop. The more complex the reasoning, the more likely it is to wobble off course.

Judgement problems. Adjacent to reasoning problems, these include issues with prioritisation, salience, and appropriateness. Models might treat every point as equally significant, failing to recognise when more information is needed, recommending the technically possible but contextually unsuitable option. LLMs are sophisticated text predictors, not people. They do not have the capacity to “judge” in the same way a human expert might.

Knowledge problems. This includes both what the model doesn’t “know”, and what it knows too much of. On one side, we have outdated information, missing local and organisational context, and the gap between general knowledge and specialist knowledge, or situated expertise. The meme-like phrase, “as a Large Language Model from OpenAI, my training data cut-off was August 2022…” seems almost nostalgic by 2026, but all GenAI models still suffer from lack of access to knowledge. The other side is overfitting on training data, where the model is exposed to the same information many times and “learns” the answer: this has lead to some of the biggest copyright issues in GenAI.

Communication problems. Simply how content is expressed for its audience. Includes generic language, repetition, verbosity, and your distinctive voice flattened into the median register of the training data. Interestingly, it seems that there are common communication issues across all models (such as the overfitting on rhetorical techniques like “it’s not x, it’s y”) and some which are unique to certain models (like Claude calling everything “the move” and claiming “the honest truth”). That suggests these are training and post-training concerns.

Aesthetic problems. Finally, the capability to judge the aesthetic quality of an output, such as form, structure, and presentation in both written text and multimodal outputs. AI writing often “looks wrong” because, aesthetically speaking, it has a distinctly inhuman rhythm or flow – a subtle, uncanny signal of slop. Of course, aesthetics are much more apparent in AI-generated image, music, and video. Although models continue to improve, and human-users continue to work around these limitations, the default position of multimodal GenAI is still generic.

The categories overlap, and there’s at least one more candidate that I considered: trust problems, or outputs that invite more confidence than they deserve. But for now I’ll stick with these seven, and using “slop” as an example explore more closely how they relate to patterns in the training and post-training processes.

Decomposing slop

On social media, “AI slop” has become the derogatory term of choice to throw at anyone suspected of using the technology. To clarify my position, I don’t really care whether people choose to use AI or not, and I have seen many examples of interesting and creative AI use. I’m interested in slop more as a cultural and communicative phenomenon than a stick to beat AI users with.

“AI slop” isn’t a thing, it’s a vibe, a feeling. It can be obvious (Shrimp Jesus) or barely noticeable; the uncanny feeling that you’ve spotted something AI-generated but can’t put your finger on why. Wikipedia already has an extensive page of AI-slop signals across media formats, but you likely know the signs yourself as a “gut feeling”.

Shrimp Jesus. Image source: Wikimedia Commons made available under CC0

In an earlier article, The Effort Economy of Slop, I borrowed the following definition originally coined on Bluesky:

Slop is something that takes more human effort to consume than it took to produce.

That clearly isn’t all AI use. If we take, for example, the recently released 13-minute short film Nightborne directed by Neill Blomkamp (District 9) you can tell that more effort went into the production of the media than my consumption of it as a viewer. TW: violence, zombies, cussin’…

But even Blomkamp’s film – which reportedly involved 32 human actors and a complete production team – has recognisable signs of “slop” which are more aspects or limitations of the technology than artistic decisions.

Here’s a breakdown of my earlier seven categories, alongside some of the technical causes of the problems and examples of what slop might look like, whether in text, image, or film:

CategoryTechnical cause of the patternWhat it looks like in slop
TruthModels predict plausible text, not verified text. There is no ground truth to check against, so a fabricated citation has the same statistical shape as a real one (e.g., Ji et al., ACM Computing Surveys, 2023).Hallucinated citations, references, quotes, and accurate-seeming but incorrect URLs, DOIs, and attributions. Confident and fluent-sounding assertions that are false.
FaithfulnessYour instructions compete with training, and over a long generation the training wins. Material in the middle of a long input is used least reliably (Liu et al., TACL, 2024).Text that starts accurate, but gradually becomes more generic the longer it gets. Output which follows instructions at first (“don’t ad lib subtitles on my slides”) and then “forgets” by the end.
ReasoningText is generated forward with no backtracking, so an early error conditions everything after it (Zhang et al., ICML, 2024). The reasoning doesn’t necessarily reflect the actual basis for the answer (Turpin et al., NeurIPS, 2023).Cause and effect asserted, but not demonstrated. Reasoning that includes logical flaws, draws on inaccurate or misinterpreted evidence, or is circular. Confident assertions that seem to come from nowhere.
JudgementPreference training rewards agreement with the user (Cheng et al., 2026) and sometimes rewards “length” almost as an end in itself (Singhal et al., COLM, 2024).Recommendations that ignore the user’s requests and seem off-topic or otherwise irrelevant. “Inhuman” decisions or errors that a human would easily avoid. Sycophantic responses.
KnowledgeOn the one hand, models have a strict “knowledge boundary” up to the training data cut-off (Pęzik et al., 2025 (preprint)). On the other, various causes can lead to memorisation of training data.Defunct or out of date assertions. “Facts” which reflect commonly held misconceptions learned through a large volume of inaccurate training data. Memorised passages of text, such as large chunks of novels or articles from training data.
CommunicationPost-training narrows the output distribution, so a small set of phrasings and structures dominates (Reinhart et al., 2025). The homogenisation might even transfer to the humans writing with it (Padmakumar and He, ICLR, 2024).“It’s not x, it’s y.”, the rule of three, AI slop vocabulary (delve, navigating…). A general flattening of authorial voice into something that “seems like AI” for some reason. Generic visual structures in media like slides.
AestheticForm reflects what raters rewarded, but the model has no inherent sense of style. Image models converge on the average of a prompt and reproduce the dominant patterns of the training set even when asked not to (Bianchi et al., FAccT, 2023).Arrhythmic or stilted prose. “Ugly” template style slides and websites. Design choices that are generic and obviously based on training methods and not human judgement. Stereotypical or otherwise problematic images, audio and video.

These are reputational hazards

Recently I ran a totally unscientific poll on LinkedIn about what people saw as the reputational hazards associated with AI (mis)use. Based on the results and the comments, hallucinations are seen as far less hazardous than obvious signs of AI slop, and the risk of flawed judgement.

People are increasingly exhausted by the overwhelming volume of AI generated content online. Publishers cranking out obviously AI-generated articles are losing readers, and authors are coming under fire for blatant AI use. Substack has recently implemented AI detection, via the Pangram platform, to “protect” users from obviously AI generated content. I don’t think that’s the answer, and AI detection tools are very problematic technologies, but it is yet another sign that public sentiment has turned against the more flagrant uses of GenAI.

I think we need more nuance in these conversations. It’s not simply a case of “did the writer/artist/director use AI”, but how, and what, and why. The cost of AI use needs to be weighed carefully against our ability as creators to catch and correct the problem patterns that are inherent to the technologies. Sometimes, that cost might be too high: to our reputations, and to the trust between ourselves and our audiences. Sometimes, AI use might not only be acceptable, but perhaps the better option. This is a trade off between cost and “catchability”, which I’ll describe in the next article.

Understanding those variables is still very much a question of human authorship.

Want to learn more about GenAI professional development and advisory services, or just have questions or comments? Get in touch:

← Back

Thank you for your response. ✨

Leave a Reply