The source data said "cotton blend." The description that came back said "crafted from premium merino wool." Nobody typed that. This is the log of why models invent things, and what finally caught them.
The sweater that promoted itself
We run LLM enrichment across 70+ brands on the SYNC PILOT. The job we scoped was almost boring: take a title, a SKU, a category, the source attributes, write a clean description. The early runs felt like cheating. Copy came out faster than anyone could read it.
Faster than anyone could read it. That was the warning, and we filed it as a feature.
Then the sweater. A blue cotton-blend sweater came back "crafted from premium merino wool," and the source data contained nothing, anywhere, about wool.
We blamed the prompt
The fix looked obvious. Tighten the prompt, add a "do not invent specifications" line, move on. (We felt clever for about one spot-check.)
Then a phone gained "wireless charging via Qi standard," a feature that appears nowhere in its spec sheet. A shoe got a size range its manufacturer has never produced. Every new rule held until it didn't, and the fabrications kept coming, politer each time.
So we blamed the prompt harder. Then we blamed the model. It wasn't the prompt, and it wasn't the model.
The pattern was worse than either theory. Give the model sparse input, "Blue Shirt, Medium," and it fills the silence: suddenly the shirt is "a versatile cotton essential ideal for layering." Plausible, unsourced. Precise language drifts too. "100% natural fibers" comes back "eco-friendly natural blend," close enough to slide past a tired reviewer, different enough to be a compliance problem. And multi-variant products, the SKU soup every apparel catalog swims in, quietly nudge the model into describing combinations that don't exist. Spot-check a batch and sooner or later a description is talking about a variant that isn't on the page.
What actually scared us was the confidence. The model states dimensions and weights with the same stylistic certainty it uses for adjectives. Right or wrong, the output reads the same. There is no tell.

Nothing was broken
The turn came when we stopped asking what was broken and accepted the answer: nothing. NIST's 2024 Generative AI Profile lists "confabulation," the confident presentation of false content, as one of 12 risks unique to or exacerbated by generative AI. A property of the architecture, not a bug in ours.
The model never looks anything up. It computes the most probable next token, and a plausible token is not the same thing as a true one. It has no idea there is a real sweater somewhere, made of actual cotton.
Nerd detail, skip freely: a 2025 analysis grounds this in computability theory. For any LLM, the set of inputs that produce incorrect output is infinite. No prompt, no fine-tune, no bigger model closes that door.
Fluency is the most dangerous feature these models have. So we stopped spending on prevention and started spending on verification.
Writer, critic, human
Catalogs turn out to be a terrible place to hallucinate, which makes them a good place to catch it. Open-ended writing has no ground truth. A catalog keeps receipts: every generated claim can be held against source data, and the gap is measurable.
The numbers told us where to aim. In the fact-checking research we leaned on, GPT-3.5 and GPT-4 averaged 63-75% accuracy working from the claim alone. Hand them reference material and they clear 80% and 89%, but only on claims with nothing ambiguous about them. Translation: on a real catalog, where sparse and ambiguous is the default, a human will have to look. The design question is how often.
Inside OKART the answer became ContentScope. One model writes. A second evaluator model hunts the output for hallucination signals: extraneous claims, confidence without grounding, format violations, tone drift. Clean output passes. Flagged output goes to a person, and every correction feeds back into the critic, so it keeps getting sharper.

The critic is the whole ballgame. Without it, humans would need to review close to everything, at which point congratulations, you have automated nothing. With it, people see only the flagged cases and the pipeline still moves 5,000 products/day, verified before publication instead of corrected after a customer files a return.
The human gate stays for a quieter reason. We have watched descriptions pass every automated check and still misrepresent what a product is for. The safety literature converges on the same stack, input checks, a judge model, output filters, human review, and our logs agree: no single layer holds.
What survived
- Hallucination is a property of the architecture. Plan for it the way you plan for returns.
- Prevention has a ceiling. Verification doesn't, so that's where the budget went.
- Autonomous means autonomous with human oversight. Machines take the confident cases; people take the ambiguity every real catalog is full of.
Somewhere in a warehouse sits that blue sweater, cotton blend, exactly what the label always said. The catalog finally agrees with the label.
It was never merino.
If you're weighing LLM enrichment for your own catalog, the Infrastructure Stress-Test includes a content-verification checkpoint that measures hallucination risk and guardrail effectiveness on your data, not a benchmark set.
Sources
-
NIST Generative AI Profile (July 2024). Defines 12 risks unique to or exacerbated by generative AI, including confabulation, with suggested actions to govern, map, measure, and manage them.
-
The perils and promises of fact-checking with large language models. Reports GPT-3.5 and GPT-4 at 63-75% average accuracy without context, improving to above 80% and 89% on non-ambiguous verdicts when context is incorporated; concludes the models "cannot completely replace human fact-checkers" because even infrequent errors carry outsized consequences.
-
LLM Guardrails: Strategies & Best Practices in 2025. Details guardrail categories spanning input, output, LLM-as-judge, and operational layers; emphasizes layered defenses and human review workflows for flagged outputs.
-
Inevitable Hallucination in LLMs. Establishes mathematical inevitability of hallucination via computability theory; notes practical error rates can be reduced to negligible levels through scaling and engineering.
-
Auto prompting without training labels: An LLM cascade for product quality assessment in e-commerce catalogs. Cascade architecture for product quality assessment; demonstrates layered LLM evaluation on e-commerce catalog data.





Share:
Your sync didn't crash. It hit a bucket of 40.
Your sync didn't crash. It hit a bucket of 40.