AI3 min read

Why AI writing tools produce slop

It is not the model. It is that the model was given nothing to work with, so it filled the gap with the average of everything it has read.

Open most AI writing tools and you get a text box. Type what you want to post about. Get a post.

The output is fluent, structurally correct, and empty. It reads like everything else, because it is the average of everything else.

People blame the model. The model is mostly not the problem.

The input is the problem

Ask for "a LinkedIn post about shipping a new feature" and you have supplied about eight words of information. The model has to produce three hundred. The remaining two hundred and ninety come from its prior, which is the average of a very large number of LinkedIn posts.

The average LinkedIn post is slop. Not because language models are bad but because most LinkedIn posts are bad, and averaging them produces something worse than any individual one.

Better prompting helps marginally. It cannot fix a fundamental shortage of information, because no amount of instruction adds facts that were never provided.

What actually fixes it

Give the model something only you have.

That is the whole insight. The difference between a generic post and a good one is specific true detail, and specific true detail has to come from somewhere. A model cannot invent your file paths, your line counts, or the four hours you spent on a one line fix.

This is why Noomachy DevTracker starts from commits rather than a prompt box. By the time the model is asked to write anything, it has:

  • the actual commit messages, parsed and humanised
  • the file paths that changed
  • lines added and removed
  • which of those changes were user facing
  • a judgment about which commits belong to the same piece of work

That is a few hundred words of real, specific, verifiable input about one afternoon. The model's job is to shape it, not to invent it.

The second problem: nothing checks the output

Even with good input, a model will reach for things that sound right. Numbers, in particular. Ask for a post about a performance improvement and there is a good chance you get a percentage that nobody supplied.

Prompting against this helps and does not solve it, because a prompt is a request. Anything that matters has to be checked.

So the output is validated deterministically. Every number in the post must appear in the source material. 40% faster fails if no 40 exists in the commits. Claims that a commit scan can never support, user feedback, revenue, popularity, get blocked by pattern regardless of phrasing.

No model checks another model's work, because a checker sharing the same instincts is not an independent check.

The third problem: no voice

Generic output also comes from having no target to aim at.

"Write a LinkedIn post" has no style constraint, so the model uses the default register, which is the average one. Supplying a voice profile, ideally derived from the person's own previous posts, gives it something specific to match.

Tone, sentence length, emoji usage, technical depth, how they tend to open. Observed from samples rather than asserted.

What is left for the model to do

Once you supply the facts, the checks and the voice, the model is doing what it is actually good at: turning structured information into readable prose, choosing an angle, deciding what to leave out.

That is a real job and models are excellent at it. The mistake is asking them to do the other job, the one where they supply the substance, because that is the job they do by producing an average.

aiwritingcontent

Keep reading