People have become good at recognising machine written text. Not perfectly, but well enough that the cost of being caught is real, especially with a technical audience that reads carefully.
The tells fall into three groups, and they are not equally important.
Vocabulary
The easy ones. Words that appear far more often in generated text than in ordinary writing:
delve, leverage, seamless, robust, cutting-edge, game-changing, transformative, revolutionary, unlock, empower, elevate, supercharge, paradigm shift, best-in-class, state-of-the-art, next-level
And the phrases:
"In today's fast-paced world", "It's not just X, it's Y", "Let that sink in", "Here's the thing", "And that's when I realized", "The future is here", "I'm thrilled to announce"
These are easy to filter, which is why we maintain a list and check the output against it rather than only instructing the model to avoid them.
They are also the least important group, because avoiding all of them still leaves text that reads as generated.
Punctuation and rhythm
More reliable than vocabulary.
The em dash. Language models use them constantly, frequently several per paragraph. Most human writers use them rarely. A short post with three em dashes is a strong signal.
Tricolons everywhere. "It's faster, cleaner, and more maintainable." Groups of three are a natural rhythm and generated text overuses them to the point where you notice.
Uniform sentence length. Human writing varies. A short sentence. Then a longer one that carries more of the argument and takes its time. Generated text tends toward a consistent medium length, which reads oddly flat.
Every paragraph the same size. Three sentences each, throughout.
Balanced closings. Every section ending with a tidy summarising sentence that restates what was just said.
Structure
The hardest to fix and the most telling.
Symmetry. Five sections of equal length, each with the same internal shape. Real writing is lumpy, because some parts of a subject deserve more room.
No dead ends. Human writing wanders slightly, mentions something and drops it, has an aside. Generated text is relentlessly on topic, which reads as smooth and strangely lifeless.
Concluding what was already said. A final paragraph that adds nothing and exists because the shape demands an ending.
Hedged everything. Every claim softened. "This can often be a useful approach in many cases." Real writers commit more.
Why the list matters even when a human is writing
Most of these are not exclusively AI tells. They are bad writing habits that AI has amplified because it learned them from the same places humans did.
Cutting em dashes, varying sentence length, letting sections be uneven and committing to claims makes writing better whoever produced it.
What we do about it
Noomachy DevTracker checks generated posts for banned vocabulary and for em dashes, and flags both. Neither blocks publishing, because they are style rather than truth, and the user is allowed to disagree.
The prompt layer is more direct than most: an explicit list of banned words, an instruction not to use em dashes, and a rule against stacking rhetorical fragments for emphasis.
The structural tells are harder to check mechanically and mostly get handled upstream. A post built from four specific commits with real file paths and a real line count has something concrete to be about, which is the best defence against the flat symmetric shape that generated text falls into when it has nothing to say.