Here is my hypothesis for why AI content can follow the style guide and still feel off. The guide never specified how conformance should be judged, and a seasoned human reviewer filled that gap invisibly for years, and we never had a reason to notice that the gap was there.
Most style guides specify the rules: tone, terminology, structure, voice. What they do not specify is the rubric: which dimensions matter, how a borderline case should be scored, what "on brand" means when a piece of content satisfies every rule and still reads wrong. A human writer or editor absorbed that judgement over time, largely without documenting it, and applied it as tacit knowledge every time they reviewed a piece of content. That tacit layer is not a minor detail sitting on top of the written rules. For most style guides, it is doing more of the work than the rules are.
An AI system has no equivalent tacit standard to fall back on. It can follow every explicit rule in the guide and still miss the thing a seasoned editor would catch instantly, because the thing the editor would catch was never written down as a rule in the first place. That is not the model failing to understand the brand. That is the brand's standard living only in a person's head, exposed the moment the person reviewing the content is not that person anymore.
Every rule set that bans certain words or constructions runs into this straight away. Banning em dashes is a rule, and a model can follow it without effort. Catching the quieter tells, the throat-clearing opener, the "actually" that adds nothing, the sentence that explains a joke instead of trusting it to land, takes the same judgement an editor has always used without ever writing it down as a rule. We ban the easy tell and miss the ones that give the writing away.
The fix, perhaps, is not a longer style guide sitting on a shelf. It is a judging method: a short, explicit statement of what separates a borderline pass from a borderline fail, written into the prompt itself rather than left in a document nobody consults mid-task. That puts the standard in front of whatever is producing the content at the moment it matters, which is also where consistency and speed come from: not a longer rulebook, but a clearer standard applied at the point of use.
I would argue that the best test of that is a practical one: build the rubric, apply it to prior work, then gauge how well it performs. It might even catch a few mistakes that happened to pass the editorial gate the first time round.
This does not mean every scrap of tacit judgement can be written down. Some of it probably cannot. We can try to capture what currently lives only in a reviewer's head and make it explicit, and closing the gap is worth doing even without closing all of it.
The irony is hard to miss once you notice it. We are already paying AI to fight the tells AI itself produces, one prompt at a time. The rubric is what turns that from a losing chase into an actual standard.