Sanitize or Fail: Two Ways to Handle Rule Violations in LLM Output
An em dash in generated text has one correct fix. Swap it for a comma, swap it for a period, or recast the sentence. The first two are a string replacement. You do not need a second model call to do it, and you do not need the model to agree.
A dollar amount you cannot verify is a different animal. If a draft says your product saves buyers some amount of money, you cannot repair that sentence, because you do not know what the true amount is or whether a true amount exists. Deleting the figure leaves a claim with a hole in it. Guessing a replacement means your code is now the one making things up.
That difference is the whole design of a post-generation check layer. Every rule you enforce on model output falls into one of two buckets, and the bucket decides the code path.
The test for which bucket a rule is in
Ask one question about the violation: can the correct output be computed from the bad output alone, without knowing what the writer meant?
If yes, it is a repair. Typography, whitespace, smart quotes, a banned unicode character, a trailing heading colon, a markdown link wrapped in stray backticks. The fix is mechanical and total. Do it in code, silently, and move on.
If no, it is a failure. An unverifiable number, a claim about results, a phrase you never want published under your name, a leftover placeholder token. There is no function from the bad text to the good text, because the information needed to write the good text is not in the bad text. Code cannot fix it, so code should refuse it.
Mixing those two buckets is where these pipelines rot. Route repairs through the model and you burn latency and money on work str.replace does perfectly. Route failures through a sanitizer and you produce text that passes your own checks while still being wrong, which is the worse outcome, because now nothing downstream knows anything was ever suspect.
The repair path
Keep it boring and deterministic. This is a pure function with no network in it.
import re
DASH_REPLACEMENTS = {
"\u2014": ", ", # em dash
"\u2013": "-", # en dash
}
def sanitize(text: str) -> str:
for bad, good in DASH_REPLACEMENTS.items():
text = text.replace(bad, good)
text = re.sub(r"[ \t]{2,}", " ", text)
text = re.sub(r"\s+([,.;:])", r"\1", text)
return text.strip()
Two properties matter here. It is idempotent, so running it twice changes nothing, which means you can call it at more than one stage without thinking hard. And it never asks the model for a second draft. In my own agent, em dashes are stripped by a sanitizer and never retried, because a retry is a chance to introduce a new problem while fixing a solved one.
Watch the blast radius of a sanitizer, though. A replacement that is correct in prose can be wrong inside a fenced code block or a URL. If your drafts contain code, split the text on fences and sanitize only the prose spans.
The failure path
Failures need an allowlist, not a pattern. Patterns tell you where the numbers are. Only a fact sheet tells you which numbers are true.
Give each product a set of figures it is allowed to state, written exactly as they may appear. Then extract every figure in the draft and check membership.
MONEY = re.compile(r"\$\s?\d[\d,]*(?:\.\d+)?")
PERCENT = re.compile(r"\d+(?:\.\d+)?\s?%")
def normalize(token: str) -> str:
return token.replace(" ", "").replace(",", "")
def check_figures(text: str, allowed: set[str]) -> list[str]:
problems = []
for pattern in (MONEY, PERCENT):
for token in pattern.findall(text):
if normalize(token) not in allowed:
problems.append(f"fabricated_number:{token}")
return problems
The allowlist should hold normalized forms too, so a figure written with a thousands separator in the draft still matches the same figure written without one in your facts. Everything that survives normalization and is still absent is, by definition, a number your source of truth does not contain. There is no judgment call left to make, which is exactly why this check belongs in code rather than in a prompt. A prompt that says "do not invent numbers" is a request. A membership test is an outcome.
Banned phrases work the same way, and they are cheaper: a substring scan for the words you refuse to publish, the placeholder tokens your template leaves behind when a field is missing, and the phrasings that read as machine filler in your house style.
Wiring the two together
Repair first, then judge. Otherwise you fail drafts for typography and lose the real signal underneath.
def review(draft: str, facts) -> dict:
text = sanitize(draft)
problems = check_figures(text, facts.allowed_figures)
problems += check_banned_phrases(text)
if problems:
return {"text": text, "status": "failed", "reasons": problems}
return {"text": text, "status": "pending_review"}
Note what a clean result is: pending_review, not published. Passing the validators means the draft earned a human's attention, nothing more.
Log the two buckets separately. Sanitizer hits are telemetry about your prompt and your model settings. Validator failures are signal about the draft itself, and they are the thing worth feeding back: you can inject the recent failure reasons for that product and that channel into the next generation, so the next draft starts from a narrower target. Never mix sanitizer counts into that feedback. If you tell the model it failed for a dash you already fixed in code, you are spending prompt tokens teaching it to avoid something that costs you nothing.
One more ordering decision worth taking deliberately: decide whether a validator failure gets a regeneration or gets dropped. Both are defensible. What is not defensible is regenerating forever. Cap the attempts, and when the cap is hit, keep the failed draft with its reasons attached rather than deleting it, because a pile of failed drafts with reasons is the most useful artifact this whole layer produces.
What this does not solve
This is a text-level guard. It has real edges, and you should know them before you trust it.
It cannot catch a false claim that contains no number. "Buyers love it" passes every check above while being unsupported. Sentences like that need a human, or a separate pass with a much narrower question.
It cannot catch figures written as words. A percentage spelled out in English slips past a digit regex. You can extend the pattern, and you will still be playing catch-up with the language.
It cannot judge tone or fit. A draft can be fully accurate and still be wrong for the venue.
And the allowlist is now a maintained artifact. Change a price and forget the fact sheet, and your validators will reject the truth. That is the cost of making the check deterministic, and it is the right cost, because the failure mode is a blocked draft rather than a published falsehood.
Who should skip this: if you publish rarely, read the draft yourself. This layer pays off when volume makes word-by-word reading the thing you quietly stop doing.
I build and run Content Agent Pro, a self-hosted Python and Next.js content agent where validators like these are the gate every draft passes before a human sees it: https://fulcrumenterprises.tech/go/content-agent-kit-pro/?c=hashnode
