My AI Said Something Wrong and It Ended Up in the Changelog
How gitaiflow checks a model's answer before saving it, and what it does with answers it doesn't trust.
CBIf you've automated anything with a language model, you've had this moment. The model returned something that looked like an answer. It wasn't. It was its own reasoning, or half a sentence, or the instructions you gave it echoed back. And because the output went straight into a file, that mess got saved, cached, and eventually showed up where people read it.
For a tool that writes commit messages and changelogs, that's the worst failure. Not an error, but a confident wrong thing sitting in your release notes. Here's how gitaiflow is built to avoid that.
1. Ask for a Format That's Hard to Fake
Plain words like TITLE: and BODY: are fragile. A diff of code that talks about titles and bodies contains exactly those words, and a model quoting that code can repeat them outside its real answer. So gitaiflow asks for three unique whole-line markers instead:
===GITAIFLOW-TITLE=== ===GITAIFLOW-BODY=== ===GITAIFLOW-END===
Reasoning models sometimes print the format before answering, so the parser uses the last title marker it finds. Anything the model wrote before the title is kept aside as analysis, and any chatter after the bullet list is kept out of the commit body.
2. Be Tolerant, but Say So
Models drift. If the markers are missing, the parser falls back to the older TITLE: and BODY: words. If those are missing too, it takes the first line as a title, and in that case it marks the result as not parsed.
That flag matters more than anything else in the pipeline. The fallback exists so nothing crashes, not so its output is trusted.
3. Reject Clearly, Don't Repair Silently
An answer is rejected for one of three reasons: the model didn't follow the format, the body is absurdly long (over 20,000 characters, which means runaway output), or the title contains leaked instruction text. The result is stored as a placeholder:
chore: update <folder name> - Summary rejected this run: <reason>. The raw model response is saved in commit.analysis of this JSON file.
The reason is stated in plain words and the model's raw response is kept, so you can see exactly what came back and nothing is lost. You can see it yourself with --show-analysis.
4. A Rejected Answer Never Gets Reused
This is where silent damage normally happens, so there are three separate guards:
- A rejected summary is not cached. You get a warning that it will be retried on the next run.
- If an older cached summary turns out to be malformed, it is regenerated instead of reused.
- A summary marked not parsed is skipped when building a changelog, so a placeholder can never become a release note.
5. Long Titles and Provider Errors Are Different Problems
A title over 72 characters isn't rejected. It's cut to fit, and the overflow moves into the body as (title continued: ...), so nothing is lost.
Provider failures are handled separately: temporary errors such as rate limits and server errors are retried up to three times with increasing delays, while a daily-quota error is not retried, because waiting can't fix it.
6. What This Can't Catch
Be honest about the limit. All of the above checks the shape of an answer, not its truth. A summary can be perfectly formatted and still get a detail wrong. That's why I read a generated summary as a draft, and why I check Breaking Changes and Security entries by hand before anything ships.
Want the Full Picture?
This post covers output validation. The full behavior reference is in the gitaiflow documentation ↗.
What I'd Tell You If You're Automating With an LLM
- Decide what a valid answer looks like before you call the model. Validation written afterwards tends to fit whatever came back.
- Mark untrusted results instead of dropping or fixing them. A visible placeholder with a reason is better than silent success.
- Keep bad output away from caches and published artifacts. The damage from a wrong answer comes from where it gets reused, not from the answer itself.