I built an OpenAPI changelog generator because writing API release notes was eating my afternoons: diff two specs, output a structured list of breaking changes, no LLM rewriting history. First real dogfood run — comparing my own API's v3 spec against v2 — it flagged exactly one breaking change. Except that change had shipped a week earlier and nothing broke.
The delta: a status field's enum went from ["queued","shipped","failed"] to ["queued","shipped"]. Textbook breaking change, the differ said. But that field lived in a response schema. A server returning fewer enum values cannot surprise a client that already handles all three — the client's switch statement still compiles. The server promised less variety, not less data.
That's when it clicked: breaking-ness has a direction, and the same textual delta flips meaning depending on which way the schema points.
number to integer — all breaking. Old clients start sending rejected payloads.
The inverse bit me earlier too: adding a value to a response enum means old clients suddenly receive data their validators reject. Same operation, opposite verdict, depending on the boundary side.
Second fix, less glamorous: I now normalize specs before diffing — expand $ref s, sort keys, canonicalize. Without that, two semantically identical specs whose properties happened to be ordered differently produced a wall of fake diffs. JSON comparison is not semantic comparison.
The direction-aware version now gates my own spec releases before anything ships. I eventually packaged it as the OpenAPI Changelog Generator — it takes a base spec URL plus the new spec and returns a human-readable changelog alongside a machine-readable change list.
If you diff specs with a naive property-level tool, check which side of the boundary the changed schema sits on. It flips the verdict more often than you'd expect.