The Petstore scores 65
The most copied OpenAPI document in the world is a bad API for an agent, and it takes a third of a second to find out. So I built the check I was doing by hand.
The Swagger Petstore is the most copied OpenAPI document in the world. It is the example in the tutorial, the fixture in the test suite, the thing you paste into a new tool to see whether it works. Here is what it looks like when you score it for the consumer everyone is now building for.
$ legible check petstore.json
PUT /user/{username} 1:13250 error arguments collide when flattened into one tool input: "username" (path and body) argument-collision 1:13250 warning no security requirement, here or at the top level: an agent cannot tell whether this needs credentials, or which security-declared 1:13315 warning described identically to 2 other operations (POST /user, PUT /user/{username}, DELETE /user/{username}): an agent cannot tell them apart duplicate-description
legible petstore.json, OpenAPI 3.0.4, 19 operations
descriptions 62 ████████████░░░░░░░░ naming 100 ████████████████████ errors 50 ██████████░░░░░░░░░░ examples 32 ██████░░░░░░░░░░░░░░ tool-schema 99 ████████████████████ safety 47 █████████░░░░░░░░░░░
score 65 agents will struggleThree operations described identically. An argument that collides with itself the moment it is flattened. No declared security, anywhere. None of that is a bug, and a validator will tell you the document is perfectly fine, because it is.
I wrote last week about what your endpoint looks like to an agent: not a set of resources and verbs, but a flattened projection where path parameters and body properties land in one argument object and the description is the only thing deciding whether your operation gets picked at all. That article ended with a list of things to change, which is easy to write and tedious to apply. I have been doing that check by hand for a while now, one engagement at a time, and the problem with being the linter yourself is that you only run when someone books you.
So legible is that check as a binary.
Why a score rather than a pass
Linters give you a list. A list of forty findings on a four-hundred operation spec tells you almost nothing, because you cannot tell whether that is a disaster or a rounding error.
Every rule here records each place it looked, not only the places it failed. Three undocumented parameters out of two hundred is a good spec with a small gap. Three out of four is a different conversation entirely, and the same list of three findings describes both. The score is passes over checks, which is what stops a large spec being punished for being large.
That decision cost more than it sounds. It means a rule cannot just scan for problems and report them; it has to walk everything in its remit and account for what it saw. It is also the only reason the bars above mean anything.
What it checks, and why those things
Twenty-three rules in six categories, and the ones that matter most are the ones that only exist because of the flattening.
argument-collision is the clearest example. Look at what the Petstore does:
put: operationId: updateUser parameters: - name: username in: path requestBody: content: application/json: schema: properties: username: { type: string }Two different username values, and HTTP has no trouble with that: one is in the path and one is in the body. Flatten it into a single argument object and one of them wins. Which one? Neither reliably, and the agent has no way to express what it meant.
The fix is usually deletion rather than renaming. The path already identifies the user, so the body does not need to. schema-depth and parameter-count measure your widest and deepest operations against the limits OpenAI and Anthropic publish for tool definitions. recursive-schema finds the structures that cannot be expressed as a tool input at all.
The rest is older advice with a new consequence. description-restates-name catches an operation whose description says “Updates a user” on an operation called updateUser, which was always a missed opportunity and is now close to invisible under deferred tool loading. error-shape wants your failures to look the same as each other. duplicate-description is how the Petstore loses most of its points.
Each rule has a page covering what it checks, why it trips an agent up, how to fix it, and when to ignore it. Those pages are built into the binary, so legible explain argument-collision works with no network.
Three things it deliberately is not
Not a validator. Run one of those as well. legible assumes the document is valid, because a spec can be perfectly valid and still be a bad API.
Not a test of the live API. Everything is read from the document. It never calls your service and it never sends your spec anywhere.
Not a constraint inventor. No rule asks you to add a pattern or a maxLength to data that does not have one. A constraint added to please a linter rejects valid input in production, and I would rather score you lower than break you.
The decision that shaped the whole thing
I tried libopenapi first and dropped it on day one.
Its high-level model is lovely for reading a spec and wrong for linting one. It hides $ref behind lazily resolved proxies, and its line numbers sit behind generic wrappers several layers deep. A linter needs the opposite of convenience: every node’s exact position, the difference between a field that is absent and one that is empty, and enough control over resolution that recursion stays visible rather than being quietly followed forever.
So it parses plain YAML nodes instead. JSON parses as YAML, so one parser covers both formats, and every finding can point at a line and a column. It checks Stripe’s eight megabyte, 594 operation spec in about a third of a second.
Using it
One static binary, no account, and your spec stays on your machine.
curl -fsSL https://raw.githubusercontent.com/JustSteveKing/legible/main/install.sh | shlegible check openapi.yamlIt follows $refs across files and URLs and caches them in .legible/, and --offline turns the fetching off entirely if you would rather it touched nothing. In CI, --fail-under 80 is the whole integration.
What I would fix first
If your score is low and you have an afternoon, do these three, because they are the cheapest points on the board. Give every operation a description that says something its name does not. Stop repeating an identifier between the path and the body. Declare your security, even if the answer is that there is none.
Then the obvious question: what did it say about mine?
The API I am building now scores 99. The one before it scores 77. That gap is not effort, it is that I designed the first one after learning what the second one taught me, which is the only honest way anyone gets a good score.
Even at 99 it found something. Four properties are camelCase in an API whose other thirty-nine are snake_case, all four in the same block, all four written on a day I was clearly not paying attention. No human consumer would have filed a bug about that. A model picking between keyLength and key_prefix has no idea which convention is real, and will guess.
That is the whole argument for running it, and it is why the score is a starting point rather than a grade. A spec at 99 still has four places where a machine will guess. Go and find yours.
legible lives here and on GitHub under MIT. If you run it and disagree with a rule, open an issue, because a rule nobody argues with is usually a rule nobody ran.
Keep Reading
Designing CLIs people want to use
API design is a user experience problem and nobody argues any more. Then we build a CLI and forget all of it. Cobra defaults, which ones are wrong, and what I do instead.
Sept 2026 · 22 min read
Shipping & ToolingBuilding Research
A desktop research workspace in Laravel and NativePHP. Streaming SSE into a queued job, distilling reports with a local model, and why cosine similarity cannot tell a paraphrase from a contradiction.
Aug 2026 · 22 min read
Shipping & ToolingControlling Code Quality When an Agent Writes Your Laravel
Specs, ADRs and path-scoped rules. The three layers of context I build around an agent so it writes Laravel the way my codebase does, not the way every tutorial does.
Aug 2026 · 14 min read