Architecture
What Actually Constrains a Coding Agent
I had fifteen agents build the same Laravel and Slim API. Written decisions mattered most, the framework second, PSRs barely at all, and Laravel already ships the first.
Oct 202612 min read
If agents write most of the code now, is a framework still worth having? The case for one used to be that it saved you typing: the router, the validation, the container, the boilerplate you did not want to write again. An agent writes a router in thirty seconds. That part of the argument is gone.
The case people make instead is that a framework constrains the agent. It makes decisions, so the agent does not have to invent them, and two sessions on Tuesday do not produce two different architectures. I think that is right, but “I think” is not much to build a team’s practices on, so I tested it.
Fifteen fresh Claude Code sessions built the same small API on Laravel and on Slim, under different instructions, and I scored what came back. It cost $15.63, and it turned up something about Laravel I was not looking for.
The experiment
The task: a team projects API. Users belong to teams, teams have projects, and there are three endpoints, to list, create and show a team’s projects. Bearer tokens, members only, errors as RFC 9457 problem details, SQLite, tests. The prompt fixed only what a script needed to score the result (the paths, two seeded users, and that composer setup, composer test and php -S work). Every architectural choice was left to the agent.
Each build started from a clean project and ran as a headless session (claude -p, Opus 5.5), isolated from my own setup and with permissions left on. There were five groups of three:
| Group | Starting point | Agent instructions |
|---|---|---|
| Slim, none | Slim 4.15 with slim/psr7 and PHPUnit, a hello-world route | None |
| Laravel, none | Fresh Laravel 13.35 skeleton | Removed (see below) |
| Laravel, as shipped | Fresh Laravel 13.35 skeleton | Whatever the skeleton ships |
| Laravel, conventions | Fresh Laravel 13.35 skeleton | My 25-line AGENTS.md |
| Slim, conventions | As Slim, none | My 28-line AGENTS.md |
Slim is the stand-in for “PSRs and nothing else”: PSR-7 messages, PSR-15 middleware and PSR-17 factories, with no opinions about where anything lives.
Every build was then probed over HTTP with 19 requests (no token, the wrong team, invalid input four ways, team_id smuggled into the body, someone else’s project, a missing team), checked for plaintext tokens in the database, run through PHPStan at level 8, and diffed. Two read-only agents then classified each build’s internals against fixed questions, citing the files, and I checked the most serious claims over HTTP myself.
The thing I was not looking for
The first Laravel build modified AGENTS.md and CLAUDE.md. I had not given it either file.
The Laravel 13 skeleton ships both. They are identical, 47 lines each, and the first thing they tell an agent to do is install Laravel Boost before touching anything. All three builds in the “as shipped” group did exactly that, and Boost replaced the files with 121 to 148 lines of Laravel’s own guidelines.
The skeleton also requires laravel/pao, described as “Agent-optimized output for PHP testing tools”. It uses laravel/agent-detector to notice that an agent is running (environment variables such as AI_AGENT and CLAUDECODE) and rewrites PHPUnit and PHPStan output into compact JSON:
{"tool":"phpunit","result":"passed","tests":29,"passed":29,"assertions":301,"duration_ms":416}I found that one because my own scoring script, also running under Claude Code, got the JSON instead of the output it expected and broke.
So a fresh Laravel app is not a neutral starting point any more. It ships instructions for agents and reshapes its tools’ output for them. That is why the experiment has a “none” group with those files deleted, and it is worth knowing before you clean up a new project and remove two Markdown files you did not write.
Drift
The question that matters most for a team is whether two sessions given the same task make the same decisions. I picked twelve decision points (authentication, token storage, where validation lives, where authorisation lives, the status a non-member gets, how a project from another team is kept out, where errors are rendered, the shape of validation errors, how responses are built, the shape of the handlers, pagination and persistence) and counted how many all three builds in a group agreed on.
| Group | Decision points agreed |
|---|---|
| Slim, none | 6 of 12 |
| Laravel, none | 9 of 12 |
| Laravel, as shipped | 11 of 12 |
| Laravel, conventions | 12 of 12 |
| Slim, conventions | 12 of 12 |
The three unconstrained Slim builds rendered errors three different ways: a custom handler on Slim’s ErrorMiddleware, problem responses returned directly from the controller, and a PSR-15 middleware catching exceptions. Two shaped validation errors as field => [messages] and the third as a list of {detail, pointer}. One put SQL straight into the controller where the other two wrote repositories.
Laravel without instructions drifted less, and where it drifted it was between things Laravel offers: two builds wrote a Policy and one a Gate::define closure; one used scopeBindings() to keep projects inside their team and two used $team->projects()->findOrFail(). Where the unconstrained builds agreed, they agreed on each stack’s default. Every Laravel build reached for Auth::viaRequest, a FormRequest, an API Resource, one resource controller, and paginate() at 15 a page. Every Slim build wrote one controller class, PSR-15 authentication middleware, and offset pagination at 20 a page.
That is the framework doing what it was supposed to: offering a default, so agents converge on it.
Twenty-five lines
Then the conventions. This is the whole Laravel file:
Conventions for this codebase. Follow them; do not invent alternatives.
## HTTP contract
- Unauthenticated requests get 401. A team the user is not a member of, or a project outside the team in the path, gets **404**, never 403: do not reveal what exists.- Every error is RFC 9457 problem details, `application/problem+json`, with `type`, `title`, `status` and `detail`. Validation failures are 422 and add `errors`: a list of `{ "field": "...", "message": "..." }`.- A single resource is wrapped: `{ "data": { ... } }`. A list is `{ "data": [ ... ], "meta": { "next_cursor": "..." | null } }`.- Lists are cursor paginated: `?limit=` defaults to 20, maximum 100; `?cursor=` takes `meta.next_cursor`. Newest first.- JSON keys are snake_case. A project is `id`, `team_id`, `name`, `description`, `created_at` (ISO 8601, UTC).
## Where things live
- **Authentication:** a request guard registered with `Auth::viaRequest('token', ...)` in `AppServiceProvider`, used as `auth:token`. Tokens are stored only as a SHA-256 hash in `users.api_token_hash`; never store or log the plain token.- **Authorisation:** a `ProjectPolicy`, and routes use scoped implicit binding (`->scopeBindings()`), so a project outside the team is a 404.- **Validation:** a `FormRequest` per write endpoint. Never validate in a controller.- **Responses:** an API Resource (`ProjectResource`).- **Errors:** rendered in one place, `withExceptions` in `bootstrap/app.php`.- **Controllers:** one `ProjectController` with `index`, `store` and `show`. Thin: no queries beyond what binding provides, no validation, no authorisation logic.- **Mass assignment:** `team_id` comes from the route, never from the request body.
## Tests
Feature tests over HTTP for every endpoint, including the unauthenticated, non-member and validation cases.The Slim version makes the same decisions in Slim’s terms: PSR-15 middleware for authentication and membership, an input class per write endpoint, one invokable action per endpoint, PDO repositories, a presenter, and one error handler on ErrorMiddleware.
Both groups went to twelve out of twelve. The classifier could not find a single deviation from the file in any of the six builds, apart from one Slim action checking that a route argument was numeric itself.
More telling is that they followed it where it went against their own defaults. Every unconstrained build, all nine of them, returned 403 to non-members. Every build with the file returned 404. Every unconstrained build used page numbers; every build with the file used cursors, newest first. The stack stopped mattering to drift the moment the decisions were written down.
Where the hypothesis was wrong
The other half of the case for frameworks is security: the framework remembers the things an agent forgets, such as authorisation, mass assignment, hashing and validation. If that held, the unconstrained Slim builds would show it.
They did not. Every one of the fifteen builds sent 401 problem details without a token and 422 problem details for every kind of invalid input. Every one refused cross-team reads, ignored team_id and id in the request body, and stored tokens as SHA-256 hashes, never in plain text. On a task this size, in October 2026, the agents did not need a framework to remember the basics.
What they got wrong was subtler, and it was in both stacks.
Eight of the nine builds without the 404 rule let anyone enumerate teams. They answer 403 for a team you are not in and 404 for a team that does not exist, so any user with a token can walk the ids and learn which teams exist. All six unconstrained Laravel builds do it, and two of the three Slim ones. Three of the Laravel builds go further: they resolve the project before checking membership, so a non-member gets 404 for a project id that does not exist and 403 for one that does.
One build misread the framework. It wrote error rendering for AuthorizationException and ModelNotFoundException, and neither branch ever runs, because Laravel’s exception handler converts both into HTTP exceptions before render callbacks see them. The fallback it ended up in passes the framework’s message straight through:
{"type":"about:blank","title":"Not Found","status":404,"detail":"No query results for model [App\\Models\\Project] 999999","instance":"/teams/1/projects/999999"}Another build hit the same behaviour and worked around it with getPrevious(). That is the one place in fifteen builds where framework magic, behaviour you cannot see from the code you are writing, cost an agent something.
None of the six builds with conventions leaked anything, because the rule “404, never 403: do not reveal what exists” removes the difference an attacker would measure. One line of written decision did what the framework did not.
This one is worth checking in your own API, agent or no agent. Request a team that does not exist and one you are not a member of, and compare the status codes. It takes two requests, and in fifteen builds only one avoided the leak without being told to.
Where PSRs fit
PSRs held the shape and decided nothing. No Slim build added a package. No Slim build made the classic PSR-7 mistake of calling $response->withHeader() and throwing the result away, which the scoring checked for twice, by search and by reading. And unconstrained Slim drifted more than anything else in the experiment.
That is what PSRs are for. They define what a request, a response and a middleware look like, so code written against them works anywhere. They do not say where validation lives or what a non-member sees. Use them at the boundaries, in SDKs, HTTP clients and middleware you want to move between projects, and do not expect them to make an architecture.
What it costs
Laravel needed less code: about 770 to 810 lines across the three Laravel groups, against 1,008 for Slim without conventions and 1,289 with them. The conventions made Slim bigger, because they asked for the layers Laravel already has: input classes, presenters, repositories, one action per endpoint. But every Slim build with conventions had the same 29 files in the same places, so a reviewer knows where everything is before opening one.
The whole batch cost $15.63 and took between two and a half and five minutes a build. That is cheap enough to run before you commit to a convention, which is the other thing I took from this: you can test your AGENTS.md the way you test code.
What this does not show
It is small. Three runs per group, one feature of about 800 lines, one model. The twelve decision points are mine, and choosing them more coarsely would raise every group’s score. I would not quote any of these numbers as a rate. What I would quote is the direction, because it held across both stacks: written decisions, then framework conventions, then interfaces.
And the security result is for a task this size. A feature with roles, background jobs and a second resource is where I would expect agents to start forgetting things, and that is the next run.
What I would do next
I would write down the decisions the framework does not make for me, before handing an agent anything. The status a non-member gets, the shape of the errors, how lists are paginated: those are exactly where the agents disagreed with each other, and once they were written down, every build followed them. If you are not sure how to write them down, how to write specs, RFCs and ADRs goes through it from a blank page.
On Laravel 13, I would read the AGENTS.md the skeleton shipped before deleting it. It is doing more than it looks.
Related
More in Architecture
The Second Pattern: The Ones I Left Out
Five patterns I know well, that come up constantly, and would not reach for in PHP. Not because they are bad ideas, but because of what the runtime does and does not give you.
Sept 2026 · 12 min read
ArchitectureThe Second Pattern: Blackboard
A pipeline works until one step both needs and improves the same piece of information. That is a cycle, and a topological sort has exactly one contract: there are no cycles.
Sept 2026 · 12 min read
ArchitectureThe Second Pattern: Event-Carried State Transfer
A consumer that receives an ID and immediately asks you for the record has not been decoupled from you. It has been given a slightly slower way to call your API.
Sept 2026 · 12 min read
Same subject