Skip to main content
WritingSpeakingCodeAboutNow

Building a home-education platform with Laravel 13, Inertia and React, and shipping it to Laravel Cloud

A Laravel 13, Inertia and React platform for home educating my son: an append-only learning record, a planner, AI kept on a short lead, and what Laravel Cloud asked of the code.

My son was about to start secondary school, and he’d been showing a real interest in tech, code in particular. When I looked at where that route would take him, the so-called computer science he’d cover was a lot less than I’d hoped. He’s also very much like his mother in that he doesn’t enjoy social situations or environments. After a long talk we offered home schooling as an option, and he warmed to it quicker than I expected. So I built a platform to run it on.

Parents set up a household, add their children and enrol each child onto curricula, and the app turns those enrolments into a daily schedule. Children work through their day, everything they finish lands in an append-only learning record, and parents produce reports from that record. A local authority can ask to see a child’s record, and the parent approves a time-limited, logged grant.

It’s a Laravel 13 app with an Inertia and React frontend, about 430 commits over two months, running on Laravel Cloud. I’m not going to tour all 113 migrations. This is the parts where I had to make a real decision, followed by what Cloud asked of the code on the way to production.

The stack

LayerPackageVersion
Frameworklaravel/framework13.33
RuntimePHP8.5
Serverlaravel/octane2.20
Frontend bridgeinertiajs/inertia-laravel, @inertiajs/react3.3, 3.6
UIReact, Tailwind CSS19.2, 4.3
Routes in TypeScriptlaravel/wayfinder0.1
Authlaravel/fortify (2FA and passkeys)1.40
Feature flagslaravel/pennant1.26
Realtimelaravel/reverb, laravel-echo1.12, 2.4
AIlaravel/ai0.10
Workflowsjuststeveking/workflow-engine1.1
Monitoringlaravel/nightwatch1.30
TestsPest with the browser plugin5.2

Production runs on Postgres. Local development started on SQLite, and the test suite still runs on both.

How the code is organised

If you’ve read much Laravel code from the last few years, app/ will look familiar.

  • Actions/ holds the domain logic as single-purpose classes, grouped by area: Schedule, Planner, Households, Oak, Topics and so on. Most expose handle().
  • Support/ holds plain domain services and value objects: LearningCalendar, Timetable, EnrolmentPlanner, CorePlan, EvidenceFile.
  • Services/Oak/ is a full API client for Oak National Academy, with typed DTOs, typed exceptions and an endpoint enum.
  • Ai/ holds the agents and the guardrails around them.
  • Workflows/ holds the multi-step research flows that sit on top of the workflow engine.
  • Controllers are thin. They resolve the user, authorise, validate, call an action and return an Inertia response or a redirect with a flash.

The opinion that shaped the codebase most is that every question the app answers has exactly one owner. Learning days come only from LearningCalendar. Curriculum content is written only by SyncCurriculumContent. The learning record is written only by RecordLearningEvent and read only by LearningRecord. A local authority’s view of a child comes only from BuildChildRecord, and that same class builds the page, the printable record and the CSV export.

That sounds like ordinary separation of concerns, and it is. I wrote it down as a rule anyway, because when the Progress page, the CSV and the planner each work out “how many lessons has this child done” for themselves, they eventually disagree. Then a parent sends their local authority a report that contradicts the screen they just looked at. One owner per question means one place to fix it.

The rules live in .ai/rules/, an index of glob patterns mapped to rule files. They started life as context for the AI agents I used while building (more on that later), but they read just as well for a human joining the project. domain-sources.md is the one I’d hand anyone first.

Households, and the URL shape that comes with them

Nearly every route in the app is scoped to a household:

/{current_household}/today
/{current_household}/schedule
/{current_household}/planner

EnsureHouseholdMembership resolves the slug against the signed-in user’s memberships in one query, and takes an argument for parent-only routes:

Route::prefix('{current_household}')
->middleware(['auth', 'verified', EnsureHouseholdMembership::class.':parent'])
->group(function () {
Route::get('activities', [ActivityController::class, 'index'])->name('activities.index');
// planner, enrolments, reports...
});

Writing the slug into every route() call would be tedious, so a middleware sets it as a URL default for the request:

URL::defaults([
'current_household' => $currentHousehold->slug,
'household' => $currentHousehold->slug,
]);

Under Octane this line is more dangerous than it looks. Octane boots the application once and reuses it across requests, so anything held in a static or a singleton outlives the request that set it. A URL default left over from the previous request would put one family’s slug into another family’s links. Octane does hand each request its own clone of the URL generator, and there’s a test that fails if the middleware ever skips writing the default because one is “already set”. That test is the guard for the whole Octane audit I did before deploying, which I come back to in the deployment section.

On the frontend, Wayfinder generates a typed function for every named route, and household-scoped routes take the slug explicitly:

import { index as scheduleIndex } from '@/routes/schedule';
// ...
href: scheduleIndex(household.slug),

From enrolments to a day of learning

A parent enrols a child onto a curriculum for an academic year at a pace, say three sessions a week. A curriculum is made of units, units are made of lessons, and each lesson has an estimated number of sessions. The planner takes every enrolment for a child and lays out the whole year as ScheduleBlock rows, one per session, on the days that child actually learns.

Which days count

LearningCalendar::learningDates() is the only place that answers this. A learning date is a term day, on one of the child’s learning weekdays, outside the half-term break, and not a household holiday. Nothing else in the app gets to decide whether Tuesday is a school day.

Pace mode

Without a timetable, EnrolmentPlanner works week by week. For each enrolment it works out how many sessions this week should get, prorating partial weeks (the first week of term might only have two learning days) and carrying the rounding error into the following week:

$dailyCap = max(1, (int) ceil($pace / $learningWeekdays));
$allowance = $pace * count($dates) / $learningWeekdays + ($carry[$enrolment->id] ?? 0);
$target = min((int) round($allowance), $dailyCap * count($dates));

Then it places sessions one at a time on the least-loaded date still under the daily cap, earliest date on a tie. Whatever is left over at the end of the week carries forward, clamped to plus or minus one:

$carry[$enrolment->id] = max(-1, min(1, $allowance - $inWeek));

Without the daily cap, a short week stacks all of its lessons onto one day. The clamp fixes a slower problem: a long run of short weeks builds up a debt, and eventually it lands as a wall of maths in a single week.

Generation is idempotent. Each lesson expands into estimated_sessions minus the sessions already placed for it, so running the planner twice places nothing the second time. Every run stamps its blocks with a plan_run_id, which lets a parent undo a run they didn’t like, and the planner only ever removes blocks still in the planned state. Finished work is never touched.

Timetable mode

Pace mode is fine until a family wants maths on Monday and Wednesday mornings specifically. A parent can bind a week template to a child, and each slot in the template is a fixed activity, a subject area, or a choice slot where the child picks what to do on the day.

allocateFromTimetable() walks the dates and, for each date, the slots in weekday order. A subject slot rotates round-robin through every enrolment in that subject area. The cursor tracking whose turn it is has to carry across dates rather than reset each day, or the first enrolment in the area takes every slot and the others never get a look in.

Two edge cases took more thought than the main loop.

When a curriculum runs out before the year does. If a child is working faster than the curriculum’s estimates, the slot’s queue empties in, say, May. Early versions quietly filled the slot with the next subject’s work, which is wrong. It was a science slot and the family planned it as science. Now an empty slot past the last planned lesson of every curriculum behind it becomes an open block of the same subject:

$ranOut = collect($ring)->every(fn (Enrolment $enrolment): bool => $key > max(
$lastPlanned[$enrolment->curriculum_id] ?? '',
$report->lastPlacedOn[$enrolment->id] ?? '',
));

A daily scheduled command, learning:notify-run-outs, reuses the same planner draft to work out when each enrolment runs out and tells the parents a fortnight in advance. It uses the same draft() as the planner preview, so the notice and the preview can’t disagree.

What the weekly target means. The planner grid shows each subject area against a weekly target. In pace mode that’s the sum of sessions_per_week across the area’s enrolments. In timetable mode it has to come from the timetable instead, or a child with three science slots and one science enrolment set at a session a week shows as “3/1”, a number that means nothing:

'target' => $timetable !== null
? Timetable::slotsForArea($timetable, (int) $area->id, $weekdays)
: (int) $enrolments->sum('sessions_per_week'),

That fix was the last commit before I started writing this. Planning bugs in this app don’t throw. They show up as a number on a page that a parent looks at and frowns.

Moving a lesson forward

A child who finishes their day early can pull the next lesson forward. Every later lesson in that subject then shuffles back a slot, each one taking the day and position of the one in front of it. Done naively that’s one UPDATE per block. PullLessonForward does it in one statement with a CASE per column, and a test pins the query count.

If you write this kind of query yourself, watch the date column. It needs a ::date cast on Postgres, and the same cast written as CAST('2026-09-22' AS date) on SQLite returns 2026, because SQLite has no date type and treats the cast as numeric. The action builds the cast per driver.

A learning record you can trust

Plans are meant to change. A parent edits an activity, archives a lesson, regenerates the term. The record of what a child actually did must not change with them, because that’s what goes to the local authority.

The record is an append-only learning_events table, enforced at three separate layers.

The application. RecordLearningEvent is the only writer. It refuses any payload carrying private writing, at any depth:

private const array FORBIDDEN_KEYS = ['body', 'note', 'journal'];
// ...
if (is_string($key) && in_array($key, self::FORBIDDEN_KEYS, true)) {
throw new InvalidArgumentException("Learning event payloads must not contain a [{$key}] key.");
}

The model. LearningEvent::booted() throws a LogicException on updating and on deleting.

The database. A migration installs triggers per driver. On Postgres the trigger refuses deletes while the household and child still exist, so erasing a household still cascades. It refuses every update except the one Postgres makes itself when a user is deleted and actor_user_id is set to null:

IF NEW.actor_user_id IS NULL
AND OLD.actor_user_id IS NOT NULL
AND NOT EXISTS (SELECT 1 FROM users WHERE id = OLD.actor_user_id)
AND (to_jsonb(NEW) - 'actor_user_id') = (to_jsonb(OLD) - 'actor_user_id') THEN
RETURN NEW;
END IF;
RAISE EXCEPTION '%s';

The to_jsonb(NEW) - 'actor_user_id' comparison says “every column except this one is unchanged” without listing the columns, so adding a column to the table later doesn’t quietly open a hole in the trigger.

Each layer catches a different mistake. The action stops a developer passing the wrong array, and the model stops someone reaching for $event->update() in a hurry. The trigger is what catches a raw query, a tinker session, or a migration nobody has written yet. Most of the time any one of them would be enough. I kept all three because this is the table a local authority ends up reading.

Finished blocks also keep a completed_snapshot of what the child actually did. Today, the schedule, the planner and the parent review all render finished work from that snapshot, so editing an activity next month doesn’t rewrite what it looked like when the child did it.

Oak National Academy as the only content source

All curriculum content comes from Oak National Academy’s API. That wasn’t the plan at the start. Early versions linked out to BBC Bitesize and IXL and could generate whole curricula with AI. On 21 September I removed all of it, purged the AI-generated curricula with a data migration, and added a test that fails if a Bitesize or IXL link ever comes back. A local authority reading a child’s record should be able to trace every lesson to a published, attributed source.

OakClient is the only thing in the app that calls Oak. Every operation is a case on an OakEndpoint enum, and a contract test checks those against a copy of Oak’s API spec vendored into resources/oak. Assets reach the browser only through a proxy controller, so the API key never leaves the server, and quiz answers are marked on the server so the answer key never reaches a child’s browser.

The hard constraint is Oak’s rate limit of 1,000 requests an hour. Importing one programme costs roughly 120 requests, and the full curriculum is around 12,000.

The first answer to that is a backlog and an hourly oak:top-up command, which reads Oak’s remaining quota and dispatches as many queued imports as will fit.

The bigger one is a record and replay archive. With OAK_ARCHIVE_MODE=record, every Oak response is written to a SQLite file as it’s fetched. That file is committed to the repository as a streamed, gzipped export of about 9 MB. With OAK_ARCHIVE_MODE=replay, OakClient answers from the archive first and only goes to Oak on a miss. Production runs in replay mode, so a fresh install can import the whole curriculum without spending any of its quota. It’s also why composer.json requires ext-pdo_sqlite on a server that otherwise runs Postgres.

Replay mode had a bug that only production found. While replaying, the client reported a made-up budget of 100,000 requests, on the reasoning that replayed imports cost nothing. That only holds for programmes the archive covers in full. Anything else fell through to live calls, and the top-up dispatched the whole backlog at once against the made-up figure. On 29 September, 208 imports expired together after six hours of releasing themselves back onto the queue. The fix costs each waiting import individually: zero when the archive answers the whole thing, a real quota reading otherwise. The top-up is now also switched off unless OAK_TOP_UP_ENABLED is set, since the curricula in use are already imported.

AI features, and how they’re kept on a short lead

The app has three AI features: researching lesson activities for a topic a child wants to study, generating lessons from a textbook, and drafting a career roadmap for an older child. All three use the laravel/ai SDK with Anthropic as the provider, and all three return structured output.

The agents

Each agent is a class with instructions and a JSON schema:

#[Provider(Lab::Anthropic)]
class LessonPlanResearcher implements Agent, HasStructuredOutput
{
use ConfiguresResearchAgent;
use Promptable;
public function schema(JsonSchema $schema): array
{
return [
'activities' => $schema->array()
->items(
$schema->object(fn (JsonSchema $schema) => [
'title' => $schema->string()->required(),
'description' => $schema->string()->required(),
'duration_minutes' => $schema->integer()->required(),
'description_ids' => $schema->array()->items($schema->integer())->required(),
'steps' => $schema->array()->items(/* learning_layer, title, body */)->required(),
])
)
->required(),
];
}
}

The shared ConfiguresResearchAgent trait reads the model, HTTP timeout and token cap from config/research.php, so changing model is an environment variable rather than a deploy. That works because the SDK checks for model(), timeout() and maxTokens() methods first and only falls back to the equivalent attributes when a method is missing or returns null.

Treating the child’s text as data

Every AI prompt here contains text the app didn’t write: a child’s description of what they want to learn, or a book title from Open Library. That text goes through UntrustedText::tag() before it reaches a prompt:

public static function tag(string $tag, string $text, int $limit = 500): string
{
$clean = Str::of(self::redact($text))
->replace(['<', '>'], '')
->squish()
->limit($limit)
->toString();
return "<{$tag}>{$clean}</{$tag}>";
}

redact() replaces email addresses, URLs and anything that looks like a phone number. Stripping the angle brackets means the text can’t close its own tag and open a new one. Every agent’s instructions end with the same block, telling the model that anything inside <child_topic>, <book_title> and friends is data to describe, never instructions to follow.

The prompt carries the child’s age, never their name or date of birth. Output goes through ResearchOutput, which drops items with no title, caps every string, limits steps per activity and throws if nothing usable came back, so a bad response fails the job and retries instead of saving an empty plan. When the model returns curriculum ids, the action intersects them with the ids it actually offered, and a hallucinated id never reaches the database.

Running the call: workflows, jobs and nested timeouts

A topic request goes past a parent before it costs anything. The topic workflow runs: await review, route the parent’s decision, research, confirm, finalise. The research step doesn’t call the model itself. It claims the topic, dispatches a queued job after the transaction commits, and parks the workflow until the job signals back:

if (! $topic->aiFeatureEnabled()) {
return StepResult::fail($topic->aiFeature()->refusalMessage());
}
if (! ResearchRun::claim($topic)) {
return StepResult::fail('The topic cannot be researched right now.');
}
ResearchTopic::dispatch($topic, $context->workflowInstanceId)->afterCommit();
return StepResult::await(ResearchRun::SIGNAL);

The workflow engine runs steps inside locked transactions, and holding a database transaction open across an HTTP call to a model is a reliable way to make everything else wait. So the AI call always happens in a job, outside any transaction, and the job reports back with a signal.

The timeouts nest deliberately:

'timeouts' => [
'agent' => (int) env('RESEARCH_AI_TIMEOUT', 150),
'job' => (int) env('RESEARCH_JOB_TIMEOUT', 180),
'workflow' => (int) env('RESEARCH_WORKFLOW_TIMEOUT', 1200),
],

And in config/queue.php:

// Must exceed the longest job timeout (research jobs run for up to
// config('research.timeouts.job') seconds), or a still-running job
// is handed to a second worker and the AI call is paid for twice.
'retry_after' => (int) env('DB_QUEUE_RETRY_AFTER', 240),

150 inside 180 inside 240. If anything still slips through, a scheduled job sweeps anything stuck in “researching” for more than 30 minutes back to its previous state.

Feature flags, checked four times

Every AI feature sits behind a Pennant flag scoped to the household, off by default, plus a master switch:

foreach (AiFeature::cases() as $feature) {
Feature::define($feature->value, fn (Household $household) => false);
}

The master switch is applied in AiFeature::enabledFor() and not inside a Pennant resolver. Pennant stores a resolver’s answer the first time it runs, so a master switch baked into a resolver would be frozen at whatever it was when each household was first checked.

public function enabledFor(?Household $household): bool
{
if ($household === null) {
return false;
}
if ($this === self::Master) {
return $this->storedFor($household);
}
return self::Master->storedFor($household) && $this->storedFor($household);
}

The flag is then checked at four points, each closing a different gap:

  1. The route, through EnsureAiFeature:topics middleware, which returns a 403. The page hides the button when the flag is off, but an old tab or a bookmarked form can still post.
  2. The workflow step, before it claims the subject. A topic approved while the feature was on mustn’t reach the model once a parent has switched it off.
  3. The job, again, as the very last thing before the model. A job can sit on the queue for a while. If the flag is now off, the job puts the subject back and reports the run as discarded, which ends the workflow instead of triggering a retry.
  4. The UI, which reads a shared aiFeatures Inertia prop and replaces the controls with a notice.

One test file exists purely to prove the app works with every flag off: Oak imports, the lesson viewer, the daily pages, child navigation and the stalled-research sweep. No scheduled command is allowed to read an AI flag.

Waiting for the result

There’s no streaming. The research takes 20 to 60 seconds and the result is a structured plan, not prose you’d want to watch arrive. The button posts with an optimistic update that flips the item to “researching”, and a small hook polls only the relevant prop until nothing is researching any more:

export function usePollWhile(active: boolean, only: string[], intervalMs = 4000): void {
const props = only.join(',');
useEffect(() => {
if (!active) {
return;
}
const poll = router.poll(intervalMs, { only: props.split(',') });
return () => poll.destroy();
}, [active, props, intervalMs]);
}
usePollWhile(textbooks.some((t) => t.status === 'researching'), ['textbooks']);

Reverb is set up and the bell notifications broadcast over it, so a websocket push would have worked too. Polling one prop every four seconds for a minute was simpler, and it can’t miss an event.

Testing it without calling a model

The SDK’s agent fakes turn the AI paths into ordinary feature tests. phpunit.xml runs the queue synchronously, so one POST drives the workflow, the job and the save:

LessonPlanResearcher::fake([['activities' => [['title' => 'A plan', 'steps' => []]]]]);
// ...post to the research route...
LessonPlanResearcher::assertPrompted(fn ($prompt) => $prompt->contains('<child_topic>Astronomy</child_topic>')
&& ! $prompt->contains('kid@example.com'));

For every gate, the assertion that matters most is the negative one:

LessonPlanResearcher::assertNotPrompted(fn () => true);

Privacy decisions that shaped the code

A child’s journal is readable by the child and their parents and nobody else. It never goes into an AI prompt, a local authority view, a CSV, a report, the data export, Inertia shared props or the learning record, and the FORBIDDEN_KEYS check above is part of how that holds. Wellbeing check-in notes get the same treatment. Anything outward-facing reads check-ins through BuildWellbeingSummary, which selects an explicit list of shareable columns and nothing else.

The admin area is aggregate-only. A platform admin sees household counts, sizes and ages, never a household’s name, slug or any member’s details. The test for it seeds “needle” values (a household name, emails, a date of birth) and asserts none of them appear in any admin response, so a new admin page that leaks something fails without anyone having to remember to write a test for it.

A local authority sees a household only through a grant the parent approves, for 7, 30 or 90 days. Every view, export, print, approval and expiry goes into the parent’s access history, and a request for a household with no grant returns 404, so slugs can’t be enumerated.

Dates of birth never leave the server. Only a derived age is serialised.

The frontend

The frontend is Inertia v3 with React 19, Tailwind 4 and the React Compiler.

Layouts are resolved in one place, app.tsx, from the page name:

layout: (name) => {
switch (true) {
case name === 'welcome':
return null;
case name.startsWith('auth/'):
return AuthLayout;
// Settings replaces the app shell rather than nesting inside it.
case name.startsWith('settings/'):
case name.startsWith('households/'):
return SettingsLayout;
// The printable household record is a document, not a screen:
// no sidebar, no page bar, nothing to hide before printing.
case name === 'local-authority/household-print':
case name === 'reports/print':
return null;
case name.startsWith('local-authority/'):
return LocalAuthorityLayout;
case name.startsWith('admin/'):
return AdminLayout;
default:
return AppLayout;
}
},

Pages pass their title and breadcrumbs to the layout as layout props, so the page bar renders exactly one h1, and a browser test asserts that across every page.

CRUD is meant to feel instant. Inertia v3’s optimistic visits do most of the work, with automatic rollback if the server refuses:

const optimisticBlock = (changes: Partial<ScheduleBlock>) => (props: Record<string, unknown>) => {
if (!Array.isArray(props.blocks)) {
return {};
}
return {
blocks: (props.blocks as ScheduleBlock[]).map((item) =>
item.id === block.id ? { ...item, ...changes } : item,
),
};
};

Child-facing pages are pitched by age band with younger: and older: Tailwind variants, and touch targets are 44 pixels.

Testing

The suite is Pest throughout, organised by domain under tests/Feature, with Playwright-backed browser tests in tests/Browser. phpunit.xml runs on in-memory SQLite with a synchronous queue, Oak switched off and Nightwatch disabled. fakeOak() turns on Http::preventStrayRequests() and serves fixtures, so no test can reach the real API by accident.

CI runs two jobs on every push:

  • SQLite: ESLint, Prettier, tsc, Pint, PHPStan at level 7, then the full Pest suite.
  • Postgres: a postgres:18 service, migrate:reset followed by migrate to prove the migrations roll both ways, then the suite again.

SQLite is the forgiving one of the pair. It ignores varchar lengths, LIKE case rules and JSON column types, so a query like where('col->key') passes on SQLite and fails on a Postgres text column. New JSON-shaped columns use json(), never text(), and test data has to fit the declared lengths. SQLite needed a fix of its own as well. Date-only columns use a custom DateOnly cast, because SQLite stores dates as text and a whereBetween was dropping the last day of a range.

Performance-sensitive paths have query-count tests. The planner went from 1,339 queries to 22, and a test keeps it there.

The browser tests aren’t in CI. They run locally with composer test:browser, and a pre-push hook runs the same checks CI does.

How it was built

The first commit is 31 July 2026: a fresh Laravel app, Pest, Laravel Boost, and renaming the starter kit’s Team concept to Household. By the end of that first day there were schedule blocks, activities, curriculum tagging and a parent dashboard.

The rough shape after that:

  • 31 July to 2 August. A fast prototype: AI topic research, a first MCP server, local authority access by PIN, evidence uploads, careers, an auto-fill planner.
  • 4 August. Logic pulled out into actions, and the research features moved onto the workflow engine.
  • 18 September. A redesign, and wellbeing records.
  • 19 September. 156 commits in one day. I wrote up every domain in docs/ with an “improvement areas” list, then worked through it. Local authority PINs were replaced by request and approval, the learning record became append-only with triggers, Postgres joined CI, and the Oak client landed.
  • 21 September. Oak became the only content source. AI flags, record and replay, db:copy, timetable-driven planning.
  • 22 and 23 September. Launch preparation and the first deploy.
  • Since then. Features driven by actually using it: choice slots, public holidays, run-out notices, replanning from today.

Most of it was written with Claude Code, using Laravel Boost’s MCP server for docs search and database inspection, and the .ai/rules files to carry decisions from one session into the next. The commit messages are long on purpose. Each one explains why, which makes git log the best documentation the project has.

Deploying to Laravel Cloud

Before the first deploy

The app had never been deployed until 23 September. I spent the day before removing things.

Passport, two MCP servers, the OAuth routes and the local authority API tokens all went. None of it was needed to launch, and /oauth/register was a public, unthrottled write endpoint. The MCP server had a quieter problem. laravel/mcp sat in require-dev, and a production install runs composer install --no-dev, so the package would have vanished and taken its routes with it, with no error at deploy time. Check this one in your own app: anything in require-dev that production code depends on disappears on deploy.

The demo seeder went too. Registration is invite-only and nothing else grants admin, so a fresh install would have had nobody able to sign in. ProductionSeeder replaces it, creating the first household and administrator from config rather than committed literals, and it’s safe to run twice.

composer.json had never declared the extensions the app needs, so it now requires ext-gd, ext-pdo_pgsql and ext-pdo_sqlite. Without GD, thumbnails fail silently and every evidence photo is served full size forever. Bandwidth would have been the only symptom.

Last came .env.production.example. It carries every variable whose absence would break a deploy quietly, each with a comment saying why:

# The framework default is "prefer", which silently falls back to an
# unencrypted connection.
DB_SSLMODE=require
# Server-side rendering needs its own always-on process. The app is entirely
# behind a login, so SSR buys nothing here.
INERTIA_SSR_ENABLED=false

The environment

The Cloud application is connected to the GitHub repository with push-to-deploy on main. The production environment runs PHP 8.5 and Node 24 in eu-west-2, with Octane and hibernation switched on.

Attached resources:

ResourceWhat it isWhy
DatabaseServerless Postgres, 0.25 compute units, suspends after five minutes idleMatches CI; the app is quiet most of the day
Object storageA private bucket (Cloudflare R2 underneath), disk named s3Evidence photos and PDFs
WebSocketsA Reverb clusterBell notifications
ComputeOne instance with the scheduler enabledRuns routes/console.php
Background processphp artisan queue:work database --sleep=10 --quietAI research, Oak imports, thumbnails

There’s no cache resource. Cache, sessions and the queue all run on the database. The site serves one family and their local authority, and a cache server would be one more thing to pay for and one more thing to fail.

Build commands:

Terminal window
composer install --no-dev --no-interaction --prefer-dist --optimize-autoloader
npm ci --audit false
npm run build:ssr

Deploy command:

Terminal window
php artisan migrate --force

Cloud injects the database, bucket, Reverb and Nightwatch variables itself. The only custom environment variables are the app key, mail settings, session driver, the Oak key and archive mode, and the Anthropic key.

What Cloud asked of the code

Four things needed changing in the code itself, and each came from how Cloud runs the app rather than from anything wrong locally.

Trusted proxies. Cloud terminates TLS in front of the app and forwards plain HTTP to the container. Without trusting the forwarded headers, every request looks insecure: generated URLs come out as http://, and a secure session cookie gets set on a connection the framework thinks isn’t secure.

// The app runs behind a platform load balancer that terminates TLS and
// forwards plain HTTP, so without this every request looks insecure.
// The container is only reachable through that load balancer, which is
// what makes trusting the forwarded headers from any address safe here.
$middleware->trustProxies(at: '*');

I wrote the comment because at: '*' looks reckless out of context. It’s only safe because Cloud runs the app on a private network that the public can only reach through its edge, and the next person to read that line needs to know it.

Object storage variable names. Cloud’s buckets are Cloudflare R2, and the variables Cloud injects for them use the AWS SDK’s global names, AWS_REGION and AWS_ENDPOINT_URL. The default Laravel config reads AWS_DEFAULT_REGION and AWS_ENDPOINT. Reading both means the disk works whichever pair is present:

'region' => env('AWS_DEFAULT_REGION', env('AWS_REGION')),
'endpoint' => env('AWS_ENDPOINT', env('AWS_ENDPOINT_URL')),
'throw' => true,

Because those are global SDK settings, any other AWS client in the app would pick them up as well. This app doesn’t talk to AWS for anything else, but yours might.

throw => true is the other half. With it off, a failed write returns false, the controller stored that as a path, and a family was told their work was saved when no file existed. Object storage fails in ways a local disk doesn’t, so the disk now throws and the controller refuses a false path.

Cloud also sets FILESYSTEM_DISK to whatever you named the bucket’s disk, so name it s3 or the app resolves to a disk it has never defined. And R2 doesn’t support per-object ACLs, so never add 'visibility' => 'public'. The write fails with a NotImplemented error, and visibility is set on the bucket instead.

You also need league/flysystem-aws-s3-v3 installed. Without it the s3 disk can’t be built at all.

Nightwatch. The fourth deploy failed with:

The [laravel/nightwatch] package was not found in the [composer.lock] file.
The Nightwatch package is required when the Nightwatch integration is enabled.

I’d switched on Cloud’s Nightwatch integration before installing the package. composer require laravel/nightwatch, commit, push, and the next deploy went through. The integration injects its own token along with LOG_CHANNEL=stack and a LOG_STACK that sends logs to both Cloud and Nightwatch. That’s why there’s no config/nightwatch.php in the repo, and why LOG_CHANNEL isn’t in my custom variables: a custom value would override the injected one.

Octane. The first deploy ran without Octane, and it went on from the second. The commit that added it records the audit behind it. I went through app/ looking for state that would leak between requests, such as mutable statics, env() calls at runtime, or a request or config captured in a constructor, and found none. The one per-request service, a cache of subject names, is bound with scoped() rather than singleton(). I ran 420 requests against a single FrankenPHP worker locally with no failures and memory flat to within 3 KB a request, then turned it on in Cloud.

Seeding and importing content

The first data went in through one-off commands from the Cloud dashboard: db:seed --class=ProductionSeeder --force, then oak:archive:import to load the committed archive, then oak:backlog and oak:top-up to import the programmes.

The --force flag is easy to forget. Seeding in production without it exits with an error, and I hit that once on 30 September.

I’d do the next bit differently. The environment template says the archive import belongs in the build commands, because build output is kept in the release and anything a one-off command writes to local disk isn’t. Cloud’s filesystem is ephemeral and every replica has its own. I ran the import as a one-off command anyway. Content already imported into Postgres is unaffected, but a later replay will miss the archive and fall back to live Oak calls.

What production taught me

None of these three would have shown up locally.

Out of memory on photo uploads. The page an upload redirects to asks for a thumbnail before it exists, and the code dispatched thumbnail generation with ->afterResponse(). That doesn’t queue the job. It runs it synchronously in the web process once the response is sent, and under Octane the web process is a long-lived worker. GD decodes every pixel, so memory follows image dimensions, not file size: a 48 megapixel phone photo passes a 10 MB upload limit and needs about 200 MB to decode. The fix is a plain dispatch() onto the queue, plus a check that reads the image header first and skips any thumbnail whose decode wouldn’t fit in the remaining memory. The Oak lesson warming had the same afterResponse() problem and got the same fix.

The tests had passed the whole time. Queue::fake() swallows after-response dispatches as well as queued ones, so it couldn’t tell the difference. The new tests use Bus::fake() and assertNotDispatchedAfterResponse(), and they fail on the old code.

Every scheduled task gets its own minute. schedule:run starts the tasks due in a minute one after another, so a slow task holds up the rest. The top of the hour had three tasks and midnight had five, all in the same hour as the Oak imports. Now each task has a minute of its own, and the nightly ones run at 03:10 and 03:30, clear of the hour the clocks change in for Europe/London, where a task can be skipped or run twice.

Schedule::job(new RecoverStalledResearch)
->cron('20,50 * * * *')
->description('Revert topic, textbook and career research stuck in progress');

The Oak quota. Covered above. Replay mode’s made-up budget sent a whole backlog of live imports at once.

What I’d check next time

The Cloud dashboard is the source of truth for this deployment. There’s no cloud.yml in the repo, and the build commands, resources and custom variables all live in Cloud. The cloud CLI makes them readable from a terminal (cloud env:get, cloud deployment:list, cloud command:list), and I’m going to make a habit of diffing what it reports against .env.production.example after any change. Writing this article was the first time I did that properly, and it turned up a few variables in the template that weren’t set in the environment.

Next up: moving the Oak archive into the build step, and whether a 512 MB instance is really enough once the archive lives on disk.

Share

XLinkedIn

Related

Keep Reading

All posts →