PROJECT 02 · 2025 – ONGOING

Self-built AI agent

I built my own operating system with Claude Code: it pulls data from the systems I work in, turns it into dashboards, and runs the repetitive administration on a schedule.

Any interface shown is a de-identified demo. All figures are illustrative and contain no real client or personal data.

29
capability modules
11
scheduled jobs, staggered
109
days of work logs (Apr–Aug 2026)
5 min
to restore onto a new machine
PROBLEM

The data was scattered; the time went into collecting it

Ad performance in one console, orders in another, tasks in Notion, schedule in a calendar. Opening each one, copying the numbers out and assembling something decision-ready took most of the morning — and it was stale again the next day.

ARCHITECTURE

Four layers, so changing one doesn't disturb the others

Capability (what it can do), workflow (when it runs) and memory (what it decides against) are stored separately. This is what keeps maintenance cost down.

Rules
Standing operating principles Verify before concluding and cite the source; propose a plan before acting; keep a human gate on anything high-risk.
Capabilities
29 modules, each owning one job Every module's spec states what it does not own. Overlapping ownership produces contradictory output — the same as it does in a real team.
Workflows
11 scheduled jobs Staggered and ordered by data dependency: the data has to land before the analysis runs.
Memory
76 structured records Project context, decisions and the reasoning behind them, and failures already encountered — so a judgement call has something to check against instead of being re-argued.

Deployment is a mounted-config setup: one version-controlled core directory, symlinked into place. New machine, pull it down, run the mount script, run the health check. A full day of rebuilding becomes five minutes.

INTERFACES

Real screens, synthetic data

These are not mockups. The dashboard was run against a structurally identical dataset with every value and label replaced by synthetic ones, then screenshotted — so the layout, fields and charts are exactly what the system renders, while nothing on screen is real. No redaction, no blurring.

● Real interface, demo data
Paid media dashboard, demo data

Paid media. Spend, orders and performance pulled from the ad platform API into the handful of numbers worth watching daily, with scale / pause / replace-creative recommendations derived from them.

Operations overview dashboard, demo data

Operations overview. Social, paid and sales in one view. Each figure carries its basis and coverage — "revenue" can mean three different things in the same report, and unlabelled it gets used to reach the wrong conclusion.

Organic search dashboard, demo data

Organic search. Search console and web analytics, used to decide whether resource goes to search or to paid. This view is what moved budget toward search: organic contributes roughly 4× per visitor what paid social does.

Daily operations dashboard, demo data

Daily brief. Calendar, tasks, priority mail and upcoming payments assembled into one page every morning. The panel top-left carries forward the three things written down the night before — closing the loop between reflection and execution rather than leaving two sets of notes that never meet.

ROLLOUT

Three stages, not one cutover

Automating something that isn't trustworthy yet only scales the error.

Stage 1
Manual validation Run every module by hand first. It only advances if the output is something I actually use to make a decision.
Stage 2
Semi-automatic Automate collection; keep judgement manual. Make "the data is fresh" reliable before anything depends on it.
Stage 3
Fully scheduled 11 staggered jobs. Enabling a schedule stays a deliberate manual step — capability and authority are two different boundaries.
OPERATIONS

The hard part of a rollout is what happens when it breaks

These three are worth more than the module count.

Failure 1

Scheduled runs failed; manual runs were fine

The scheduler reported "not logged in" while the identical command worked in my own terminal. The cause was the operating system deliberately isolating background processes from the user keychain — a security design, not a fault, so "fix it" was never an option.

Resolution: moved to a long-lived authorisation token, and specifically chose the variety that cannot trip metered API billing. That choice carried real financial exposure, so I verified the billing model before making it.

Failure 2

Unattended modules stopped to ask a question

There is nobody in a scheduled environment to answer "save this file?", so the whole run stalled. This class of error never appears in manual testing, because in manual testing I am sitting there.

Resolution: ran a full end-to-end simulation, which surfaced two environment-specific faults — one waiting on a save prompt, one silently renaming its output so the downstream job couldn't find it. Turned it into a standing rule: modules that run on a schedule must specify three things — write directly, never prompt, fixed filename prefix.

Failure 3

One constraint I could not fix

An external service expires its authorisation every seven days. I believed a configuration change had solved it permanently; a week later reality disagreed. Checking the provider's actual policy confirmed there was no fix available for an account of my type.

Resolution: stopped trying to eliminate it and reduced its cost instead — a one-command re-authorisation script, plus monitoring that tells me which script to run when it lapses.
A second lesson came with it: monitoring file timestamps is unreliable, because the file still gets written when the service behind it is dead. Checking the error records inside the file is what actually catches it.

STACK

Built, not bought

Claude Code Python JavaScript / Node.js Meta Graph API Google Search Console & GA4 APIs Notion API n8n Cloudflare Pages Scheduled jobs
WHY IT TRANSFERS

Same work, different object

The expensive part of running projects is rarely the decision. It's assembling, reconciling and aligning information before the decision can be made. That is the part I automate.

Which is the same thing systems implementation is: taking a process that only works because someone is holding it together, and making it run on its own.

Back to
Portfolio