14 May → 7 Oct 2026. The last hundred are the v3 rebuild, from 18 September.
1,427 files tracked in the repository.
London·BTech — West Herts College·Available now
I build AI systems end to end — model routing, retrieval, safety gates, interface and deployment. Eight systems below — three of them built for the company I work for.
Summary
A desktop AI assistant, rebuilt on its own v3 core. A family app running on four phones. A trading terminal with a nine-agent research desk. An apprenticeship tracker that never presses submit. A backtesting engine. For the company I work for: an internal platform, a lead finder, and a catalogue whose website cannot read its own submissions. Dated commit history for all of it.
Every system on this page in one film. There is no editor and no timeline: each frame is a function of time in a web page, rendered by headless Chrome at 3840 × 2160 and 120 frames a second, with a score synthesised in code. Shown here at 1080p.
01 — The ledger · recounted 8 October 2026
Counted from the working trees on 8 October 2026. The command that produced each figure is printed with it.
Sourced from eight git repositories, a backtesting archive and a CV. No stock imagery and no mock-ups.
14 May → 7 Oct 2026. The last hundred are the v3 rebuild, from 18 September.
1,427 files tracked in the repository.
Plus 56,964 more in the Electron HUD.
Twenty-seven of them for the v3 core alone, plus live probes and a benchmark.
Sixty modules. A test fails the build if any of them passes 600 lines.
Competition maths, graduate science, calendar arithmetic. It was 22 until problems were solved by running a program.
1 → 26 Aug 2026. Four weeks, one app, four phones.
No framework, no build step, nothing to keep updated.
Light and dark, by a screenshot harness I wrote myself.
Equities, FX, futures and crypto. Costs modelled, one position at a time.
Plus 271 checks driven against the running application over its debug port.
Electron located rather than installed. There is no node_modules to have.
Read from GOV.UK by Education Partners, and re-checked against their closing dates.
Across Kohli, VANTAGE, Education Partners, Lead Finder, the questionnaires and this site.
6 → 15 Sep 2026. Ten days, eight phases, one branch.
Over 78 tables and 35 migrations. Generated schema snapshots are not counted.
5 Oct 2026. Six public sources read on a schedule, every lead shown with its evidence.
Node built-ins only, with 53 tests that never touch the network.
Live with the catalogue, each with its own questions and its own rules.
Sealed with RSA-OAEP on arrival. The website holds only the sealed copy.
chrome-devtools, memory, git-mcp, defillama over stdio; Microsoft Learn over HTTP.
Across Groq, NVIDIA and OpenRouter, each with its measured behaviour written down, and a local model for anything private.
The AI assistant on my desktop, used every day — and in September rebuilt on a core of its own. One loop, native tool calling, and three ways to think: Spark 1.1 for fast answers, Zenith 1.1 for deep ones, and Auto choosing per question. It researches live, reads documents and pictures, talks back through a voice orb, launches the rest of my apps, and has a Private mode in which nothing leaves the PC.
On 19 September the old runtime was put through forty questions in one pass, and the result was committed before a line of v3 was written. Asked the capital of France, it took sixty seconds and four model calls and ended with "That took longer than I'm willing to wait." The first v3 turn engine was 1,800 lines against the 62,000 it replaced.
Default to CONCISE.
Both after-figures are from the runs that closed the fixes, on 23 September. A later run of the benchmark that evening read 34 of 40, and that report is the one tracked in the repository: the lanes are free and live, and the variance is real.
The same question in a new chat each time, recorded off the running HUD in real time. Spark answers in about fourteen seconds with the recommendation and its reasons. Zenith takes about forty and writes a comparison table, when the other choice wins, and a build plan — the same loop with four times the evidence and no clock.
Should a small startup building a social app use PostgreSQL or MongoDB? Compare them and recommend one.
Spark, Zenith and Auto are not three engines. They are one turn loop and a table of how much time, how many rounds and how much evidence each may spend. Choose one.
Voice lives over the open chat. It listens with an echo-cancelled microphone, ends a turn after 720 ms of silence, transcribes on this PC with faster-whisper, answers on Spark, and speaks sentence by sentence with Kokoro while the answer is still being written. Talking over it interrupts it. Every voice turn is saved as an ordinary chat line.
The orb above is the HUD's own drawing routine and state table. The microphone level it reacts to is generated here — a page should not ask for a microphone.
The microphone is paused until the dictate button or the voice orb asks for it — no wake word, no clap detector, and fifteen ambient behaviours switched off. Speech is transcribed and spoken on this machine, sentence by sentence while the answer is still streaming, and speaking over it interrupts it. Only the language-model call goes out, and Private mode keeps even that local.
Every name above is a file in syrix_v3 — 60 modules, 15,259 lines, none over 600.
An architecture test fails the build if a module passes that, or if anything from the runtime it
replaced is imported. The old code was deleted rather than wrapped: fourteen modules nothing
imported, found by a reachability tool before anything went.
An answer is not sent to one model and waited on. The answer role's lanes are tried in a fixed order, each with a hedge timer: if the first has not started streaming when its timer fires, the next starts beside it, the first to stream wins, and the other is cancelled. Every lane on the list was called before it was trusted.
The modes are a table, not three engines. budget.py holds the whole difference as data: rounds, sources, time, and which lane answers.
Spark answers in sentences on the balanced tier. Zenith gets more rounds and more sources, and splits work between sub-agents that report back on the fast lanes. Auto decides once, and the decision can be explained.
Groq gpt-oss-120b and gpt-oss-20b; NVIDIA nemotron-3-super-120b, nemotron-3-ultra-550b and glm-5.3; a free Nemotron on OpenRouter. qwen2.5:7b runs locally, started on demand and stopped with SYRIX.
Every model there was called, not assumed. The catalogue records the IDs that had been retired, the provider that answered 402 to everything, and the ones listed as available that never answered.
It can go to a named website or profile and follow its links. Pages are numbered, and only an address already seen in the turn may be opened.
Official sites are resolved through Wikidata and tickers checked on the company's own listing. LinkedIn's robots.txt disallows everything, so a profile there is reported as not opened rather than guessed at.
The program a model writes is untrusted: it must parse, may import only the standard maths and text modules, may touch no underscore attribute and none of eval, exec or open, and runs in an isolated interpreter from an empty folder.
Anything that fails a check is not run. It works on every lane, because the program is plain text.
Markets and weather read live. A quote uses the price's own decimals, and a closed market says so rather than presenting the last trade as now.
Charts, pictures, product cards and related questions arrive with the answer rather than after it, and a stock widget sits beside the text.
PDFs, Word files and images stay in the conversation as things the model can go back to, not a summary taken once.
Vision runs on NVIDIA's llama-3.2-11b-vision with an omni model behind it — both measured reading the colour and the text in a test image before either was trusted.
Voice is an orb over the open chat, not a separate screen: hands-free turn-taking, typing still works, and every voice turn is saved as an ordinary chat line.
Kokoro runs in a child process, so the speech model is loaded only while voice is in use and fully released after.
Off at every launch. While it is on, answers come from the local model, the chat is never written to history or memory, and the process refuses any connection that leaves the machine.
Supercomputer will not run while it is on, because its job is to browse.
An agentic workspace with its own embedded browser, a permission gate and a resource governor that yields while I am gaming.
Since 27 September it plans and answers on the v3 core, on a thread of its own, so neither mode's context reaches the other.
A launch pad for the apps of the ecosystem: VANTAGE, Kohli, Education Partners, Lead Finder, GMC Hub and this site.
The interface only ever sends an app's id; what that id launches lives in a file the main process reads, so nothing typed into a chat can become a command line.
Settings was rebuilt and every control on it works — the ones wired to nothing were removed rather than left as decoration.
Identity, loyalty and tone live in SYRIX.md, an editable file read on every turn. How it researches and formats stays in code, so editing its personality cannot change its answers.
Gmail, Outlook, Hostinger, Slack, Discord, Teams, Google Calendar, Notion, Todoist, Asana and Trello in one directory, each behind its own permission switches and a token store encrypted with Windows DPAPI.
Send and delete are blocked at the connector, not by asking the model nicely. MCP runs both ways: five external servers in, SYRIX's own tools out.
Questions about SYRIX are answered from SYRIX and never from the web — a search once described an unrelated company with the same name as if it were this one.
Who built it is answered one way, in any language and any word order, and the owner's name is never sent to a search engine.
One teardown for every way out — window, tray or a typed request — and the local model is stopped with it.
No reply from the pre-v3 assistant can reach the screen: a failed turn says it failed rather than falling back to the old chain.
A launch pad in the HUD's sidebar. Each key is an application SYRIX lives inside, in its own colour, with a live dot when it is already running. The interface only ever sends an id: what that id starts is held in a file the main process reads, so nothing typed into a chat can ever become a command line.
The same core, cut down differently for each application: what it may read, what it may change, who it asks first. Each scene below is that app's own SYRIX surface, playing the flow its code actually runs.
Hello
Ask me anything at all — I can also read and change what your family keeps in this app, and search the web when something needs to be current.
The widest fitting. It reads the family's own lists, calendar and bills through the same store the screens are drawn from, and writes back to it. An ordinary write — a list item, an event, a job — just happens. Money, health, messages, location and anything destructive always stop at a card first, in plain words.
add_item needs no cardmark_bill_paid always asksNine specialists, run in waves of three, each reading its own slice of a dossier gathered in code before any of them is asked. A chief writes one view that has to carry the price that would prove it wrong, and a reviewer checks it against the dossier rather than against the other members.
SYRIXKabirr Education PartnersThe narrowest fitting, on purpose. Seventeen tools that only search, read and check, in a loop bounded at eight rounds. What can be proved is worked out first and handed in as evidence, so a course is never suggested on grades nobody has. It drafts an application; a test asserts it never presses submit.
GMC AI✓Task created.Open it
Seven routes — conversation, itself, the Hub, reasoning, an action, research, a draft — decided in code before the model runs. A change to the Hub is only ever a card that shows every value it would write, with "Don't do it" focused by default, and the browser approves an id: the action itself lives on the server.
SYRIXLead Finder edition
I can read every lead in this app. Ask me who needs what, where they are, how to reach them, or whether an approach is allowed.
Which companies need help with their sponsor licence? Which tenders close in the next two weeks? Is it legal to email these companies about our services?Nine tools, and not one of them writes. It reads every lead, Companies House and official guidance from GOV.UK, legislation.gov.uk and the ICO, and every lead carries an AI overview cached against a hash of its record — redone only when the evidence changes.
Each scene is rebuilt from that application's own markup, colours, step labels, tool names and card wording. The questions and the details inside them are examples; nothing in them is a record from the app.



SYRIX.md, opened from the SYRIX page..env at boot and never enter config.jsonA 3D, scroll-driven site for the project — Three.js, a boot sequence and a raymarched core — rewritten in October for the v3 core, with a build journey that runs from a clap script to the ecosystem.
Open the SYRIX site →
A private family hub — lists, jobs, dates, meals, money, documents, markets and an assistant, on my family's actual phones. Plain HTML, CSS and JavaScript. No framework, no build step, no packages, nothing to install. Twenty-five screens, and it installs to a home screen like an app.
Real captures · light and dark
















Light and dark — both real captures, not recolours.
The same assistant, built for one household and given the run of it. It is not a separate chat surface: it holds the app's own context — who is using this phone, which screen they are on, what is on it — and it can act on the data through the same store the interface uses. One controller drives both the chat tab and the per-screen overview cards, so there is no second pipeline to drift out of step.
The whole surface, named. They run one at a time rather than in parallel, and every write passes a permission check before it is allowed to touch the store.
/* Six model providers, none of them reliable. A lane is not picked
because it is preferred — it is picked because it has been
answering, and answering quickly. */
const BACKOFF_429 = [30, 60, 120, 240, 480];
function reward(state, ms) {
state.fails = 0; state.until = 0;
state.score = state.score * 0.7 + 0.3; // EWMA
state.latency = state.latency * 0.7 + (ms / 1000) * 0.3;
}
function candidates(env, tier) {
/* A dead key or a retired model is quarantined for a day, not
retried every request for the rest of the afternoon. */
out.push({ provider, model, state,
rank: state.score / (1 + state.latency / 10) });
return out.sort((a, b) => b.rank - a.rank);
}
Exponential backoff, an exponentially weighted moving average of score and latency, and a latency-discounted rank. One model was left out of the table entirely: it is fast and it advertises tool calling, but handed a tool schema it returns an empty array and a page of garbage. Every lane in there was checked against a real tool call before it was listed.
A multi-asset research and trading terminal, on my own machine, against free keyless endpoints. A hand-written canvas chart, a portfolio across three books that are never added together, a newsroom, a screener, and a research desk that takes a position and then has that position checked against the evidence it was built from.
One real capture — NVDA, live, 28 August — taken apart along the application's own panel boundaries.
Every figure on it was fetched, checked and given a freshness state before it was drawn.
/* A filed revenue and a filed profit from different years produced a
profit margin of four hundred per cent, and each figure was
individually correct. `fy` and `fp` name the REPORT a fact appeared
in — a 10-K carries two years of comparatives — so one fiscal-year
label is attached to three different revenues. Periods are keyed on
their real start and end dates instead. */
function periodKind(row) {
const days = (row.periodEnd - row.periodStart) / 86400000;
if (!row.periodStart) return 'instant'; // a balance-sheet moment
if (days >= 300) return 'annual';
if (days >= 60) return 'quarterly';
return 'other';
}
Two figures are only divided by each other when they cover the same period. A balance-sheet figure with no instant at the anchor's year end returns null rather than the nearest one to hand.
One question, split nine ways, each pass reading only its own slice of the evidence.
A guard that removes the conclusion cannot distinguish a sound one from an unsound one. A check against the evidence can — so the sentence that takes a judgement stays, and every figure in it has to appear in evidence that was actually fetched, with the price that would prove it wrong.
A UK apprenticeship and university tracker that runs on my own PC. It finds vacancies across every free source that legitimately publishes them, reads the adverts, checks them against published entry requirements, and drafts the application. It never submits one — a test asserts it stays that way.
The line below is every apprenticeship the application is holding, and how many are left after each check it makes. It is written left to right, in the order the checks run.
Reading it left to right: every vacancy the application holds, and how many survive each check it makes. The distance between the first figure and the second is the point — a search listing publishes no entry requirements, so nothing in it has been compared against anything. The application says so on its own screen rather than letting 3,877 read as 3,877 opportunities.
Nothing on this sheet is typed. Every line is read out of something, and says what.
It opens the employer’s own form and stops. GOV.UK applications go through One Login with two-factor, so the sending stays with the candidate — and a test asserts the words that would break that promise never appear.
Ten more sources are named on the Sources screen as deliberately not read, each with the published reason. LinkedIn's User Agreement prohibits automated collection. Workday is public and unauthenticated but undocumented, and not offered for syndication.
A third build of the same assistant, cut down to one job. It has no household to run and no markets to read — it can see the profile, the grades, the uploaded files and whatever is on the screen in front of it, and its tools do nothing but search, read and check. Narrowing what an assistant can reach is the cheapest way to make it reliable.
Five trading strategies, run against equities, FX, futures and crypto on one engine, ranked by the metrics that survive contact with real money. I built it because I wanted to know whether the setups I read about actually hold — and the honest answer needed a machine.
SMC_BOS on a synthetic series — the engine's wiring-test mode, which the README labels clearly because these numbers carry no meaning for live trading. The mechanics are the production ones: position size derived from stop distance, ATR stop, fixed 2R target, one position at a time.
class BaseRisk(bt.Strategy):
"""Shared: ATR stop, RR target, % risk sizing, one position at a time."""
params = dict(risk_pct=0.01, atr_period=14, atr_mult=2.0, rr=2.0)
def size_for_risk(self, stop_dist):
if stop_dist <= 0: return 0
risk_cash = self.broker.getvalue() * self.p.risk_pct
return max(int(risk_cash / stop_dist), 0)
Position size is derived from the stop distance, never guessed. Commission is modelled at 0.05% a side. One position at a time, so results cannot be inflated by stacking entries.
The engine also ships a synthetic mode for wiring tests, clearly marked so those results are never mistaken for evidence. An over-fitted backtest is worse than none.
The internal operating platform for Golden Management Consultancy, an immigration consultancy. Thirty-five screens over one Postgres database: clients and cases, projects, tasks, calendar, email, team chat, documents, a full HR module, automations, reports, and an assistant that answers immigration questions from published sources and can act on the Hub when a person approves the card. Built to a written specification I was given, phase by phase, on a stack I did not choose.
Every screen shows a person only what their role allows, and it is decided on the server — hiding a button is not a control. The mark fills to how much of the platform each role can actually reach.
The stack was locked before the first commit and the work was sequenced into ten phases, each with a written bar for being finished. Changing a locked choice needs a written reason in the dependency file, which is where the two changes that happened are recorded.
The same assistant again, tuned for immigration casework and for this company's own data. It answers from published sources with the figures cited, reads the Hub when the person asking is allowed to, and proposes an action as a card that shows the real parameters it would run with. It is free to use and holds no provider key: the Hub pairs with a gateway I own and is handed a revocable device token.
A local app that finds organisations and people with a current, evidenced need for the consultancy's services, shows how to reach them, and retires a lead only when a source shows the need has passed. It reads six public sources on a schedule and every lead carries the published words that make it one. Lead data never leaves the PC.


Every lead shows when the source first published the evidence, not when this app found it. Leads read from a register were showing "today" for companies incorporated in 2022 — so each kind of lead has its own date and its own limit.
The same core in five applications, and a Deck in the HUD that launches every one of them. It is not five models and nothing here is retrained — what changes is what it can reach, who it asks, and what it is not allowed to say.
The parts worth building once. Every fitting inherits them, which is the only reason five assistants were affordable to build at all.
The pattern across all five is the same: the model is the least trusted part of the system, and every build spends most of its code deciding what it is allowed to see and checking what it gave back.
480 commits on SYRIX, 60 on Kohli, 22 on VANTAGE, 172 on GMC Hub, 18 on Lead Finder, 6 on the film, 3 on the catalogue, and one squashed import for Education Partners.
762 commits across eight repositories, May to October 2026 — counted with
git rev-list on each one. Twelve milestones, then the raw log.
A 40,000-line single file split into modules behind a routing layer.
Model lanes measured before the router is allowed to trust them.
Isolated browser, permission gate, resource governor, evidence store.
v1, an in-app assistant, and a gateway needing no key or open port.
Two-lane runtime complete, with a threat model and a recorded rollback.
An apprenticeship desk that reads the adverts and never presses submit.
22 commits, each with its own test, and a handoff that records what is not finished.
Eight phases to a written specification: 35 screens, 78 tables, and an assistant that asks before it acts.
A baseline committed first, then one loop with native tool calls — 1,800 lines against the 62,000 it replaced.
Three cuts in two days, each frame a function of time, rendered at 4K and 120 fps.
A voice orb in the chat, a mode in which nothing leaves the PC, and a launch pad for the other apps.
A lead finder that dates every lead by its source, and questionnaires the website itself cannot read.
git logPrepare Jarvis/SYRIX for safe GitHub sync
Complete safe modularization routing phase 1
tools: free-brain benchmark — prove which lane is best before routing
feat(supercomputer): Phase 1–2 — safety baseline + isolated mode with fail-closed dispatch
feat(supercomputer): Phase 3 — isolated embedded browser sandbox
fix(supercomputer): the real cause — 21k tokens per step was starving the pool
fix(supercomputer): 16 bugs from the first live drive — safety, scope, answers
feat(supercomputer): a resource governor, so a task can run while Boss games
feat(supercomputer): ask for what only Boss knows, never invent it
fix(chat): repair an answer from the web instead of shipping the gap
fix(tests): the suite was calling live search APIs, and I raised the ratchet
perf(spark): 16.4s → 10.5s, and stop refusing before actually looking
fix: ground self knowledge and long browser tasks
docs: threat model Spark Zenith 1.1
feat: complete Spark Zenith 1.1 Vanguard runtime
fix: make secret guard Unicode-safe
Baseline: 40 questions against the runtime being replaced
SYRIX v3: the turn engine
v3 memory: one store, with provenance, that ages itself
v3 live: the cutover, and the local lane actually working
Delete what nothing imports, and show what is actually running
v3 brain: benchmark eval 30/40 -> 40/40, and the fallback chain no longer fails open
v3 compute: problems are solved by running a program; hard eval 22/25 -> 25/25
v3: go to a named website or profile, and follow its links
Private mode, and a spoken-turn pipeline for the voice orb
HUD: voice orb in the chat, Private mode, account menu, Connectors and Settings
SYRIX.md: identity, loyalty and tone in an editable file
HUD: Deck, a launch pad for the apps of the SYRIX ecosystem
Questions about SYRIX are answered from SYRIX, never from the web
Close tears down everything at once; pre-v3 replies can no longer leak
Kohli v1: private family hub
Make SYRIX work with no API key, no PC left on and no port opened
Deploy Kohli Core, and fix what only a live deploy could reveal
The passcode has to cover the price cache too
Stop a bio being wiped by somebody else walking down the road
Fix the More panel putting you back where you started
Deploy from a staged allow-list, not from the whole folder
Drop 'unsafe-inline' from style-src, and fix the boot flag it exposed
Finance: a real charting screen, and the family section earns its place
Keypad keys pop on press, and stop the lock screen sliding on a swipe
Rework the lock screen: pill instead of boxes, Enter in the pad, less on it
Education/opportunity screen: eligibility, universities, live driver
SYRIX gateway: display-advert spec, transport, and deploy notes
Chat media, photo editor, and household deck/finance refinements
Kabirr Education Partners: apprenticeship discovery and application tracker
A private multi-asset research and trading terminal
Fix TradingView widget chart rendering nothing
Make alerts actually fire
Fix Discover heatmap 'Colour by' going flat on two of six options
Remove the redundant Watchlist rail tab; give its header a real switcher
Un-block the price feed from the bid/ask enhancement, fix a stuck splitter
Charts and paper fills keep moving into pre/post-market
Fix real logos never showing, and a slow news front page
Show a plain-English gist for every instrument, free, before any button
Replace the SYRIX depth cycle button with an explanatory menu
Fix the AI overview reading fields that have never existed
Correct STATUS.md: the research panel's vendor-summary path was never true
Revert a same-day misdiagnosis of the AI overview card's data shape
Stop the crosshair freezing mid-drag, add a crosshair line-style setting
Align Performance holdings table with the paper blotter, reset tab scroll
Let a symbol be added to the watchlist by search, not just the chart in view
feat(web): Phase 1 — identity & access
fix(web): add team column to tasks so TEAM_LEAD can see unassigned team work
ci(web): add lint/typecheck/test/dependency-scan/secret-scan/SAST workflow
fix(web): remove leaked Resend key from tracked .env.example
feat(web): add step-up auth ("Approvals") for sensitive actions
feat(web): add GMC AI research pipeline — retrieval, ranking, SSRF guard
fix(web): route by subject, not by the words that happen to be in a question
fix(web): stop the assistant answering from memory, and let it read pictures
fix(web): search for the rule, not for the case
test(web): prove one employee's GMC AI cannot reach another's
feat(db): flag seed rows in the database instead of in their names
feat(web): let one person hold both platform and HR administration
fix(web): close an IDOR on expense line items, and drop dead starter schema
fix(web): bind step-up auth to the session that passed the challenge
fix(web): keep error logs in production, and close the approval bypass
Showreel v1: 2:38 film of six systems, 4K 120fps
Showreel v2: 3:20 film, rebuilt sections, new score and sound design
Bound the encoder's memory: cap decode threads, filter threads, input queue and lookahead
Showreel v3: glass renderer, rebuilt intro and finale, new sections and drops
v3 revision: darker glass, SYRIX reaching into the apps, synced handover, new close
Lead engine: public sources, enrichment, scoring and evidence-based retirement
Stricter matching, verified websites without search engines, retries
Community leads: exclude sellers, US routes and existing-visa paperwork
SYRIX: chat panel and an AI overview on every lead
Reliable overviews, services that match GMC's offer, side panel restored
Catalogue website as published before the October 2026 simplification
Simplify the catalogue and its navigation
Regenerate the questionnaire bundle for the 5 October 2026 deployment
A single squashed import on 26 August 2026, so there is no history to show. Saying that is better than padding a log with something. What can be counted is the tree.
src/web/server.jscloud/All of this started as programs on this PC, reachable from nowhere else. That is safe, and no use on a phone. Six of them are hosted so they can be opened from anywhere — and the work was doing that without giving up what was true while they were local. One of the six is a read-only copy, because the real application cannot run on the free plan, and it says so.
A household planner nobody can open on their phone is a file. Putting these on the internet is the easy half; the half worth writing down is that nothing which was true on a loopback address stopped being true once there was a public URL — no key in anything shipped, nothing listening at home, and every device holding only a credential that can be taken back.
wrangler.toml rather than clicked together in a dashboardnpx — nothing installed to publish any of itThe family app is a Worker and a Pages project, and forgetting the second deploy leaves the fixes live on this machine and nowhere else. The apprenticeship app puts its API inside the Pages output instead, so there is only ever one thing to ship and it cannot be half-shipped.
Every one of these was found the same way — the deploy reported success and the thing it published was wrong. None of them can be caught on this machine.
main and the repository is on master. Without --branch=main wrangler uploads a preview URL and production carries on serving the old build — which looks exactly like a deploy that worked..gitignoreA whole-folder deploy publishes whatever is sitting in the folder, at a public address, including anything holding keys. The deploy stages an allow-list instead and refuses to run if a staged file is shaped like a credential.Six of these are reachable from somewhere other than my desk, one can read email, open pages and run code, one holds another company's staff and client records, and one takes in strangers' answers about themselves. The controls below are the ones that exist in the repositories, written up with the gaps they do not close.
It reads the web, opens documents, holds connector tokens and can be given a task to run on its own. That is the whole attack surface in one program, so it has a written threat model — ten invariants, eleven trust boundaries, and a list of the gaps it does not close, dated and kept with the code.
127.0.0.1; LAN needs an explicit config flag127.0.0.1 with an origin check on every state-changing request, so nothing else on the network can reach it. A credential scan over the whole codebase runs as a test, and a separate test asserts it never presses submit..gitignore and a whole-folder deploy publishes whatever is sitting in it.The threat model is a release gate, not a certificate. It records controls that exist, gaps that are open, and what should happen when a control fails — and the gaps are the part worth reading.
None of this is a claim that SYRIX is secure in absolute terms — the document says so in its own second paragraph. Writing down what is not covered is the only version of this worth showing to somebody who will check.
Eight decisions that shape everything above, each one traceable to code in these repositories.
A figure in an answer is worked out by code that ran — a calculator, or a program the model wrote and an isolated interpreter executed. The model phrases the result and nothing else.
Ambiguous command: ask. Unknown domain: deny. Missing capability: degrade and report it. The default is never a best guess.
Speech recognition and synthesis run on the machine. Private mode answers on a local model and refuses every connection off the PC. Lead data, and every questionnaire answer, is only ever readable on one computer.
Gmail can read, search and draft. Send and delete are blocked at the connector, with a DPAPI-encrypted token and an audit trail — not enforced by prompting.
Every model in the lane catalogue was called before it was trusted, and the dead ones are written down with how they failed. Provider choice is data, not preference.
Kohli, Lead Finder and the questionnaires have zero packages and no build step. The HUD transpiles JSX at runtime. None can break from an update I did not make.
Screenshot harnesses, overflow checks and live acceptance probes. A passing unit test and a working screen are separate claims.
SYRIX runs on my desktop, Kohli on four phones in my family, and the GMC tools at work. Defects surface in use rather than in review.
Each of these replaced an assumption that had already cost something. The cost is written next to the rule.
Every request carries a user agent naming the application and a route back to me — never a copied browser string. There is no CAPTCHA handling, no cookie replay and no headless browser to get past a challenge. A host that puts one up is recorded as blocked and the run continues without it.
The difference between reading a public page and evading a control is whether you would be willing to say which you were doing.
A cache that refuses to write on a miss leaves every caller retrying a failing source at full speed. It took the family app down twice — a price cache that would not write while the passcode was set turned one render into an unbounded fetch loop — so a remembered failure now gets a short life of its own, ninety seconds against the hour a success gets.
Negative caching is not an optimisation. It is the thing that stops a retry becoming a denial of service against somebody else.
A guard that deleted any sentence naming an action and a direction meant the assistant could read every filing a company had published and then remove the one sentence that took a judgement — while still printing the readings that implied it. It was replaced by a pass that checks the finished answer against the evidence that was actually fetched, and repairs it.
A guard that deletes the conclusion cannot tell a good one from a bad one. A check against the evidence can.
Each of these left a directory, a config entry or a commit on my own machine.
Having the tools is not the skill. This is the routing table I actually use, and the same reasoning is hard-coded into SYRIX's model router.
Everything listed appears in code I have written and run. Nothing here is from a tutorial I watched.
Python · TypeScript · JavaScript (ES5 → modern) · HTML · CSS · GLSL · SQL · PowerShell · Bash
Prompt engineering · context-window and cost/latency management · model routing and fallback · quota ledgers · function/tool calling · structured output validation · evaluation harnesses · streaming
Retrieval-augmented generation · vector databases (Chroma) · embeddings · chunking · hybrid ranking · query rewriting · semantic verification · freshness gating · citation grounding
Native tool-calling loops with concurrent tool rounds · sub-agents · multi-agent councils · sandboxed code execution · MCP as both server and client · permission gating · sandboxed browser automation · evaluation sets auto-graded without a model judge
Groq · NVIDIA NIM · OpenRouter · Ollama (local) · gpt-oss 120b/20b · Nemotron 3 Super and Ultra · GLM-5.3 · Qwen 2.5 · Llama 3.2 Vision · lane racing, hedging and rate-limit admission
Faster-Whisper STT · Kokoro TTS · voice-activity detection with barge-in · echo cancellation · multimodal and vision-language models · PaddleOCR · MediaPipe · YOLO
Electron · React · Next.js (App Router, Server Actions) · Tailwind · shadcn/ui · TanStack Query and Table · Canvas 2D · WebGL and raymarched shaders · CSS 3D · Progressive Web Apps · service workers · responsive and accessible UI · light/dark theming
FastAPI · WebSocket state bridges · PostgreSQL with Drizzle · schema migrations · BullMQ on Redis · Clerk · Cloudflare Workers, Pages, KV and R2 · Wrangler · zero-dependency Node servers · OpenTelemetry · Git
pandas · numpy · DuckDB · backtrader · yfinance · ccxt · Finnhub · Alpha Vantage · Stooq · TradingView · IBKR · technical indicators and risk sizing
AES-256-GCM · RSA-OAEP key wrapping · sealed-inbox design · PBKDF2 key derivation · CSP without unsafe-inline · OAuth token handling · DPAPI-encrypted stores · secret redaction · deny-by-default allowlists · per-action permissions · role and capability models · server-side scoping · step-up authentication · audit trails · threat modelling
pytest · Vitest · golden regression suites · Playwright · CDP automation · screenshot and layout-overflow harnesses · server-render smoke checks · live acceptance probes · CI with secret scanning and static analysis · cross-device verification
Deterministic frame rendering in headless Chrome · score and sound design synthesised in an OfflineAudioContext · ffmpeg · AV1 NVENC at 4K 120 fps · fragmented MP4 streamed through MediaSource
OCDS procurement feeds · Companies House · daily CSV register diffing · RSS and Atom · per-host rate limiting with backoff · evidence-dated records with automatic retirement
Formal evaluation harnesses for agent output · walk-forward validation for strategies · real-time graphics and shader work · distributed edge deployment
Thirty-four things, routed into ten builds. Most of them went into more than one, which is the point — a technique is only learned once it has survived a second problem.
The AI assistant on my desktop, on its own v3 core.
0 of 0A private family app, running on four phones.
0 of 0A multi-asset research and trading terminal.
0 of 0An apprenticeship desk that never presses submit.
0 of 0A multi-market backtesting engine.
0 of 0An internal platform for the company I work for.
0 of 0Leads dated by their source, retired only with evidence.
0 of 0Questionnaires the website itself cannot read.
0 of 0Four minutes, every frame rendered from code.
0 of 0Hand-written, no framework, no build step.
0 of 0
Everything on this page was built and run on my own desk. Every figure below was read off the machine itself.
Knowing the parts is not the same as writing software that respects them. Everything below exists because this is one desktop with one card in it.
Eight gigabytes of video memory is the constraint the whole assistant is designed around. A vision model, a speech model and a 7B language model do not fit in it at once, so the runtime has to decide what is resident, what is evicted, and what it is prepared to give back the moment the machine is wanted for something else.
nvidia-smi and psutil, and the hardware panel reads Win32_Processor, Win32_PhysicalMemory, Win32_DiskDrive and Win32_VideoController. The figures beside the photograph came from those calls, not from a spec sheet.
I'm in London, studying BTech at West Herts College. Before that, GCSEs at Northwood School in Pinner, and school up to Grade 8 at Amity International in India.
Every project on this page was built outside coursework, self-directed and unassessed. One of them is now relied on daily by my family.
What pulled me into computing was not the code, it was the leverage. A rule written once runs a million times without getting tired or bored or slightly wrong on a Friday, and the machine will tell you exactly where you were mistaken if you build it so that it can. That is a very unusual thing to be handed at seventeen. Most of what I have learned since came from being wrong in a way the system could prove.
Language models sharpened that rather than replaced it. They are the first tool I have used that is genuinely capable and genuinely unreliable at the same time, which makes the engineering question the interesting one: not can it answer, but how do you know. Almost everything I have built since is some version of that — gather the evidence in code so it can be checked, compute the figures rather than letting the model say them, test the finished answer against what was actually retrieved, and report the failure rather than papering over it. The verifier that fails on screen in the recording above is the whole point, not an embarrassment.
The part I want to go further into is evaluation: how you measure whether an agent is actually getting better, rather than whether the last demonstration went well. That is where I think the hard, unglamorous, useful work is.
I have also managed a personal investment portfolio for 18 months — 22 holdings, reviewed monthly, with 20+ written reviews of my own decisions and roughly 27% in one annual period. The discipline carries into the engineering: establish what is known, record what is not, and size the risk accordingly.
Previously a competitive footballer: Player of the Month at the LaLiga Football School in Delhi, and invited to train with Real Madrid's youth programme in Spain.
Four Springpod work-experience programmes, each completed with a partner company and issued 14 November 2025. Every card opens the certificate PDF exactly as it was issued, certificate ID included.
Kabirr Kohli · AI Systems Engineer · London
Based in London, open to relocation. Happy to walk through any of the above in detail, including the parts that did not work first time.