London·BTech — West Herts College·Available now

Kabirr Kohli

I build AI systems end to end — model routing, retrieval, safety gates, interface and deployment. Eight systems below — three of them built for the company I work for.

8shipped systems
762commits across eight repos
634klines written

Summary

A desktop AI assistant, rebuilt on its own v3 core. A family app running on four phones. A trading terminal with a nine-agent research desk. An apprenticeship tracker that never presses submit. A backtesting engine. For the company I work for: an internal platform, a lead finder, and a catalogue whose website cannot read its own submissions. Dated commit history for all of it.

The film

Four minutes, rendered from code

Every system on this page in one film. There is no editor and no timeline: each frame is a function of time in a web page, rendered by headless Chrome at 3840 × 2160 and 120 frames a second, with a score synthesised in code. Shown here at 1080p.

4:00 running time 4K · 120 fps master, AV1 0 samples — every sound is synthesised 6 commits, 26–27 Sep 2026

01 — The ledger · recounted 8 October 2026

Figures,
and where
each one
came from.

Counted from the working trees on 8 October 2026. The command that produced each figure is printed with it.

100%
verifiable

Sourced from eight git repositories, a backtesting archive and a CV. No stock imagery and no mock-ups.

Kabirr Kohli · London
0
commits on SYRIX

14 May → 7 Oct 2026. The last hundred are the v3 rebuild, from 18 September.

git rev-list --count HEAD
0
Python files

1,427 files tracked in the repository.

git ls-files '*.py' | wc -l
0
lines of Python

Plus 56,964 more in the Electron HUD.

git ls-files '*.py' | xargs wc -l
0
test files

Twenty-seven of them for the v3 core alone, plus live probes and a benchmark.

ls tests/**/*.py | wc -l
0
lines in the v3 core

Sixty modules. A test fails the build if any of them passes 600 lines.

git ls-files 'syrix_v3/*.py' | xargs wc -l
0
of 25 on the hard set

Competition maths, graduate science, calendar arithmetic. It was 22 until problems were solved by running a program.

python tools/v3_eval.py
0
commits on Kohli

1 → 26 Aug 2026. Four weeks, one app, four phones.

git rev-list --count HEAD
0
lines, zero packages

No framework, no build step, nothing to keep updated.

cat package.json → there isn't one
0
screens captured

Light and dark, by a screenshot harness I wrote myself.

node tests/shots.js
0
strategies backtested

Equities, FX, futures and crypto. Costs modelled, one position at a time.

python run.py
0
assertions in VANTAGE

Plus 271 checks driven against the running application over its debug port.

node tests/run-all.js
0
lines, no lockfile

Electron located rather than installed. There is no node_modules to have.

git ls-files | xargs wc -l
0
vacancies held locally

Read from GOV.UK by Education Partners, and re-checked against their closing dates.

wc -c data/vacancies.json
0
packages installed

Across Kohli, VANTAGE, Education Partners, Lead Finder, the questionnaires and this site.

find . -name package.json
0
commits on GMC Hub

6 → 15 Sep 2026. Ten days, eight phases, one branch.

git rev-list --count HEAD
0
lines of TypeScript

Over 78 tables and 35 migrations. Generated schema snapshots are not counted.

git ls-files '*.ts' '*.tsx' | xargs wc -l
0
commits on Lead Finder

5 Oct 2026. Six public sources read on a schedule, every lead shown with its evidence.

git rev-list --count HEAD
0
lines, zero packages

Node built-ins only, with 53 tests that never touch the network.

npm test
0
eligibility questionnaires

Live with the catalogue, each with its own questions and its own rules.

ls eligibility/config/services
0
bit key on every answer

Sealed with RSA-OAEP on arrival. The website holds only the sealed copy.

server/sealed.js
0
MCP servers wired in

chrome-devtools, memory, git-mcp, defillama over stdio; Microsoft Learn over HTTP.

config.json → syrix_mcp_servers
0
model lanes, raced

Across Groq, NVIDIA and OpenRouter, each with its measured behaviour written down, and a local model for anything private.

syrix_v3/lanes/catalogue.py
Project 01 · the centre of everything else Core v3 · Spark 1.1 · Zenith 1.1 · Supercomputer · Voice · Private · Deck
SYRIX project mark
Project 01

SYRIX

The AI assistant on my desktop, used every day — and in September rebuilt on a core of its own. One loop, native tool calling, and three ways to think: Spark 1.1 for fast answers, Zenith 1.1 for deep ones, and Auto choosing per question. It researches live, reads documents and pictures, talks back through a voice orb, launches the rest of my apps, and has a Private mode in which nothing leaves the PC.

  • Python 3.11
  • Electron + React
  • gpt-oss-120b
  • Nemotron 3
  • GLM-5.3
  • faster-whisper
  • Kokoro TTS
  • Ollama
  • MCP
  • WebSocket bridge
SYRIX v3 answering how a transformer model works: a seven-part numbered answer, five named sources with three more, three related questions, and the composer set to Auto
SYRIX v3, answering. Captured from the running application on 26 September 2026. The reply is the one it gave — its sources numbered under it and three follow-up questions offered — with the composer set to Auto.

Why it was rebuilt, and what it scores now

On 19 September the old runtime was put through forty questions in one pass, and the result was committed before a line of v3 was written. Asked the capital of France, it took sixty seconds and four model calls and ended with "That took longer than I'm willing to wait." The first v3 turn engine was 1,800 lines against the 62,000 it replaced.

The old runtime · 40 questions 124 / 164 checks 4.65 model calls a turn, worst Spark turn 60.4 s, and 18 of 40 replies at 90 characters or fewer — its prompt said Default to CONCISE.
v3 · benchmark set 30 → 40 / 40 Word problems, logic traps, multiple choice, facts, small-company lookups, code run against tests, live data and a follow-up chain. No model grades another.
v3 · hard set 22 → 25 / 25 All three misses were long exact computations a model had predicted rather than worked out. Now a model writes a short program, it runs, and the printed result is the evidence.

Both after-figures are from the runs that closed the fixes, on 23 September. A later run of the benchmark that evening read 34 of 40, and that report is the one tracked in the repository: the lanes are free and live, and the variance is real.

Spark 1.1 and Zenith 1.1

One question, two depths

The same question in a new chat each time, recorded off the running HUD in real time. Spark answers in about fourteen seconds with the recommendation and its reasons. Zenith takes about forty and writes a comparison table, when the other choice wins, and a build plan — the same loop with four times the evidence and no clock.

Should a small startup building a social app use PostgreSQL or MongoDB? Compare them and recommend one.

SPARK 1.1real time · 18 s
Recommends PostgreSQL and says why in four paragraphs, with eight sources and three follow-ups. One line is out of date: it says MongoDB lacks multi-statement transactions, which it has had since version 4.0 — the kind of thing the deeper pass exists to catch.
ZENITH 1.1real time · 43 s
A seven-row comparison — data model, transactions, JSON, queries, operations, hiring, cost at scale — three cases where MongoDB is the better call, and how each feature of a social app maps onto Postgres. It gets transactions right: they exist in MongoDB, and are slower.
Modes

One loop, given different room

Spark, Zenith and Auto are not three engines. They are one turn loop and a table of how much time, how many rounds and how much evidence each may spend. Choose one.

SupercomputerAgentic tasks and automation
Tool rounds
Search queries
Sources read
Prompt, characters
Time limit8 s
Sub-agentsno

ListeningSpeak, or type below
Voice

An orb in the conversation, not a separate screen

Voice lives over the open chat. It listens with an echo-cancelled microphone, ends a turn after 720 ms of silence, transcribes on this PC with faster-whisper, answers on Spark, and speaks sentence by sentence with Kokoro while the answer is still being written. Talking over it interrupts it. Every voice turn is saved as an ordinary chat line.

  • 0.74 sto the first spoken sentence once the speech model is warm
  • child processKokoro is loaded only while voice is in use, and fully released after
  • mic idleuntil the orb or the dictate button asks for it — no wake word, no clap

The orb above is the HUD's own drawing routine and state table. The microphone level it reacts to is generated here — a page should not ask for a microphone.

Voice, from microphone to orb

MICecho-cancelled VAD720 ms pause STTfaster-whisper BRAINSpark 1.1 TTSKokoro · local ORBin the chat

The microphone is paused until the dictate button or the voice orb asks for it — no wake word, no clap detector, and fifteen ambient behaviours switched off. Speech is transcribed and spoken on this machine, sentence by sentence while the answer is still streaming, and speaking over it interrupts it. Only the language-model call goes out, and Private mode keeps even that local.

What a question passes through before it is answered

  1. turn.pyThe one entry point. Auto is decided once, before any model call, and nothing in the loop can start another mode's work — the old runtime ran the deep engine after the fast one had already answered.
  2. research/triage.pyDecides whether the question needs evidence at all. A question about SYRIX, a follow-up that leans on the thread, and text pasted into the message are never sent to a search engine.
  3. research/gather.py · rank.pySearch lanes raced in parallel, sponsored results dropped, and a headline is never accepted as an answer to the question it is a headline about.
  4. context.pyHistory compacted rather than cut from the end, and compacted after the answer so it costs the turn nothing.
  5. persona.py · SYRIX.mdIdentity and tone read from an editable file on every turn, capped at 1,580 characters so the conversation always keeps room in the prompt.
  6. lanes/The model call. Lanes are raced behind hedge timers, and a lane whose minute refills in a second is waited for rather than swapped for a weaker one.
  7. loop.pyRounds of native tool calls. Every call in a round runs at once, and out of time means answer from what arrived and say what is missing — never an apology instead.
  8. compute.pyA problem to work out is solved by a short program in an isolated interpreter with a six-second limit. The printed result reaches the answer as evidence, with the program beside it.
  9. check.pyThe finished answer is read back against the question and against the evidence before it is shown.
  10. polish.pyCitations normalised, labels and closing offers removed — with inline code held aside, so a rule written for prose cannot corrupt a command.

Every name above is a file in syrix_v3 — 60 modules, 15,259 lines, none over 600. An architecture test fails the build if a module passes that, or if anything from the runtime it replaced is imported. The old code was deleted rather than wrapped: fourteen modules nothing imported, found by a reachability tool before anything went.

Under the hood

Six free lanes, raced rather than queued

An answer is not sent to one model and waited on. The answer role's lanes are tried in a fixed order, each with a hedge timer: if the first has not started streaming when its timer fires, the next starts beside it, the first to stream wins, and the other is cancelled. Every lane on the list was called before it was trusted.

 

  • prompt ÷ 3.5 + replyis what a request costs against a lane's minute. A reply's tokens are reserved whether used or not, so harder questions buy rounds and sources, never a bigger reply.
  • empty is not brokenA lane that returns nothing once is retried without tools and costs no cooldown, so one hiccup does not lose it for five minutes.
  • qwen2.5:7b, locallyis the floor under every cloud lane, started on demand and stopped with SYRIX.

What is actually in there

Fourteen parts
01

Spark 1.1 · Zenith 1.1 · Auto

The modes are a table, not three engines. budget.py holds the whole difference as data: rounds, sources, time, and which lane answers.

Spark answers in sentences on the balanced tier. Zenith gets more rounds and more sources, and splits work between sub-agents that report back on the fast lanes. Auto decides once, and the decision can be explained.

02

The lanes

Groq gpt-oss-120b and gpt-oss-20b; NVIDIA nemotron-3-super-120b, nemotron-3-ultra-550b and glm-5.3; a free Nemotron on OpenRouter. qwen2.5:7b runs locally, started on demand and stopped with SYRIX.

Every model there was called, not assumed. The catalogue records the IDs that had been retired, the provider that answered 402 to everything, and the ones listed as available that never answered.

03

Research and browsing

It can go to a named website or profile and follow its links. Pages are numbered, and only an address already seen in the turn may be opened.

Official sites are resolved through Wikidata and tickers checked on the company's own listing. LinkedIn's robots.txt disallows everything, so a profile there is reported as not opened rather than guessed at.

04

Problems solved by running them

The program a model writes is untrusted: it must parse, may import only the standard maths and text modules, may touch no underscore attribute and none of eval, exec or open, and runs in an isolated interpreter from an empty folder.

Anything that fails a check is not run. It works on every lane, because the program is plain text.

05

Live answers

Markets and weather read live. A quote uses the price's own decimals, and a closed market says so rather than presenting the last trade as now.

Charts, pictures, product cards and related questions arrive with the answer rather than after it, and a stock widget sits beside the text.

06

Documents and pictures

PDFs, Word files and images stay in the conversation as things the model can go back to, not a summary taken once.

Vision runs on NVIDIA's llama-3.2-11b-vision with an omni model behind it — both measured reading the colour and the text in a test image before either was trusted.

07

The voice orb

Voice is an orb over the open chat, not a separate screen: hands-free turn-taking, typing still works, and every voice turn is saved as an ordinary chat line.

Kokoro runs in a child process, so the speech model is loaded only while voice is in use and fully released after.

08

Private mode

Off at every launch. While it is on, answers come from the local model, the chat is never written to history or memory, and the process refuses any connection that leaves the machine.

Supercomputer will not run while it is on, because its job is to browse.

09

Supercomputer

An agentic workspace with its own embedded browser, a permission gate and a resource governor that yields while I am gaming.

Since 27 September it plans and answers on the v3 core, on a thread of its own, so neither mode's context reaches the other.

10

The Deck

A launch pad for the apps of the ecosystem: VANTAGE, Kohli, Education Partners, Lead Finder, GMC Hub and this site.

The interface only ever sends an app's id; what that id launches lives in a file the main process reads, so nothing typed into a chat can become a command line.

11

Settings, and SYRIX.md

Settings was rebuilt and every control on it works — the ones wired to nothing were removed rather than left as decoration.

Identity, loyalty and tone live in SYRIX.md, an editable file read on every turn. How it researches and formats stays in code, so editing its personality cannot change its answers.

12

Connectors + MCP

Gmail, Outlook, Hostinger, Slack, Discord, Teams, Google Calendar, Notion, Todoist, Asana and Trello in one directory, each behind its own permission switches and a token store encrypted with Windows DPAPI.

Send and delete are blocked at the connector, not by asking the model nicely. MCP runs both ways: five external servers in, SYRIX's own tools out.

13

It knows what it is

Questions about SYRIX are answered from SYRIX and never from the web — a search once described an unrelated company with the same name as if it were this one.

Who built it is answered one way, in any language and any word order, and the owner's name is never sent to a search engine.

14

It closes cleanly

One teardown for every way out — window, tray or a typed request — and the local model is stopped with it.

No reply from the pre-v3 assistant can reach the screen: a failed turn says it failed rather than falling back to the old chain.

The Deck

Every app of the ecosystem, one key away

A launch pad in the HUD's sidebar. Each key is an application SYRIX lives inside, in its own colour, with a live dot when it is already running. The interface only ever sends an id: what that id starts is held in a file the main process reads, so nothing typed into a chat can ever become a command line.

DeckThe SYRIX ecosystem
SYRIX, inside every app

Five fittings, five different jobs

The same core, cut down differently for each application: what it may read, what it may change, who it asks first. Each scene below is that app's own SYRIX surface, playing the flow its code actually runs.

‹ BackSYRIXFamily Assistant
◌ Spark ›+ New◷ History
Hello

Ask me anything at all — I can also read and change what your family keeps in this app, and search the web when something needs to be current.

Your family What is happening in the family this week? What bills are still unpaid? Add milk to the groceries
Checking the app
Checking the app
⛉Just checking
Mark "Water" as paid?
CancelYes, do it
⌁Message SYRIX…◉➤
Kohli · the household

The widest fitting. It reads the family's own lists, calendar and bills through the same store the screens are drawn from, and writes back to it. An ordinary write — a list item, an event, a job — just happens. Money, health, messages, location and anything destructive always stop at a card first, in plain words.

  • add_item needs no card
  • mark_bill_paid always asks
  • 26 tools · 10 read · 16 write
SYRIXAutomatic ▾
Working
  1. Gathering the evidence on NVDA
  2. Planning the council
  3. Reading the chart
  4. Checking momentum
  5. Reading the news
  6. Reading the accounts
  7. Checking the earnings record
  8. Reading insider filings
  9. Weighing the consensus
  10. Reading the sector
  11. Building the case against
  12. Writing the answer
  13. Checking it against the evidence
The desk — 9 specialists on NVDA
Price action
Momentum and indicators
The story
The business
Earnings and events
Insiders and ownership
Consensus and positioning
The sector and the tape
The case against
VANTAGE · the desk

Nine specialists, run in waves of three, each reading its own slice of a dossier gathered in code before any of them is asked. A chief writes one view that has to carry the price that would prove it wrong, and a reviewer checks it against the dossier rather than against the other members.

  • 9 specialists · waves of 3
  • reads the desk, writes nothing to it
  • views drawn blank here: no invented calls
SYRIXKabirr Education Partners
Working out what you can actually prove
Reading your profile and grades
Checking what your grades open and what they block
Finding courses you could actually get onto
Checking what they actually ask for
Within reach
A stretch
Blocked — and why
Drafted, never submitted. Everything waits for you to send it.
Education Partners · the application

The narrowest fitting, on purpose. Seventeen tools that only search, read and check, in a loop bounded at eight rounds. What can be proved is worked out first and handed in as evidence, so a course is never suggested on grades nobody has. It drafts an application; a test asserts it never presses submit.

  • 17 tools · read and check only
  • 8 rounds at most
  • never submits
GMC AI
routeaction
PlanningWriting the answer
⚠
Create a taskNothing has changed yet. This will touch Tasks.
Title
Due
Assignee
Client
Don't do itGo ahead

✓Task created.Open it

GMC Hub · the casework

Seven routes — conversation, itself, the Hub, reasoning, an action, research, a draft — decided in code before the model runs. A change to the Hub is only ever a card that shows every value it would write, with "Don't do it" focused by default, and the browser approves an id: the action itself lives on the server.

  • 7 routes · 7 specialists in the deep lane
  • deny is the default
  • values drawn blank: they are client records
SYRIXLead Finder edition

Hi, I’m SYRIX.

I can read every lead in this app. Ask me who needs what, where they are, how to reach them, or whether an approach is allowed.

Which companies need help with their sponsor licence? Which tenders close in the next two weeks? Is it legal to email these companies about our services?
Searching the leads: closing within 14 days
Working out the answer
Ask about your leads, a company or the rules...↑
Lead Finder · the pipeline

Nine tools, and not one of them writes. It reads every lead, Companies House and official guidance from GOV.UK, legislation.gov.uk and the ICO, and every lead carries an AI overview cached against a hash of its record — redone only when the evidence changes.

  • 9 tools · none of them writes
  • official guidance before the open web
  • an overview per lead, cached by hash

Each scene is rebuilt from that application's own markup, colours, step labels, tool names and card wording. The questions and the details inside them are examples; nothing in them is a record from the app.

The SYRIX Connectors directory: Gmail connected, then Google Calendar, Google Drive, GitHub, Notion, Outlook, Hostinger, Slack, Discord, Microsoft Teams, Todoist, Asana, Trello and Jira, each with its status
Connectors, as a directory. Sign-ins are encrypted on the PC with Windows DPAPI, and SYRIX asks before anything that changes something. The connected account's address is masked.
The Gmail connector's permissions: read, search, read full email, read drafts and create drafts allowed; send and delete blocked
Gmail, permission by permission. Read, search and draft are allowed. Send and delete are blocked at the connector.
SYRIX Settings: General — default mode Auto, Spark or Zenith; sleep when idle after 15 minutes; Enter to send — with sections for Appearance, Voice, Privacy, SYRIX, Supercomputer, Connectors and Advanced
Settings, rebuilt. Every control on it is wired to something — the ones that were not were removed rather than left as decoration. Identity and tone live in SYRIX.md, opened from the SYRIX page.

The rules underneath it

  • Speech is recognised and spoken on my own machine, not on somebody else's API
  • In Private mode a question is answered by a local model or not at all
  • The microphone is idle until dictation or the voice orb asks for it
  • Every log line goes through the redactor before it is written
  • Keys load from .env at boot and never enter config.json
  • 1,427 settings, and not one of them holds a credential
  • Arithmetic is not the model's jobA figure worked out in an answer comes from a program that ran. The model phrases the result and nothing else.
  • Effort buys time, never a bigger replyA request is priced at its prompt plus its reserved reply, and the free lanes refuse what their minute cannot cover — so a hard question that reserved more was the likeliest to land on the weakest model. Hard effort now buys rounds and sources, and a test asserts the price against the real caps.
  • A fallback must not read like oneFree lanes rotate as each one runs dry, and the answer must not reveal which one it came from. An empty reply costs a retry, not a cooldown, so a lane that hiccups once is not lost for five minutes.
  • Fail closedAmbiguous command: ask. Unknown domain: deny. A lane that cannot call tools is told so, rather than writing the call out as its answer.
  • It yieldsA resource governor hands CPU, GPU and RAM back the moment the machine is wanted for something else, so a long task can run while I am using the PC.
0commits, 14 May → 7 Oct 2026
0lines of Python
0lines in the Electron HUD
0lines in the v3 core

SYRIX has its own website.

A 3D, scroll-driven site for the project — Three.js, a boot sequence and a raymarched core — rewritten in October for the v3 core, with a build journey that runs from a clap script to the ecosystem.

Open the SYRIX site →
The SYRIX website: the word SYRIX in gold over a dark sphere
Project 02

Kohli

A private family hub — lists, jobs, dates, meals, money, documents, markets and an assistant, on my family's actual phones. Plain HTML, CSS and JavaScript. No framework, no build step, no packages, nothing to install. Twenty-five screens, and it installs to a home screen like an app.

  • Vanilla JS
  • PWA + service worker
  • Cloudflare Workers
  • AES-256-GCM
  • PBKDF2 · 310k
  • Canvas charts
  • Open-Meteo
  • OpenStreetMap
  • Zero dependencies

Twelve screens, one turn

Real captures · light and dark

Kohli — the in-app SYRIX family assistant screen
Assistant
Kohli — the finance screen with UK indices and a sector heatmap
Markets
Kohli — a full stock chart for AAPL with volume and indicators
Charting
Kohli — the planner showing a week of household jobs
Planner
Kohli — the encrypted family-code lock screen
Lock screen
Kohli — weekly meal planning
Meals
Kohli — documents sorted by expiry
Documents
Kohli — month calendar
Calendar
Kohli — Simple Mode, four large buttons
Simple Mode
Kohli — shared shopping lists
Lists
Kohli — household jobs
Jobs
Kohli — search across everything on the device
Search

What it does

  • 25 screens · one codebase · light and dark
  • Installs to a home screen · works offline
  • Encrypted on device · syncs across four phones
  • Built-in AI assistant that reads and edits the app
  • Live markets — London, New York, Mumbai
  • Charting with indicators, drawn on canvas
  • Live map, location sharing that expires itself
  • Meals → shopping list in one tap
  • Documents sorted by what expires first
  • Simple Mode — four buttons, for grandparents
  • Voice input · photo recognition
  • Zero dependencies · zero build step
  • AES-256-GCMThe whole store — messages, photos, locations, prices — is ciphertext until the passcode is entered.
  • 310,000 roundsPBKDF2 stretching from passcode to key. Five wrong codes and the wait doubles up to a minute.
  • Re-locks itselfFive minutes in the background and it closes again. A passcode that only guards opening isn't a passcode.
  • Keys never in the appAnything shipped to a phone can be read off it. The model keys live in a Cloudflare Worker gateway instead.
  • Sync it can't readWhat goes up is encrypted with a key derived from the family code, and that code never leaves the device.
  • Never coordinatesThe assistant is told “at home” or “1.2 mi away”. It is never given a latitude and longitude.
Kohli lists screen in light mode
Light
Kohli planner screen in light mode
Light
Kohli lists screen in dark mode
Dark
Kohli planner screen in dark mode
Dark

Light and dark — both real captures, not recolours.

What it grew

  • An education desk: apprenticeships, degree apprenticeships and universities
  • An eight-stage application tracker, from saved to offer
  • Entry requirements read from GOV.UK and checked against real grades
  • Photos, GIFs and a photo editor in the family thread
  • The family hub rebuilt as a deck of live keys, each carrying a figure
  • A lock screen that does not tell anyone how long the code is
  • Paths are signed, not encryptedThe merge has to match records across phones, so the same record must produce the same name every time. The name is an HMAC; the real path travels inside the ciphertext.
  • 210,000 roundsPBKDF2 from the family passcode to two separate keys, one for AES-GCM with a fresh IV per write and one for the HMAC above. Neither is ever written to disk.
  • Written only when it changedThe gateway compares the merged state with the held one and skips the write if they match — four phones with the map open would otherwise spend a tenth of a day's free-tier writes on nothing.
  • A streamed reply that drops its tool call is re-askedMeasured at nine failures in ten on a question that answers reliably without streaming. An empty reply is the worst failure available: the app looks broken and there is nothing to retry from.

SYRIX, inside this one

The same assistant, built for one household and given the run of it. It is not a separate chat surface: it holds the app's own context — who is using this phone, which screen they are on, what is on it — and it can act on the data through the same store the interface uses. One controller drives both the chat tab and the per-screen overview cards, so there is no second pipeline to drift out of step.

Reads10

screen_facts list_goals list_tasks list_items list_events list_bills list_budgets health_summary read_messages recall

Writes16

add_goal update_goal add_task update_task complete_task add_item tick_item add_event add_bill mark_bill_paid log_food log_water send_message set_sharing delete_record remember

Reaches outside7

search_web read_url compare_prices get_weather market_find market_quote market_overview

The whole surface, named. They run one at a time rather than in parallel, and every write passes a permission check before it is allowed to touch the store.

  • It knows who it is talking toIdentity questions are answered in code from a table of twelve patterns, before any model call — and it addresses each person by their own relation, so the same assistant calls me by name and calls my mother's son by his.
  • Never handed a coordinateLocation reaches it as “at home” or “1.2 miles away”. A latitude and a longitude are never put in a prompt, because a prompt is the one place data goes that you cannot take back.
  • A write asks firstEvery tool that changes something passes a permission check, and anything irreversible raises a confirmation card. One at a time, because a second card appearing over the first is a way to get somebody to agree to something they did not read.
  • Checked against its own toolsA grounding pass tests the finished reply against what the tools actually returned, so a figure on a family screen cannot be one the assistant made up.
syrix/gateway.js
/* Six model providers, none of them reliable. A lane is not picked
   because it is preferred — it is picked because it has been
   answering, and answering quickly. */
const BACKOFF_429 = [30, 60, 120, 240, 480];

function reward(state, ms) {
  state.fails = 0; state.until = 0;
  state.score   = state.score * 0.7 + 0.3;              // EWMA
  state.latency = state.latency * 0.7 + (ms / 1000) * 0.3;
}

function candidates(env, tier) {
  /* A dead key or a retired model is quarantined for a day, not
     retried every request for the rest of the afternoon. */
  out.push({ provider, model, state,
             rank: state.score / (1 + state.latency / 10) });
  return out.sort((a, b) => b.rank - a.rank);
}

Exponential backoff, an exponentially weighted moving average of score and latency, and a latency-discounted rank. One model was left out of the table entirely: it is fast and it advertises tool calling, but handed a tool schema it returns an empty array and a page of garbage. Every lane in there was checked against a real tool call before it was listed.

0screens
0lines in the repository
0test files · 12,784 lines
0packages installed
VANTAGE project mark
Project 03

VANTAGE

A multi-asset research and trading terminal, on my own machine, against free keyless endpoints. A hand-written canvas chart, a portfolio across three books that are never added together, a newsroom, a screener, and a research desk that takes a position and then has that position checked against the evidence it was built from.

  • Electron
  • Canvas 2D
  • Node built-ins only
  • Yahoo · Cboe · Binance
  • SEC EDGAR XBRL
  • DevTools Protocol
  • 19 indicators
  • Zero dependencies

What the window is made of

One real capture — NVDA, live, 28 August — taken apart along the application's own panel boundaries.

The VANTAGE terminal: a one-year NVDA candle chart with volume, a drawing rail, a watchlist with company marks, the research column, the paper blotter and the order ticket

Every figure on it was fetched, checked and given a freshness state before it was drawn.

What it does

  • Four canvas layers, each repainted only when its own inputs change
  • Either chart in the same frame — this engine, or TradingView’s, switched from the toolbar
  • Bar sizes from one second to one month, on a fractional-index viewport
  • Candle · OHLC · line · area · Heikin Ashi · Renko · Range
  • 19 indicators, checked against hand-computed and published values
  • 14 drawing tools, anchored in time and price, not pixels
  • Lot-based positions from a ledger — FIFO, LIFO, average, shorts
  • Risk: exposure, concentration, beta, correlation, drawdown
  • Paper broker with five order types and stated fill assumptions
  • Newsroom, treemap heatmap, screener, economic calendar
  • Every request in the main process — the page has no network at all
  • Nothing is live on its own authorityA provider has to declare real-time and zero delay. Silence means delayed, and the badge says so.
  • A blocked figure is not publishedShowing a number with a warning still puts the number on screen. It is withheld instead, and the reason is named.
  • Three books, never summedReal holdings entered by hand, paper, imported. Every total carries its own book's caveat.
  • Positions are derived, never storedThey come out of the transaction ledger, so an average entry cannot go stale against its own lots.
  • No total without a real ratePer-currency subtotals and the reason, rather than one number built on an assumed exchange rate.
  • Worst freshness winsA portfolio total inherits the least trustworthy figure inside it, and unpriced holdings are named rather than dropped.
The VANTAGE Discover screen: a squarified treemap of the United States market, tiles sized by market value and coloured by the day's move, with company marks on the larger tiles
Discover — a squarified treemap, laid out in code
engine.js
/* A filed revenue and a filed profit from different years produced a
   profit margin of four hundred per cent, and each figure was
   individually correct. `fy` and `fp` name the REPORT a fact appeared
   in — a 10-K carries two years of comparatives — so one fiscal-year
   label is attached to three different revenues. Periods are keyed on
   their real start and end dates instead. */
function periodKind(row) {
  const days = (row.periodEnd - row.periodStart) / 86400000;
  if (!row.periodStart) return 'instant';   // a balance-sheet moment
  if (days >= 300) return 'annual';
  if (days >= 60) return 'quarterly';
  return 'other';
}

Two figures are only divided by each other when they cover the same period. A balance-sheet figure with no instant at the anchor's year end returns null rather than the nearest one to hand.

SYRIX

The council

One question, split nine ways, each pass reading only its own slice of the evidence.

pricePrice action
technicalsMomentum
newsThe story
fundamentalsThe business
earningsEvents
ownershipInsiders
positioningConsensus
sectorThe tape
riskThe case against
chief …
  1. The dossierNine sections gathered in code, not chosen by the model — price with swing levels, consensus, accounts, news bodies, insider filings, sector, peers, holdings.
  2. The conductorPicks the roster on the fast lane. A full council is eleven requests and the gateway allows twenty a minute, so a trimmed roster is stated in the answer.
  3. Nine specialistsDifferent questions over different evidence, in waves of three, so the gateway sees a queue rather than a burst.
  4. The chiefWrites the answer from their reports. Every dated fact carries the gap between then and now, worked out rather than eyeballed.
  5. The reviewerChecks the answer against the dossier and triggers one repair round. A view is only accepted with the price that would prove it wrong.

Why it is built this way

  • One model asked eight things at once hedges“Would you buy this, and where would you get in” is not one question — it is price structure, the tape, the accounts, the calendar, who is selling, the sector, and the case against. Asked together you get four paragraphs that mention every consideration and commit to none. Asked separately, each pass reads a small slice closely.
  • The members are specialists, not copiesThe published pattern runs the same question past several models to average out their weaknesses. There is one model behind this gateway, so averaging it against itself buys nothing. Splitting the question buys the thing that matters.
  • The evidence is gathered before anyone is askedA model that chooses its own tools fetches what it happens to think of — it reads the news and forgets the insider filings. A checklist gathered in code is complete every time, and what was unavailable is named rather than silently missing.
  • And a figure nobody fetched cannot be checkedBecause the dossier is a plain object built before a word is written, the reviewer at the end can compare the answer against it mechanically instead of taking its word for anything.
  • The roster is trimmed, out loudA full council is eleven gateway requests and the allowance is twenty a minute, so the conductor takes what is affordable and the answer states that it ran short-handed. A degradation nobody is told about is a feature that silently does not work.
  • Waves of threeThe specialists run three at a time, so the gateway sees a queue rather than a burst — and the token budget is spent on reasoning rather than on transport, because nothing has to be re-fetched per round.
  • The review is against the evidence, not the other membersRanking answers by peer vote rewards whichever one sounds most certain. The reviewer is handed the dossier and asked what in the answer is not supported by it, and one repair round follows.
  • Read the stance looselyThe brief asks for one exact form and models write it a dozen ways. A parser insisting on the exact phrasing found one stance in seven and reported the rest as having no view — a parsing failure presented as a desk that would not commit.

A guard that removes the conclusion cannot distinguish a sound one from an unsound one. A check against the evidence can — so the sentence that takes a judgement stays, and every figure in it has to appear in evidence that was actually fetched, with the price that would prove it wrong.

0lines, no lockfile
0unit assertions
0checks against the running app
0packages installed
Kabirr Education Partners mark
Project 04

Kabirr Education Partners

A UK apprenticeship and university tracker that runs on my own PC. It finds vacancies across every free source that legitimately publishes them, reads the adverts, checks them against published entry requirements, and drafts the application. It never submits one — a test asserts it stays that way.

  • Node built-ins only
  • 127.0.0.1 only
  • Cloudflare Pages · KV
  • Hand-written PDF + DOCX
  • GOV.UK Find an Apprenticeship
  • UCAS tariff
  • Service worker
  • Zero dependencies
KEP

What is left after each check

The line below is every apprenticeship the application is holding, and how many are left after each check it makes. It is written left to right, in the order the checks run.

  1. 3,877held on my own PC
  2. 127advert read in full
  3. 30not ruled out
  4. 12requirements published, and met

Reading it left to right: every vacancy the application holds, and how many survive each check it makes. The distance between the first figure and the second is the point — a search listing publishes no entry requirements, so nothing in it has been compared against anything. The application says so on its own screen rather than letting 3,877 read as 3,877 opportunities.

KEP

The application fills itself in

Nothing on this sheet is typed. Every line is read out of something, and says what.

application pack 0%
  1. GradesGCSE maths 5 · English 5 · BTEC Level 3 Computer Scienceprofile · confirmed by the candidate, never assumed
  2. Scalevocational — Distinction = 4, not the letter it looks likechosen by the published qualification type
  3. CV2 pages · 1,140 words recovered · confidence: goodzlib inflate → object streams → Tj/TJ → /ToUnicode
  4. Evidence11 skills, each carrying its document and lineshown in work, never inferred from a job title
  5. VacancyLevel 4 · closes in 19 days · requirements publishedGOV.UK Find an Apprenticeship, read keyless
  6. Checkmeets the published requirementsgreen · amber · red — amber means read it yourself
  7. Relevance8 of the advert’s own terms coveredtoken overlap, counted — not a judgement about you
  8. Packassembled · every block traced to a profile fieldno language model writes a claim about a person
Not submitted

It opens the employer’s own form and stops. GOV.UK applications go through One Login with two-factor, so the sending stays with the candidate — and a test asserts the words that would break that promise never appear.

The Apprenticeships screen: a count of what can be applied for, split by whether requirements were published and met, and category filters across every sector
Counted three ways, because the honest answer needs three. What can be applied for, what published no requirements to check against, and what asks for a grade that is not held — kept apart instead of averaged into one number.
The Career pathways screen: what can be proved from uploaded files, a warning that most of the stored database has not been read yet, and pathways ordered by openings not ruled out
Pathways ordered by openings, not by opinion. And it says on the screen that most of the database has not been read yet, so the counts underneath are not mistaken for the whole picture. Real captures, driven against a stand-in candidate — the interface is the real one, the grades on it are not.

Sources, and the ones it refuses

  • GOV.UK Find an Apprenticeship — read keyless, from the public pages
  • DfE Display Advert API v2 — a free self-registered key, OpenAPI 3.0.1
  • Employer ATS boards — the endpoint their own careers page calls
  • Reed and Adzuna — documented APIs, free tiers
  • Every UK university, and course pages where they are server-rendered
  • DuckDuckGo, then Bing RSS, then Wikipedia — in order, stopping at the first that answers

Ten more sources are named on the Sources screen as deliberately not read, each with the published reason. LinkedIn's User Agreement prohibits automated collection. Workday is public and unauthenticated but undocumented, and not offered for syndication.

  • It identifies itselfA real user agent naming the application, with a contact route. Never a copied browser string.
  • No bot-defence handling, everNo CAPTCHA solving, no cookie replay, no headless browser to get past a challenge. A host that puts one up is marked blocked and the app carries on without it.
  • A cache that can miss records the missOtherwise the retry loop runs at network speed. Found twice in the family app; encoded here.
  • Polite per host, measuredThe public service caps its search at ten results a page, so 4,453 live vacancies are 446 pages. One page every 700ms is seven and a half minutes; three at a time is forty seconds.
  • Loopback onlyBound to 127.0.0.1, with an Origin check on every state-changing request. Nothing else on the network can reach it.
  • Ten milliseconds of CPUThe phone copy streams the database straight out of KV without parsing it — parsing costs nineteen milliseconds and the free plan allows ten.

SYRIX, inside this one

A third build of the same assistant, cut down to one job. It has no household to run and no markets to read — it can see the profile, the grades, the uploaded files and whatever is on the screen in front of it, and its tools do nothing but search, read and check. Narrowing what an assistant can reach is the cheapest way to make it reliable.

  • Its own device token against the same gateway — and no provider key at all
  • 17 tools, an eight-round bounded loop
  • A credential scan over the whole codebase, as a test
  • Forbidden phrases, pinned: a good chance · you are likely · should be fine
  • Tokens are held backUntil a round commits to answering. All eight rounds narrate before asking for a tool, so text appeared and was withdrawn eight times — streaming nobody can read is worse than none.
  • Older evidence is squeezed firstThe free lanes cap context and output together, so six results at full size leave a few hundred tokens to answer with. The oldest are trimmed to 1,200 characters; the newest two keep 4,000.
  • Seeded, then disarmedVerified rows are injected as a tool result and the research tools are then withdrawn — handed real rows, a lane would still go and invent universities to fill out the list.
  • A promise is not an answerA reply announcing a lookup it never made is refused, and the loop runs again. Bounded, because a model that will not stop promising must not be allowed to spin.
0lines, hand-written
0vacancies stored
0assertions
0packages installed
Project 05

Backtest Lab

Five trading strategies, run against equities, FX, futures and crypto on one engine, ranked by the metrics that survive contact with real money. I built it because I wanted to know whether the setups I read about actually hold — and the honest answer needed a machine.

  • Python
  • backtrader
  • pandas · numpy
  • yfinance
  • ccxt
  • ATR risk sizing
Entry ATR stop 2R target —

SMC_BOS on a synthetic series — the engine's wiring-test mode, which the README labels clearly because these numbers carry no meaning for live trading. The mechanics are the production ones: position size derived from stop distance, ATR stop, fixed 2R target, one position at a time.

SMC_BOSBreak of structure — swing high/low breakout
RSI_TrendRSI pullback filtered by the 200 EMA
Boll_MRBollinger mean reversion
EMA_Ribbon21/55 cross, trend following
Donchian20-bar channel breakout, turtle-style
strats.py
class BaseRisk(bt.Strategy):
    """Shared: ATR stop, RR target, % risk sizing, one position at a time."""
    params = dict(risk_pct=0.01, atr_period=14, atr_mult=2.0, rr=2.0)

    def size_for_risk(self, stop_dist):
        if stop_dist <= 0: return 0
        risk_cash = self.broker.getvalue() * self.p.risk_pct
        return max(int(risk_cash / stop_dist), 0)

Position size is derived from the stop distance, never guessed. Commission is modelled at 0.05% a side. One position at a time, so results cannot be inflated by stacking entries.

What the README says before it says anything else

  1. Out-of-sampleBacktest 2020–2023, then check 2024–2025 separately.
  2. Walk-forwardDoes the edge hold in chunks, or was it one lucky run?
  3. Forward paper trade100+ live paper fills before a penny of real money.
  4. Real costsSpread and slippage, not just commission.
  5. Beware overfittingTune the parameters until it looks great and it will break live.

The engine also ships a synthetic mode for wiring tests, clearly marked so those results are never mistaken for evidence. An over-fitted backtest is worse than none.

Project 06

GMC Hub

GMC Hub project mark

The internal operating platform for Golden Management Consultancy, an immigration consultancy. Thirty-five screens over one Postgres database: clients and cases, projects, tasks, calendar, email, team chat, documents, a full HR module, automations, reports, and an assistant that answers immigration questions from published sources and can act on the Hub when a person approves the card. Built to a written specification I was given, phase by phase, on a stack I did not choose.

  • Next.js 16
  • TypeScript
  • Drizzle · Postgres
  • Clerk · MFA
  • Tailwind v4 · shadcn/ui
  • TanStack Query
  • Upstash Redis
  • Cloudflare R2
  • Vitest
GMC

Who can reach what

Every screen shows a person only what their role allows, and it is decided on the server — hiding a button is not a control. The mark fills to how much of the platform each role can actually reach.

The Golden Management Consultancy mark
the platform 30 resources it protects
one database Every read scoped before it runs, on the server
  1. EMPLOYEEOwn work. The tasks assigned to them and the clients they are staffed on — a membership row, not a guess.
  2. TEAM_LEADTheir team. Scope compares the caller's team against the resource's own team, so an unassigned task is still theirs to hand out.
  3. MANAGERAcross teams, plus reports. The first role that sees work it is not staffed on.
  4. ADMINThe platform: people, roles, integrations, connectors, automations, the kill switch. Everything except the record of what it did.
  5. SECURITY_ADMINSessions, the audit log, security events — and deliberately no business data. Kept apart from ADMIN so one account cannot both act and erase its trail.
  6. HR_ADMINIdentity, pay, leave, absence, documents, rotas, hiring, safety. Deep rather than wide: ordinary business scope underneath.

Built to a specification, not to taste

The stack was locked before the first commit and the work was sequenced into ten phases, each with a written bar for being finished. Changing a locked choice needs a written reason in the dependency file, which is where the two changes that happened are recorded.

  • Phases 0 to 7 done: foundations, identity, core work, the queue, integrations, the full workflows, reports, and the assistant
  • One branch. Typecheck, lint and tests gate every commit, because there is no second person to review a merge
  • Continuous integration on every push: lint, types, tests against a real Postgres, dependency audit, secret scanning and static analysis
  • Free tier throughout, verified rather than assumed — the host that bans commercial use on its free plan is named in the specification as the trap it is
  • Every check is server-sideHiding a button is not a control, so a hand-written request gets the same answer as the interface. Permissions are resource-level and action-level, and the scoped query is the one the screen already uses.
  • A capability sits beside a role, not above itA person holds exactly one role. Finance access and HR administration are granted alongside it, because an administrator who took the HR role to configure leave would lose the screen that grants roles, with no way back. Neither capability widens business-data scope.
  • A dead integration degrades a featureIt never takes the app down. External systems stay authoritative for their own data and the Hub stores a reference and a sync state, so a mail outage costs the mail screen and nothing else.
  • Thirty-five screens are rendered as a real employee, on every runOnly the session is faked; every permission predicate and scoped query runs for real against the database. It catches the screen that typechecks, passes its tests and still throws when something actually renders it.
  • Test rows are identified by a reserved addressNever by a name pattern. A suite that died before its cleanup had put six accounts into the live staff directory, and a pattern match would eventually delete somebody real.

SYRIX, inside this one

The same assistant again, tuned for immigration casework and for this company's own data. It answers from published sources with the figures cited, reads the Hub when the person asking is allowed to, and proposes an action as a card that shows the real parameters it would run with. It is free to use and holds no provider key: the Hub pairs with a gateway I own and is handed a revocable device token.

  • Seven routes — conversation, itself, the Hub, reasoning, an action, research, a draft
  • Seven immigration specialists in the deep lane, and a reviewer that reads the answer back against the evidence
  • Nothing is stored: the transcript lives in the page, and a new chat destroys the old one
  • No prompt or reply reaches the database, the logs, or the error tracker
  • A prompt is advice; a branch is a controlThree attempts to fix routing by sharpening the instructions all failed — a sum of billable hours kept going to the Home Office and citing a visitor visa page. It took a rule: an instruction to compute, three or more quantities, and no word that can mean immigration. Hub access is a branch before the broker for the same reason, so nothing the model reads can argue it out of.
  • Search for the rule, not for the caseThe applicant's own figures are the most distinctive words in the question and the least useful for finding a page. No page is about 210 days; every page about the rule is about long residence. A short rule-name query runs alongside the detailed one, and ranking scores the URL path — a slug is the one part of a page that cannot be padded with the question's words.
  • A question that turns on a published rule never falls back to memoryHowever badly retrieval went. A backstop that measured coverage against the raw question read a misspelling as "found nothing" and answered from memory: a threshold three years out of date, with a retrieval date it had invented. That is exactly when memory is most confident and most stale.
  • The approval card shows the parameters, not the intentFully rendered, as they would run. The proposal is held on the server and the client approves an identifier, so what executes is the Hub's own action with the arguments that were on the card.
  • It is not allowed to be neutered by its own safety systemThe guards fix structural failures — narration leaking into an answer, a figure nothing supports. None of them may refuse to answer, because a guard that fires on a correct answer is a guard that gets switched off. Two that did were found by probing and removed.
0commits, 6 → 15 Sep 2026
0lines of TypeScript
0tables, 35 migrations
0lines of tests
Project 07

GMC Lead Finder

GMC Lead Finder icon

A local app that finds organisations and people with a current, evidenced need for the consultancy's services, shows how to reach them, and retires a lead only when a source shows the need has passed. It reads six public sources on a schedule and every lead carries the published words that make it one. Lead data never leaves the PC.

  • Node.js · built-ins only
  • OCDS procurement data
  • Companies House
  • CSV register diffing
  • RSS · Atom
  • SYRIX
  • node:test
  1. 01 · readSix sources Tenders every three hours from Find a Tender and Contracts Finder, the sponsor register every six, the penalties report daily, and two communities hourly.
  2. 02 · matchOne service, with the words A lead names which service fits and prints the wording that matched. Sellers, other countries' routes and paperwork the firm does not do are rejected before scoring.
  3. 03 · proveWho they are Companies House for status, directors and registered office. A website is accepted only if it shows the company number, the postcode, or the name with the town.
  4. 04 · rankA score with its reasons A sorting aid, not a verdict — every point it adds is printed beside it.
  5. 05 · retireOnly with evidence A closed or awarded tender, a restored rating, a dissolved company, a deleted post. A lead someone has contacted is never aged out.
GMC Lead Finder: a list of tender leads with SYRIX scores on the left, and an open lead on the right with its SYRIX AI overview and the published words that make it a lead
A lead, open. The score with its reasons, a SYRIX overview of the lead, and the tender's own words as the evidence. Public tender notices only; contact details removed.
The Lead Finder with its SYRIX panel docked on the right, offering four starting questions
SYRIX, docked beside it. The Lead Finder edition: it reads every lead and can change none of them.

Dated by the source, never by the app

Every lead shows when the source first published the evidence, not when this app found it. Leads read from a register were showing "today" for companies incorporated in 2022 — so each kind of lead has its own date and its own limit.

LeadDated byToo late after
PostWhen it was posted30 days
TenderFirst publication, read from the procurement's own historyThe deadline; with none, 90 days
PenaltyThe period the penalties were issued inA year after it ends
New sponsorBetween two daily registers compared90 days
Expansion sponsorBounded by the company's incorporation dateTwo years — the route's maximum
  • A notice that was amended is not newFind a Tender re-publishes old frameworks every time they are amended, and one first published in 2021 arrived looking current. The first publication is now read from the procurement's own record history and kept with the lead.
  • Polite to every source, by hostOne request every four seconds to the tender service and every thirty to the community feed, an honest user agent, and a long pause on a refusal. A search engine that answers with an "anomaly" page is detected and left alone for six hours rather than retried.
  • Names collide unless the key keeps the detail"Group", "Holdings", "UK" and a bracketed place are kept in a company's key, or two different firms merge into one lead — and a live company is preferred over a dissolved namesake.
  • Quality was measured, not tuned by feelThe community rules were written against 98 real posts and took accepted leads from 40 down to the 12 that were genuinely asking for help.
  • SYRIX, read-onlyA side panel with nine tools: totals, search, counts, one lead in full, the service list, a live Companies House lookup, official guidance from GOV.UK, legislation.gov.uk and the ICO, the web, and one page read. None of them can change a lead.
0commits, 5 Oct 2026
0lines, zero packages
0offline tests
0read-only SYRIX tools
SYRIX

One assistant, built five times

The same core in five applications, and a Deck in the HUD that launches every one of them. It is not five models and nothing here is retrained — what changes is what it can reach, who it asks, and what it is not allowed to say.

The SYRIX mark
one core 5 applications carry it
unchanged One router, one gateway, one grounding check — and five different jobs
  1. Kohli · the householdThe widest. It holds the app's own context — who is on this phone, which screen, what is on it — and writes to the same store the interface uses. Ten tools read and sixteen write, one at a time.
  2. VANTAGE · the deskNine specialists over a dossier gathered in code before any of them was asked. It reads everything on the desk and writes nothing to it: the output is a view, with the price that would prove it wrong.
  3. GMC Hub · the caseworkSeven routes and seven immigration specialists. It reads Hub records only if that employee's own setting allows it, and every action is a card showing the parameters it would run with.
  4. Education Partners · the applicationThe narrowest, deliberately. Seventeen tools that only search, read and check, in a loop bounded at eight rounds. It drafts the application and a test asserts it never presses submit.
  5. Lead Finder · the pipelineNine tools and not one of them writes. It reads the leads, Companies House and official guidance, and a per-lead overview is cached against a hash of the record, so it is redone only when the evidence changes.

What is the same in all five

The parts worth building once. Every fitting inherits them, which is the only reason five assistants were affordable to build at all.

  • A figure in an answer is computed in code; the model is allowed to phrase it and nothing else
  • The answer is checked against what was actually retrieved before it is shown
  • Questions about what it is are answered from a registry, so it cannot be talked out of the answer
  • External text is evidence, never an instruction — detection annotates, it never permits
  • No provider key in any application: each pairs with the gateway and holds a revocable device token
  • Free lanes throughout, degraded in depth rather than refused when one runs dry

And what each one had to be cut down to

  • Kohli — a prompt is the one place data cannot be taken back fromSo location reaches it as "at home" or "1.2 miles away" and never as a coordinate, and identity is answered in code from a table of relations rather than guessed, so it addresses each person as what they actually are to the person asking.
  • VANTAGE — the members are specialists, not copiesThe published pattern averages several models against each other. There is one model behind this gateway, so averaging it against itself buys nothing; splitting the question does. The reviewer is pointed at the dossier rather than at the other members, because ranking by peer vote rewards whichever answer sounds most certain.
  • GMC Hub — a prompt is advice, a branch is a controlWhether it may read Hub records at all is decided before the broker runs, so nothing the model reads can argue it out of it. Its routing rules are code for the same reason: three attempts to fix a misrouted question by sharpening the instructions all failed.
  • Lead Finder — a stream can end without an errorSome lanes closed an overview half way through a sentence and reported success. Overviews are requested whole, checked for completeness and retried on the other tier, and a chat answer that stops mid-sentence is fetched again in full.
  • Education Partners — seeded, then disarmedVerified rows are handed in as a tool result and the research tools are withdrawn, because given real rows a lane will still go and invent more to round out the list. Four phrases are pinned as forbidden — a good chance · you are likely · should be fine — since this one is read by somebody deciding what to do next.

The pattern across all five is the same: the model is the least trusted part of the system, and every build spends most of its code deciding what it is allowed to see and checking what it gave back.

git rev-list

Eight repositories

480 commits on SYRIX, 60 on Kohli, 22 on VANTAGE, 172 on GMC Hub, 18 on Lead Finder, 6 on the film, 3 on the catalogue, and one squashed import for Education Partners.

total 762 commits, eight repositories
  1. SYRIX48014 May → 7 Oct 2026
  2. Kohli601 → 26 Aug 2026
  3. VANTAGE2226 → 27 Aug 2026
  4. Education Partners126 Aug 2026 · squashed import
  5. GMC Hub1726 → 15 Sep 2026
  6. The film626 → 27 Sep 2026
  7. Lead Finder185 Oct 2026
  8. Catalogue35 Oct 2026
02

Build history

762 commits across eight repositories, May to October 2026 — counted with git rev-list on each one. Twelve milestones, then the raw log.

14 May 2026 · SYRIXModularisation

A 40,000-line single file split into modules behind a routing layer.

23 Jul 2026 · SYRIXBenchmarked routing

Model lanes measured before the router is allowed to trust them.

24–27 Jul 2026 · SYRIXSupercomputer mode

Isolated browser, permission gate, resource governor, evidence store.

1–2 Aug 2026 · KohliFamily app shipped

v1, an in-app assistant, and a gateway needing no key or open port.

10 Aug 2026 · SYRIXSpark Zenith 1.1

Two-lane runtime complete, with a threat model and a recorded rollback.

26 Aug 2026 · EducationEducation Partners

An apprenticeship desk that reads the adverts and never presses submit.

26–27 Aug 2026 · VANTAGEA terminal, in two days

22 commits, each with its own test, and a handoff that records what is not finished.

6–15 Sep 2026 · GMC HubAn employer's platform

Eight phases to a written specification: 35 screens, 78 tables, and an assistant that asks before it acts.

19–24 Sep 2026 · SYRIXThe v3 core

A baseline committed first, then one loop with native tool calls — 1,800 lines against the 62,000 it replaced.

26–27 Sep 2026 · FilmFour minutes from code

Three cuts in two days, each frame a function of time, rendered at 4K and 120 fps.

27 Sep – 7 Oct 2026 · SYRIXVoice, Private, Deck

A voice orb in the chat, a mode in which nothing leaves the PC, and a launch pad for the other apps.

1–7 Oct 2026 · GMCLeads and a catalogue

A lead finder that dates every lead by its source, and questionnaires the website itself cannot read.

From git log

unedited · 89 of 762
  1. SYRIX

    Prepare Jarvis/SYRIX for safe GitHub sync

  2. SYRIX

    Complete safe modularization routing phase 1

  3. SYRIX

    tools: free-brain benchmark — prove which lane is best before routing

  4. SYRIX

    feat(supercomputer): Phase 1–2 — safety baseline + isolated mode with fail-closed dispatch

  5. SYRIX

    feat(supercomputer): Phase 3 — isolated embedded browser sandbox

  6. SYRIX

    fix(supercomputer): the real cause — 21k tokens per step was starving the pool

  7. SYRIX

    fix(supercomputer): 16 bugs from the first live drive — safety, scope, answers

  8. SYRIX

    feat(supercomputer): a resource governor, so a task can run while Boss games

  9. SYRIX

    feat(supercomputer): ask for what only Boss knows, never invent it

  10. SYRIX

    fix(chat): repair an answer from the web instead of shipping the gap

  11. SYRIX

    fix(tests): the suite was calling live search APIs, and I raised the ratchet

  12. SYRIX

    perf(spark): 16.4s → 10.5s, and stop refusing before actually looking

  13. SYRIX

    fix: ground self knowledge and long browser tasks

  14. SYRIX

    docs: threat model Spark Zenith 1.1

  15. SYRIX

    feat: complete Spark Zenith 1.1 Vanguard runtime

  16. SYRIX

    fix: make secret guard Unicode-safe

  17. SYRIX

    Baseline: 40 questions against the runtime being replaced

  18. SYRIX

    SYRIX v3: the turn engine

  19. SYRIX

    v3 memory: one store, with provenance, that ages itself

  20. SYRIX

    v3 live: the cutover, and the local lane actually working

  21. SYRIX

    Delete what nothing imports, and show what is actually running

  22. SYRIX

    v3 brain: benchmark eval 30/40 -> 40/40, and the fallback chain no longer fails open

  23. SYRIX

    v3 compute: problems are solved by running a program; hard eval 22/25 -> 25/25

  24. SYRIX

    v3: go to a named website or profile, and follow its links

  25. SYRIX

    Private mode, and a spoken-turn pipeline for the voice orb

  26. SYRIX

    HUD: voice orb in the chat, Private mode, account menu, Connectors and Settings

  27. SYRIX

    SYRIX.md: identity, loyalty and tone in an editable file

  28. SYRIX

    HUD: Deck, a launch pad for the apps of the SYRIX ecosystem

  29. SYRIX

    Questions about SYRIX are answered from SYRIX, never from the web

  30. SYRIX

    Close tears down everything at once; pre-v3 replies can no longer leak

  31. Kohli

    Kohli v1: private family hub

  32. Kohli

    Make SYRIX work with no API key, no PC left on and no port opened

  33. Kohli

    Deploy Kohli Core, and fix what only a live deploy could reveal

  34. Kohli

    The passcode has to cover the price cache too

  35. Kohli

    Stop a bio being wiped by somebody else walking down the road

  36. Kohli

    Fix the More panel putting you back where you started

  37. Kohli

    Deploy from a staged allow-list, not from the whole folder

  38. Kohli

    Drop 'unsafe-inline' from style-src, and fix the boot flag it exposed

  39. Kohli

    Finance: a real charting screen, and the family section earns its place

  40. Kohli

    Keypad keys pop on press, and stop the lock screen sliding on a swipe

  41. Kohli

    Rework the lock screen: pill instead of boxes, Enter in the pad, less on it

  42. Kohli

    Education/opportunity screen: eligibility, universities, live driver

  43. Kohli

    SYRIX gateway: display-advert spec, transport, and deploy notes

  44. Kohli

    Chat media, photo editor, and household deck/finance refinements

  45. KEP

    Kabirr Education Partners: apprenticeship discovery and application tracker

  46. VANTAGE

    A private multi-asset research and trading terminal

  47. VANTAGE

    Fix TradingView widget chart rendering nothing

  48. VANTAGE

    Make alerts actually fire

  49. VANTAGE

    Fix Discover heatmap 'Colour by' going flat on two of six options

  50. VANTAGE

    Remove the redundant Watchlist rail tab; give its header a real switcher

  51. VANTAGE

    Un-block the price feed from the bid/ask enhancement, fix a stuck splitter

  52. VANTAGE

    Charts and paper fills keep moving into pre/post-market

  53. VANTAGE

    Fix real logos never showing, and a slow news front page

  54. VANTAGE

    Show a plain-English gist for every instrument, free, before any button

  55. VANTAGE

    Replace the SYRIX depth cycle button with an explanatory menu

  56. VANTAGE

    Fix the AI overview reading fields that have never existed

  57. VANTAGE

    Correct STATUS.md: the research panel's vendor-summary path was never true

  58. VANTAGE

    Revert a same-day misdiagnosis of the AI overview card's data shape

  59. VANTAGE

    Stop the crosshair freezing mid-drag, add a crosshair line-style setting

  60. VANTAGE

    Align Performance holdings table with the paper blotter, reset tab scroll

  61. VANTAGE

    Let a symbol be added to the watchlist by search, not just the chart in view

  62. GMC Hub

    feat(web): Phase 1 — identity & access

  63. GMC Hub

    fix(web): add team column to tasks so TEAM_LEAD can see unassigned team work

  64. GMC Hub

    ci(web): add lint/typecheck/test/dependency-scan/secret-scan/SAST workflow

  65. GMC Hub

    fix(web): remove leaked Resend key from tracked .env.example

  66. GMC Hub

    feat(web): add step-up auth ("Approvals") for sensitive actions

  67. GMC Hub

    feat(web): add GMC AI research pipeline — retrieval, ranking, SSRF guard

  68. GMC Hub

    fix(web): route by subject, not by the words that happen to be in a question

  69. GMC Hub

    fix(web): stop the assistant answering from memory, and let it read pictures

  70. GMC Hub

    fix(web): search for the rule, not for the case

  71. GMC Hub

    test(web): prove one employee's GMC AI cannot reach another's

  72. GMC Hub

    feat(db): flag seed rows in the database instead of in their names

  73. GMC Hub

    feat(web): let one person hold both platform and HR administration

  74. GMC Hub

    fix(web): close an IDOR on expense line items, and drop dead starter schema

  75. GMC Hub

    fix(web): bind step-up auth to the session that passed the challenge

  76. GMC Hub

    fix(web): keep error logs in production, and close the approval bypass

  77. Film

    Showreel v1: 2:38 film of six systems, 4K 120fps

  78. Film

    Showreel v2: 3:20 film, rebuilt sections, new score and sound design

  79. Film

    Bound the encoder's memory: cap decode threads, filter threads, input queue and lookahead

  80. Film

    Showreel v3: glass renderer, rebuilt intro and finale, new sections and drops

  81. Film

    v3 revision: darker glass, SYRIX reaching into the apps, synced handover, new close

  82. Lead Finder

    Lead engine: public sources, enrichment, scoring and evidence-based retirement

  83. Lead Finder

    Stricter matching, verified websites without search engines, retries

  84. Lead Finder

    Community leads: exclude sellers, US routes and existing-visa paperwork

  85. Lead Finder

    SYRIX: chat panel and an AI overview on every lead

  86. Lead Finder

    Reliable overviews, services that match GMC's offer, side panel restored

  87. Catalogue

    Catalogue website as published before the October 2026 simplification

  88. Catalogue

    Simplify the catalogue and its navigation

  89. Catalogue

    Regenerate the questionnaire bundle for the 5 October 2026 deployment

Education Partners has one commit

A single squashed import on 26 August 2026, so there is no history to show. Saying that is better than padding a log with something. What can be counted is the tree.

88files tracked
13,981lines in src/
5,936lines in web/
2,442lines of tests
1,809lines in server.js
1,598lines in cloud/
02 · continued

Where it all runs

All of this started as programs on this PC, reachable from nowhere else. That is safe, and no use on a phone. Six of them are hosted so they can be opened from anywhere — and the work was doing that without giving up what was true while they were local. One of the six is a read-only copy, because the real application cannot run on the free plan, and it says so.

one account Six projects, all inside the free plan
syrixWorkerModel access and paired devices. The model arrives as a runtime binding, so there is no provider key to obtain, rotate or leak, and KV holds the device tokens and the day's counters.
kohliPagesThe family app. Deployed from a staged allow-list rather than the folder it is developed in, and the stager refuses to run if anything staged is shaped like a credential.
kabirr-educationPages · KVApp and API in one deploy — a worker inside the output directory — so it cannot be half-shipped. Its own namespace and its own passcode, with no reach into the others.
kabirr-kohliPagesThis page, and the sites it links to, served under a per-page content policy. The policy is enforced only on the live deploy, so it has its own test.
gmc-hubWorker · read-only copyA captured copy of every screen behind an access code, served with the PC off. Any attempt to save is refused with a plain message — a Worker is a small isolate and cannot host the real application, so it does not pretend to.
gmc-cataloguePages · Functions · D1The catalogue and thirteen questionnaires in one deploy. Its database holds sealed envelopes and nothing it can open, and each deploy deletes the ones before it.

From this PC to a phone

A household planner nobody can open on their phone is a file. Putting these on the internet is the easy half; the half worth writing down is that nothing which was true on a loopback address stopped being true once there was a public URL — no key in anything shipped, nothing listening at home, and every device holding only a credential that can be taken back.

  • One account, six projects, all inside the free plan
  • Workers AI as a binding — no provider key in anything shipped
  • Separate KV namespaces and a D1 database, none able to read another
  • Every binding declared in wrangler.toml rather than clicked together in a dashboard
  • Wrangler run through npx — nothing installed to publish any of it

The family app is a Worker and a Pages project, and forgetting the second deploy leaves the fixes live on this machine and nowhere else. The apprenticeship app puts its API inside the Pages output instead, so there is only ever one thing to ship and it cannot be half-shipped.

What deploying it taught me

Every one of these was found the same way — the deploy reported success and the thing it published was wrong. None of them can be caught on this machine.

  • The production branch is not the defaultPages publishes to main and the repository is on master. Without --branch=main wrangler uploads a preview URL and production carries on serving the old build — which looks exactly like a deploy that worked.
  • Pages does not read .gitignoreA whole-folder deploy publishes whatever is sitting in the folder, at a public address, including anything holding keys. The deploy stages an allow-list instead and refuses to run if a staged file is shaped like a credential.
  • The data goes up before the site doesThe obvious order leaves a window where the page is live against an empty store, and a reader cannot tell that apart from a broken deploy.
  • A big value cannot go on a command lineWindows caps one at about 32,000 characters and the vacancy snapshot is six million, so it is written to a file and uploaded from there.
  • The edge serves the old file for a minuteA deploy is verified with a cache-buster. Stale assets read as a failed deploy and invite a second one on top of the first.
  • A proxied app is served at a host it does not think it hasThe framework refuses a form submission whose origin disagrees with the host it believes it is serving, and trusts the forwarded header over the real one — so the front door has to tell it what it is. Redirects naming the tunnel are rewritten too, or the first one hands the visitor a disposable address.
  • Rebuilding to change an address is a taxThe public URL is compiled into the build, so every new tunnel meant a full rebuild. Pointing it at the permanent Worker instead moved the moving part into a key-value entry, and a new tunnel became one command.
  • A read-only copy must refuse, not failThe captured Hub answers every write with "This is a read-only preview. Changes are not saved." — and lets requests to other origins through untouched, or it breaks the sign-in provider's own setup.
0Cloudflare projects
0D1 database, sealed
0servers of my own
0provider keys shipped
Security

What holds it shut

Six of these are reachable from somewhere other than my desk, one can read email, open pages and run code, one holds another company's staff and client records, and one takes in strangers' answers about themselves. The controls below are the ones that exist in the repositories, written up with the gaps they do not close.

all seven 7 of 7 aligned
closed Nothing turns until every lever is home
  1. SYRIXA written threat model, a nine-step gate, and grants that die with the task
  2. KohliPaired devices, ceilings, and a log that never holds a conversation
  3. VANTAGEA loopback route that serves a registry, not an address handed to it
  4. EducationLoopback only, an origin check on every write, and a credential scan as a test
  5. GMC HubEvery check server-side, step-up authentication bound to the session that passed it, and an audit trail no administrator can reach
  6. CatalogueAnswers sealed on arrival with a key the website does not hold
  7. This siteAn allow-list deploy that refuses anything shaped like a key

SYRIX, which needed the most

It reads the web, opens documents, holds connector tokens and can be given a task to run on its own. That is the whole attack surface in one program, so it has a written threat model — ten invariants, eleven trust boundaries, and a list of the gaps it does not close, dated and kept with the code.

  • External content is evidence, never an instruction source
  • Spark and Zenith share one safety boundary — the deeper mode gets more research, never more permission
  • Private mode refuses every connection that leaves the machine, at the socket
  • Connector tokens in a DPAPI vault, per user, per machine — and no plaintext fallback
  • Every local service binds 127.0.0.1; LAN needs an explicit config flag
  • No wildcard CORS — the loopback origin and port, or an explicit list
  • Diagnostics bounded, local and redacted before write
  • Nine gates, in that orderKill switch, known action, capability, task state, declined-before, scope, permission, timeout, audit. The capability check precedes the permission check so a disabled capability can never even raise a card, and the scope check precedes it so an out-of-scope target is refused rather than asked about.
  • A grant is for one action, one scope, one taskHeld in memory and never persisted. It dies when the task ends, when the mode is left, and on restart — enforced by a boot stamp, so a grant that somehow reached disk could not be honoured anyway. There is no "always allow" and no setting that pre-approves anything.
  • Declining is stronger than not grantingA refused action and scope goes into the task's deny set, so the same effect cannot be reached by asking again in different words or routing through a different verb.
  • A string check cannot catch DNS rebindingAn ordinary-looking name can resolve to a private address, so the host is resolved and refused if any record comes back private — and that is why a fetch is validated again on every redirect rather than once at the start.
  • A policy with a switch is a policy that will be switched offThe URL policy reads no configuration at all. Its only tunable is a cache lifetime. Reads are deny-listed rather than allow-listed, because a research assistant permitted sixteen websites cannot research; the interactive verbs keep their allow-list and their permission gate.
  • Injection detection annotates, it never permitsPage text is marked as data and flagged when it tries to speak to the model, but no score it produces can unlock anything — if a "safe enough" number could, a sufficiently persuasive page would eventually reach it. The real defence is structural: a grant can only be created by a person answering a card.
  • Deny unknown, everywhereAn undeclared verb fails closed at the gate. A connector action with no explicit entry is blocked, safe verbs seed as allowed and destructive ones as blocked, and an action marked not-implemented can never be flipped on.

The other six

  • Kohli — a phone is paired, not logged inIt swaps a pairing code for a random 32-byte token of its own, and with no code set the gateway refuses to issue tokens at all rather than falling back to an open door. Twenty requests a minute per device, four hundred a day, twelve hundred across the household. Any token can be revoked, one or all.
  • Kohli — the log records that something happened, never what was saidTime, endpoint, decision, size, and the device as a six-byte hash: enough to spot one phone hammering the gateway, not enough to say whose it is. A family assistant's usage log is not a place to keep family conversation.
  • Kohli — the passcode covers the cacheThe app opens on a keypad and the data it had already drawn sits behind the same passcode. A lock over the screen but not over what is cached behind it is decoration.
  • VANTAGE — a registry, not an addressThe image route on the loopback server will not fetch a URL handed to it, only one the application has already decided to publish. That is the difference between a cache and an open proxy running on the machine.
  • Education Partners — loopback, and an origin check on every writeBound to 127.0.0.1 with an origin check on every state-changing request, so nothing else on the network can reach it. A credential scan over the whole codebase runs as a test, and a separate test asserts it never presses submit.
  • GMC Hub — the administrator cannot reach the audit trailSessions, the audit log and security events belong to a security role that holds no business data at all, and the platform administrator is not given them. That is the point: one compromised account should not be able to act and then erase the record of having acted.
  • GMC Hub — step-up authentication is bound to the session that passed itA sensitive action asks for a second factor, and the approval that grants is tied to the session that answered the challenge rather than to the account — otherwise passing it once in one place authorises a request made from somewhere else. Found and fixed, along with an identifier on an expense line that was trusted without a scope check.
  • Catalogue — the website cannot read what it receivesAnswers are encrypted with AES-256-GCM on arrival and the key wrapped with a 3072-bit RSA-OAEP public key. The private key is on one computer, which collects every two minutes and deletes the website's sealed copy only after its own copy is saved and read back.
  • This site — an allow-list, not a folderThe deploy is assembled from a list of what may ship and scanned for credential-shaped strings before it goes, because Pages does not read .gitignore and a whole-folder deploy publishes whatever is sitting in it.

And what it does not do

The threat model is a release gate, not a certificate. It records controls that exist, gaps that are open, and what should happen when a control fails — and the gaps are the part worth reading.

  • Two read paths still bypass the canonical URL policy, and are named as doing so
  • The injection wrapper covers one path; other tool output can still reach a prompt as raw text
  • Child processes inherit more environment than they need
  • Nothing here has been audited by anyone but me

None of this is a claim that SYRIX is secure in absolute terms — the document says so in its own second paragraph. Writing down what is not covered is the only version of this worth showing to somebody who will check.

03

Engineering principles

Eight decisions that shape everything above, each one traceable to code in these repositories.

01Computed, not predicted

A figure in an answer is worked out by code that ran — a calculator, or a program the model wrote and an isolated interpreter executed. The model phrases the result and nothing else.

02Fail closed

Ambiguous command: ask. Unknown domain: deny. Missing capability: degrade and report it. The default is never a best guess.

03Local by default

Speech recognition and synthesis run on the machine. Private mode answers on a local model and refuses every connection off the PC. Lead data, and every questionnaire answer, is only ever readable on one computer.

04Permissions are per-action

Gmail can read, search and draft. Send and delete are blocked at the connector, with a DPAPI-encrypted token and an audit trail — not enforced by prompting.

05Measure before routing

Every model in the lane catalogue was called before it was trusted, and the dead ones are written down with how they failed. Provider choice is data, not preference.

06Minimal dependencies

Kohli, Lead Finder and the questionnaires have zero packages and no build step. The HUD transpiles JSX at runtime. None can break from an update I did not make.

07Verify in a real browser

Screenshot harnesses, overflow checks and live acceptance probes. A passing unit test and a working screen are separate claims.

08Run it daily

SYRIX runs on my desktop, Kohli on four phones in my family, and the GMC tools at work. Defects surface in use rather than in review.

03 · continued

Rules that came from defects

Each of these replaced an assumption that had already cost something. The cost is written next to the rule.

09

Identify honestly

Every request carries a user agent naming the application and a route back to me — never a copied browser string. There is no CAPTCHA handling, no cookie replay and no headless browser to get past a challenge. A host that puts one up is recorded as blocked and the run continues without it.

The difference between reading a public page and evading a control is whether you would be willing to say which you were doing.

10

A cache that can miss must record the miss

A cache that refuses to write on a miss leaves every caller retrying a failing source at full speed. It took the family app down twice — a price cache that would not write while the passcode was set turned one render into an unbounded fetch loop — so a remembered failure now gets a short life of its own, ninety seconds against the hour a success gets.

Negative caching is not an optimisation. It is the thing that stops a retry becoming a denial of service against somebody else.

11

Verify the answer rather than suppress it

A guard that deleted any sentence naming an action and a direction meant the assistant could read every filing a company had published and then remove the one sentence that took a judgement — while still printing the readings that implied it. It was replaced by a pass that checks the finished answer against the evidence that was actually fetched, and repairs it.

A guard that deletes the conclusion cannot tell a good one from a bad one. A check against the evidence can.

04

AI tooling

Each of these left a directory, a config entry or a commit on my own machine.

Agentic CLI Claude Code My main driver. Skills, hooks, subagents, MCP servers, custom slash commands. 3,523 files · ~/.claude
Agentic CLI Codex Second opinion and long refactors. SYRIX keeps a handover document written for it specifically. 8,547 files · ~/.codex
Agentic CLI OpenClaw A third opinion when the first two agree with each other too easily. openclaw@2026.4.22 · installed globally
IDE agent Cursor Where SYRIX's early modularisation happened. avan_cursor_mode.py
IDE agent Antigravity Used through the Gemini-backed agent flow. 2,874 files
IDE agent Kiro · Copilot · Continue Spec-driven work, inline completion, and a local-model IDE loop. Continue → Ollama
Protocol MCP — both directions SYRIX is an MCP server over stdio, and a client of five. chrome-devtools · memory · git-mcp · defillama · MS Learn
Local models Ollama The floor under every cloud lane, and the only model Private mode will use. Started on demand and stopped with SYRIX. qwen2.5:7b-instruct
Inference Groq · NVIDIA NIM · OpenRouter Six lanes, raced behind hedge timers, each with its measured latency and failure modes written beside it. gpt-oss-120b · 20b · Nemotron 3 Super · Ultra · GLM-5.3
Orchestration Native tool calling One OpenAI-style tool protocol across every lane. Orchestration frameworks were tried, kept behind flags, and deleted with the old core when nothing imported them. syrix_v3/gateway.py · loop.py
Memory Chroma · DuckDB · plain files v3 keeps one memory store with provenance that ages itself, over a Markdown vault I can read without any of it. syrix_v3 memory · Obsidian vault
Automation Playwright · crawl4ai Real browsers driven headlessly — for SYRIX's research mode, for testing this page, and for rendering the film one frame at a time. five widths, every build
Research Perplexity · Gemini · GPT · Grok For the things a search engine is bad at. never in the critical path of anything I ship
Hosted agents Manus · NemoClaw Long research runs I would otherwise sit and watch. Used alongside the work, never wired into it. outside SYRIX · nothing they touch ships
Media Higgsfield Video and advertising cuts, driven from the same agent surface as everything else. MCP server on my Claude account

Knowing which one to reach for

Having the tools is not the skill. This is the routing table I actually use, and the same reasoning is hard-coded into SYRIX's model router.

The jobWhat I reach forWhy that one
Anything private — money, health, family, keysPrivate mode, locallyIt physically cannot leave the machine: the local model answers, nothing is saved, and the process refuses any connection off the PC.
A spoken replySpark 1.1Latency beats depth when someone is waiting mid-sentence, so voice turns are always Spark, and the first sentence is spoken while the rest is still being written.
Multi-step research with sourcesZenith 1.1More rounds, more sources and sub-agents that each take part of the question, with the answer read back against the evidence before it is shown.
An exact computationA program, not a modelA model predicts digits; a program works them out. The three misses on the hard set were all predictions, and running a program fixed all three.
Long refactor across a big unfamiliar repoCodexIt holds a wide file set without losing the thread. I keep a handover document in SYRIX written for it specifically.
Architecture and debuggingClaude CodeIt challenges assumptions and asks for evidence rather than agreeing, which is what a design decision needs.
Reading a picturellama-3.2-11b-vision, an omni model behind itBoth were measured reading a test image's colour and its text before either was trusted, and the candidates that could not were dropped.
A number in an answerNo model at allEvery figure SYRIX and Kohli quote is computed in code. The model is only allowed to phrase it. That is the single biggest reason they don't invent numbers.
05

Skills

Everything listed appears in code I have written and run. Nothing here is from a tutorial I watched.

Languages

Python · TypeScript · JavaScript (ES5 → modern) · HTML · CSS · GLSL · SQL · PowerShell · Bash

LLM engineering

Prompt engineering · context-window and cost/latency management · model routing and fallback · quota ledgers · function/tool calling · structured output validation · evaluation harnesses · streaming

RAG & retrieval

Retrieval-augmented generation · vector databases (Chroma) · embeddings · chunking · hybrid ranking · query rewriting · semantic verification · freshness gating · citation grounding

Agents

Native tool-calling loops with concurrent tool rounds · sub-agents · multi-agent councils · sandboxed code execution · MCP as both server and client · permission gating · sandboxed browser automation · evaluation sets auto-graded without a model judge

Models & inference

Groq · NVIDIA NIM · OpenRouter · Ollama (local) · gpt-oss 120b/20b · Nemotron 3 Super and Ultra · GLM-5.3 · Qwen 2.5 · Llama 3.2 Vision · lane racing, hedging and rate-limit admission

Speech & vision

Faster-Whisper STT · Kokoro TTS · voice-activity detection with barge-in · echo cancellation · multimodal and vision-language models · PaddleOCR · MediaPipe · YOLO

Frontend

Electron · React · Next.js (App Router, Server Actions) · Tailwind · shadcn/ui · TanStack Query and Table · Canvas 2D · WebGL and raymarched shaders · CSS 3D · Progressive Web Apps · service workers · responsive and accessible UI · light/dark theming

Backend & infra

FastAPI · WebSocket state bridges · PostgreSQL with Drizzle · schema migrations · BullMQ on Redis · Clerk · Cloudflare Workers, Pages, KV and R2 · Wrangler · zero-dependency Node servers · OpenTelemetry · Git

Data & markets

pandas · numpy · DuckDB · backtrader · yfinance · ccxt · Finnhub · Alpha Vantage · Stooq · TradingView · IBKR · technical indicators and risk sizing

Security

AES-256-GCM · RSA-OAEP key wrapping · sealed-inbox design · PBKDF2 key derivation · CSP without unsafe-inline · OAuth token handling · DPAPI-encrypted stores · secret redaction · deny-by-default allowlists · per-action permissions · role and capability models · server-side scoping · step-up authentication · audit trails · threat modelling

Testing & QA

pytest · Vitest · golden regression suites · Playwright · CDP automation · screenshot and layout-overflow harnesses · server-render smoke checks · live acceptance probes · CI with secret scanning and static analysis · cross-device verification

Video & motion

Deterministic frame rendering in headless Chrome · score and sound design synthesised in an OfflineAudioContext · ffmpeg · AV1 NVENC at 4K 120 fps · fragmented MP4 streamed through MediaSource

Data pipelines

OCDS procurement feeds · Companies House · daily CSV register diffing · RSS and Atom · per-host rate limiting with backoff · evidence-dated records with automatic retirement

Currently going deeper on

Formal evaluation harnesses for agent output · walk-forward validation for strategies · real-time graphics and shader work · distributed edge deployment

06

Where each skill went

Thirty-four things, routed into ten builds. Most of them went into more than one, which is the point — a technique is only learned once it has survived a second problem.

  • Python 3.11
  • JavaScript · ES5 → modern
  • Electron · main/renderer split
  • WebGL · GLSL raymarching
  • CSS 3D
  • Canvas 2D
  • FastAPI · WebSocket bridge
  • Cloudflare Workers · Pages · KV
  • PWA · service workers
  • Faster-Whisper · Silero VAD
  • Kokoro TTS · voice activity
  • YOLO · PaddleOCR · Florence-2
  • Chroma · embeddings · RAG
  • DuckDB
  • Model routing · quota ledgers
  • MCP — server and client
  • AES-256-GCM · PBKDF2 · HMAC
  • CSP without unsafe-inline
  • DPAPI · secret redaction · audit logs
  • pandas · numpy · backtrader
  • yfinance · ccxt · Finnhub
  • Playwright · CDP · pytest
  • Multi-agent orchestration
  • Sandboxed code execution
  • SEC EDGAR · XBRL periods
  • PDF and DOCX parsed by hand
  • Rate limiting · backoff · caching
  • RSA-OAEP · AES-GCM sealing
  • Cloudflare D1 · OCDS · Companies House
  • Headless rendering · Web Audio · ffmpeg
  • TypeScript · Next.js · Server Actions
  • PostgreSQL · Drizzle · migrations
  • Role and capability permissions
  • BullMQ · Redis · background workers
Project 01SYRIX

The AI assistant on my desktop, on its own v3 core.

0 of 0
Project 02Kohli

A private family app, running on four phones.

0 of 0
Project 03VANTAGE

A multi-asset research and trading terminal.

0 of 0
Project 04Education Partners

An apprenticeship desk that never presses submit.

0 of 0
Project 05Backtest Lab

A multi-market backtesting engine.

0 of 0
Project 06GMC Hub

An internal platform for the company I work for.

0 of 0
Project 07GMC Lead Finder

Leads dated by their source, retired only with evidence.

0 of 0
—GMC Catalogue

Questionnaires the website itself cannot read.

0 of 0
—The film

Four minutes, every frame rendered from code.

0 of 0
—This page

Hand-written, no framework, no build step.

0 of 0
The desk this was built on: a white water-cooled PC with an RTX 4060, an MSI monitor, a condenser microphone and a mechanical keyboard, lit blue
07

The machine

Everything on this page was built and run on my own desk. Every figure below was read off the machine itself.

  1. RTX 40608,188 MiB VRAM The local model loads here, the film's frames were encoded here, and this page’s shader renders here.
  2. Core i7-12700K12 cores · 20 threads Water-cooled, on a Z690 board. The block in the photo sits on it.
  3. 32 GB RAM2 × 16 GB Corsair Holds a 7–8B model resident with the HUD, a browser and Python all up.
  4. Crucial P31 TB NVMe Every repository on this page, a model cache, a vector store, a backtesting archive and a 4K film master.
  5. 1920 × 1080144 Hz Why the HUD is built to hold a frame budget, not to look right in a screenshot.
  6. The microphonewhere SYRIX listens Idle until the voice orb or dictation asks for it, then transcribed on this machine. None of it leaves the desk.
08

What the software knows about it

Knowing the parts is not the same as writing software that respects them. Everything below exists because this is one desktop with one card in it.

Designed around eight gigabytes

Eight gigabytes of video memory is the constraint the whole assistant is designed around. A vision model, a speech model and a 7B language model do not fit in it at once, so the runtime has to decide what is resident, what is evicted, and what it is prepared to give back the moment the machine is wanted for something else.

  • Whisper transcribes on CUDA in fp16 — half the memory, and the accuracy holds
  • Vision models are unloaded on a cooldown once they stop being asked for
  • A low-VRAM path that changes which models are eligible, not just their batch size
  • Cold storage: weights pushed out of VRAM rather than reloaded from disk
  • Thread affinity, so vision work does not land on the cores the game wants
  • CPU, GPU and RAM ceilings for background work, set as percentages
  • It yields the machineA resource governor watches memory pressure against a threshold and hands capacity back rather than competing for it. A long research task can run while I am gaming, because it is written to lose that argument.
  • Measured, not assumedUtilisation comes from nvidia-smi and psutil, and the hardware panel reads Win32_Processor, Win32_PhysicalMemory, Win32_DiskDrive and Win32_VideoController. The figures beside the photograph came from those calls, not from a spec sheet.
  • Read-only by designThe monitor reports and does nothing else — no process kills, no background loops. A tool that can end processes on the machine it is diagnosing is a tool that can make the fault worse.
  • Storage has a floorA guardian refuses to start work that would take the disk below a set headroom. Filling the drive a model is caching to is a failure that looks like a hundred unrelated ones.
Kabirr Kohli as a small child, absorbed in a tablet
09

About

I'm in London, studying BTech at West Herts College. Before that, GCSEs at Northwood School in Pinner, and school up to Grade 8 at Amity International in India.

Every project on this page was built outside coursework, self-directed and unassessed. One of them is now relied on daily by my family.

What pulled me into computing was not the code, it was the leverage. A rule written once runs a million times without getting tired or bored or slightly wrong on a Friday, and the machine will tell you exactly where you were mistaken if you build it so that it can. That is a very unusual thing to be handed at seventeen. Most of what I have learned since came from being wrong in a way the system could prove.

Language models sharpened that rather than replaced it. They are the first tool I have used that is genuinely capable and genuinely unreliable at the same time, which makes the engineering question the interesting one: not can it answer, but how do you know. Almost everything I have built since is some version of that — gather the evidence in code so it can be checked, compute the figures rather than letting the model say them, test the finished answer against what was actually retrieved, and report the failure rather than papering over it. The verifier that fails on screen in the recording above is the whole point, not an embarrassment.

The part I want to go further into is evaluation: how you measure whether an agent is actually getting better, rather than whether the last demonstration went well. That is where I think the hard, unglamorous, useful work is.

I have also managed a personal investment portfolio for 18 months — 22 holdings, reviewed monthly, with 20+ written reviews of my own decisions and roughly 27% in one annual period. The discipline carries into the engineering: establish what is known, record what is not, and size the risk accordingly.

Previously a competitive footballer: Player of the Month at the LaLiga Football School in Delhi, and invited to train with Real Madrid's youth programme in Spain.

  • EducationBTech, West Herts College, Watford — 2025 to present
  • BasedLondon, UK
  • Investing18 months · 22 holdings · monthly reviews · ~27% in one period
  • HonoursPlayer of the Month, LaLiga Football School Delhi · Real Madrid youth invitation
11

Some of my other work

Six more sites, and the SYRIX site again. Five are hosted here and the client one opens their own deploy — each in a new tab.

The SYRIX website: the word SYRIX in gold over a dark sphere, with a nav bar and four figures beneath it
01 SYRIX

The site for the runtime above — the core, the pipeline and the capabilities, walked through on scroll.

Two pages · three.js · WebGL Open site
The Avananka site: the SYRIX wordmark in blue over a dark hexagonal grid, with a full navigation bar
02 Avananka

The product site for SYRIX: features, a HUD preview, the trading desk, pricing, access and the build journey.

13 pages · no framework Open site
The Literal Humans site: heavy black type reading Resilient growth for mission-driven brands on cream, above a line of drawn animals
03 Literal Humans

A London agency site rebuilt as one continuous scroll, with a raymarched centrepiece running the length of it.

One page · raw WebGL · zero dependencies

An independent redesign concept, built unprompted. Not affiliated with Literal Humans — the brand, the copy and the illustration are theirs. literalhumans.com ↗

Open site
The Rangkar site: cream and gold serif type reading Where tradition meets tailored elegance on a dark brown ground
04 Rangkar

A bespoke Indian tailoring house in Britain. Markup, styles and script in one document.

One file · 22 KB Open site
The Wife’s Tongue site: the bar’s name in a heavy orange-to-pink gradient over a dark neon-lit background
05 Wife’s Tongue

A cocktail bar in London — menu, events and the room. Same shape: one self-contained file.

One file · 26 KB Open site
The GMC Immersion site: the company name in gold over an alpine lake at dusk, with the line Sport, Culture, Growth and two buttons
06 GMC Immersion

My first paid client work: a thirteen-slide PDF rebuilt as an interactive site, with the deck's own dark-to-light turn kept as the page's structure.

Client work · deck copy verbatim · palette sampled from the PDF

A full redesign of a client deck for Golden Management Consultancy. The copy, the photography and the brand are theirs; the layout, the motion and the code are mine. goldenmc.co.uk ↗

Open site
The GMC catalogue: the firm's name, twelve services for businesses, nine for individuals and four destinations, beside a dotted globe
07 GMC Catalogue

Every service deck as one catalogue, with an eligibility questionnaire per service. The website gives the result at once and cannot read a single answer: each is sealed with a public key on arrival and opened only on one computer.

Client work · 13 questionnaires · RSA-OAEP sealed inbox · zero packages

Built for Golden Management Consultancy. The decks' wording, the photography and the brand are theirs; the site, the questionnaires and the sealing are mine. goldenmc.co.uk ↗

Open site

Kabirr Kohli · AI Systems Engineer · London

Made by Kabirr Kohli

Open to
university, degree apprenticeships
and engineering roles.

Based in London, open to relocation. Happy to walk through any of the above in detail, including the parts that did not work first time.