Under the hood

From spoken life to useful action, in one system.

This page shows how recordings, communications, files, and live requests become memory, decisions, actions, approvals, and new software. The tables, models, budgets, and limits are left in.

The loop, in plain words

Living Code OS is a living operating system for one person. It listens through the sources you choose to connect, holds them as one living context, and reacts in near real time. Everything further down this page is this same loop, with the numbers left in.

Fieldyspoken life Omispoken life Other wearablesspoken life live or near-live memory
Capture→ Connect→ Reason→ Apply policy→ Act or ask→ Record the outcome
  • Capture. A recording, a wearable conversation, a mail thread, a calendar event, a file, a text, a call or a question you just asked lands on whichever surface you used. It goes in through one ingestion path and the original is kept whole.
  • Connect. The same pass that files it also notices who and what it mentions, and links those people, companies, projects and topics to everything else that mentions them. There is one context, not one per app.
  • Reason. A question fires four searches at once over that context, fuses them by rank, and hands the result to a model with a tool belt. Background workers do the same on their own schedules.
  • Apply policy. Every proposed effect is checked against the authority you configured: observe, recommend, ask first, or act automatically within defined rules.
  • Act or ask. Inside your boundary it runs. Outside it, the exact action that would run is filed to a human approval area, with the payload in editable fields.
  • Record the outcome. Every packet, request and resolution is journaled with a sequence number. Nothing is claimed that was not seen to succeed.

You reach it through the visual interface, a connected wearable, a messaging bot, SMS, and two phone lines. Every channel reaches the same resident agent and the same memory, so the visual interface is one surface, not the whole system. One person, one instance, with no user accounts inside it.

Self-extension: asking it for a capability it does not have

You can ask it for software that does not exist yet. It plans, builds, tests and stages the work, then hands it back to you. Production merges still require human approval.

Ask for a feature→ Plan→ Build→ Sandbox test→ Staging→ Human merge→ New capability

Ask in words. A planner picks the files to study and writes a plan. A builder writes them. A sandbox on a real machine runs a syntax check and the source-contract assertions before anything opens a pull request to staging. A headless browser then drives staging against the planner's own test script, judged by a model from a different lineage than the one that wrote the code. The run ends by filing a merge request and waiting for you. The Shipwright below carries the budgets, the token ceilings and the protected paths.

One memory

Everything you work on, and everything about any company you point it at, goes in through one ingestion path and is kept whole. The original record is the truth; everything else is derived from it and can be rebuilt.

Search returns 700-character snippets for ranking. Once an agent knows what it wants, read_original reconstructs the whole record from the raw segments in 14,000-character pages.

Ingest, chunk, embed

Every source item, from a recording to a mail thread to a day of conversation, is normalised into one item shape and passed to a single ingestion function, whatever the source.

The pipeline, with the numbers

The raw object is written to R2 at raw/<source>/<external_id>.json and registered as a row in the D1 meetings table, keyed for idempotency by a source-prefixed external id. The row moves through four states: stored, embedded, extracted or failed. Items that carry a newer edit timestamp are rebuilt in place, so an edited page or a growing thread stays one document.

From the original come chunks of 1,100 characters with 150 of overlap, held in a chunks table and a mirrored FTS5 table with porter unicode61 tokenisation. Each chunk is embedded with bge-m3 on Workers AI (1,024 dimensions, twenty texts per call) and upserted into Vectorize under the chunk's own id, so a vector hit joins back to its row without reading metadata.

Four-leg retrieval

A question is asked four ways at once, and the answers are merged by rank. Whoever asks, and from whichever surface, every answer comes back through one system.

Retrieval is one function that every surface calls. Four legs run in parallel: the title leg matches query terms against titles and pulls chunks from the top four; the vector leg embeds the query and queries Vectorize with room and source filters pushed down; the FTS leg runs BM25 over the full-text table; the anchor leg resolves up to three entities from the query and reads the chunks that literally contain their names.

Fusion

Each list is filtered against the caller's room access, then fused into one ranked list. An ask function runs the search beside a dossier lookup and has Llama 3.3 70B on Workers AI write a cited answer, with graph summaries labelled background only.

The weights and the caps

Reciprocal rank fusion, constant 60, weighted 1.0 vector, 0.9 FTS, 0.85 title, 0.8 anchor. A small recency boost favours items under 7, 30 and 120 days old. Results are capped at three excerpts per document and k plus four in total.

query k, rooms titleterms · top 4 · w 0.85 vectorbge-m3 → Vectorize · w 1.0 FTSBM25 · porter unicode61 · w 0.9 anchor≤3 entities · literal names · w 0.8 room filter RRF constant 60 recency 7 / 30 / 120 d ≤3 per document k + 4
FOUR LEGS IN PARALLEL · ROOM-FILTERED · FUSED BY RECIPROCAL RANK · ONE LIST

The graph, built at ingest

As each document arrives, the system notes who and what it mentions and links them. Your map of people, companies, projects and topics grows as a side effect of reading, and it is the same map every surface reasons over.

Extraction, anchors, dossiers, communities

An extraction pass asks Llama for people, companies, projects and topics, creates entities with deterministic ids and a uniqueness constraint on kind plus canonical name, and writes co-mention edges from each entity to the next six in list order, with the source id as evidence. Merges never delete: losers go to a merge table and an alias map and are skipped on read. Anchors record which chunks contain an entity's name and feed both the anchor leg and the entity timeline.

An event pass asks a Haiku-class model for up to eight dated, past-tense one-liners per document. Dossiers are Haiku prose of at most 130 words per entity and room, rebuilt from the latest twenty events and ten anchored excerpts, never rolled forward, skipped when evidence is under three items. A five-minute cron drains the dirty queue eight at a time. Communities come from label propagation over normalised co-mention weights among entities with three or more mentions, with Haiku briefs of up to 120 words.

The 3D display

The constellation behind every window reads exactly two tables, entities and edges: unmerged entities ordered by mention count (default 800, maximum 25,000), then the edges among exactly those nodes.

How a node is tinted

Hue is kind: person blue, company green, project amber, topic violet, contact pink, skill teal, document gold. Brightness is recency, a glow from 0.55 to 1.0 over 180 days. Size is mention count, clamped to 24. Dossiers, events, communities and anchors serve answer time only and never reach the display.

The resident agent

The agent you talk to is a loop: read the question, pick tools, run them, read the results, answer in a sentence or three. Jeffery is this deployment's resident agent; every deployment names its own.

It carries a belt of about eighty tools, a curated slice of the board's single MCP vocabulary, and it remembers the conversation across every surface you use.

The loop, the ceilings and the three memory tiers

The loop runs on a Sonnet-class model with up to eight rounds, a 60-second ceiling per model call and 25 seconds per tool. Tool calls in a round run in parallel; results are capped at 6,000 characters. The tool list is built once, frozen, and tagged for prompt caching; the system prompt is a cached block placed before the uncached summary and facts. A quick-acknowledge pass on a Haiku-class model answers in under fourteen words before the real loop runs.

Conversation memory has three tiers. Tier one is the working window: raw turns, trimmed to the last twenty. Tier two is one rolling summary row, refreshed by Haiku every thirty turns to at most 150 words. Tier three is a daily rollup that ingests each day's turns back into the brain as a document. Standing facts are different: short sentences you asked the agent to keep, never chunked or embedded, pushed into every prompt as a block of the newest forty.

Model routing

Three tiers: Haiku-class, Sonnet-class, Opus-class. You can pick; otherwise an auto-router decides with three regex families: tool-ish, deep-ish, code-ish.

The routing rules

Under 90 characters and no match: Haiku. Tool-ish: Sonnet. Deep-ish, or over 900 characters: Opus. When torn, a Haiku classification with eight output tokens and a 3.5-second timeout answers fast, tools or deep.

By default the resident agent sits at the ask-first end of the authority ladder, so its own real-world actions are filed rather than fired. Its system prompt is a block of labelled rules. Condensed:

Replies are one to three spoken sentences, no markdown. The graph is yours to steer. Real-world actions are requested, never executed; say they are filed for approval. Everything you make is ledgered, one artifact per request. Model choice belongs to the operator. Never report an action you did not see succeed.

The fleet and the gateway

Background workers are separate programs with their own schedule, their own memory and a daily spending ceiling. They never hold the model key; a gateway holds it for them.

Each worker runs the same five-step loop, recall → sense → reason → act → record, and wakes on its own cron, hourly by default. Its directive is edited live, so behaviour changes without a redeploy.

The registry row, the loop template and the gateway

Each worker is its own Cloudflare Worker. Its registry row holds name, purpose, spec, directive, brain (default Haiku-class), daily budget (default fifty cents) and state (live, paused or killed). Every run writes a flight-recorder row (trigger, sense results, recall, reasoning summary, packet statuses, error, cost), always, in a finally block.

The loop template in full: fetch its own row and stand down if not live, read the live directive; run up to four declared board tools; make one gateway call for a JSON summary at 500 tokens; file one signal packet; record.

The gateway refuses unregistered or non-live agents, prepends the worker's working memory (capped at 4,000 characters), returns 429 when today's spend meets the budget, caps output at 4,096 tokens, meters the call under the agent's surface and updates spend. Budget means a cost ceiling and nothing else.

A Foundry birth

A hire begins with a spec. The Foundry is a separate Worker with a sandbox container: a Cloudflare sandbox image with wrangler preinstalled, lite instance type, at most three instances. Every step streams to the Hive as build packets, so the birth is watchable.

The sequence

Normalise the name. Open a sandbox. Design a directive on a Sonnet-class model if none was supplied. Materialise the loop template. Syntax check. Deploy from inside the container. Health-poll the new Worker up to ten times, five seconds apart. Commit the source. Register the child with brain and budget from the spec.

Authority, and the approval rail outside it

You set the authority boundary. For each kind of effect you choose one of four postures: observe, recommend, ask first, or act automatically within defined rules. Inside that boundary the system acts. Outside it, the exact action that would run waits for a human, in editable fields. What is in the fields at approval is what runs.

The boundary is not a preference toggle. It is carried by the key an actor holds: every Control deck key matches one of four profiles, and the profile decides which scopes that actor can reach at all.

ProfileScopesThe posture it encodes
observerreadObserve. Sees tasks, agenda, costs, board status, the fleet, memory health. Changes nothing.
operatorread, approve, memory, writeRecommend and ask first. Requests, approves and denies actions, searches and writes memory, files proposals. Cannot fire commands itself.
executoradds execute, connectorAct within defined rules. Runs machine commands, starts builds, sends texts, schedules, generates media, calls connectors, under the money line.
ownereverything, plus adminFull authority. Mints and revokes keys. Minting an owner or executor key itself demands the PIN.

whoami returns a key's name, matched profile, scopes, rooms, callable tool names and the money line, so any actor can be asked what it is allowed to do before it does it. Your own site login is promoted to owner on the deck.

The short version: routine work runs, consequential work waits for you, and you decide which is which.

(a) Authorised to run automatically

Inside the boundary, no card and no wait:

  • Opening a pull request.
  • Sending a text.
  • Blocking time on your own calendar, with no one else on it.
  • Adding a task.
  • Generating an image.
  • A direct Shipwright build under twenty dollars, started by an execute-scope key.
  • A machine command you issued yourself.
  • Anything an execute-scope deck key does under the money line.

(b) Must wait for human approval

Outside the boundary, filed as a card with the payload that would fire:

  • Email sends.
  • Calendar moves.
  • Shipwright builds.
  • Agent births.
  • Production merges.
  • Machine shell, input and app commands requested by an agent rather than by you.
  • Destructive-looking shell commands (recursive removes, filesystem formats, raw disk writes, shutdown, reboot, fork bombs, mode 777, killall) regardless of who asked.

Every surface that wants an effect in category (b) calls one function that inserts a pending row with the exact JSON payload, journals a request packet through the board Durable Object, and pushes an Approve and Deny card to the paired messaging chat. Nothing executes until a resolve function runs with approval, reachable only from the board tray, the Control deck approve tool (approve scope, plus PIN over the money line), or a button press in the paired chat.

Pending · forge agent · filed by quartermaster
name tray-watcher purpose report failed actions and anything pending over thirty minutes brain haiku · cron every fifteen minutes budget $0.50 / day · first brief the originating task cost under the $25 line · no PIN required
ApproveDenyEdit fields

Above the twenty-five-dollar PIN line the approve tool demands a PIN; starting a build over $20 and minting an owner or executor key demand it too. The PIN is a second factor layered on scope, so a leaked operator key can work the tray but cannot spend.

Resolution, failure and the journal

Resolve refuses a second resolution, applies tray edits into the stored parameters, and dispatches by type. A thrown executor becomes a failed status carrying the error text, and the board raises an interrupting card that reads "Approved, but it did NOT execute" with an open-approvals action.

Every packet is journaled with a sequence number, source, actor, kind, payload and status. Request and resolution are both packets. Card runs file decision packets that create receipts (evidence, confidence, expected upside, cost of delay, proposed action, reversibility, success test), and outcome packets grade them. A tray-watching fleet worker reports any failed action with its exact error and anything pending more than thirty minutes.

The Shipwright

The Shipwright plans, builds, tests and stages a new capability on request. It goes as far as a tested pull request from staging to main, and no further: production merges still require human approval, every time.

It is a Workflow with a budget on every run, a guard that stops on overrun, and a list of paths it is refused: workflow files, the schema, auth, the Worker config, anything matching secret.

The phases, with the budgets and the ceilings

Start inserts a run with a budget: default $20 from the API, $10 from the Quartermaster. Plan resets staging to main's head, then an Opus-class planner picks three to eight files to study and writes a JSON plan: approach, conventions, files, a test script, risks.

Build has the same lineage write each file whole at up to 30,000 tokens (streamed, so truncation is visible) or as exact find-and-replace patches when a file exceeds 20,000 characters, checkpointing each in R2 so a retry resumes past finished files. At most twelve files.

The machine gate clones staging into a sandbox, runs a syntax check and 377 source-contract assertions, and allows two repair rounds before a pull request to staging is opened and auto-merged. Await deploy polls CI up to ten minutes. Test drives a headless browser against staging with the planner's own script and has the newest balanced model from a different provider judge the evidence. Up to two fix cycles. Handoff opens or reuses a staging-to-main pull request and files a merge action; the run ends awaiting you.

Two MCP doors

Outside software connects through one of two doors: a small one that reads and remembers, and a large one that commands. They share one keyring and never share a key.

The Brain Stem is hard-capped to sixteen memory tools, twelve read and four write, narrowed further by the key's scopes and tool profile. It is the door for another person's instance: a collaborator's resident agent reading a room you chose to share, seeing that region of memory and nothing that acts on the world.

The Control deck is the command surface, 106 tools across seven scopes. Which of them a given key can call is decided by the four profiles in Authority above: observer, operator, executor, owner.

The sixteen memory tools and the seven deck scopes

Brain Stem: search, ask, remember, graph neighbours, entity brief, entity timeline, community brief, find people, read document, list documents, read original, note fact, recall facts, retract fact, inbox, notify. A call outside that set returns an error stating that command surfaces live on the Control deck.

Control deck: read 50, approve 8, execute 30, write 9, memory 5, connector 1, admin 3.

Keys, rooms, audit

There are no user accounts inside an instance. You are the only login, and keys are how other people reach in. People who work together each run their own instance, and instances share parts of the same brain through the rooms you choose to share. Nothing is shared until you mint a key for it.

Hashing, rooms and the wildcard

Tokens are shown once; only the SHA-256 hash is stored, and revocation takes effect on the next request. Every key carries a rooms list, an optional read filter and a tool profile; the wildcard widens reads to every room, present and future, never writes. Both doors accept the token in the URL path for clients that cannot set headers. Every deck side effect files an audit packet under the key's actor.

On Cloudflare

One Worker is the server. No build step, no bundler, no transpile step. Everything else is a Cloudflare primitive bound to it.

PrimitiveBindingWhat it holds
WorkermainThe whole server: board API, both MCP doors, webhooks, the auth gate, cron dispatch.
Static assetsASSETSThe board front end, the graph page, the face mesh, the 3D bundle.
D1DB60 tables plus the FTS5 virtual table: the board spine, memory (meetings, chunks, entities, edges, anchors, events, dossiers, facts), keyring, agents and runs, actions, turns, model calls, settings.
R2MEMORYOriginals under raw/, uploads, generated media, Shipwright checkpoints, the Hand binaries.
Vectorize ×2VECTORS, PEOPLEChunk embeddings (1,024 dimensions) and the people index.
Workers AIAIbge-m3 embeddings; Llama 3.3 70B extraction and cited answers; FLUX schnell images; Whisper large v3 turbo transcription; two text-to-speech models; document-to-markdown.
Browser RenderingBROWSERHeadless browsing for the Shipwright's test phase and the Porthole's page watch.
Durable Objects ×4BOARD, PHONE_BRIDGE, MACHINE_HUB, COMPANION_LIVEThe board's single writer and WebSocket fan-out; one per phone call; one per paired machine; one per live companion capture. SQLite-backed.
Workflows ×5BUILD_CYCLE, INGEST, PEOPLE_FLOW, FEATURE_FORGE, DEEP_THINKThe daily card build, paged ingestion, people import, the Shipwright, Deep Think. Each has a staging twin.
Cron triggerssevenThe heartbeat; schedule below.
Service bindingsemail tracker, FOUNDRYThe email tracking Worker and the agent Foundry.
ContainersSandboxA sandbox image with wrangler preinstalled; lite instance type; at most three instances.
Satellite Workersagents/*Foundry-born fleet workers, each with its own cron and a service binding back to the board.
AccessedgeIdentity login on the human host, verified in process with a six-hour key cache.

Seven cron triggers are the heartbeat: the five-minute tick keeps ingest, reschedules, media jobs and approval pushes moving, and the rest carry the daily passes.

The seven cron triggers, in full

Every five minutes: ingest poll, reschedule ticks, media jobs, approval pushes, the Quartermaster if dirty, eight dirty dossiers and four community briefs.

Every two hours: fast-cadence cards.

Daily 09:00 UTC: the build-cycle Workflow, catalogue refreshes, the Librarian sweep, the conversation rollup, the daily graph pass, the community build.

09:20 UTC: the daily creative. 10:20 and 16:20 UTC: Pulse and task re-derivation. 10:50 UTC: the Quartermaster's full pass.

The request path

health · login · logout→ MCP mounts→ pre-gate capability routes→ the gate→ board WebSocket→ API layer→ assets
What each hop checks

The MCP mounts and each pre-gate webhook carry a credential of their own shape. At the gate, anything denied stops, assets included. The API layer resolves the actor (master bearer, keyring token, Access JWT, site cookie), enforces scopes per route, and records a timing row.

Lines of code: the board front end about 5,900; the graph page about 1,900; the Worker entry file and source directory together about 18,500; the Hand daemon about 1,800 lines of Go; an 811-line schema, applied idempotently on every deploy.

What it costs to run

Every figure here describes what the system costs to run, not what anyone is charged. The platform is tens of dollars a month at a single operator's volume. Model spend is the larger line, metered call by call into one table with surface, provider, model, cache reads, cache writes and cost.

$0.50
per worker · per day
Default fleet budget. The gateway returns 429 when it is spent.
$10–20
per Shipwright build
$20 from the API, $10 from the Quartermaster. A guard stops the run on overrun.
$5
per Deep Think
Two to five research passes, a synthesis, and a contrarian pass.
$25
the PIN money line
Above it, approving asks for a PIN. Scope alone cannot spend.
Model tierIn · out, per million tokensWhere it runs
Haiku-class$1 · $5Fleet workers, summaries, events, dossiers.
Sonnet-class$3 · $15The resident agent's loop, the Quartermaster.
Opus-class$15 · $75The Shipwright's planner and builder.
Real-time speechest. $32 · $64 per million audio tokensThe voice rail over WebRTC.

Prompt caching on the resident agent's frozen tool block holds the per-turn cost down; cache reads are priced at a tenth of input. A typical day at solo scale is a few dollars of model spend. A day with several builds and briefs on the deep model is tens.

Early access

Want to see it on your own data?

It isn't available yet. The early-access list hears first. The home page has the live agent if you'd rather ask him.