Follow the build

What is running, what changed, and what it cost.

A dated record from a living operating system for one person. It opens with what already runs in the private deployment, then the site you are reading, then what gets demonstrated next. Entries are written the night they happen and stay as written. Dates are ISO. Costs are measured, not estimated.

The log

Three dated entries. The undated plan at the end is not one of them.

Living Code OS is not open for signup yet. What is open is this log, the agent on the home page, and a list that hears first when either changes.

2026-08-23entry 001 · running

What already runs in the private deployment.

The system this site describes was running for one person before the site existed. This entry is the inventory: what it listens to, what it notices, what it does inside the authority you set, and where you can reach it. One person, one instance. Every line below is drawn from the overview and verified against the code.

It listens, into one memory. Voice recordings from wearables, meeting recordings, mail threads, calendar entries, files, uploads, text messages, phone calls and the day's own conversation all pass through a single ingestion function. The original of each is kept whole in object storage; chunks, embeddings, entities and edges are derived from it and can be rebuilt at any time. One retrieval function serves every surface, so whoever asks and by whatever channel, they read the same memory.

How one question reads the whole memory

Every query fires four legs in parallel: a title match, a vector search, a BM25 full-text match, and an entity-anchor pass that reads the chunks literally containing a named person, company or project. Each list is filtered by the caller's room access, then fused with reciprocal rank fusion at constant 60, weighted 1.0 for vector, 0.9 for full text, 0.85 for title and 0.8 for anchor, with a small boost for items under 7, 30 and 120 days old. Search returns short snippets for ranking; once the agent knows what it wants, it reads the whole original in 14,000-character pages.

It notices, and keeps the relationships. A language pass at ingest extracts people, companies, projects and topics, writes co-mention edges with the source item as evidence, and marks what changed. Dossiers, dated events and community briefs are rebuilt from that graph on a five-minute cron, so the picture of who and what you are dealing with stays current without anyone maintaining it.

It reacts, through a resident agent with a tool belt. A tool loop of up to eight rounds carries about eighty tools: read the memory, steer the graph, read a repository, draft, file, propose. Calls inside a round run in parallel, the tool block is frozen and cached, and a fast acknowledgment lands before the real loop starts. The agent reads and proposes; it does not fire real-world actions on its own.

It acts in the background, on a budget. The fleet is a set of workers, each its own deployment with its own schedule, a directive it re-reads on every wake, working memory, and a daily budget of fifty cents by default. The budget is enforced by a gateway that holds the key and refuses the call when the day is spent. Every run writes a flight-recorder row with its trigger, what it read, what it filed, its error if any, and its cost.

It builds new software on request. The Shipwright plans a change, picks the files to study, writes them whole, checkpoints each one so a retry resumes where it stopped, verifies the result on a machine in a sandbox, opens a pull request to staging, waits for CI, drives a headless browser test against staging, and has a different model lineage judge the evidence. Then it files a merge request and stops. Production merges always wait for a person.

You set the authority boundary. Anything with a real-world effect, an email, a calendar move, a build, a new worker, a production merge, lands in the approval area as a card holding the exact payload that will fire. You can edit the fields, and what is in the fields at approval is what runs. Spend above a set money line asks for a PIN. Nothing executes on a promise; it executes on the payload you can read.

It is reachable where you already are. The visual interface with its chat and voice windows, a messaging bot that carries the approval cards with their buttons, two phone lines, SMS threads any worker can hold under its own name, an iMessage bridge running on your own machine, and an iOS companion that records audio and posts health samples. Same memory behind every one of them.

Every call is metered. One ledger row per model call with its surface, provider, model, cache reads and writes, and cost, alongside per-worker, per-build and per-brief spend. That is why this log can say what a thing cost instead of estimating it.

running in the private deployment · documented here next
2026-08-23entry 002 · shipped

Day zero: this site, and a resident agent you can talk to.

Tonight the site went live with a working agent on the front page. It knows one document and a short note, it cannot do anything in the world, and every answer it gives is metered and visible. Below is how each piece is put together.

The site. The whole thing is one Cloudflare Worker serving static assets, with no build step. That is the same shape as the product. It was built in a single overnight session by an AI coding agent working from the product overview and a walkthrough of the live system.

The agent. The window on the home page is Jeffery, this deployment's resident agent; every deployment names its own. He runs on Workers AI, on Llama 3.3 70B Instruct FP8. His memory is the product overview plus a short note on who the system is for, and nothing else. He cannot act: no tools, no approval area, no email, no calendar. He runs on a separate, read-only deployment, not the instance its builder runs on.

How a question is answered

The overview was chunked the way the product chunks: 1,100 characters with 150 characters of overlap, which came to 89 chunks, plus five more from the positioning note. Each chunk was embedded with bge-m3 on Workers AI and upserted into Vectorize, and mirrored into a D1 FTS5 table. Every question runs two legs in parallel, a vector search against Vectorize and a BM25 match against the full-text table, and the two lists are fused with reciprocal rank fusion before the model sees anything. The product runs four legs over the whole of one person's memory; two documents need these two.

Voice in, voice out

Voice in is the browser's own speech recognition, with Whisper large v3 turbo on Workers AI as the fallback when the browser has none. Voice out is the Deepgram Aura-1 text-to-speech model on Workers AI. Both are optional; typing works the same.

The graph behind every page. It is a hand-written WebGL point cloud, about 2,600 nodes on desktop, coloured with the same hue grammar the real system uses for entity kinds. Newcomers streak in as meteors. A focus channel lets the agent steer it after every answer, the way the product's show_on_graph tool steers the real constellation. It is decorative: the page reads fine without it.

The list. The early-access form stores every submission in a D1 table first, so nothing is lost even if mail delivery fails, and then sends through Cloudflare Email Sending. A first submission sends two emails: a notification to one Duval Software inbox, and a confirmation to the address that signed up. A repeat submission sends the notification only, and updates the row it already has. Turnstile guards the form.

The meter. Every model call the demo makes is recorded in the model_calls ledger with its surface, model and sizes, the same way the product meters its own calls. The demo carries daily caps of 800 chat turns and 300 voice replies; when a cap is reached the demo says so and points to the list.

What one turn costs. Measured tonight: about 1,700 prompt tokens and roughly 45 neurons to the first token. That is fractions of a cent per answer, which is the point of running it on Workers AI.

first measured turn · ~45 neurons · < $0.01
2026-08-23entry 003 · shipped

The overview, verified against the code.

The document Jeffery remembers was written and checked against the repository today. It covers what the system is, the memory and the graph behind it, the resident agent and the fleet, every channel it is reachable on and every source it listens to, the approval area and what passes through it, the two machine doors, the daemon that reaches your computers, the architecture underneath, and what it costs to run. It is the source for everything on How it works, and it is what the agent on the home page reads before it answers you.

upcomingno date yet

What we'll demonstrate next

These capabilities already run in the private deployment. Next, we will document them end to end so you can inspect the evidence, not just the claim. Nothing below is promised as a date; it is what goes on this page as each one is captured.

  • The Shipwright building a feature end to end: plan, build, a machine gate in a sandbox, a pull request to staging, CI, a browser test judged by a different model lineage, and the merge request left for a person.
  • A background worker being born: its specification, its budget, its first scheduled run, and every step as it happens.
  • A machine paired through the Hand, from the single pasted line to the first command it runs.
  • The face first speaking: the constellation rearranged into a head, the jaw tracking live audio, and the detonation back to the constellation.
  • An action moving through the approval area: the exact payload, an edit to the fields, and the receipt after it fires.
  • A weekly cadence: an entry here, a demonstration you can watch, and a ledger of what that week's model calls cost.

Each of those becomes a dated entry above when it is documented, with its cost chip.

Early access

Hear about each entry first.

Living Code OS is not open for signup yet. When a new entry lands here, the early-access list hears before anyone else. Your details stay with Duval Software. No drip, no resale.

Join the early-access list