Brad SebastianAgentic AI Operating Record
Built and operated since July 2026

A commercial operation run by agents,
governed by a person.

This is the record of a production multi-agent AI system I designed, built and run every day. It manages a live commercial workload end to end, reaches my email, files, browser and cloud storage through Model Context Protocol, and cannot send anything to another human being until I have read that item and said go. It is the same architecture a company would use to put AI inside its commercial engine, at one-person scale, with me as platform owner, product manager and governance function.

6
Layers, one record of truth
100%
Outbound actions gated by approval
4
Gates on every outbound document
~2,500
Lines in the presentation builder

Every claim on this page can be shown live from the files that produce it.

Architecture

Six layers, one record of truth.

The design principle is that files, not conversations, hold the state. A chat is disposable. The record is not.

Data

One record of truth

Every engagement is a structured record: organization, status, a step ladder with dates, contacts with tier and send state, an activity log, signals and flags. Exports are regenerated from the record, never kept as a second copy of the truth.

Generation

Documents from code

Outbound documents are produced by generator scripts built on a master, one per engagement, from a single verified source file. A fact corrected once propagates to every document on rebuild.

Gates

Nothing ships unchecked

Four gates run on every outbound document before I see it: a written quality gate, a formatting and parser gate, a mechanical provenance gate that fails a document on any unverified tool, figure or claim, and a cold read as the recipient. Details below.

Presentation

A self-rebuilding dashboard

A 2,500-line builder renders the record into a single self-contained application: stages, momentum-based hot detection from the last five days of activity, weighted scoring, next actions phrased as people and meetings, embedded documents, and one-click commands that hand a task back to an agent. Nobody edits it by hand.

Skills

Written procedures the agents load by name

Each skill is a procedure with hard rules: what to do, in what order, and what never to do. The agent loads one by name and follows it. The library is described below.

Integration and memory

MCP, and a rulebook written when the rule is learned

Agents reach Gmail, Google Drive, Chrome, Microsoft 365 and the local file system through Model Context Protocol. A standing rulebook records every ruling the day it is made, with the incident that caused it; a session-state file is what any new session reads first.

Governance

The part that separates running a platform from using a chatbot.

Approval on every outbound action

No message, document or email leaves without me reading that specific item and approving it in the moment. No batch approvals, no standing permission. The dashboard can queue a send; only my approval of that item executes it.

A written trust boundary

The agents do not act on third-party systems autonomously; I drive every submission myself. Full-agentic runs are tested only on work I designate as expendable, at my initiation. Trust is earned first; the record has to show the system deserves more room before it gets it.

Rules come from incidents

The rulebook was not written in advance. It is a log of real failures converted into standing rules, each with the incident attached.

A tool name inferred from context reached a dozen documents before I caught that I had never used it. Rule: every tool comes from a verified list that only I can edit.
An email built with embedded images rendered blank in Gmail. Rule: build for the delivery platform and verify in the destination before sending.
A phrase from an unrelated thread contaminated a strategy brief. Rule: high-stakes work gets a clean session with only its own material in it.

A challenged term is a removed term

If I question where a claim came from, it comes off immediately and returns only if I say keep. Leaving it in pending a ruling is how errors ship twice.

Optimize for the meeting

My own evidence says live conversations convert and cold submissions stall, so the system is tuned to earn meetings: every engagement carries the shortest path to a human, and unsent outreach outranks finished paperwork. That is a strategic decision made by the person, from evidence, and no prompt made it.

Accuracy over warmth

An unverified personal connection is never used to make a note warmer. A clean professional note beats a warm note with a false claim in it.

Skills library

Reusable agent procedures I wrote.

Described here in general terms; the procedures themselves are working files, not marketing.

Engagement kit

Captures a new engagement, runs the analysis, produces the tailored documents from the generators, logs the record, updates the dashboard, then runs the gates.

Targeted outreach

A tiered contact list for an organization, verified profile by profile, only real connection points, drafted notes in my voice. Hard rules: warm paths before cold notes, and never an unverified affinity claim.

Rehearsal and scoring

Compresses everything known about a high-stakes conversation into a roleplay brief for a separate session, then scores the transcript against a fixed rubric so repeat rehearsals are comparable.

monte-carlo-document-check

A review-and-revise loop for anything going to a third party: a cold-read agent scores the document as the recipient would, an editor agent applies the edit list under hard rules, the outputs rebuild and verify, and the next round starts from the new version until the scores plateau.

update-dashboard

Rebuilds every document whose generator changed, re-runs the gates, re-embeds the outputs and refreshes the dashboard.

Research and share

Researches a question and sends the findings as a formatted email, behind a send gate, with a disclosure footer. Plus skill-creator and the standard document skills for Word, PDF, PowerPoint and Excel output.

Gates

What runs before I see a document.

Quality gate

A written truth check: every tool, company, number and date must trace to me or to the verified source file. Tense and date check, American spelling, voice rules. Each rule cites the incident that caused it.

Formatting and parser gate

An in-house parser check: single column, standard headings, no tables or text boxes, consistent dates, text-layer PDF, metadata. House rule: fix the detector, never the truth.

Provenance gate

A mechanical fabrication check that fails a document on banned or false terms, tools not on the verified list, paraphrase drift in descriptors, and stale figures.

Recipient pass

The rendered document read cold as the person who will receive it, confirming that everything it needs to answer is visibly answered on the page.

Findings

What I have learned about where these systems fail.

Fabrication happens by paraphrase, not invention

The model rewords a verified fact into nearby language that echoes what the reader wants. The defense is a verified vocabulary and a mechanical check, not a reminder to be careful.

Confidence is not evidence

Word counts stated without counting; figures that trace to nothing. Anything that matters gets measured.

Design from description does not work

Taste cannot be specified in prose. Give the tool the real assets and a written brief, then let a purpose-built tool do the visual work.

The platform is part of the deliverable

An artifact that renders where it was made and fails where it lands is a failure. Test in the destination.

Long contexts contaminate

A session carrying weeks of unrelated work leaks vocabulary across threads. High-stakes work gets a clean room and a handoff document.

The human owns the facts, the strategy and the go decision

Every substantive correction in this system has come from me knowing my own story better than the model does, and the system is built so that my correction wins immediately.