GoodLeads ships product through an organization of AI agents — general managers, an engineering manager, specialists, rituals, memory, and budgets. One rule holds all of it together.
For the operator running outbound: records delivered before Apollo has them indexed, contacts who can actually authorize a purchase, and a quality floor that rises without a corresponding rise in cost. For anyone evaluating the architecture: what follows is how.
For the operator: fewer dead dials, more first conversations, leads nobody else has called yet. For the business: a cost structure that doesn't multiply with markets.
Every agent must name the variable it moves before work begins. If it can't, we don't take the work. That's how companies align thousands of employees they can't individually supervise — a P&L, not a rulebook. We applied it to agents from day one.
The equation isn't a poster on the wall. It's the admission gate. A proposal that can't name its variable gets sent back — by other agents. The moves that pass, by variable:
"Which variable does this move?" is the first question every agent asks — and the last gate before work ships.
Business units across the top — one General Manager per market, owning how often its records sell. Departments down the side — one craft per row, serving every unit. A specialist sits at each intersection, answering to both. Every seat is an identity: a system prompt, a skill library, tool permissions, a memory scope — not whichever model happens to back it.
The entire matrix reaches him through a single authored briefing line — and that line doesn't grow when the matrix does.
Filled dot = the department serving that market · open dot = the posted seat.
Departments carry the craft across units; units carry the market. Specialists answer to both — the classic matrix trade, made cheap.
Add a market: a column. Add a craft: a row. Neither adds a line to the founder's day.
A new market — or a new craft — starts as a hiring spec, written against shared context infrastructure: the memory stores, the wiki, the toolkit ladder, the harness itself. Because all of that already exists, the spec is short and the start date is immediate.
We didn't invent new ceremonies for agents. We gave them the ones that already work — every ritual a software company would recognize, run by agents, on a cadence.
Every ritual is a write to memory. Standups write what happened. Retros write how we work. Reviews write what we know. Cadence is how experience becomes an institution.
It's also why the founder can be out of the loop without being out of touch — synthesis reaches him on a schedule, instead of him going looking for it.
Because that's how organizations — and people — actually retain knowledge. Each layer has its own persistence, its own write cadence, and its own ritual that writes it.
What we know is a live table, not a build artifact. Edit the wiki and the agent's next session works from the new knowledge — versioned, auditable, no redeploy. Changing an agent's mind takes seconds, not a release cycle.
Every agent session runs in a container that vanishes when the work ends. The agent's memory doesn't. Each seat carries its own persistent store — market archetype, baselines, patterns, every finding it ever filed — and it compounds. Fire the session, keep the employee.
Every week it works, an agent gets cheaper to run — and harder to replace.
Any agent can propose an experiment against its variable — and spend real vendor dollars running it, inside a monthly envelope enforced by the server, not by trust.
The agent proposes, spends, ships, verifies — and retreats when production disagrees.
The envelope is real money against real vendor credits — skip-trace, gap-fill, validation — not simulated cost. An agent that wants to test a hypothesis pays for it like anyone else.
Variants forward-replay through the production context-builder — test and production are one code path, so experiments predict production. There is no separate test harness to drift.
A win writes an append-only config row. A daily verifier watches production metrics and auto-reverts regressions, filing an auditable finding for each retreat. No one has to notice.
When we use a model as a pipeline worker, the model call is a cost. What we sweep out of it is an asset.
Deterministic rules handle what's known, free. The model gets only what they can't resolve.
Every resolution is captured — recurring patterns become owned rules that handle those cases forever.
Name intelligence and industry classification route only their uncertain tail to a model. Resolutions are captured; recurring patterns are swept into rules, crosswalks, and per-state mappings.
Each pattern should cost roughly one model call, ever. The call rate on known patterns falls; per-record cost declines while quality holds. That's a curve we watch, not a slogan.
Every brain gets the pattern — and so does the harness itself. Every hardcoded rule in the org is a marker for an answer an agent will eventually learn to discover on its own.
The model is rented. The library is owned.
The founder never opens an engineering tool. His entire interface is email and Google Docs — the tools he'd be using anyway.
The EM authors a daily narrative in a VP's voice — today, what moved, the calls it needs, what it's watching. Decisions arrive as one-click buttons. A click records a durable, auditable decision; the agent executes the instruction it authored for that option. One click. The agent handles the rest — and auto-reverts if production disagrees.
He comments on docs. That's it. The org mines those comments, clusters them, and promotes the recurring ones into formal judge scorers that grade future agent work. Management judgment, compiled.
The org meets the founder in his tools — and learns from what he says there.
Overnight standups clean across markets. The enrichment experiment cleared simulation — projected fill-rate lift holds within the envelope.
Ship the winning vendor variant to production? Simulation says yes; spend fits the envelope; the daily verifier will watch it and retreat on regression.
One market's fill rate drifting near its floor — an experiment proposal is coming if it holds through Friday.
Who decides what the agents work on? Strategy lives in a stack of toolkits, laddering from durable commitments down to this week's evidence. An idea descends only by earning it — and every level answers to the equation above it.
Thirteen durable commitments — speed from public availability, confidence not certainty, unit economics. The constitution's articles.
The strategy: what we're building, why it wins, what we're explicitly not doing.
One per bet — the learning loop, commerce, the intelligence surfaces. Each carries its own outcomes and kill criteria.
Scoped, reviewed, retroed. The rung where ideas become shipped work.
The evidence layer. What the org learned today — feeding everything above it.
Every layer carries the same anatomy: the job to be done · the outcomes it targets · the KPIs that say it's working
Agents are first-class readers of the ladder. A proposal must cite the tenet or toolkit it serves — the equation gate, one level more specific. And new seats inherit the ladder on day one, which is how the org absorbs more agents without absorbing more chaos.
Every rung of autonomy the org climbs safely deletes a category of human cost from the equation. Here's how far it goes — and where everyone else stops.
Agents write the code.
Agents review each other — two model families, both green to merge. No same-family blind spots.
Agents merge without a human — tier-gated by what the change touches.
Merged code deploys itself to production. The merge is the only gate — keyless, immutable releases, self-healing rollback.
Agents verify their own shipments against production metrics, daily.
Agents revert themselves on regression — and file the finding that says why.
The org compiles the founder's judgment into formal evaluation criteria.
Agents propose changes to their own instructions — a structured case (why · evidence · target outcome), approved on the trust ladder.
Tiers are computed mechanically from what a change touches. Fixable mechanical gaps don't block — the EM repairs them itself and re-reviews. The founder sees only what's worth his judgment — in plain English, in a doc.
Auto-merge when both reviewers agree. The founder never sees these unless he asks.
Auto-merge, logged in the daily sweep.
Auto-merge with ownership gates satisfied.
Routes to the founder as a plain-English risk summary. A comment approves it. Rare by design.
Everything a licensee's contract covers — pipeline, brains, API. The agent layer mounts behind a single feature flag: switched off, the org disappears and the data plane runs as if it never existed.
backend/pipeline/classification-service/The GMs, the EM, the rituals, the memory, the learning loop. A licensee build literally subtracts this directory. A CI test fails the build on any import that crosses the line.
agents/The commercial architecture: the deliverable is licensable because the org was never fused into it.
Expansion means posting a seat req and filling it the same morning: a seeded memory store, a state configuration, a budget envelope. The horizontal layer — review, deploy, verification, the founder's briefing — is already amortized. It doesn't grow with markets.
Not aspirations. The live-market count below is fetched from the production API on every load; the rest is stamped from the repository at publish time — never hand-typed history.
Repository figures as of 2026-07-10 · market count live from /api/v1/leads/states
A competitor copying the surface ships in a sprint. The rows below each presuppose a commitment made months earlier.