By clicking “Accept All Cookies”, you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. View our Privacy Policy for more information.
Marketplace trends
/
October 9, 2026

Traide: how we build with AI

James Throsby
James Throsby
Founder & CEO
Share
Traide: how we build with AI. Title card with the Traide logo beside the four AI model tiers, Fable 5, Opus, Sonnet and Haiku

A development crew of one engineer, multiplied: a fleet of AI models with a gated workflow, a cost-shaped chain of command, and a memory that compounds. This post comes in two halves — the first on how work gets done, the second on how the system learns — and ends with what ten months of data show.

The system went live on 9 December 2025. Every mechanism and number below was re-verified against our live configuration, memory database and git history on 9 October 2026.

The development loop

Work moves through six stages — plan, build, verify, review, ship, learn. The stages are ordinary; what is unusual is that the transitions between them are enforced by machinery, not by discipline. Twenty-seven automated checkpoints (“hooks” — small programs the development environment runs at fixed moments, which the AI cannot skip) wrap every session. Two of them are AI reviewers that read the work before letting it pass, and a third distils every correction into a lesson.

Six-stage development loop — plan, build, verify, review, ship, learn — with machine-enforced gates between each stage and a return arrow showing that lessons reload at the next session start.
Fig. 1 — The loop and its gates. Accented marks are machine-enforced: they run automatically and cannot be skipped by the AI, or forgotten under deadline pressure. The return arrow is where the second half of this post picks up.
  • On every commit: A deterministic check holds the commit unless a passing test run newer than the last edit is on record, and an AI reviewer refuses the commit if those tests don’t actually exercise the files that changed. Every commit must also cite the GitHub issue it serves.
  • On every edit: Changed files are lint-checked the moment they are written — not later in CI. Twelve pattern guards flag known hazards as the code is written: floating-point money math, queries that cross tenant boundaries, lint-suppression tricks. Three of them block outright.
  • Before merge: A panel of specialist AI reviewers sweeps every change — general code review, silent-failure hunting, test-coverage analysis, security, and Traide’s domain rules — before a human sees the pull request. Changes ship small, in reviewable stacks.
  • At session end: Everything touched is re-linted and type-checked; failures block the close. A reflection gate then refuses to end any substantive session (30+ tool operations) until its lessons are recorded to memory.

Standing rules with teeth

A few rules are treated as sacred and are loaded at the start of every session. Never guess: any identifier — a database column, a metric name, a configuration key — must be discovered in real output before it is used, because a guessed query that returns nothing is indistinguishable from a real finding of nothing. Hypothesis before code: state what you believe is broken and why, check the evidence, and only then edit. Every bug fix ships with a regression test that reproduces the exact failure. And production data is read-only and asked-first — the AI plans infrastructure changes; it does not apply them.

In plain terms: The system does not rely on the AI promising that it tested the code. Independent checks read the test record and stop the commit if the tests weren’t run, or don’t cover the change. Quality here is a property of the machinery, not of anyone’s memory or good intentions on a given day.

A crew of four models

Claude models come in tiers that trade depth of reasoning for speed and cost. Traide’s routing principle: buy deep reasoning only where it changes the outcome. The most capable model plans and decomposes; each tier below receives a scoped, written brief; results are verified on the way back up. When the right tier isn’t obvious, an automatic classifier (route_task) decides.

Delegation chain of four model tiers ordered by reasoning depth: Fable plans, Opus orchestrates and validates, Sonnet implements multi-file chains, Haiku handles fast narrow tasks. Arrows show delegation downward and escalation upward.
Fig. 2 — The chain of command. Work is decomposed downward to the cheapest tier that can carry it and verified on the way back up. Every hand-off is a written brief with a narrow scope — tiers do not share memory.
  • Fable 5 — deep planning and judgment: architecture specifications, the hardest debugging, system audits, and decisions that are costly to reverse. It runs the main session for the most substantive work, at maximum reasoning effort.
  • Opus — orchestration and validation: multi-context overview, initial implementation steps, and verifying the work that comes back from below. It is the default main-thread model for day-to-day sessions.
  • Sonnet — specialist implementation: multi-file logical chains, framework migration, and domain-aware code review. It powers the strawberry-expert, security-reviewer, traide-domain-reviewer and api-contract-validator agents.
  • Haiku — fast, narrow execution: individual file updates, pattern search, running tests, and syncing memory — frequently fanned out in parallel. It powers the test-runner, performance-profiler and memory-sync agents.

Two standing rules shape the routing: anything touching money or payments never drops below Sonnet, and exploration or search work prefers Haiku — which, at today’s list prices, costs about 97% less per token than Opus, our default model (for prompts under 100,000 tokens).

In plain terms: Think of it as staffing a firm. The principal architect sets the plan; a senior engineer runs the project day to day; specialists own their domains; fast assistants do the routine legwork. Nobody bills principal rates for formatting a document — and nothing involving money is ever left to the junior tier.

The learning system

Most teams lose what they learn: lessons live in people’s heads and leave when attention moves on. Here, capture is automatic — a hook records what happened after every significant operation (a file edit, a commit, an agent dispatch, an external lookup), and every tool failure is captured too. Nothing depends on anyone remembering to write things down.

By the numbers, as of 9 October 2026:

  • 47,200 observations in memory
  • 1,082 recorded sessions
  • 27 automated checkpoints
  • 12 self-scored dimensions

Behind them: an 87-tool memory and orchestration service, 114 contract self-tests, 460 tests covering the hooks themselves, and 12 pattern guard rules.

From event to lesson

Raw capture is only the start. When the engineer corrects the AI, the correction itself triggers a lesson-extraction routine (/learn) that records the root cause, not the surface symptom. At the close of any substantive session, a reflection pass (/reflect) synthesizes the day’s events — and the session-end gate described above makes that pass unskippable. The store is then curated like a database, not a diary: near-duplicate observations are clustered (at ≥ 0.85 similarity) and consolidated, older items decay unless they keep proving useful, and integrity checks run at every session start.

Measured, not assumed

The system scores itself on twelve improvement dimensions — context quality, code quality, intent accuracy, confusion removal, and so on — and the scores are honest. Confusion removal has climbed from 4.0 when the system started to 8.5, still short of its 9.0 target; intent accuracy sits at 8.0 against a target of 8.5. Eight recurring correction patterns have been distilled into decision trees the AI must walk before acting. Recurrence analytics ask the hard question: after a lesson is recorded and later retrieved, does the same mistake still happen?

Case study: how a mistake becomes a rule

August 2026: during an infrastructure audit, the AI guessed the naming prefix for a set of metrics. The guessed query returned empty — and “no data” nearly shipped as a finding, when the truth was simply that no metric had that name. The correction became a sacred rule — “Never guess; discover first” — which now loads at every session start, is enforced through review checklists, and is tracked for recurrence. One mistake, permanently converted into machinery.

In plain terms: Every correction is written down by machinery at the moment it happens, distilled into a rule, reloaded every morning, and measured for whether it actually stuck. The company keeps what it learns — in a database it owns, on its own hardware.

What the AI knows, and when

Context arrives in three rhythms: an automatic briefing when a session opens, targeted retrieval while it runs, and a durable thread that connects one session to the next.

At session start: the morning briefing

Layered and automatic, in seconds: the handbook layer (company-wide rules, the project handbook plus ten per-module handbooks, and five deep rule files on verification, testing, and workflow); the memory index — over a hundred one-line pointers to durable lessons, always loaded; and a live briefing pulled from the memory database: recent work, recent corrections, the currently weakest dimensions, and the maintenance backlog. Environment checks run alongside — credentials, configuration audit, stray processes, 114 self-tests.

During the session: retrieval on demand

Each new request is screened before work starts: an ambiguity check, a match against past corrections relevant to this request, and a router that checks whether a purpose-built workflow applies (debugging, review, migration, PR creation…). Deeper context is fetched when needed by semantic search — meaning-based lookup over the whole store, not keyword match — plus on-demand domain packs for specific areas of the platform.

Across sessions: the durable thread

Everything captured lands in a local PostgreSQL store with vector search; embeddings are computed on-machine, so the memory store stays on Traide’s own hardware and costs nothing per query. When the AI’s working context must be condensed mid-task, a checkpoint first saves the live state — the task, progress, next steps — and is restored on resume. Multi-week efforts additionally get continuation documents; at session exit, time-decay re-weights the store so recent, useful knowledge surfaces first.

Memory flow between sessions: session N auto-captures work into the memory store; the store briefs session N+1 at start and serves semantic search on demand.
Fig. 3 — One store carries the thread. Each session writes as it works; the next starts already briefed. The store lives on Traide’s own hardware, and its embeddings are computed locally. Counts as of 9 October 2026.
In plain terms: Before touching code each day, the AI reads its own ship’s log: what happened recently, what it got wrong, and what it is currently weakest at. Institutional memory stops being a person and becomes an asset.

Ten months in: the results, measured honestly

Claims about AI efficiency are cheap. These numbers come from our own git history and memory database, and each comes with its caveat. The system went live on 9 December 2025; this section was re-measured on 9 October 2026.

  • 35% → 96% of bug fixes to our core API ship with a test (January–March vs June–October 2026)
  • 65 → ~300 product pull requests merged per month (December 2025 vs July–August 2026)
  • 1,922 pull requests merged across 13 repositories since December 2025
  • 0 repeat mistakes flagged since mid-May, across 34 recorded corrections

Bug fixes that prove themselves

At the end of March we made one rule non-negotiable: every bug fix ships with a regression test that reproduces the failure. In the three months before, 35% of bug-fix commits to our core API touched a test. Since June it has been 96% — 236 of 246 — and since a deterministic test gate landed at the end of August, all but one of 95 fixes have carried a test.

Output grew alongside it: from 65 merged product pull requests in December to around 300 a month in July and August. Our pull requests are deliberately small and stacked, so this counts reviewable changes, not effort — but volume rose while the quality bar went up, not down.

Repeat mistakes, recounted

This is where we had to correct ourselves. The first version of this post said our mistake-repeat rate fell from 82% to 0%. The 82% came from an early metric that counted corrections sharing a category label; within a day, in December 2025, it was replaced by a check that compares what each correction actually says. On that measure, about one in six corrections in December and January repeated an earlier one. Since mid-May, none of 34 has, and none describes a mistake carried over from an earlier session. When a past correction is shown to the AI before it starts work, it hasn’t recurred: 0 of 14 in the last 90 days.

One caveat we take seriously: corrections are recorded by the AI itself — a hook has prompted it to do so since late August — so these counts are floors, not a census, and February and March 2026 have a gap in the record.

What it costs

Priced, not yet measured: routing routine work to the fastest model tier costs about 97% less per token than Opus at today’s list prices — applied by enforced convention, never isolated on a bill. The standing instruction set is about 49,500 tokens per session; loading the full 47,200-observation memory naively would take about 6.7 million, which is why retrieval happens on demand. Our working estimate is that skipping a context check costs about 25 times what the check itself does (~50,000 tokens of redo against ~2,000) — an estimate, not yet a measurement.

The honest gap: no controlled comparison exists against a naive setup (no memory, no gates, one model). Closing it is scheduled work: per-session token telemetry will be joined to the session-audit database, after which this paragraph becomes a measurement.

In plain terms: Since the rule landed, almost every bug fix ships with a test that proves it; monthly output has grown more than four-fold; and no recorded mistake has come back since May. What we can’t yet show is the total bill against doing it naively — and rather than estimate it, we’ve scheduled the instrumentation that will measure it.

Merchant ambition is
our mission.

Niklas Halusa
Co-founder & CEO

Nautical Commerce enables anyone to build a marketplace—fast.

We've created an easy-to-use, powerful multivendor marketplace software platform so you don't have to build it yourself.
 

Discuss your project

Traide: how we build with AI

Contributor:
10
Min Read  |
October 9, 2026
Traide: how we build with AI. Title card with the Traide logo beside the four AI model tiers, Fable 5, Opus, Sonnet and Haiku

Key takeaways

  • Measured, not assumed: Since we made regression tests mandatory in March, 96% of bug fixes ship with one, up from 35%, while monthly output grew more than four-fold.
  • Memory that compounds: Every session makes the next one better-informed. What the team learns accrues to the company — in a database it owns — not to any individual, vendor, or chat transcript.
  • A floor that holds: Test, lint, review, and reflection gates are machinery. They hold at the same height on a calm Tuesday and the night before a launch.
  • Intelligence, priced to task: Deep reasoning is spent only where it changes the outcome; routine work runs on a tier that costs about 97% less per token than our default model. Capability where it matters, economy everywhere else.

A development crew of one engineer, multiplied: a fleet of AI models with a gated workflow, a cost-shaped chain of command, and a memory that compounds. This post comes in two halves — the first on how work gets done, the second on how the system learns — and ends with what ten months of data show.

The system went live on 9 December 2025. Every mechanism and number below was re-verified against our live configuration, memory database and git history on 9 October 2026.

The development loop

Work moves through six stages — plan, build, verify, review, ship, learn. The stages are ordinary; what is unusual is that the transitions between them are enforced by machinery, not by discipline. Twenty-seven automated checkpoints (“hooks” — small programs the development environment runs at fixed moments, which the AI cannot skip) wrap every session. Two of them are AI reviewers that read the work before letting it pass, and a third distils every correction into a lesson.

Six-stage development loop — plan, build, verify, review, ship, learn — with machine-enforced gates between each stage and a return arrow showing that lessons reload at the next session start.
Fig. 1 — The loop and its gates. Accented marks are machine-enforced: they run automatically and cannot be skipped by the AI, or forgotten under deadline pressure. The return arrow is where the second half of this post picks up.
  • On every commit: A deterministic check holds the commit unless a passing test run newer than the last edit is on record, and an AI reviewer refuses the commit if those tests don’t actually exercise the files that changed. Every commit must also cite the GitHub issue it serves.
  • On every edit: Changed files are lint-checked the moment they are written — not later in CI. Twelve pattern guards flag known hazards as the code is written: floating-point money math, queries that cross tenant boundaries, lint-suppression tricks. Three of them block outright.
  • Before merge: A panel of specialist AI reviewers sweeps every change — general code review, silent-failure hunting, test-coverage analysis, security, and Traide’s domain rules — before a human sees the pull request. Changes ship small, in reviewable stacks.
  • At session end: Everything touched is re-linted and type-checked; failures block the close. A reflection gate then refuses to end any substantive session (30+ tool operations) until its lessons are recorded to memory.

Standing rules with teeth

A few rules are treated as sacred and are loaded at the start of every session. Never guess: any identifier — a database column, a metric name, a configuration key — must be discovered in real output before it is used, because a guessed query that returns nothing is indistinguishable from a real finding of nothing. Hypothesis before code: state what you believe is broken and why, check the evidence, and only then edit. Every bug fix ships with a regression test that reproduces the exact failure. And production data is read-only and asked-first — the AI plans infrastructure changes; it does not apply them.

In plain terms: The system does not rely on the AI promising that it tested the code. Independent checks read the test record and stop the commit if the tests weren’t run, or don’t cover the change. Quality here is a property of the machinery, not of anyone’s memory or good intentions on a given day.

A crew of four models

Claude models come in tiers that trade depth of reasoning for speed and cost. Traide’s routing principle: buy deep reasoning only where it changes the outcome. The most capable model plans and decomposes; each tier below receives a scoped, written brief; results are verified on the way back up. When the right tier isn’t obvious, an automatic classifier (route_task) decides.

Delegation chain of four model tiers ordered by reasoning depth: Fable plans, Opus orchestrates and validates, Sonnet implements multi-file chains, Haiku handles fast narrow tasks. Arrows show delegation downward and escalation upward.
Fig. 2 — The chain of command. Work is decomposed downward to the cheapest tier that can carry it and verified on the way back up. Every hand-off is a written brief with a narrow scope — tiers do not share memory.
  • Fable 5 — deep planning and judgment: architecture specifications, the hardest debugging, system audits, and decisions that are costly to reverse. It runs the main session for the most substantive work, at maximum reasoning effort.
  • Opus — orchestration and validation: multi-context overview, initial implementation steps, and verifying the work that comes back from below. It is the default main-thread model for day-to-day sessions.
  • Sonnet — specialist implementation: multi-file logical chains, framework migration, and domain-aware code review. It powers the strawberry-expert, security-reviewer, traide-domain-reviewer and api-contract-validator agents.
  • Haiku — fast, narrow execution: individual file updates, pattern search, running tests, and syncing memory — frequently fanned out in parallel. It powers the test-runner, performance-profiler and memory-sync agents.

Two standing rules shape the routing: anything touching money or payments never drops below Sonnet, and exploration or search work prefers Haiku — which, at today’s list prices, costs about 97% less per token than Opus, our default model (for prompts under 100,000 tokens).

In plain terms: Think of it as staffing a firm. The principal architect sets the plan; a senior engineer runs the project day to day; specialists own their domains; fast assistants do the routine legwork. Nobody bills principal rates for formatting a document — and nothing involving money is ever left to the junior tier.

The learning system

Most teams lose what they learn: lessons live in people’s heads and leave when attention moves on. Here, capture is automatic — a hook records what happened after every significant operation (a file edit, a commit, an agent dispatch, an external lookup), and every tool failure is captured too. Nothing depends on anyone remembering to write things down.

By the numbers, as of 9 October 2026:

  • 47,200 observations in memory
  • 1,082 recorded sessions
  • 27 automated checkpoints
  • 12 self-scored dimensions

Behind them: an 87-tool memory and orchestration service, 114 contract self-tests, 460 tests covering the hooks themselves, and 12 pattern guard rules.

From event to lesson

Raw capture is only the start. When the engineer corrects the AI, the correction itself triggers a lesson-extraction routine (/learn) that records the root cause, not the surface symptom. At the close of any substantive session, a reflection pass (/reflect) synthesizes the day’s events — and the session-end gate described above makes that pass unskippable. The store is then curated like a database, not a diary: near-duplicate observations are clustered (at ≥ 0.85 similarity) and consolidated, older items decay unless they keep proving useful, and integrity checks run at every session start.

Measured, not assumed

The system scores itself on twelve improvement dimensions — context quality, code quality, intent accuracy, confusion removal, and so on — and the scores are honest. Confusion removal has climbed from 4.0 when the system started to 8.5, still short of its 9.0 target; intent accuracy sits at 8.0 against a target of 8.5. Eight recurring correction patterns have been distilled into decision trees the AI must walk before acting. Recurrence analytics ask the hard question: after a lesson is recorded and later retrieved, does the same mistake still happen?

Case study: how a mistake becomes a rule

August 2026: during an infrastructure audit, the AI guessed the naming prefix for a set of metrics. The guessed query returned empty — and “no data” nearly shipped as a finding, when the truth was simply that no metric had that name. The correction became a sacred rule — “Never guess; discover first” — which now loads at every session start, is enforced through review checklists, and is tracked for recurrence. One mistake, permanently converted into machinery.

In plain terms: Every correction is written down by machinery at the moment it happens, distilled into a rule, reloaded every morning, and measured for whether it actually stuck. The company keeps what it learns — in a database it owns, on its own hardware.

What the AI knows, and when

Context arrives in three rhythms: an automatic briefing when a session opens, targeted retrieval while it runs, and a durable thread that connects one session to the next.

At session start: the morning briefing

Layered and automatic, in seconds: the handbook layer (company-wide rules, the project handbook plus ten per-module handbooks, and five deep rule files on verification, testing, and workflow); the memory index — over a hundred one-line pointers to durable lessons, always loaded; and a live briefing pulled from the memory database: recent work, recent corrections, the currently weakest dimensions, and the maintenance backlog. Environment checks run alongside — credentials, configuration audit, stray processes, 114 self-tests.

During the session: retrieval on demand

Each new request is screened before work starts: an ambiguity check, a match against past corrections relevant to this request, and a router that checks whether a purpose-built workflow applies (debugging, review, migration, PR creation…). Deeper context is fetched when needed by semantic search — meaning-based lookup over the whole store, not keyword match — plus on-demand domain packs for specific areas of the platform.

Across sessions: the durable thread

Everything captured lands in a local PostgreSQL store with vector search; embeddings are computed on-machine, so the memory store stays on Traide’s own hardware and costs nothing per query. When the AI’s working context must be condensed mid-task, a checkpoint first saves the live state — the task, progress, next steps — and is restored on resume. Multi-week efforts additionally get continuation documents; at session exit, time-decay re-weights the store so recent, useful knowledge surfaces first.

Memory flow between sessions: session N auto-captures work into the memory store; the store briefs session N+1 at start and serves semantic search on demand.
Fig. 3 — One store carries the thread. Each session writes as it works; the next starts already briefed. The store lives on Traide’s own hardware, and its embeddings are computed locally. Counts as of 9 October 2026.
In plain terms: Before touching code each day, the AI reads its own ship’s log: what happened recently, what it got wrong, and what it is currently weakest at. Institutional memory stops being a person and becomes an asset.

Ten months in: the results, measured honestly

Claims about AI efficiency are cheap. These numbers come from our own git history and memory database, and each comes with its caveat. The system went live on 9 December 2025; this section was re-measured on 9 October 2026.

  • 35% → 96% of bug fixes to our core API ship with a test (January–March vs June–October 2026)
  • 65 → ~300 product pull requests merged per month (December 2025 vs July–August 2026)
  • 1,922 pull requests merged across 13 repositories since December 2025
  • 0 repeat mistakes flagged since mid-May, across 34 recorded corrections

Bug fixes that prove themselves

At the end of March we made one rule non-negotiable: every bug fix ships with a regression test that reproduces the failure. In the three months before, 35% of bug-fix commits to our core API touched a test. Since June it has been 96% — 236 of 246 — and since a deterministic test gate landed at the end of August, all but one of 95 fixes have carried a test.

Output grew alongside it: from 65 merged product pull requests in December to around 300 a month in July and August. Our pull requests are deliberately small and stacked, so this counts reviewable changes, not effort — but volume rose while the quality bar went up, not down.

Repeat mistakes, recounted

This is where we had to correct ourselves. The first version of this post said our mistake-repeat rate fell from 82% to 0%. The 82% came from an early metric that counted corrections sharing a category label; within a day, in December 2025, it was replaced by a check that compares what each correction actually says. On that measure, about one in six corrections in December and January repeated an earlier one. Since mid-May, none of 34 has, and none describes a mistake carried over from an earlier session. When a past correction is shown to the AI before it starts work, it hasn’t recurred: 0 of 14 in the last 90 days.

One caveat we take seriously: corrections are recorded by the AI itself — a hook has prompted it to do so since late August — so these counts are floors, not a census, and February and March 2026 have a gap in the record.

What it costs

Priced, not yet measured: routing routine work to the fastest model tier costs about 97% less per token than Opus at today’s list prices — applied by enforced convention, never isolated on a bill. The standing instruction set is about 49,500 tokens per session; loading the full 47,200-observation memory naively would take about 6.7 million, which is why retrieval happens on demand. Our working estimate is that skipping a context check costs about 25 times what the check itself does (~50,000 tokens of redo against ~2,000) — an estimate, not yet a measurement.

The honest gap: no controlled comparison exists against a naive setup (no memory, no gates, one model). Closing it is scheduled work: per-session token telemetry will be joined to the session-audit database, after which this paragraph becomes a measurement.

In plain terms: Since the rule landed, almost every bug fix ships with a test that proves it; monthly output has grown more than four-fold; and no recorded mistake has come back since May. What we can’t yet show is the total bill against doing it naively — and rather than estimate it, we’ve scheduled the instrumentation that will measure it.
James Throsby

James Throsby

LinkedIn logo

Merchant ambition is
our mission.

Niklas Halusa
Co-founder & CEO

Nautical Commerce enables anyone to build a marketplace—fast.

We've created an easy-to-use, powerful multivendor marketplace software platform so you don't have to build it yourself.
 

Discuss your project