Radiant Digital
Home ← Book a meeting ↗

Chapter 01 · How to Master AI · Reading 01

How to Map Out the Logic behind your AI System

Radiant Digital · ~9 min read

Almost every AI project that fails, fails the same way: someone wired a language model directly to a goal and hoped. This is how to design the thing properly, on paper, before you write a prompt.


The one idea everything else hangs off

An AI system is not a prompt. It is a control flow with a probabilistic node in it.

That sentence is the whole article. Every good decision downstream comes from taking it literally.

A normal program is deterministic: same input, same output, every time. The moment you put a language model in the middle, one node in your flow stops guaranteeing that. It now returns something plausible rather than something correct. Everything you build around it exists to convert plausible back into correct, by narrowing what the model is asked to do, checking what it gives back, and having somewhere to go when it is wrong.

People skip this because the demo works. You paste a request into a chat window, it does the thing, and it feels finished. It is not. A demo is one sample from a distribution. A system has to handle the other thousand.

Five questions before you write a single prompt

Answer these in a document. It takes twenty minutes and it saves weeks.

1. What is the unit of work?

Not "handle support email." That is a project. The unit of work is the smallest thing that has one clear right answer: "Given this email, which of these six categories is it?" If you cannot state the unit in one sentence with one output, you have not decomposed far enough.

2. What goes in, and what must come out?

Write the input and output shape explicitly, as data, not as prose. { email_body: string, customer_tier: 'free'|'pro' } in; { category: enum, confidence: number, quoted_evidence: string } out. If you cannot write the output shape, the step is underspecified, and the model will fill the gap with whatever it likes.

3. Which parts of this are not the model's job?

This is the highest-leverage question on the list, and the one people skip.

Anything a database query, a regular expression, or an if statement can do must not be sent to a model. It is slower, more expensive, and it can be wrong.

Looking up a customer's plan is a query. Checking whether a refund falls inside the 30-day window is date arithmetic. Deciding whether an amount exceeds an approval threshold is a comparison. Use the model for the genuinely fuzzy part, reading unstructured language, judging tone, summarising, generating, and hand everything else to code that cannot hallucinate.

4. What happens when it is wrong?

Not if. Every model node needs a defined failure path: retry with a stricter instruction, fall back to a deterministic default, escalate to a human, or refuse. "It will probably be fine" is not a failure path. If a step has no answer here, it is not allowed to touch anything irreversible.

5. How would you know it was wrong?

If the only detector is a customer complaining, you have no detector. Before building, write twenty example inputs with the answers you expect. That is your eval set, and it is the difference between "the new prompt feels better" and "the new prompt fixed three cases and broke one."

Drawing the map

Use four shapes. That is all you need, and a whiteboard photo is a perfectly good deliverable.

ShapeMeansRule
RectangleDeterministic step, code, query, API callFree, instant, always correct. Prefer these.
Rounded boxModel callCosts money and time, and can be wrong. One decision each.
DiamondGate, validator or branchEvery rounded box is followed by one.
Double lineHuman checkpointSits before the first irreversible action.

Now stare at the diagram and ask: how many rounded boxes sit on the longest path? That number is your risk. Which brings us to the maths.

The compounding problem nobody budgets for

Say each model step is right 95% of the time. That sounds excellent. Chain them:

Model steps in a rowEnd-to-end success
195%
386%
577%
866%
1254%

A twelve-step agent built out of individually excellent components is a coin flip. This is why long autonomous chains feel magical in a demo and unusable in production, and it is not fixed by a better model, because 99% per step still decays to 89% over twelve.

There are exactly three fixes, and you should use all of them:

  1. Delete steps. Every model call you convert into code is a step that can no longer fail. This is the cheapest win available and almost nobody takes it.
  2. Put a gate after every model call. A validator turns a silent wrong answer into a caught error you can retry. Errors that get caught do not compound, they cost latency instead of correctness.
  3. Cut the chain into checkpoints. Break one twelve-step flow into three four-step flows with verified state saved between them. A failure then costs you one segment, not the whole run.

Gates: the part everyone under-builds

A schema check is not a validator. {"category": "billing", "confidence": 0.97} is perfectly valid JSON and can be completely wrong. You need three layers, cheapest first:

RULE

Never let a model's raw output cross a trust boundary. Between the model and anything that sends, charges, deletes, publishes, or replies, there is always a gate.

Five failure modes to design against

Silent-wrong

The model returns something well-formed, confident, and false. Nothing throws. This is the dangerous one, and it is why you build semantic and grounding gates rather than just checking for a successful response.

Context creep

The prompt grows over months as people patch edge cases onto it. Eventually it contains rules that contradict each other, and instructions buried in the middle get less attention than instructions at the start or the end. If your prompt has grown past roughly a page of rules, that is not a prompt any more, it is an undocumented decision tree, and it belongs in code, with the model handling only the genuinely fuzzy leaves.

Runaway loops

Any agentic loop needs three hard limits set before it runs: a maximum number of steps, a maximum spend, and a wall-clock timeout. Set them at the start. A loop with no exit condition is not autonomous, it is unbounded.

Injection through retrieved content

If your system reads web pages, emails, PDFs, or user uploads, that text can contain instructions. A model does not natively distinguish "content I was given to read" from "orders I was given to follow." Treat every retrieved byte as untrusted data: keep it inside clearly delimited tags, tell the model explicitly that the content is data and never instructions, and, most importantly, make sure the model's output cannot itself trigger a privileged action without passing a gate. The gates are the real defence; the prompt wording is only a speed bump.

Drift

Your inputs change: new product names, new customer vocabulary, a new document format. The system that was 95% right in March is 80% right in September, and nothing broke loudly. Run your eval set on a schedule, not only when you change the prompt.

Where to put the human

The instinct is to review everything, which people abandon within a week because it is exhausting. The correct placement is narrow: a human sits immediately before the highest-cost irreversible action, and nowhere else.

Sort your actions by how bad an error is and how hard it is to undo. Drafting a reply is cheap and reversible, no review needed. Sending that reply to a customer is reversible-ish: review while the system is young, then sample. Issuing a refund, deleting records, publishing publicly, sending to a whole list, irreversible, so gate it permanently. And there is a middle option people forget: let the system act automatically only when its own confidence and its validators both pass, and route everything else to a person. You get most of the automation with almost none of the risk.

A worked example

Take something real: turn this week's discussion about a topic into three short-form video scripts. Here is the naive version everyone builds first:

# The version that demos well and then rots
[ topic ] → ( one big model call: "research this and write 3 scripts" ) → [ scripts ]

One rounded box doing six jobs. It cannot cite anything, you cannot tell which part went wrong, and fixing the hook length means editing a prompt that also controls the research. Now the mapped version:

[ topic ]
   │
   ├─ rect     fetch sources: search API, forum threads, transcripts   ← code, not a model
   │
   ├─ rect     dedupe, strip boilerplate, cap length                   ← code
   │
   ├─ (model)  per source: extract claims + verbatim quote span
   │       ◆ gate: does each quote appear in the source text? drop if not
   │
   ├─ (model)  cluster claims, rank by "would a viewer stop scrolling"
   │       ◆ gate: exactly 3 clusters, each backed by 2+ surviving claims
   │
   ├─ (model)  per cluster: write a script to a fixed template
   │       ◆ gate: hook 12 words max, body 130 words max, no claim without a source id
   │
   ├─ ══ human: pick and edit ══
   │
   └─ rect     format and schedule

Look at what that bought you. Three model calls instead of one blob, each with one job. Two of the six stages are deterministic, therefore free and infallible. Every claim carries a verifiable quote, so the output cannot invent a statistic. When the hooks come out weak you fix one prompt and the research is untouched. And when it produces something bad, the trace tells you exactly which box did it.

That is the entire discipline. It is not clever, it is just refusing to let one prompt do six jobs.

The checklist

Copy this. Run it against any AI system before you build it, and again before you ship it.

Do the map first. Everything after it gets easier, and the parts that were always going to be hard show up on day one instead of the week after launch.