Chapter 1 of 7 · The model on its own
1aA model only writes text
A model only writes words, by guessing what comes next. It cannot act for you.
The AI model at the centre of all this is a text predictor. Given some text, it guesses a likely next token, a small piece of text, adds it, and guesses again. It has no hands, no phone and no way to reach a restaurant.
Chapter 1 of 7 · The model on its own
1bAsking it to book
You ask your AI model to book you a table for Friday.
It says your table is booked. No restaurant ever heard about it.
Go deeper
In more detail
- Tokens, not words: in English one token is about three quarters of a word, so 100 words is roughly 130 to 150 tokens, and the count of tokens going in and coming out is how a model's work is billed.
- Models do not all cut text the same way: some split it into whole words, some into fragments of words, and some into single characters.
- Token size varies with the writing system: in Chinese or Japanese a single character may be a token, and Bitdeer's guide says this variability is why understanding tokens matters for performance and cost.
- Writing one token at a time is not a law of nature: a research model trained with Bitdeer AI's graphics chips learned to write separate parts of an answer at the same time and then merge them, up to twice as fast per token without losing accuracy.
- Underneath the tokens is a great deal of arithmetic, run on graphics chips (GPUs): processors built to do many calculations at the same time.
For engineers
Mainstream large language models are autoregressive: each token is generated one after another, conditioned on everything before it.
Chapter 1 of 7 · The model on its own
1cAnswers are guesses
Because it is guessing, the same question can get different answers, and sometimes a confident wrong one.
Each next token is picked from a range of likely options, so asking twice can give different answers. Now and then the likeliest-sounding answer is simply false, and it arrives in the same sure tone as a true one.
Chapter 1 of 7 · The model on its own
1dAsking three times
You ask your AI model the same question, three times.
Three different answers, and the last was confidently wrong.
Go deeper
In more detail
- Variation is normal, not a bug: AI agents work by probabilities, so they may not always produce identical results, and people building with them are told to expect that.
- A system can be built to say “I'm not sure” when the evidence is thin, rather than make something up: one Bitdeer example tells a support assistant to answer only from the text it was handed, and to say it does not know if that text lacks the answer.
- The other fix comes in part 3a: let it look things up instead of guessing. Bitdeer's articles call this retrieval, fetching the right documents when the question arrives so the reply rests on them, which cuts down on made-up answers.
- Whether a system makes up a plausible answer or cites the right policy is not a question of which model it uses. It is a question of how the knowledge behind it is set up.
- Wrong answers can hide inside step-by-step reasoning too: a model can invent a step or a conclusion and carry it forward with high confidence.
For engineers
Output is sampled from a probability distribution over next tokens. Confidence estimation lets a model abstain when evidence is insufficient.
Chapter 2 of 7 · The program around it
2aThe model has no memory
The model remembers nothing between messages. The app you type into is a program around it, and it sends the whole chat again with every new message.
The program (engineers say agent harness) is the ordinary app between you and the model: it takes what you type, sends it to the model and shows you the reply. Every call, one trip to the model, starts from a blank page, so the program sends everything said so far again. That bundle is the context window: everything the model reads in one call.
Chapter 2 of 7 · The program around it
2bHow the chat is sent again
You mention your partner's nut allergy, then chat on.
It seems to remember the allergy. Really, the program sent it again, every time.
Go deeper
In more detail
- Every call starts from nothing, so the context window is the model's only memory of the chat. Bitdeer's own chatbot example shows it: the program keeps the chat as a list, adds each new message to the end, and sends the whole list with every call.
- That is why a long chat gets slower and costs more with every message: longer prompts cost more, and processing more tokens slows the reply.
- Reading and writing are separate jobs for the model. In an agent's long task, reading is often the bigger job, so one model Bitdeer AI hosts gives reading and writing their own separate shares of computing power.
- A reply appears piece by piece because the service can send each token as soon as it is written, instead of waiting for the whole answer.
For engineers
Model APIs are stateless. Each request carries the full message list: the system message, then every user and assistant turn so far.
Chapter 2 of 7 · The program around it
2cThe context window size limit
The context window has a size limit. When a chat outgrows it, the simple way drops the oldest part without telling anyone; the better way squashes it into a summary.
The context window, everything the program sends the model in one call, can only hold so much. When the conversation outgrows it, something has to go. Dropping the oldest messages is simple and silent. Summarising them keeps the facts that matter in far less space.
Chapter 2 of 7 · The program around it
2dDropping or summarising
Your chat has grown longer than the context window holds.
Dropped, the allergy vanished without a word. Summarised, it survived in one line.
Go deeper
In more detail
- The limit covers what goes in and what comes out together: it is the most text the model can handle in one request, the question plus the reply.
- A builder can also cap how long a reply may be, because the reply uses up room in the window and adds cost and waiting.
- Managing context (what the model is given), memory design and how the work fits together is described as becoming as important as the model itself.
- A long window has a physical cost: the longer the chat and the more people using the model at once, the less chip memory is left for everything else.
- Some models Bitdeer AI hosts accept windows of up to a million tokens. Bitdeer's own article notes that most models with large advertised windows become uneconomical well before the limit, and says its hybrid design keeps a million-token window usable in production.
For engineers
Context window overflow: truncation drops the oldest tokens; compaction summarises them.
Chapter 2 of 7 · The program around it
2eStanding instructions
Before your request, the program gives the model a short instruction saying what it is for.
This instruction rides at the top of every context window, ahead of anything you type. It sets the job: what the model is for and what to aim for. Rules you give once, like no outdoor seating, are added to it and ride along too. Engineers call it the system message.
Chapter 2 of 7 · The program around it
2fRules on every message
You give your rules once, then ask for Friday night.
You gave your rules once. They rode in on top of every message after.
Go deeper
In more detail
- Instructions guide the model, but they are words, and a rule that only lives in words is only as firm as the wording. Bitdeer's write-up on OpenShell describes enforcing rules at the runtime level, in the program that runs the agent, instead of relying solely on prompts or guardrails inside the agent (part 7i).
- An instruction can be short: Bitdeer's sample call sets the whole job with one line, “You are a knowledgeable assistant. Provide concise and clear explanations to scientific questions.” A rule about outdoor seating would be one more line like it.
- Bitdeer's guardrails article lists “intent alignment”, keeping an agent within its defined objectives, ethical limits and domain rules, as a key function of guardrails, and lists prompt templates (a set layout for requests) as a way to reduce unpredictability.
- Because the instruction rides along on every call, a long one is paid for again each time: longer prompts cost more, and trimming unnecessary words helps.
For engineers
The system message is sent as the first message in every request, before the user's turn.
Chapter 3 of 7 · Letting it act
3aUsing tools to act
Words become actions: the program shows the model a menu of tools, the model writes a request, the program carries it out.
The model only writes. The program, the ordinary app around the model, can reach the outside world. So it adds a menu of tools to the context window (everything the model reads in one call): search restaurants, check availability, make a reservation. The model writes a tool call, a filled-in request for one, and the program does it and brings back the answer.
Chapter 3 of 7 · Letting it act
3bA real booking system
You ask whether Paolo's has a table, then book it.
This time a real booking system answered, and really booked it.
Go deeper
In more detail
- Requests are filled-in forms, so the program can check them before acting: in Bitdeer's developer guide the model returns structured arguments, the program carries them out, and a strict setting makes the model follow the form exactly. A table for 12? Ask first.
- Looking up documents is a tool too, and it is the other fix for part 1d's confident wrong answers: instead of guessing whether Nonna's opens on Fridays, the model looks it up. Bitdeer's own example is a currency rate, where the agent calls a financial service to find the latest figure instead of guessing.
- Each tool on the menu comes with a plain description and an exact list of what it needs, the way a booking at Paolo's needs a date, a time and a party size. The description is what helps the model decide when to use the tool.
- Whatever a tool sends back becomes more text the model has to read, and read text is counted and billed, so fetching a long document for one small fact costs the whole document.
- A menu can be long: at a developer event Bitdeer AI co-hosted, one team gave an agent more than 300 specialist tools and let it decide which to use.
For engineers
A tool reaches the model as a name, a description and an input schema; the model replies with a structured call the program executes.
Chapter 3 of 7 · Letting it act
3cThe loop: act, then look
The program goes round and round: ask, act, look at the result, go again, until the job is done. Looking is how problems get caught.
One request is rarely enough. The program runs a cycle called the agent loop: it puts each result back into the context window, the model reads it and decides the next step, and the cycle repeats. If the program never showed the model the result, it would never notice that something failed.
Chapter 3 of 7 · Letting it act
3dActing without looking
You ask for 7pm at Paolo's, which is already full.
Without looking, it booked nothing. Looking, it saw “full” and got 7:30.
Go deeper
In more detail
- It can check its own work before replying, a pattern Bitdeer's blog calls reflection: the agent produces a first answer, critiques it, then revises it.
- Looking can repeat: some agents reflect on their output several times, improving it with each pass, and adapt to what they find along the way.
- Each look can shape the next step: in a multi-step search, each step has its own question, informed by what the previous step found, like asking about 7:30 once 7pm comes back full.
- Looking is also how an agent recovers: one Bitdeer write-up points to a public test that measures how well a model can chain commands, read the output and recover from errors across many steps.
- Retrying belongs to the loop too, when the fault is a busy service: Bitdeer's error guide marks a “too many requests” reply as safe to retry after a short wait, and a malformed request as not worth repeating.
For engineers
The agent loop is the ReAct pattern (Reason plus Act), a feedback loop: reason, act, observe, reflect, repeat, with a cap on iterations.
Chapter 3 of 7 · Letting it act
3eMemory in a notes file
Lasting memory is an ordinary file, not the model.
When the chat ends, the context window is thrown away and the model has learned nothing. What carries over is a plain notes file that the program writes to, and reads back in at the start of the next chat. You could open it and edit it yourself.
Chapter 3 of 7 · Letting it act
3fA new chat, the same notes
A week later, you open a new chat to book again.
The model learned nothing. The notes file remembered for it.
Go deeper
In more detail
- Reusable skill files are written the same way: in the agent framework Bitdeer's guide describes, a skill is plain text, and more than 100 ready-made ones are available.
- In the same framework, memory is a set of ordinary text files kept locally by default, so you can open them and read what the agent remembers.
- True long-term memory, where an agent remembers people and keeps learning over time, is described in Bitdeer's article on the agent economy as a complex hurdle that still needs advances in how information is stored and retrieved.
- Once the notes grow to years of past bookings, a program can search them and bring in only the few that matter tonight, such as the usual table size and any allergy, instead of reading every past dinner. Bitdeer describes this as plug-in storage for memory and retrieval that lets agents build up context over time.
For engineers
Persistent memory lives outside the model: local Markdown files, or a vector store for larger collections.
Chapter 4 of 7 · An agent
4aWhat makes an agent
All of that together is an agent: you give a goal once, it plans the steps, and it asks before anything that matters.
Model, program, context window, instruction, tools, the loop and the notes file: put them together and you no longer spell out each step. You give a goal; it works out a plan, carries it out, and stops to check with you where it counts, such as before spending your money.
Chapter 4 of 7 · An agent
4bFrom goal to plan
You give it one goal: plan Friday date night.
One goal and one yes from you. It planned, booked, and asked before paying.
Go deeper
In more detail
- A workflow starts when an event happens; an agent is given the goal once and works out the steps itself. Bitdeer's blog puts it that way.
- Bitdeer's example of a workflow: an invoice arrives, the system pulls out the details, sorts the expense, checks it against a budget and routes it for approval, with people brought in at decision points rather than at every step.
- Bitdeer's blog describes four stages, each riskier than the last: an assistant that answers when asked, a workflow that starts on an event, an agent that is given a goal, and a team of specialised agents.
- How much a person checks depends on the stakes: for routine tasks the AI drafts and a human approves, and for decisions with medical, financial or legal exposure a person signs off before any action. The safety chapter comes back to it.
- General-purpose models are flexible but often lack the precision complex workflows need, while specialised agents carry deep knowledge of one field and fit into the processes around it.
For engineers
Planning pattern: decompose the goal into steps, execute each with tools, gate consequential actions on human approval.
Chapter 4 of 7 · An agent
4cRunning on a schedule
An agent can work on its own schedule, and that needs a home that never switches off.
Once it can act alone, it can also act without being asked: every morning, every hour, whenever something changes. A heartbeat, a timer, wakes it up to check whether there is work to do. That only works if the computer it runs on is always on, such as a rented one in the cloud.
Chapter 4 of 7 · An agent
4dLaptop or always-on cloud
You ask it to check your calendar every day at 7am.
On the laptop, the 7am check never ran. In the cloud, Saturday got its table.
Go deeper
In more detail
- A heartbeat is a timer that wakes the agent regularly to check whether there is work to do: in Bitdeer's guide, a built-in heartbeat lets the agent monitor tasks and act without anyone prompting it.
- Running an agent on your own laptop means it stops whenever the machine shuts down, and Bitdeer's guide adds two more risks: fiddly setup and keys stored on the device.
- An always-on home need not cost much: for the small computer an agent needs, Bitdeer's guide lists pay-as-you-go pricing with no charge while it is stopped.
- The model's side is pay-per-use as well: Bitdeer AI's Model Studio is a service you call over the internet and pay for by the token, so a check that runs once a day pays for that day's text, not for chips sitting idle.
- A scheduled job in Bitdeer's guide is an ordinary chore: read the inbox, pull out the action items and send a daily digest at a set time.
For engineers
A scheduler or heartbeat loop triggers the agent on an interval; it needs an always-online host.
Chapter 4 of 7 · An agent
4eWhere each step runs
The program is cheap everyday software; only the model's thinking needs graphics chips, paid for by the amount of text.
Reading notes, checking a calendar and calling a booking system are jobs for an ordinary computer. The expensive part is the model, which runs on specialised graphics chips in a data centre and is paid for by how much text goes in and out.
Chapter 4 of 7 · An agent
4fOrdinary computer or chips
You book one table: five steps, two kinds of computer.
Most steps ran on an ordinary computer. Only the model's turns needed graphics chips.
Go deeper
In more detail
- Bitdeer's example: a basic cloud computer with 2 cores and 4 GB of memory is enough to run an agent, at about $0.0363 an hour, because running the agent does not need graphics chips.
- The model is billed separately, per million tokens. Those are two different units, an hour of computer time and a count of tokens, so the two numbers cannot be compared directly.
- Bitdeer's guidance keeps the two apart: ordinary computers can handle scheduling, the front doors that route requests (gateways), databases and the agent's runtime, keeping the graphics chips focused on the model's own work.
For engineers
The agent runtime is a CPU-only virtual machine; inference is an API call billed per token.
Chapter 4 of 7 · An agent
4gThe loop, counted in calls
Every turn of the agent loop is another call to the model, so one request from you becomes dozens of calls.
This is the agent loop from part 3c, counted in trips to the model. Every turn of it is a new call, and every call carries the whole context window again. An agent that takes ten turns to finish sends ten calls, each with a fuller context window than the last.
Chapter 4 of 7 · An agent
4hModel calls in a week
You ask once for a daily check, then a week passes.
You asked once. The model was called 40 times.
Go deeper
In more detail
- One request from you can mean many calls: agent work may trigger several model calls, lookups and tool actions for a single request, so Bitdeer's planning guide says capacity cannot be worked out from one call.
- Bitdeer's own worked example: an agent making 60 tool calls in a row adds only a few hundred tokens of new text at each step, while carrying tens of thousands of tokens of history into every one.
- Text the model has already read can be kept on hand for reuse, and Bitdeer AI's price list charges much less for it: one model lists $0.15 per million tokens of fresh input against $0.003 for cached input.
- One way to cut the cost: a smaller model for easy steps and a bigger one for hard steps. Bitdeer's write-up on always-on agents puts it as sending each step to the right model, since many steps are high-volume and repetitive.
- Cost is described as the most significant constraint: token use grows with usage, so an expense that looks manageable in a trial can become unsustainable at scale.
For engineers
Each turn of the ReAct loop is a separate inference call carrying the full context. KV cache (key-value cache) and prefix caching reuse the computation for a repeated prompt prefix; model routing sends simple steps to smaller models.
Further reading
- How Should Enterprises Configure AI Training and Inference Resources?
- Smarter, Faster, with Native Visual Understanding: DeepSeek-V4.1-Flash Is Now Live on Bitdeer AI Model Studio
- Day 0 Availability: Power Always-On Agents with NVIDIA Nemotron 3.5 Lightning on Bitdeer AI Model Studio
- From AI Hype to Real Deployment: What Enterprises Are Actually Struggling With
Chapter 4 of 7 · An agent
4iSplitting work to helpers
An agent can break a big task into pieces and start helper agents for them, which work at the same time and report back.
Checking five restaurants one after another is slow, and every result piles into the main context window. Instead, the agent can start five helpers, one per restaurant. They work at the same time and each sends back one line. Engineers call these subagents.
Chapter 4 of 7 · An agent
4jAlone or with helpers
You ask it to check five restaurants for Friday.
Alone: five slow rounds and a stuffed context window. With helpers: one round and five lines.
Go deeper
In more detail
- Helpers suit jobs that split into independent pieces: in the agent-to-agent standard Bitdeer describes, agents can finish subtasks independently and at the same time, like five restaurants checked at once.
- A helper can start helpers of its own: in Bitdeer's trip-planning example, the flight agent asks a currency agent to work out exchange rates, and agents can be created on demand.
- The split-and-merge idea works inside a single model too: a research model that Bitdeer AI supplied chips for splits a hard question into parts, works each part on its own path, then merges the paths into one answer.
- What a helper sends back is a structured answer, filled in like a form rather than written as a loose message, in a format built so that each step can be traced.
For engineers
Sub-agents run in parallel with isolated context and return a compressed result to the orchestrator.
Chapter 5 of 7 · Many agents
5aShared standards
Two shared standards: one plugs any tool into an agent, the other lets one agent ask another.
A standard is an agreed way of connecting, like one shape of plug for every socket. Without one, each tool the agent uses needs its own custom wiring. With one, your calendar, a restaurant search and your messages all plug in the same way. A second standard covers agent to agent: your agent can ask a restaurant's own booking agent directly, the way you would ask a person at the front desk.
Chapter 5 of 7 · Many agents
5bTools and another agent
You ask it to book a place that runs its own agent.
Three tools plugged in the same way, and one agent asked another.
Go deeper
In more detail
- The tool standard is MCP, the Model Context Protocol, and it has three roles. The host is the app that starts the job, the client carries the conversation, and the server offers one capability, such as a restaurant search.
- In Bitdeer's worked example, the tool standard records the request, status and results of a tool call, which helps when something goes wrong (part 6a).
- One app can connect to several tools at once, a calendar, a restaurant search and a messaging tool, and the people who build each tool only have to build their own side.
- The agent-to-agent standard is A2A. Each agent keeps its own memory and workspace, and a task sent between agents has named boxes for the goal, the details and the background, plus a signal for success or failure.
- The two standards can work together: the tool standard inside each agent, to manage its tools and safety limits, and the agent-to-agent standard between agents, to share out the work. Your agent uses the first to read your calendar and the second to ask Paolo's own booking agent.
For engineers
MCP: host, client and server over JSON-RPC 2.0. A2A: agents discover each other and delegate tasks.
Chapter 5 of 7 · Many agents
5cTeams of specialist agents
A team of specialists with a coordinator does more, but every conversation between them is one more place for a mix-up.
Split the job between specialist agents: one reads your calendar, one finds restaurants, one books, one sends messages. A coordinator, the agent that hands out the work and collects the results, keeps them in step. The catch is the conversations between them. Four specialists and a coordinator, all talking directly, can pair up in 10 ways, and each pairing is a place where a message can be misread.
Chapter 5 of 7 · Many agents
5dDirect notes or coordinator
You ask a team of four agents to book Friday dinner.
Agents passing notes turned Friday into Saturday. A coordinator kept it Friday.
Go deeper
In more detail
- Unlike helpers (part 4i), specialists each have their own expertise or role, as in Bitdeer's example of a marketing agent team.
- A team can include a checker. In Bitdeer's product-launch example, one agent runs quality checks on the others' work for factual accuracy, tone and brand alignment, the way one agent here could check that the booking says Friday and stays under $150.
- The number of pairings is every two agents that could talk to each other: 1 for two agents, 3 for three, 6 for four, 10 for five. It counts places a mix-up could happen, not how likely one is.
- In the agent-to-agent style, each agent acts on its own and can start subtasks of its own, so one request can fan out into several pieces of work.
- A team can also send each step to whichever model suits it best: a large model for coordinating the work and hard reasoning, and smaller specialised models for high-volume, repetitive work.
For engineers
Coordinator pattern: a lead agent distributes tasks, aggregates results and keeps alignment.
Chapter 5 of 7 · Many agents
5eShared systems under load
When everyone's agents rush at once, the systems behind them all strain together: the model, the search and the booking systems.
Your agent does not run alone. On Valentine's Day morning, thousands of other agents call the same model service, where the model runs on graphics chips, and the same restaurant search and booking systems. When any one of them runs short of room, every agent waiting on it slows down, and a slow agent can arrive after the table has gone.
Chapter 5 of 7 · Many agents
5fA normal day or a rush
Valentine's Day: everyone's agents book at 7am.
On a normal Friday it got 7pm. On Valentine's Day, the rush took it.
Go deeper
In more detail
- Companies that run models for others add graphics chips as demand rises and release them when traffic drops, which is how a rush stays affordable. That is the business Bitdeer AI Cloud is in.
- A rush hits three things at once: the model, the search databases behind lookups and outside services such as the booking system. If the system behind an agent cannot adapt in real time, the result is slow responses or even outages.
- Bitdeer's guidance for a service that must always be up is to set ahead of time a minimum number of copies running, some warm spare capacity, the level at which to add more, a limit on the waiting line and a fallback for when that is still not enough.
- Electricity is the slow part. How fast a data centre can secure and switch on stable power often decides how fast it can add capacity, and building new supply can take years while demand for AI grows quarter by quarter.
For engineers
Load spikes hit LLM inference, vector databases and third-party APIs together; elastic GPU capacity absorbs the inference share.
Chapter 6 of 7 · When it goes wrong
6aFinding the broken part
When an agent fails, find which part broke before you blame the model.
Besides the model, an agent has five parts working together: the screen you see, the steps it runs in order, the knowledge it reads, the tools it uses, and the human checks where a person approves. A failure can start in any of them. Replaying what happened, one step at a time, usually points at one part, and it is often not the model.
Chapter 6 of 7 · When it goes wrong
6bTracing one wrong booking
It booked a steakhouse for your vegetarian dinner.
The model did exactly what the notes said. The notes were out of date.
Go deeper
In more detail
- The five parts come from Bitdeer's blog: the screen (what you see), the steps (what runs, in what order), the knowledge (what it reads), the tools (what it can do) and the human checks.
- The steps part is the decision engine: it decides whether a message needs a search, whether earlier context matters or whether a tool should be called, and it is invisible in normal use.
- Out-of-date or badly organised knowledge is the most common cause they name, and reliably the last place teams look.
- A replay is only possible if every step was written down as it happened. Bitdeer's Agent Builder captures every interaction, and Bitdeer's worked example of the tool standard records the request, status and results of a tool call.
- Blaming the model and switching to another is usually the wrong diagnosis, and an expensive one. In the steakhouse case, a better model would have booked the same steakhouse from the same old notes.
For engineers
Interface, orchestration, knowledge, tools and human oversight layers. Traceable actions make the replay possible.
Chapter 6 of 7 · When it goes wrong
6cLevels of independence
How much to let it do alone: more freedom gets more done, and lets a mistake travel further before anyone sees it.
You choose where the agent sits on a ladder of freedom: answering only when asked, suggesting and waiting for your tap, booking on its own, or running as a team. Each rung up saves you effort. Each rung up also lets a mistake travel further before a person sees it, so there is more to undo.
Chapter 6 of 7 · When it goes wrong
6dHow far a mistake reaches
You ask for next Friday; it picks the wrong Friday.
Waiting for your tap: nothing to undo. On its own: two things. As a team: four.
Go deeper
In more detail
- Bitdeer's blog draws the ladder in four stages. An assistant answers only when asked. A workflow starts when an event happens, such as an invoice arriving. An agent is given a goal and finds its own steps. A multi-agent system is several specialists working together.
- A flawed summary costs minutes to fix. An agent that cancels a booking without authorization creates a business incident with financial, legal and reputational cost.
- Bitdeer's blog asks five questions first: is its reasoning reliable enough for this task; what data will it touch, and does it need all of it; which systems can it reach; which decisions need a person first; and could you rebuild what happened afterwards?
- The rung can differ by task: booking alone for a regular Friday table, waiting for your tap for anything with a deposit.
- Bitdeer's blog calls this leverage with less certainty: more freedom gets more done, and carries more risk, together. Its advice is to manage an agent like a team member, saying what it can and cannot do, setting clear boundaries and watching how it performs.
For engineers
Autonomy stages from assistant to multi-agent system, with risk escalating at each step.
Chapter 7 of 7 · Keeping an agent safe
7aAsking before acting
Ask before anything that costs money or cannot be undone.
Most steps are cheap to get wrong: a bad suggestion costs you a second. A payment or a cancellation is not. Those steps get a gate, a point where the agent stops and waits for a person to say yes before it goes on.
Chapter 7 of 7 · Keeping an agent safe
7bPausing before it pays
Paolo's wants a deposit before it will hold the table.
Without the gate, you heard about the $20 from your bank. With it, you were asked first.
Go deeper
In more detail
- Bitdeer's blog says where humans stay in the loop is a design decision with real consequences, not a default to leave unexamined.
- Bitdeer's blog sets the gate by the stakes. For routine tasks the AI drafts and a person approves, and for decisions with medical, financial or legal exposure a person must sign off before any action is taken.
- A gate does not mean a person approves every step. In Bitdeer's workflow stage, humans are involved at decision points, not at every step.
- The human check is the right design for high-stakes work, not a limitation to remove later, and good controls are what let an AI be given more authority over time without more risk.
- A gate can also depend on the situation: Bitdeer's guardrails article describes turning a tool on only under set conditions, such as the context or a confidence level.
For engineers
Human-in-the-loop approval on consequential tool calls, with a threshold policy.
Chapter 7 of 7 · Keeping an agent safe
7cHard limits and a stop
Hard limits and a stop button, so one mistake cannot multiply.
An agent works fast and never tires of repeating itself. If something goes wrong inside the loop, the ask, act, look cycle from part 3c, it can go wrong again and again before anyone looks. The program enforces limits, such as one table per evening, and a stop button halts everything. Both cap the damage.
Chapter 7 of 7 · Keeping an agent safe
7dCapping a repeating glitch
A glitch makes it book the same Friday over and over.
Without limits: five tables for one Friday. With them: one, and a stop.
Go deeper
In more detail
- A spending ceiling that switches the whole job off is a different thing from a cap on how many tries one step gets (part 3c).
- Bitdeer's guardrails article names isolation and rate limits for high-risk tool calls in particular, so the booking tool can be held to one table an evening while harmless lookups run freely.
- It also names a supervisor: agents that watch a running agent and can override its behaviour.
- Bitdeer's own model service works this way. By default it allows 100 requests a minute and answers anything above that with a retry-later message, and it has a separate limit on how much text can be used.
- A repeating glitch also costs real money, because the model is paid for by the amount of text. Bitdeer's blog notes that what begins as a manageable expense in testing can become unsustainable at scale.
For engineers
Rate limits on high-risk tool calls, budget ceilings and a supervisory kill switch.
Further reading
Chapter 7 of 7 · Keeping an agent safe
7eChecking facts first
Check the facts that matter before committing; never assume.
The model fills gaps with whatever sounds likely. For small things that is fine. For a fact that costs money when it is wrong, the agent should go and read it, and say so when it cannot find out.
Chapter 7 of 7 · Keeping an agent safe
7fGuessing or reading
You only want a table you can cancel for free.
Cancelling on Thursday: guessing cost $50, reading cost nothing.
Go deeper
In more detail
- Whether an answer invents a plausible policy or quotes the real one is not decided by which model you use. Bitdeer's blog calls it a data architecture decision, and describes the fix as fetching the relevant documents at the moment of the question rather than relying on what the model learned in training.
- A system can be told to say “I'm not sure” rather than guess when the evidence is thin.
- The lookup can be told to cite. Bitdeer's article puts the retrieved documents into the prompt with an explicit instruction to cite or synthesise them, so the answer can point back to the cancellation policy it read.
- Documents can carry labels such as date, author and type, and the search uses them to narrow the pile before it starts.
- Reading is only as good as what is read: if the source is biased, out of date or badly organised, the agent's answers suffer too. Bitdeer's blog adds that results may be influenced by the data sources, so businesses need regular audits to keep answers accurate and fair.
For engineers
Retrieval grounds the answer in a source; confidence estimation lets the model abstain.
Chapter 7 of 7 · Keeping an agent safe
7gOnly the access it needs
Give it only the access the job needs.
Access is what an agent is allowed to do. With wide access, an agent asked for one thing can do many others by mistake. If its job is booking, it needs to add bookings, not cancel them. Giving it only what the job needs means a misunderstanding stays small.
Chapter 7 of 7 · Keeping an agent safe
7hFull or book-only access
You ask it to “tidy up the calendar”.
With full access, three bookings gone. With book-only access, none.
Go deeper
In more detail
- People already get this on a cloud. In Bitdeer's guide, finance can only view billing and developers can only operate servers inside the test area. An agent's book-only access is the same idea.
- The same questions apply to data and reach: what will it read, does it need all of it, and which systems can it reach?
- Access can be set by the agent's role or by the situation it is in, not one setting for everything.
- Bitdeer's Agent Builder sets strict boundaries for the data and outside services an agent can reach.
For engineers
Least privilege: scope tool permissions by role or scenario; lock some at creation time.
Chapter 7 of 7 · Keeping an agent safe
7iHidden orders in text
It can be tricked by text it reads, so the rules that matter belong in walls around it, not in its instructions.
To the model, everything in the context window, all the program sends it in one call, is just text: your request, its instruction and whatever it reads along the way. A cleverly written line in an email or an invite can talk it into things. A rule in the instruction can be argued with; a wall the program enforces around the agent cannot.
Chapter 7 of 7 · Keeping an agent safe
7jInstruction or hard wall
A fake email tells your agent to send $400 away.
With the rule in its instruction, $400 went. With a wall, it held.
Go deeper
In more detail
- Engineers call this prompt injection. Bitdeer's guardrails article lists defences that stack: check what comes in for tricks and unclear intent, send each request down a path that fits its risk, and build walls behind that.
- A wall works because of where it sits. OpenShell, the runtime Bitdeer's blog describes, enforces its rules in the program that runs the agent instead of relying only on prompts or guardrails inside the agent.
- OpenShell has four walls: two fixed when the agent is created (its files and its processes) and two adjustable while it runs (its network and where its model requests go).
- A privacy router decides where each request to a model goes, which helps keep sensitive context such as calendar details under tighter control.
- Bitdeer's setup guide runs the agent in a sandbox, a sealed-off area, and lets you write your own network rules for what the agent may reach.
For engineers
Runtime policy enforcement outside the model: filesystem, process, network and inference controls; outbound traffic denied unless allowed.
Chapter 7 of 7 · Keeping an agent safe
7kSlow drift over time
The longer an agent runs, the more it can drift from what you wanted. A weekly summary and readable notes keep that visible.
An agent that runs for weeks keeps adding to its own notes and its own picture of what you want. Small wrong guesses pile up quietly: that is drift. If you can read what it did and what it now believes, through a weekly summary and a notes file in plain words, you catch it early.
Chapter 7 of 7 · Keeping an agent safe
7lWith or without a summary
You skip one Friday; it decides you've stopped Fridays.
With no summary, three Fridays lost. With one, the drift lasted a week.
Go deeper
In more detail
- Two plain ways to watch an agent over time: a log of what it did, and notes kept as ordinary text files a person can open.
- Bitdeer's blog ties drift to long tasks, where agents keep exchanging conversation history, tool results and intermediate steps turn after turn. That is also what makes long agent work cost more.
- Watching does not stop at launch. Bitdeer's guardrails article says the controls must carry on into production, so someone can step in, see what happened and be held to account.
- Memory design matters as much as the model. In Bitdeer's account of teams running AI in production, context management, memory design and workflow integration become as important as the model itself.
For engineers
Context drift over long-running tasks; mitigated by real-time logging, audit trails and human-readable memory.
Further reading
Chapter 7 of 7 · Keeping an agent safe
7mLayers of checkpoints
Safety is several checkpoints, each catching something different.
No single switch makes an agent safe. Four checkpoints each stop a different kind of mistake: what it reads, what the model decides, what it writes and which tools it may use. Two more block nothing themselves: a log lets you see what it did, and the rules you write set the limits the other four enforce. Together they cover what any one of them would miss.
Chapter 7 of 7 · Keeping an agent safe
7nOne booking, six checks
One booking passes through all six checkpoints.
Six checkpoints, one booking, and nothing slipped through.
Go deeper
In more detail
- The six layers come from Bitdeer's guardrails article: input, model, output, tool access, monitoring and governance.
- Only the first four stop something directly. Monitoring and governance watch and set the rules, so they are drawn as a log and a rulebook, not as gates.
- Each layer has its own techniques. The input layer uses prompt templates to keep requests in a set shape, and the output layer scans replies for harmful, biased or personal content.
- The last layer is written policy: rules for how the agent may be used, a safety check during development, testing and rollout, and a plan for what to do when something fails.
- Bitdeer builds guardrails in from the outset instead of after an incident. Its Agent Builder combines model-level safeguards with platform controls: strict access limits, a record of every interaction, and scanning of inputs and outputs.
For engineers
Defence in depth: input filtering, model alignment, output screening, tool permissioning, monitoring with override, and governance policy.
The whole agent
That is the whole agent: a model that only writes words, and everything built around it.
Around the model: a program that carries the context window and its instructions, a menu of tools, a loop that looks at every result, a notes file, and a home that never sleeps. Then checkpoints on every way out. The safety does not come from the model being clever. It comes from what is built around it.
A portfolio project by Kum Wai, with concepts from Bitdeer.