Agentic systems engineering, with a person on every irreversible step
Saffron Systems builds agentic software: programs that use a language model to plan, call tools and take actions inside your systems, under explicit permissions, measured evaluation and human approval for anything that cannot be undone. The goal is software that finishes real work and leaves a record of what it did, not a chat window.
- Division
- Saffron Systems, the software and application engineering division of Saffron
- Principle
- Automate everything that dulls the hand, never the hand itself
- Default rule
- Irreversible actions require a person's approval
- Sister division
- Saffron Automations runs operational workflows
What an agentic system is
An agentic system gives a language model a goal, a set of tools it may call, the data it may read, and rules about what it may change. The model decides which step to take next; the surrounding software decides whether that step is allowed, records it, and stops when a person must decide. The engineering is almost entirely in that surrounding software.
A working agentic system has six parts, and a missing part is where most failures come from:
- Tools with narrow, typed inputs, each one doing one thing an engineer can reason about.
- Permissions that separate reading, reversible writes and irreversible actions, with credentials scoped to exactly what each tool needs.
- Approval gates that pause the run and ask a named person before sending money, messages to third parties, deleting data or changing access.
- An action log recording every tool call, its inputs, its result and the reason the model gave, so any outcome can be traced.
- Evaluation: a suite of real tasks with known correct outcomes, run before every change to prompts, tools or models, so quality is measured rather than assumed.
- Limits on cost, time and number of steps, with a defined fallback when a limit is reached.
Who builds what: Systems and Automations
Saffron has two divisions that work with agents, and they do different jobs:
Runs operational workflows for a business: follow-up, intake, scheduling, revenue recovery and similar work, delivered as a service with reporting on what was done. Start there if the goal is an outcome in your operations. saffronautomations.com
Engineers custom agentic software when the agent has to live inside your own product or infrastructure, integrate with systems no standard tool reaches, or meet requirements a managed service cannot, such as your own hosting, audit or data residency.
When an agent is the wrong answer
Many problems described as needing an agent are better solved with ordinary software. Saffron will say so in discovery:
- The steps never change. If the workflow is the same every time, a deterministic program is cheaper, faster and easier to test than a model deciding each step.
- Every action is high-stakes and irreversible. If a person must approve every single step, the agent adds cost without removing work.
- The data it needs is not reachable. An agent cannot act on information it cannot read. The integration work comes first.
- Nobody can say what a correct outcome looks like. Without that, there is nothing to evaluate against, and the system cannot be trusted to improve.
How Saffron engineers for safety
Treat everything the agent reads as data, never as instructions
Emails, web pages, documents and tool results can contain text written to redirect an agent. This is known as prompt injection, and it is the first risk listed in the OWASP Top 10 for LLM Applications. Saffron's designs keep instructions and retrieved content in separate channels, restrict which tools can be triggered by content the agent has read, and require approval for any action that sends information out of your systems.
Least privilege, enforced outside the model
A model's judgement is not a security boundary. Permissions are enforced by the tool layer and by the credentials each tool holds, so that even a confused or manipulated model cannot take an action it was never given the means to take.
Evaluation before every change
Prompts, tools and models all change behaviour. Each change runs against the evaluation suite first, and the results are kept, so a regression is caught by a test rather than by a customer.
Risk management that names the harms
For systems with material consequences, the design documents what could go wrong, who would be affected and which control addresses it, following the structure of the NIST AI Risk Management Framework.
How a build runs
- Discovery. The task, the systems involved, the actions the agent would take, and which of them are irreversible.
- Evaluation first. A set of real tasks with correct outcomes is assembled before the agent is built, so there is a way to know when it works.
- Tools and permissions. The narrowest set of tools and credentials that can do the job.
- The agent loop. Planning, tool use, approval gates, limits and the action log.
- Shadow running. The agent proposes actions while people still do the work, and the two are compared.
- Supervised launch. Approval gates on, monitoring on, with a documented way to stop the agent immediately.
- Handoff. Source, evaluation suite, runbook and credentials transferred to you.
What you own
- The source code, prompts and tool definitions, in your repository.
- The evaluation suite and its history of results.
- The model provider accounts and keys, in your name, so you can change provider without asking anyone.
- The action log, stored in your systems.
- Reusable Saffron components, under a permanent licence to use them.
What Saffron can show you today
Client agent systems are private. What is public, and exactly what it demonstrates:
| Evidence | What it demonstrates | What it does not |
|---|---|---|
| How a Saffron release is proven | How Saffron governs work drafted by software agents on its own properties: the same worktree rule, checks, visual evidence and signed release as any engineer, with a person holding the key. You can verify a published signature yourself. | It governs releases of Saffron's own sites. It is not a client agent system. |
| saffron-quality-gate | Saffron's engineering standard written down as a runnable command-line check, the same bar whether code is drafted by a person or an agent. | A pattern scanner flags known failure patterns. It does not prove correctness. |
| Saffron Automations | The operational side of Saffron's agent work, described as services with their own terms and proof. | Those pages describe Automations' services, not Systems engineering engagements. |
Questions buyers ask
What is the difference between an agent and an automation?
An automation follows the same steps every time. An agent decides which step to take next based on what it finds. Agents are useful when the path varies from case to case; automations are better when it does not.
Can an agent take actions without a person approving them?
Reversible, low-risk actions can run on their own within limits you set. Actions that cannot be undone, such as sending money, messaging people outside your company, deleting data or changing access, require approval by default. You decide where that line sits, and the decision is written down.
How do you know whether an agent is working?
By measuring it against an evaluation suite of real tasks with known correct outcomes, before launch and before every change, and by reviewing the action log in production.
Which language model do you use?
The model is chosen per task, based on quality on your evaluation suite, cost, speed and the provider's data terms. The system is built so the model can be replaced without rewriting it.
Where does our data go?
Only to the systems named in the design: your own infrastructure and the model provider you have approved, under that provider's business terms. The design lists every place data travels.
Describe the work you want an agent to do
Write with the task, the systems it touches, and which actions in it cannot be undone. The reply will tell you whether an agent is the right tool, and if it is not, what is.