Inside OutcomeCI: built for trust and learning
How the vault, Outcome runner, and Outcome spec make agent execution auditable, with training and distillation in early beta.
An agent’s final answer tells you what it says it did. To trust the work, you need more: what instructions it ran, what access it had, which actions it requested, and what actually happened.
That is the thinking behind OutcomeCI’s architecture. We organize it around three pillars: the vault, the Outcome runner, and the Outcome spec. Each answers a different question about trust. Together, they make agent work something you can inspect and learn from.
The vault: give access, keep the keys
A workflow names a credential. The vault holds its value. Agents use connected APIs through the Outcome runner’s broker, without reading the underlying secret.
This separates permission to use a service from possession of its credentials. In the workspace vault, a workflow can use a credential only after you grant it access. Each cloud run receives a short-lived lease scoped to those grants. Cancelling the run revokes its lease.
Credentials can change without changing the workflow. Rotate a secret and the next run uses the new value. Revoke an entry to disable it for every workflow. The spec continues to refer to the same name.
Local runs resolve those names from an encrypted local vault. Cloud runs resolve them from the workspace vault. That keeps credential storage separate from the file that describes the work.
The Outcome runner: a traced execution environment
The Outcome runner is a traced execution environment. It executes the workflow’s steps, mediates connected API calls, and records their outcomes. That Trace package is central to why we built it, with two priorities in a deliberate order.
First: auditability
You should be able to follow the work from its instructions to its result. Each run records the workflow revision, step results, and API call outcomes. The broker’s journal holds the proposed requests, their purpose, their review, and their results. Agent transcripts and token usage are captured when the agent runner exposes them.
The broker enforces each step’s grants. When a step has a policy, an independent reviewer checks proposed writes against it before the broker adds credentials and sends the call. The API response returns through the broker, which gives the agent the projected result and records the call’s outcome. The reviewer cannot expand the step’s permissions. Missing, malformed, or non-allow decisions fail closed.
A useful audit trail includes more than successful actions. Refused requests remain visible with their reasons. A confirmed effect is distinguished from an uncertain one. An identical confirmed request returns its saved receipt; an uncertain request is not automatically replayed.
That distinction matters: a step finishing does not, by itself, prove every intended effect happened. The Trace package gives you the evidence to check. Policy review adds judgment, but it is not a proof that an agent’s work is correct.
Second: training and distillation (early beta)
Training and model distillation are in early beta.
The same Trace package has value beyond reviewing a single run. Inputs, step outputs, call results, and available transcripts give you source material for understanding which approaches worked and building better models from that evidence.
Our second priority is making execution data useful for training and model distillation. A reviewed trace can show how a capable agent approached a task. A checked result can help identify examples worth learning from. Connecting those examples to the workflow revision preserves the context that produced them.
The Trace package is the foundation. Selecting examples, checking outcomes, removing sensitive data, and preparing datasets happen in your curation and training pipeline. Capturing a trace does not automatically make it a good training example.
Auditability comes first because you need to understand the work before deciding what to teach a model from it.
The Outcome spec: make intent explicit
The spec makes the work readable before it runs. A workflow declares its steps, connected APIs, and permissions in a versioned file. Each step says what it is trying to do and which capabilities it may use.
For example, this step excerpt permits posting to one Slack channel and adds a policy for the reviewer:
steps:
- announce:
reason: >
Post a summary to the team.
can:
- slack.post: {channel: updates}
policy: >
Share the summary only.
No customer details.reason describes the task. can defines the grant the broker enforces.
policy gives the reviewer instructions for assessing the proposed write.
Those roles are distinct and visible in the same file.
Compilation includes referenced instruction files in the workflow revision. A run records that revision, connecting what you reviewed beforehand to what actually executed. Local and cloud runs use the same runner image, so moving a workflow online preserves its spec and execution model.
Know what ran. Learn from what worked.
The vault controls access. The spec defines the work. The Outcome runner leaves a Trace package you can audit and build on.
These three pillars let us start with the question that matters most: can you inspect what the agent did? From there, the same evidence becomes a foundation for evaluating results, improving workflows, and, in early beta, training or distilling models from work you have actually checked.