Thoughts
Thoughts

Inside OutcomeCI: built for trust and learning

How the vault, Outcome runner, and Outcome spec make agent execution auditable, with training and distillation in early beta.

An agent’s final answer tells you what it says it did. To trust the work, you need more: what instructions it ran, what access it had, which actions it requested, and what actually happened.

That is the thinking behind OutcomeCI’s architecture. We organize it around three pillars: the vault, the Outcome runner, and the Outcome spec. Each answers a different question about trust. Together, they make agent work something you can inspect and learn from.

Three pillars: the vault controls credential access, the Outcome runner traces execution, and the Outcome spec makes intent and permissions explicit.
Three pillars: the vault controls credential access, the Outcome runner traces execution, and the Outcome spec makes intent and permissions explicit.

The vault: give access, keep the keys

A workflow names a credential. The vault holds its value. Agents use connected APIs through the Outcome runner’s broker, without reading the underlying secret.

This separates permission to use a service from possession of its credentials. In the workspace vault, a workflow can use a credential only after you grant it access. Each cloud run receives a short-lived lease scoped to those grants. Cancelling the run revokes its lease.

A workflow names vault. The workspace vault checks its grant and issues a short-lived run lease. The Outcome runner broker authenticates the API call, keeping the secret out of agent context.
A workflow names vault. The workspace vault checks its grant and issues a short-lived run lease. The Outcome runner broker authenticates the API call, keeping the secret out of agent context.

Credentials can change without changing the workflow. Rotate a secret and the next run uses the new value. Revoke an entry to disable it for every workflow. The spec continues to refer to the same name.

Local runs resolve those names from an encrypted local vault. Cloud runs resolve them from the workspace vault. That keeps credential storage separate from the file that describes the work.

Read about the vault.

The Outcome runner: a traced execution environment

The Outcome runner is a traced execution environment. It executes the workflow’s steps, mediates connected API calls, and records their outcomes. That Trace package is central to why we built it, with two priorities in a deliberate order.

First: auditability

You should be able to follow the work from its instructions to its result. Each run records the workflow revision, step results, and API call outcomes. The broker’s journal holds the proposed requests, their purpose, their review, and their results. Agent transcripts and token usage are captured when the agent runner exposes them.

The broker enforces each step’s grants. When a step has a policy, an independent reviewer checks proposed writes against it before the broker adds credentials and sends the call. The API response returns through the broker, which gives the agent the projected result and records the call’s outcome. The reviewer cannot expand the step’s permissions. Missing, malformed, or non-allow decisions fail closed.

An agent proposes a call. The broker checks grants and applies policy review when configured. A permitted call reaches the connected API, and its response returns through the broker to the agent. Agent transcripts feed directly into the Trace package alongside API results and policy decisions. The package supports audit first, then curated data for training and distillation in early beta.
An agent proposes a call. The broker checks grants and applies policy review when configured. A permitted call reaches the connected API, and its response returns through the broker to the agent. Agent transcripts feed directly into the Trace package alongside API results and policy decisions. The package supports audit first, then curated data for training and distillation in early beta.

A useful audit trail includes more than successful actions. Refused requests remain visible with their reasons. A confirmed effect is distinguished from an uncertain one. An identical confirmed request returns its saved receipt; an uncertain request is not automatically replayed.

That distinction matters: a step finishing does not, by itself, prove every intended effect happened. The Trace package gives you the evidence to check. Policy review adds judgment, but it is not a proof that an agent’s work is correct.

Explore what a run records.

Second: training and distillation (early beta)

Training and model distillation are in early beta.

The same Trace package has value beyond reviewing a single run. Inputs, step outputs, call results, and available transcripts give you source material for understanding which approaches worked and building better models from that evidence.

Our second priority is making execution data useful for training and model distillation. A reviewed trace can show how a capable agent approached a task. A checked result can help identify examples worth learning from. Connecting those examples to the workflow revision preserves the context that produced them.

The Trace package is the foundation. Selecting examples, checking outcomes, removing sensitive data, and preparing datasets happen in your curation and training pipeline. Capturing a trace does not automatically make it a good training example.

Auditability comes first because you need to understand the work before deciding what to teach a model from it.

The Outcome spec: make intent explicit

The spec makes the work readable before it runs. A workflow declares its steps, connected APIs, and permissions in a versioned file. Each step says what it is trying to do and which capabilities it may use.

For example, this step excerpt permits posting to one Slack channel and adds a policy for the reviewer:

steps:
  - announce:
      reason: >
        Post a summary to the team.
      can:
        - slack.post: {channel: updates}
      policy: >
        Share the summary only.
        No customer details.

reason describes the task. can defines the grant the broker enforces. policy gives the reviewer instructions for assessing the proposed write. Those roles are distinct and visible in the same file.

Compilation includes referenced instruction files in the workflow revision. A run records that revision, connecting what you reviewed beforehand to what actually executed. Local and cloud runs use the same runner image, so moving a workflow online preserves its spec and execution model.

Explore the Outcome spec.

Know what ran. Learn from what worked.

The vault controls access. The spec defines the work. The Outcome runner leaves a Trace package you can audit and build on.

These three pillars let us start with the question that matters most: can you inspect what the agent did? From there, the same evidence becomes a foundation for evaluating results, improving workflows, and, in early beta, training or distilling models from work you have actually checked.

Run your first workflow, or see how local execution works.