Bring your own model key

An ordinary model step uses a named reasoning profile with model: <provider>/<model> and key: secrets.<name>. Store your provider's API key in Vault and grant it to the workflow. The provider bills your account for the model calls; OutcomeCI records the call's usage with the run.

This provider registry and its capability checks were introduced in CLI 0.56.0. Run these examples with CLI 0.56.0 or newer and a compatible managed runtime. Earlier documentation remains available in the version selector.

Supported ordinary model providers

Use the exact prefix below before the model ID. Store the Vault entry's provider using that same prefix; matching is case-insensitive.

Provider Model prefix and Vault provider Key requirement
OpenAI openai Your Vault key, or the cloud's metered key
Anthropic anthropic Your Vault key, or the cloud's metered key
Google Gemini gemini Your granted Vault key
Groq groq Your granted Vault key
Mistral mistral Your granted Vault key
DeepSeek deepseek Your granted Vault key
xAI xai Your granted Vault key
Together AI together_ai Your granted Vault key
Fireworks AI fireworks_ai Your granted Vault key
Cerebras cerebras Your granted Vault key
OpenRouter openrouter Your granted Vault key
Cohere cohere_chat Your granted Vault key
SambaNova sambanova Your granted Vault key
Perplexity perplexity Your granted Vault key; tool-free calls only
NVIDIA NIM nvidia_nim Your granted Vault key
DeepInfra deepinfra Your granted Vault key

The provider registry identifies supported services. Capability checks use the installed LiteLLM package's bundled model metadata snapshot; they do not query a provider's live model catalog. The provider still decides whether a model exists, is available to your account, and accepts the request.

Ordinary steps with can, typed returns, or converse require tool calling. Both the primary model and its fallback are checked. A model known to lack tool support, or known to use a non-chat interface, is rejected with a diagnostic that names the step and model. Perplexity currently supports tool-free calls through this integration; required tools are never silently dropped.

When the installed metadata does not establish a capability, validation reports it as unverified and permits the workflow. Unverified does not mean supported; the provider performs the final check. See Capability checks for the validation report and runtime checks.

Model IDs can contain provider namespaces and colon variants, for example openrouter/anthropic/claude-sonnet-4.5 or openrouter/anthropic/claude-sonnet-4.5:exacto. Such syntax is accepted, but availability of a specific model or variant is determined by the provider.

Each registry entry uses a fixed provider endpoint. Workflow profiles cannot set api_base or arbitrary URLs. Azure, Bedrock, Vertex AI, and self-hosted endpoints require additional configuration and are outside this registry.

Capability checks

oci validate, MCP validate_workflow, and API workflow create/upload responses include model_capabilities and capability_warnings. These reports describe the ordinary model profiles used by the workflow, including fallbacks. A historical workflow list/get response can have null capability reports, meaning not evaluated, rather than verified support.

Known incompatibilities stop validation. Missing metadata or an unavailable local LiteLLM package produces an explicit unverified warning instead. Model names outside the metadata snapshot can still be published and run with a supported provider and the required Vault key.

Before an ordinary model call, the runtime checks the actual messages, images, and tools for both the primary and fallback models, before resolving provider credentials or contacting the provider. Known lack of tool or image support, known non-chat mode, and a known input-context overflow stop the call. Input token counts use LiteLLM's token counter with a packaged portable tokenizer. They are estimates, so runtime reports remain unverified and include an input_tokens_estimated warning even when model metadata is known. Checks do not fetch remote images or tokenizer files. If metadata or a safe token estimate is unavailable, the report remains unverified and the provider decides whether to accept the request.

Output limits come from the installed LiteLLM metadata, with a conservative default where no limit is known. The runtime also bounds the output reservation when checking available context. These checks do not guarantee that every provider-specific limit or model feature is satisfied.

Runtime reports and warnings are retained with model-turn responses and captures, and in local run artifacts. Inspect them alongside the model turns and usage when investigating a failed or unverified request. They record capability information without including credential values.

These checks apply to ordinary model calls. Typed dispatcher decisions use their separate Decisions contract described below.

Store and grant the key

  1. Choose the provider prefix from the table and a Vault path, such as models/gemini.
  2. Store a user-owned API key credential at that path, with provider gemini and the key in its api_key secret field.
  3. Grant the credential to the workflow that will use it. A fallback key needs its own grant to the same workflow.
  4. Declare the Vault path under secrets, then reference its name with the profile's key.

For example, store and grant the two keys used by the first workflow below:

printf %s "$GEMINI_API_KEY" | oci vault put models/gemini \
  --workspace-id WORKSPACE_ID --provider gemini --credential-type api_key \
  --value-stdin --workflow-id WORKFLOW_ID
 
printf %s "$GROQ_API_KEY" | oci vault put models/groq \
  --workspace-id WORKSPACE_ID --provider groq --credential-type api_key \
  --value-stdin --workflow-id WORKFLOW_ID

In the dashboard, use Vault → Add credential, choose an API key, set its provider and path, and grant workflow access. For a local run, store the same paths in the local Vault with oci vault local put and --credential-type api_key --value-stdin.

The key reference is required for all registry providers except OpenAI and Anthropic. Omitting it for those two providers uses the cloud's metered key; omitting it for another provider fails validation. A declared key that is missing, revoked, granted to another workflow, or stored under the wrong provider never falls back to the cloud's key.

Gemini with a Groq fallback

This complete ordinary workflow uses a model to summarize a manually supplied request. It does not require type: dispatcher.

apiVersion: outcomeci.workflow/v1
name: summarize-with-byok
trigger: manual
 
secrets:
  gemini: vault:models/gemini
  groq: vault:models/groq
 
reasoning:
  summarize:
    model: gemini/gemini-2.5-flash
    key: secrets.gemini
    fallback:
      model: groq/llama-3.3-70b-versatile
      key: secrets.groq
 
steps:
  - summary:
      using: summarize
      from: trigger
      reason: Summarize the request in one sentence and identify its main topic.
      returns:
        summary: string
        topic: string

The profile's fallback is one alternative {model, key?} profile. It cannot have another fallback. It is tried when the primary provider is rate limited or unavailable; it is not a way to ignore invalid requests or missing Vault access. Each provider uses its own declared key. This is distinct from the top-level agent reasoning.fallback list.

Namespaced models with an independent fallback key

This second complete workflow chooses models hosted through OpenRouter and DeepInfra. Keep the provider prefix before each model's own namespace.

apiVersion: outcomeci.workflow/v1
name: categorize-with-byok
trigger: manual
 
secrets:
  openrouter: vault:models/openrouter
  deepinfra: vault:models/deepinfra
 
reasoning:
  categorize:
    model: openrouter/anthropic/claude-sonnet-4.5
    key: secrets.openrouter
    fallback:
      model: deepinfra/meta-llama/Llama-3.3-70B-Instruct
      key: secrets.deepinfra
 
steps:
  - category:
      using: categorize
      from: trigger
      reason: Classify the request as a question, a bug report, or another request.
      returns:
        category: enum[question, bug, other]
        summary: string

Store the first key with Vault provider openrouter, even though the model namespace is anthropic. Store the fallback key with provider deepinfra. Grant both to categorize-with-byok. The model names in these examples show the accepted syntax; select available models with tool support from your provider account before running them.

Policy reviewers and typed decisions

This registry expands ordinary model profiles, including their use in conversation steps. It does not expand the policy reviewer or Decisions endpoint:

  • reasoning.review remains limited to OpenAI and Anthropic.
  • A decision step requires a type: dispatcher workflow and uses openai/gpt-6-luna or a TypeSafe Jev model. TypeSafe requires its own granted Vault key under service typesafeai (or the accepted alias typesafe) and is not an ordinary chat provider. Its model IDs use typesafe/.
  • Decision profiles do not support fallback profiles.

See Dispatchers for typed routing decisions and Policy review for permission-review configuration.