Skip to content

Intelligence (LLM Runtime)

The intelligence layer is Athena’s LLM dispatch system. It owns provider routing, circuit breaking, fallback, the tool loop, and budget tracking. All AI node executions pass through it.


Provider Matrix

Athena supports 15 LLM / media providers:

ProviderModalities
OpenAIchat, vision, tools, streaming
Anthropicchat, vision, tools, streaming, prompt caching
Google (Gemini)chat, vision, tools, streaming
Groqchat, tools, streaming
Coherechat, tools, streaming
Perplexitychat, streaming
xAI (Grok)chat, tools, streaming
OpenRouterchat, tools, streaming (model-dependent)
DeepSeekchat, reasoning, tools, streaming
Mistralchat, tools, streaming
MiniMaxchat, tools, streaming
Moonshot (Kimi)chat, tools, streaming
ZAi (GLM)chat, tools, streaming
XiaomiMimochat, tools, streaming
Fal.aimedia generation only (image/video)

Provider API keys are stored as workspace secrets and resolved at execution time. Key lookup order: group config → workspace secret → 401 Unauthorized.


Circuit Breaker

Each provider has an independent circuit breaker:

ParameterValue
Failure threshold to open5 failures in 60 s
Open duration before probe30 s
Probe successes to close2

States: Closed (normal) → Open (failures exceeded) → HalfOpen (probe) → Closed (2 successes) or back to Open (probe fails).

When a circuit is open, callers receive an ExternalService error immediately (no LLM call made). This error is in the fallback-trigger set, so node-level fallback still fires if configured.


Fallback Chain

An AI node whose primary model produces text can carry an ordered fallback chain: the primary model plus up to three fallback models (four tiers total). In the builder’s “Model & Fallback Chain” section, “Add fallback” opens a model chooser for the next tier; each fallback row can be reordered or deleted. Save-time validation rejects a broken chain (a tier that repeats another, exceeds three fallbacks, doesn’t exist, is retired, produces a different output type than the primary, or is missing a capability the node needs).

Image and video generation nodes have no chain. Only a text-output AI node dispatches through the LLM chain described below; an image or video node dispatches through media generation instead, which carries no tiers at all — there is nothing for a saved fallback to run against. The builder hides the chain editor when the node’s primary model is not text-output, and if a node somehow saves a non-empty fallback_registry_ids there anyway, save-time validation rejects it.

Wizard-bound (model_field) nodes get no chain either: no builder affordance, and a saved fallback_registry_ids there is rejected — the legacy single-slot field below keeps working on such a node, unaffected.

Decision nodes have no chain for a different reason: the builder shows none, and a chain saved on one is simply ignored, not rejected — save-time validation skips every non-AI node, and the engine never reads a fallback chain off a Decision node’s config.

A node’s authored chain lives in model_config.fallback_registry_ids, in order. Registry ids are UUIDs (model-registry ids); the values below are placeholders:

{
"model_config": {
"registry_id": "019f902b-db73-7ea4-8ef2-60a8ff5747cd",
"fallback_registry_ids": [
"019f902b-e21c-7d3a-9b1f-2a6c4e9f10ab",
"019f902b-f430-7c5e-8a02-7b1d3f6e40cd"
]
}
}

The older single-slot field, model_config.fallback_registry_id, still works on a text-output node — it behaves as a one-entry chain — but a node with both set uses fallback_registry_ids and ignores the legacy field. On an image or video node the legacy field is inert instead: media dispatch never reads it, the builder hides the fallback editor there, and it isn’t validated at all.

Errors that advance the chain to the next tier:

Error typeAdvances?
RateLimit (429)Yes
ExternalService (500 / 502 / 503 / 504, timeout, connection error)Yes
Circuit breaker open (skipped without a call)Yes
Validation (400, prompt too long)No
Unauthorized / Forbidden / NotFoundNo
Policy or spend-cap denial, security gateNo
Structured-output validation failureNo — retries the same tier
Any error once output has been streamed to the callerNo

A tier only advances on one of the eligible errors above and only when nothing from that tier reached the caller yet. Anything else — including an eligible error that arrives after streamed content — stops the chain where it is and surfaces.

Streaming gets one exception: the stream-open retry (the node’s retry_config) applies to the primary only. Every tier, primary or fallback, is dispatched once past that point; the first committed chunk from any tier ends the chain — no later tier is tried after that, win or lose.

Chain-wide deadline: no new tier starts once 60 seconds have elapsed since the node’s first dispatch. Tiers that never got a chance are recorded as skipped for the deadline. If every tier that was tried fails, the node returns the last tried tier’s error — earlier tiers’ errors don’t mask it, and every tier (dispatched or skipped) is recorded in the output’s fallback_chain.

Each tier dispatches with its own provider and model. Its fixed_params are the node’s authored model_config.fixed_params — merged with the same exposed params the primary uses — then validated and stripped against that tier’s own input schema; resolution defines no per-tier registry defaults. per_model_params[<tier>] and per_model_prompts[<tier>].system_prompt apply on top for a tier that sets them. Web search stays on for a tier only when the node’s own web-search setting is on and that tier’s model supports web search — independent of whether the primary supports it. The user’s message itself is never rewritten per tier — a fallback answers the same request on a different model.

Run view: the execution chronology shows the chain as a fallback_chain flow — one line per tier with rank, provider · model, outcome, a reason, and latency, with the tier that answered marked.


Tool Loop Limits

LimitDefaultEnv override
Rounds (LLM call ↔ tool result cycles)8ATHENA_TOOL_MAX_ROUNDS
Tool calls per round8ATHENA_TOOL_MAX_CALLS_PER_ROUND
Total tool calls across all rounds20ATHENA_TOOL_MAX_TOTAL_CALLS

Exceeding any limit terminates the loop with the last assistant message as output. The execution is marked Completed, not Failed.


Budget Tracking

Every execution tracks token and cost accumulation in execution variables:

VariableTypeUpdated
__accumulated_tokensu64After every LLM round
__accumulated_cost_microdollarsu64After every LLM round

Cost is calculated in microdollars from the model_pricing table (admin-managed, 5-minute in-process cache).

Budget enforcement happens before every AI node:

ActionEffect
EndExecution fails with TokenLimitExceeded or CostLimitExceeded
WarnLogged once via a latch variable; execution continues
HandoffExecution fails with TerminationReason::Handoff

Pre-call token estimation uses a 1.2x safety buffer: estimated = (prompt_chars / 4 × 1.2) + 4096 output tokens.


Anthropic Prompt Caching

Athena uses a direct Anthropic HTTP client (bypassing the SDK) to inject cache_control breakpoints into prompts. This is enabled by default and can be disabled via ATHENA_ANTHROPIC_PROMPT_CACHING=off.

Cache breakpoints are placed at:

  1. System prompt — if 4096+ characters
  2. Last tool definition — if more than one tool is defined
  3. Last assistant message — if 4096+ characters

Maximum 4 breakpoints per request. Cache hit rates are tracked in cost accounting (cache_read_input_tokens vs cache_creation_input_tokens).


Streaming vs Sync

Both execution paths share the same intelligence layer. The streaming path emits token chunks via SSE as the LLM generates them. The sync path waits for the complete response.

SSE token stream event types:

EventMeaning
textLLM text chunk
reasoningThinking / reasoning chunk (DeepSeek, Claude extended thinking)
tool_callLLM invoked a server-side tool
tool_resultServer-side tool returned a result
llm_finalLLM call complete with token usage

  • Architecture — how the execution pipeline works
  • Workflows — BFS loop that calls the intelligence layer
  • Credits — cost accounting and credit pre-deduction
  • Error Codes — external_service_error, rate_limit codes