Intelligence (LLM Runtime)
The intelligence layer is Athena’s LLM dispatch system. It owns provider routing, circuit breaking, fallback, the tool loop, and budget tracking. All AI node executions pass through it.
Provider Matrix
Athena supports 15 LLM / media providers:
| Provider | Modalities |
|---|---|
| OpenAI | chat, vision, tools, streaming |
| Anthropic | chat, vision, tools, streaming, prompt caching |
| Google (Gemini) | chat, vision, tools, streaming |
| Groq | chat, tools, streaming |
| Cohere | chat, tools, streaming |
| Perplexity | chat, streaming |
| xAI (Grok) | chat, tools, streaming |
| OpenRouter | chat, tools, streaming (model-dependent) |
| DeepSeek | chat, reasoning, tools, streaming |
| Mistral | chat, tools, streaming |
| MiniMax | chat, tools, streaming |
| Moonshot (Kimi) | chat, tools, streaming |
| ZAi (GLM) | chat, tools, streaming |
| XiaomiMimo | chat, tools, streaming |
| Fal.ai | media generation only (image/video) |
Provider API keys are stored as workspace secrets and resolved at execution time.
Key lookup order: group config → workspace secret → 401 Unauthorized.
Circuit Breaker
Each provider has an independent circuit breaker:
| Parameter | Value |
|---|---|
| Failure threshold to open | 5 failures in 60 s |
| Open duration before probe | 30 s |
| Probe successes to close | 2 |
States: Closed (normal) → Open (failures exceeded) → HalfOpen (probe)
→ Closed (2 successes) or back to Open (probe fails).
When a circuit is open, callers receive an ExternalService error immediately
(no LLM call made). This error is in the fallback-trigger set, so node-level
fallback still fires if configured.
Fallback Chain
An AI node whose primary model produces text can carry an ordered fallback chain: the primary model plus up to three fallback models (four tiers total). In the builder’s “Model & Fallback Chain” section, “Add fallback” opens a model chooser for the next tier; each fallback row can be reordered or deleted. Save-time validation rejects a broken chain (a tier that repeats another, exceeds three fallbacks, doesn’t exist, is retired, produces a different output type than the primary, or is missing a capability the node needs).
Image and video generation nodes have no chain. Only a text-output AI
node dispatches through the LLM chain described below; an image or video node
dispatches through media generation instead, which carries no tiers at all —
there is nothing for a saved fallback to run against. The builder hides the
chain editor when the node’s primary model is not text-output, and if a node
somehow saves a non-empty fallback_registry_ids there anyway, save-time
validation rejects it.
Wizard-bound (model_field) nodes get no chain either: no builder
affordance, and a saved fallback_registry_ids there is rejected — the
legacy single-slot field below keeps working on such a node, unaffected.
Decision nodes have no chain for a different reason: the builder shows none, and a chain saved on one is simply ignored, not rejected — save-time validation skips every non-AI node, and the engine never reads a fallback chain off a Decision node’s config.
A node’s authored chain lives in model_config.fallback_registry_ids, in
order. Registry ids are UUIDs (model-registry ids); the values below are
placeholders:
{ "model_config": { "registry_id": "019f902b-db73-7ea4-8ef2-60a8ff5747cd", "fallback_registry_ids": [ "019f902b-e21c-7d3a-9b1f-2a6c4e9f10ab", "019f902b-f430-7c5e-8a02-7b1d3f6e40cd" ] }}The older single-slot field, model_config.fallback_registry_id, still works
on a text-output node — it behaves as a one-entry chain — but a node with
both set uses fallback_registry_ids and ignores the legacy field. On an
image or video node the legacy field is inert instead: media dispatch never
reads it, the builder hides the fallback editor there, and it isn’t validated
at all.
Errors that advance the chain to the next tier:
| Error type | Advances? |
|---|---|
RateLimit (429) | Yes |
ExternalService (500 / 502 / 503 / 504, timeout, connection error) | Yes |
| Circuit breaker open (skipped without a call) | Yes |
Validation (400, prompt too long) | No |
Unauthorized / Forbidden / NotFound | No |
| Policy or spend-cap denial, security gate | No |
| Structured-output validation failure | No — retries the same tier |
| Any error once output has been streamed to the caller | No |
A tier only advances on one of the eligible errors above and only when nothing from that tier reached the caller yet. Anything else — including an eligible error that arrives after streamed content — stops the chain where it is and surfaces.
Streaming gets one exception: the stream-open retry (the node’s
retry_config) applies to the primary only. Every tier, primary or fallback,
is dispatched once past that point; the first committed chunk from any tier
ends the chain — no later tier is tried after that, win or lose.
Chain-wide deadline: no new tier starts once 60 seconds have elapsed since
the node’s first dispatch. Tiers that never got a chance are recorded as
skipped for the deadline. If every tier that was tried fails, the node returns
the last tried tier’s error — earlier tiers’ errors don’t mask it, and
every tier (dispatched or skipped) is recorded in the output’s
fallback_chain.
Each tier dispatches with its own provider and model. Its fixed_params are
the node’s authored model_config.fixed_params — merged with the same
exposed params the primary uses — then validated and stripped against that
tier’s own input schema; resolution defines no per-tier registry defaults.
per_model_params[<tier>] and per_model_prompts[<tier>].system_prompt apply
on top for a tier that sets them. Web search stays on for a tier only when
the node’s own web-search setting is on and that tier’s model supports
web search — independent of whether the primary supports it. The user’s
message itself is never rewritten per tier — a fallback answers the same
request on a different model.
Run view: the execution chronology shows the chain as a fallback_chain
flow — one line per tier with rank, provider · model, outcome, a reason, and
latency, with the tier that answered marked.
Tool Loop Limits
| Limit | Default | Env override |
|---|---|---|
| Rounds (LLM call ↔ tool result cycles) | 8 | ATHENA_TOOL_MAX_ROUNDS |
| Tool calls per round | 8 | ATHENA_TOOL_MAX_CALLS_PER_ROUND |
| Total tool calls across all rounds | 20 | ATHENA_TOOL_MAX_TOTAL_CALLS |
Exceeding any limit terminates the loop with the last assistant message as
output. The execution is marked Completed, not Failed.
Budget Tracking
Every execution tracks token and cost accumulation in execution variables:
| Variable | Type | Updated |
|---|---|---|
__accumulated_tokens | u64 | After every LLM round |
__accumulated_cost_microdollars | u64 | After every LLM round |
Cost is calculated in microdollars from the model_pricing table (admin-managed,
5-minute in-process cache).
Budget enforcement happens before every AI node:
| Action | Effect |
|---|---|
End | Execution fails with TokenLimitExceeded or CostLimitExceeded |
Warn | Logged once via a latch variable; execution continues |
Handoff | Execution fails with TerminationReason::Handoff |
Pre-call token estimation uses a 1.2x safety buffer:
estimated = (prompt_chars / 4 × 1.2) + 4096 output tokens.
Anthropic Prompt Caching
Athena uses a direct Anthropic HTTP client (bypassing the SDK) to inject
cache_control breakpoints into prompts. This is enabled by default and can
be disabled via ATHENA_ANTHROPIC_PROMPT_CACHING=off.
Cache breakpoints are placed at:
- System prompt — if 4096+ characters
- Last tool definition — if more than one tool is defined
- Last assistant message — if 4096+ characters
Maximum 4 breakpoints per request. Cache hit rates are tracked in cost accounting
(cache_read_input_tokens vs cache_creation_input_tokens).
Streaming vs Sync
Both execution paths share the same intelligence layer. The streaming path emits token chunks via SSE as the LLM generates them. The sync path waits for the complete response.
SSE token stream event types:
| Event | Meaning |
|---|---|
text | LLM text chunk |
reasoning | Thinking / reasoning chunk (DeepSeek, Claude extended thinking) |
tool_call | LLM invoked a server-side tool |
tool_result | Server-side tool returned a result |
llm_final | LLM call complete with token usage |
Related Docs
- Architecture — how the execution pipeline works
- Workflows — BFS loop that calls the intelligence layer
- Credits — cost accounting and credit pre-deduction
- Error Codes —
external_service_error,rate_limitcodes