AI layer

Router, not bindings

llm@1 is routed: after the grant check, the kernel’s router picks provider and model from Settings → AI (tiers fast/smart/best, fallback, embeddings). App inputs are strict objects — a model or provider id in params is InvalidParams.

Budgets and ledger

Per app per calendar month, tokens and/or micro-USD (default 2M tokens, 5 USD). Each call reserves before running and settles to reported usage (gateway cost, else catalog price; local costs 0). Over budget is BudgetExceeded, by hand and in jobs. Spend per app and model is shown in Settings → AI over Activity’s date ranges.

Model picker

With llm.choose, an app can offer a picker among the models the user made available. Handles are opaque per-app hashes — no provider ids leak and handles can’t be compared between apps. Local-only, pins and budgets still win.

Tools

llm@1 1.2 carries tool calls; tools@1 (Iro’s own, under tools.use) lists what the app declares, was granted and has a provider for, and runs a model’s call by tool name as the app. The app runs the loop (list → chat → call → results, ≤ 10 turns). Sensitive tools always ask, saying a model asked. Jobs can’t prompt.

Model apps

API-key gateways (Anthropic, OpenAI, OpenAI-compatible with a user-entered origin) and credential-free local Ollama (http://127.0.0.1:11434 only). Keys are typed into Iro’s page, kept in the vault, and sent only to the account’s origins. Streams are scrubbed with a held-back tail so split secrets are still caught.