Skip to content

Models And Modes

Models And Modes

One picker above the message box controls three things: Model — which AI answers you; Mode — whether one agent or a coordinated swarm does the work; and Effort — how hard Alfrada OS double-checks its own output before you see it. Leave all three on Auto and you never have to think about any of this. This page is for when you want control.

Model and Effort are picker controls, not chat commands: the agent can't switch its own model or Effort mid-conversation. Change them in the picker above the message box, then send the next message — for example, pick Claude Sonnet 5 - Reasoning and send:

After switching model in the picker
Redo the last answer with deeper reasoning — same question, but show your working and challenge your own first draft.

The new model takes over from the next message; earlier turns stay as they were.

Or leave the model on Auto, move Effort to Beast in the same picker, and send:

After turning Effort up to Beast
Produce the final version of this deliverable. Review it against the brief before you show it to me and fix anything that doesn't hold.

Beast keeps reviewing and revising until the rubric clears; on Auto it also reaches for a top-tier model.

The Picker At A Glance

Open the model selector in the chat input. At the top you get a controls strip with two rows:

  • a Mode row — Auto / Agent / Swarm
  • an Effort row — Auto / Fast / Medium / Beast, with the helper line: "How thoroughly Alfrada OS double-checks its own work — separate from the model you choose."

Below that sits a search box, then category tabs: Recommended (the default shortlist), plus Fast, Open SOTA, and Closed SOTA for browsing the full catalog.

One special case: pick Swarm and the model and Effort controls grey out, with a note explaining why — "Swarm orchestrates its own models per agent — model and performance settings don't apply here."

Choosing A Model

The Recommended tab is a six-model shortlist that covers most jobs:

  • GPT-5.6 Luna — fast and cost-efficient for high-volume everyday work: chat, sorting, classification, lightweight workflows.
  • Gemini 3.5 Flash-Lite — Community Auto's default. Very fast, reads huge documents (1M context), and handles images and tools.
  • Claude Sonnet 5 - Reasoning — near-flagship intelligence for coding and agent work, with the same huge 1M context.
  • Claude Sonnet 4.6 - Reasoning — the best-value reasoner: near-flagship quality for coding and agent work at a lower price, 1M context.
  • Nemotron 3 Ultra — a heavyweight open-weight model built for multi-step planning and long-running agent projects, pinned to Zero Data Retention hosting.
  • GPT-5.4 Mini — a small, budget-friendly reasoner that punches above its weight at coding and delegated sub-tasks. (It appears on the Recommended shortlist only — you won't find it under the three browse tabs.)

A simple rule of thumb:

  • use faster models for iteration, triage, and lightweight drafting
  • use stronger models for synthesis, reasoning-heavy analysis, and polished deliverables
  • switch models mid-session whenever the work changes — you keep the whole conversation

Model Notes: Nemotron 3 Ultra

One shortlist pick fills a niche worth naming: long, organized, multi-step projects that don't justify flagship pricing.

Nemotron 3 Ultra is built for planning-heavy work — coordinating research, analysis, and deliverables while keeping track of a moving plan across many tool runs. It's a 550-billion-parameter open-weight model (a hybrid Transformer-Mamba mixture-of-experts) that activates only 55B of those parameters per request — which is how it delivers that planning depth at the workhorse (3×) band instead of a 12× flagship rate. In Alfrada OS it reads up to 256K tokens of context — enough for long briefs, big source sets, and a plan that evolves as the session runs — it can lead or work inside a Swarm, and it's pinned to Zero Data Retention hosting.

If EU data residency is your requirement, pick Nemotron 3 Super 120B (the Bedrock route with the blue EU dot) instead — the smaller, EU-hosted sibling.

Pick Nemotron 3 Ultra in the picker, then:

Long-running project plan
Build a step-by-step launch plan for [product]. Research the market, map the buyer, identify risks, then produce a memo and a slide outline. Keep the plan visible as you work.

Nemotron is a good fit when the session needs planning, research, synthesis, and multiple deliverables.

Auto — What It Actually Does

When your model is set to Auto, Alfrada OS reads each message — how long it is, whether files are attached, which tools the task needs, how complex it looks — and decides how capable a model that turn deserves. It then picks the cheapest model above that quality bar whose provider is currently up and healthy. And it never downgrades you mid-conversation: once a session has used a stronger model, Auto won't drop below that level for later turns.

On Community, everyday asks can still land on a budget model such as Gemini 3.5 Flash-Lite. On Pro, Pro Max, and Black, Auto never uses a budget model — it picks the cheapest healthy model whose quality score is 60 or above.

Why you should care: you get a capable model on paid Auto without touching the picker, and Community still keeps the cheap default for small asks.

Effort — How Hard Alfrada OS Checks Its Own Work

Effort is separate from the model. It controls how thoroughly Alfrada OS reviews and revises its own output before showing it to you. The four options, exactly as the picker describes them:

OptionWhat the picker says
Auto"Alfrada OS picks the best model and effort for each turn"
Fast"No self-review — fastest. On Auto, the best budget model for the job"
Medium"One self-review-and-revise pass. On Auto, a workhorse model"
Beast"Reviews and revises until it passes (capped). On Auto, a top model, and grades any charts/images"

Effort does two independent things:

  1. On any model, it sets how deep the self-review loop goes — from none (Fast) to review-and-revise-until-it-passes (Beast).
  2. On Auto, it also nudges which model gets picked: Fast leans budget, Medium leans workhorse, Beast reaches for a top model.

Your Effort choice is remembered per session, so a deliverable-grade conversation stays on Beast until you change it.

Agent vs Swarm

  • Agent — one conversation thread does the work. It can still delegate: within a turn it can hand pieces of the job to parallel sub-agents through Smith, then fold their results back into the same thread.
  • Swarm — a planner drafts a multi-worker plan for you to review, then coordinated workers run in parallel, each assigned its own model by the planner. What makes Swarm different isn't parallelism — Agent mode has that too — it's the planner, the plan review, and the per-worker model choices.

Rule of thumb:

  • If you can describe the task as one main line of work, start with Agent.
  • If it naturally breaks into research, critique, synthesis, and output roles, use Swarm.

The agent itself can suggest Swarm mid-task when it sees the shape of the work. Nothing switches without a consent banner — you click, or it doesn't happen. You can change how that works at Settings → Agent Behavior → Safety → Switching to Swarm mode: Ask (the default), Always allow, or Never.

For the full Swarm story — plans, workers, costs — see Swarm.

What Models Cost You

Every model carries a multiplier — a token band — applied to the tokens it burns from your monthly budget:

BandMultiplier
budget
light
workhorse
plus
premium
sota12×
apex18×
overkill24×
background services

A 12× model isn't 12× better — pick it when the task deserves it. Background services (like titling and memory upkeep) cost nothing. Your current balance and what each turn cost are in Settings → Account → Usage & Plan.

Plans And Auto's Range

On Community and Pro, Auto picks from a curated pool of cost-efficient models — nearly all of it in the 1×–3× bands, with one premium-band pick (GLM 5.3) included. On Pro Max and Black, Auto is unrestricted and can range over the entire catalog. Paid plans still refuse budget models and anything below quality 60, so Pro Auto will not drop to Flash-Lite even on a short prompt.

This only shapes what Auto chooses. Manual picks are unaffected on every plan — if you select a model yourself, you get it. The Fast tab includes GLM 5.3 Flash (Budget 1×, native image input, direct Z.ai). Community Auto can pick it; paid Auto skips it because quality 58 sits under the paid floor of 60, same as Gemini 3.7 Flash. Pin it when you want the GLM family without Premium 6×.

Pinned Models Never Strand You

When a model is retired, sessions pinned to it roll forward automatically to its successor — no re-pin, no migration step. Real examples:

  • Gemini 3.5 Flash and Gemini 3.6 Flash → Gemini 3.7 Flash
  • DeepSeek V4 Pro → DeepSeek V4 Flash
  • Claude Opus 4.8 → Claude Opus 5
  • GLM 5.2 → GLM 5.3

Retired entries even say so in the picker: "Retired — sessions pinned here roll forward to …".

Labels You'll See On Models

  • EU (blue residency dot) — an EU-hosted route, e.g. MiniMax M2.5, Nemotron 3 Super 120B, and Kimi K2.5 on Bedrock (London), or Qwen3.8 Max on Alibaba's EU endpoint. The dot is the residency signal; the name no longer carries an [EU] suffix.
  • Global (grey residency dot) — the worldwide route for the same model family, e.g. MiniMax M2.5 via OpenRouter, Nemotron 3 Ultra. Where a family has both routes they share a name and are told apart by the dot.
  • Zero Data Retention (ZDR) — every request is pinned to endpoints that never retain your data. All live OpenRouter-routed models are ZDR-pinned with tool support.

One deliberate exception: Claude Fable 5 is explicitly labeled non-ZDR — the picker warns that "the provider may retain inputs and outputs" — so avoid sending sensitive data to that one model.

Wondering how each model actually performs on real Alfrada OS jobs? See the public benchmark study: Can It Run Alfrada?

For The Curious

Under the hood: identifiers and tiers

The six Recommended models, by internal model path:

Picker labelModel path
GPT-5.6 Lunaopenai:gpt-5.6-luna
Gemini 3.5 Flash-Litegoogle_genai:gemini-3.5-flash-lite
Claude Sonnet 5 - Reasoninganthropic:claude-sonnet-5
Claude Sonnet 4.6 - Reasoninganthropic:claude-sonnet-4-6
Nemotron 3 Ultraopenrouter:nvidia/nemotron-3-ultra-550b-a55b
GPT-5.4 Miniopenai:gpt-5.4-mini

Auto's internal capability tiers are named budget, workhorse, and sota. When your model is Auto, the Effort setting biases which tier the router targets: Fast → budget, Medium → workhorse, Beast → the top tier.

Built for Alfrada OS.