Skip to content

Can It Run Alfrada? — Inspect The Work, Not Just The Score

July 20, 2026

You can now inspect how nine frontier models perform inside the same production Alfrada OS harness. The public Can It Run Alfrada? study compares complete knowledge-work packages — including sources, calculations, memos, decks, charts, and schedules — rather than judging an isolated answer.

What you can do

  • Read the public LLM benchmark from the new Benchmark link in the Strategize Labs website navigation.
  • Compare nine models across two end-to-end jobs: a UAE market-entry decision package and a probabilistic semiconductor-capex forecast.
  • Inspect all 126 judge verdicts from Kimi K3, GLM-5.2, Gemini 3.5 Flash, GPT-5.6 Sol, GPT-5.6 Luna, Claude Fable 5, and Claude Opus 4.8.
  • Open the evidence behind each score — judge rationales, catalogued defects, charts, and downloadable work-product packages.
  • Compare quality, decision risk, cost, and speed together. Sol leads audited quality at 8.84; K3 offers a strong quality-to-cost balance; Luna delivers the lowest-cost high-quality result.
  • Select Inkling manually in Alfrada OS. Thinking Machines' multimodal model supports a 1M-token context window, reasoning, tool use, and text, image, and audio understanding through OpenRouter.

Where this shows up

  • You are choosing a model for consequential analysis. Instead of relying on a single leaderboard number, you can inspect which model's evidence and calculations survived an adversarial audit.
  • You manage analysts, scientists, reporters, or other knowledge workers. The study shows where autonomous work reduces production time and where accountable expert review remains load-bearing.
  • You want to try Inkling on a suitable workload. It is available in the model picker for deliberate testing, while Auto continues routing to models with more consistent audited quality.

Try it

  • "Compare the benchmark's Sol, K3, Luna, and GLM results for quality, cost, and load-bearing defects. Recommend the right model for a recurring strategy-analysis workflow."
  • "Open the evidence packages for the top three models and identify which calculations were independently reproduced by the judges."
  • Pick Inkling in the picker, then: "Produce a first-pass research brief, clearly separate sourced facts from assumptions, and include a reconciliation table for every headline number."

Heads up

  • Inkling is intentionally not in Auto routing or Swarm. It scored 5.60 in this study: 4.79 on the decision package and 6.41 on the forecast. The uneven result included recommendation-bearing arithmetic errors, conflicting source-of-truth files, and a forecast the judges could not reproduce.
  • A mechanical pass is not the same as trustworthy work. Inkling shipped every required artifact but fell below the study's practical 6.0 usability cutoff.
  • The study uses one run per model per case. Treat small score gaps as ties and use the published evidence to judge fit for your own work.

Built for Alfrada OS.