Model rankings

What models real apps run in production. Measured from every metered run across hosted SkillSafe apps over the last 30 days — 243 runs — not benchmarks. Updated several times a day.

# Model Provider Runs (30d) Apps Success
1 Gemma 4 26B (Workers AI) @cf/google/gemma-4-26b-a4b-it Workers AI 161 2 45.3%
2 GPT-5.6 Terra gpt-5.6-terra OpenAI 59 14 88.1%
3 GPT Image 2 gpt-image-2 OpenAI 12 6 100.0%
4 FLUX.2 Klein 4B (Workers AI) @cf/black-forest-labs/flux-2-klein-4b Workers AI 9 2 0.0%
5 GPT-5.6 Sol gpt-5.6-sol OpenAI 1 1 100.0%
6 Claude Haiku 4.5 claude-haiku-4-5 Anthropic 1 1 0.0%

Aggregates only — no per-app or per-user figures. Data as of 2026-08-27.

How this is measured

  • Production runs, not benchmarks. Every row is built from metered runs that real apps executed and real users paid for over a trailing 30-day window.
  • Runs count toward the model that actually executed them. An app's configured model gets the credit unless the user overrode it — model aliases resolve to the concrete model that ran.
  • Success rate is the share of terminal runs that completed rather than failed (timeouts and provider errors count as failures; in-flight runs are excluded).