Model rankings
What models real apps run in production. Measured from every metered run across hosted SkillSafe apps over the last 30 days — 243 runs — not benchmarks. Updated several times a day.
| # | Model | Provider | Runs (30d) | Apps | Success | Share |
|---|---|---|---|---|---|---|
| 1 | Gemma 4 26B (Workers AI) @cf/google/gemma-4-26b-a4b-it | Workers AI | 161 | 2 | 45.3% | |
| 2 | GPT-5.6 Terra gpt-5.6-terra | OpenAI | 59 | 14 | 88.1% | |
| 3 | GPT Image 2 gpt-image-2 | OpenAI | 12 | 6 | 100.0% | |
| 4 | FLUX.2 Klein 4B (Workers AI) @cf/black-forest-labs/flux-2-klein-4b | Workers AI | 9 | 2 | 0.0% | |
| 5 | GPT-5.6 Sol gpt-5.6-sol | OpenAI | 1 | 1 | 100.0% | |
| 6 | Claude Haiku 4.5 claude-haiku-4-5 | Anthropic | 1 | 1 | 0.0% |
Aggregates only — no per-app or per-user figures. Data as of 2026-08-27.
How this is measured
- Production runs, not benchmarks. Every row is built from metered runs that real apps executed and real users paid for over a trailing 30-day window.
- Runs count toward the model that actually executed them. An app's configured model gets the credit unless the user overrode it — model aliases resolve to the concrete model that ran.
- Success rate is the share of terminal runs that completed rather than failed (timeouts and provider errors count as failures; in-flight runs are excluded).