Clean

Read benchmark and evaluation results and see what the measurement could actually detect. A difference you can see is not a difference you can prove: the same 3% headline is decisive at one spread and meaningless at another, and the spread is the one number the report does not carry. The smallest effect a suite can resolve is fixed before it runs - 5 runs a side at 4% noise can only see 7.09%, so a 3% regression is invisible by construction. 20 benchmarks at alpha 0.05 turn a clean tree red 64.2% of the time. A p99 from 50 runs IS the largest run. And 80% against 70% is not separable at 100 tasks each. Five lanes over one results sheet: design the measurement, read whether the difference is real, work out the detection floor, read the tail as the order statistic it is, and decide what to change - plus a free browser-side engine computing Welch's t, exact p-values, intervals, minimum detectable effects, Holm thresholds, binomial rank intervals and Wilson intervals to double precision. Derived from the benchmark and agent-eval skills in affaan-m/everything-claude-code (https://github.com/affaan-m/everything-claude-code). Not affiliated with or endorsed by affaan-m.

Share

Details

PricingUsage-based + 10% creator margin
Billed model rate$2.20 in / $13.20 out per 1M tokens
Creator margin+10%
Effective rate$2.40 in / $14.40 out per 1M tokens
Security scanClean — skill and frontend scanned
Model gpt-terra
Created2026-09-01
Updated2026-09-03

Every public app is built from a security-scanned skill and must pass a clean scan — skill and frontend — before it can be listed. Have a skill of your own? Turn it into an app — or read the step-by-step walkthrough.