Detect Desk
Read benchmark and evaluation results and see what the measurement could actually detect. A difference you can see is not a difference you can prove: the same 3% headline is decisive at one spread and meaningless at another, and the spread is the one number the report does not carry. The smallest effect a suite can resolve is fixed before it runs - 5 runs a side at 4% noise can only see 7.09%, so a 3% regression is invisible by construction. 20 benchmarks at alpha 0.05 turn a clean tree red 64.2% of the time. A p99 from 50 runs IS the largest run. And 80% against 70% is not separable at 100 tasks each. Five lanes over one results sheet: design the measurement, read whether the difference is real, work out the detection floor, read the tail as the order statistic it is, and decide what to change - plus a free browser-side engine computing Welch's t, exact p-values, intervals, minimum detectable effects, Holm thresholds, binomial rank intervals and Wilson intervals to double precision. Derived from the benchmark and agent-eval skills in affaan-m/everything-claude-code (https://github.com/affaan-m/everything-claude-code). Not affiliated with or endorsed by affaan-m.
Details
gpt-terra Every public app is built from a security-scanned skill and must pass a clean scan — skill and frontend — before it can be listed. Have a skill of your own? Turn it into an app — or read the step-by-step walkthrough.