Evalset Desk - score an LLM eval run, decide if it ships, then design the judge

evalset-desk.skillsafe.ai

Clean

Paste the export of one LLM eval run (CSV, TSV, JSONL or a Markdown table): your browser scores every row with 95% intervals, pairs the candidate against the baseline (wins, regressions, an exact sign test), measures Cohen's kappa, splits by slice and checks the claim and your ship bar. A verdict lane says ship, hold or fix the eval first; a judge lane designs the LLM-as-judge. Derived from the agent skill @wshobson/llm-evaluation (wshobson/agents, MIT).

Share

Details

PricingUsage-based + 10% creator margin
Billed model rate$2.20 in / $13.20 out per 1M tokens
Creator margin+10%
Effective rate$2.40 in / $14.40 out per 1M tokens
Security scanClean — skill and frontend scanned
Model gpt-terra
Created2026-10-02
Updated2026-10-10

Every public app is built from a security-scanned skill and must pass a clean scan — skill and frontend — before it can be listed. Have a skill of your own? Turn it into an app — or read the step-by-step walkthrough.