Eval Forge — grade any AI output against the task it was given

eval-forge.skillsafe.ai

Clean

Paste the task an AI agent or model was given and the output it produced — code, a report, an analysis, structured data — and get a rigorous evaluation: a ship/revise/rework verdict with a 0-100 weighted score, a five-dimension scorecard, a rubric scored criterion by criterion (yours, or derived from the task), severity-ranked findings each quoting the output they concern, and a full refined version that actually satisfies the task. A free instant prescan flags placeholders, boilerplate, broken JSON and truncated endings before you spend anything. Derived from the @github/agentic-eval skill (MIT).

Share

Details

PricingUsage-based + 10% creator margin
Billed model rate$2.75 in / $16.50 out per 1M tokens
Creator margin+10%
Effective rate$3.00 in / $18.00 out per 1M tokens
Security scanClean — skill and frontend scanned
Model gpt-terra
Created2026-08-10
Updated2026-08-11

Every public app is built from a security-scanned skill and must pass a clean scan — skill and frontend — before it can be listed. Have a skill of your own? Turn it into an app — or read the step-by-step walkthrough.