Eval Forge — grade any AI output against the task it was given
Paste the task an AI agent or model was given and the output it produced — code, a report, an analysis, structured data — and get a rigorous evaluation: a ship/revise/rework verdict with a 0-100 weighted score, a five-dimension scorecard, a rubric scored criterion by criterion (yours, or derived from the task), severity-ranked findings each quoting the output they concern, and a full refined version that actually satisfies the task. A free instant prescan flags placeholders, boilerplate, broken JSON and truncated endings before you spend anything. Derived from the @github/agentic-eval skill (MIT).
Details
gpt-terra Every public app is built from a security-scanned skill and must pass a clean scan — skill and frontend — before it can be listed. Have a skill of your own? Turn it into an app — or read the step-by-step walkthrough.