Compare like with like.
Different tasks, tools, budgets, and human assistance can change the outcome. A shared task set matters more than a single headline score.
AI evaluations / Public results
A closer look at how AI setups handle real tasks. Published results, the conditions behind them, and room for an honest comparison.
0
No completed results published
—
Setups with verified completed runs
—
Distinct case / suite pairs evaluated
—
No evidence to calculate a rate
The record
The first reviewed run will start the record. Until then, there is no leaderboard, success rate, or winning setup to report.
Example records and setup checks do not count as model evaluations.
How to read future resultsA useful comparison explains both the result and the conditions that produced it.
Different tasks, tools, budgets, and human assistance can change the outcome. A shared task set matters more than a single headline score.
When reviewed results are published, the run date, setup, coverage, and available metrics should make it clear what changed between runs.
Unavailable measurements stay unavailable. A blocked run, illustrative example, or setup check cannot establish a model's performance.
Publication boundary
Only reviewed, sanitized result summaries will be published here. Task prompts, private source, raw logs, and account details stay outside this public record.
For the owner: run and review evaluations in your private local workspace, then publish the approved export through the website's normal release workflow.
Explore Edudojo