Skip to content
All projects

AI evaluations / Public results

Show the evidence.

A closer look at how AI setups handle real tasks. Published results, the conditions behind them, and room for an honest comparison.

Read-only results. Evaluations run privately; this page never starts a run.

Verified completed runs

0

No completed results published

Evaluated setups

—

Setups with verified completed runs

Task coverage

—

Distinct case / suite pairs evaluated

Success rate

—

No evidence to calculate a rate

The record

Results & comparisons

Awaiting first publication

No verified results published yet.

The first reviewed run will start the record. Until then, there is no leaderboard, success rate, or winning setup to report.

Example records and setup checks do not count as model evaluations.

How to read future results

Evidence needs context.

A useful comparison explains both the result and the conditions that produced it.

01 / Comparable conditions

Compare like with like.

Different tasks, tools, budgets, and human assistance can change the outcome. A shared task set matters more than a single headline score.

02 / A traceable record

Keep the history in view.

When reviewed results are published, the run date, setup, coverage, and available metrics should make it clear what changed between runs.

03 / Honest limits

A missing value is not zero.

Unavailable measurements stay unavailable. A blocked run, illustrative example, or setup check cannot establish a model's performance.

Publication boundary

Public evidence. Private execution.

Only reviewed, sanitized result summaries will be published here. Task prompts, private source, raw logs, and account details stay outside this public record.

For the owner: run and review evaluations in your private local workspace, then publish the approved export through the website's normal release workflow.

Explore Edudojo