Stable 0.2.1
Reviewed reusable cases
Prepare historical Ledger tasks or supplied prompts from an explicitly reviewed pre-solution base.
Evaluation evidence / longitudinal comparison · stable v0.2.1
YYLO Benchmark prepares reviewed cases, delegates independent task or workflow attempts, and evaluates retained outputs with new checks and judges. Results remain separate, without an automatic winner.
The thin runner uses trusted-host workspace/context hygiene, not enforced filesystem/account/network isolation. Retained prompts do not prove literal provider-message delivery. Workflow prefixes include upstream effects and stop at the selected step.
Read the Benchmark version changelog →Install / first run
Install the registry-verified stable release, then begin with a change small enough to validate and review. Compare installation paths or review changes by version.
Version-verified capabilities
Stable 0.2.1
Prepare historical Ledger tasks or supplied prompts from an explicitly reviewed pre-solution base.
Stable 0.2.1
Compare models, harnesses and configurations in independent fresh starting repositories.
Stable 0.2.1
New checks, judges and human assessments consume retained outputs without rerunning candidates.
Stable 0.2.1
Display partial results, individual assessments, cost coverage, integrity errors and disqualifications honestly.
Stable 0.2.1
Evaluate retained outputs against explicit project and task checklists.
Product direction / clearly separated
Benchmark evidence is designed to help teams learn which agent and workflow works best for a repository and change type. Automatic production workflow selection requires connected real-world outcome history and is not a current-release claim.