YYLO Benchmark · stable 0.2.1
YYLO Benchmark documentation
Prepare reviewed cases, delegate independent attempts, and evaluate retained outputs with frozen checklists and separate result rows.
On this page
Published thin runner. The v3 record format is not a package version. Benchmark 0.2.0 introduced the breaking lifecycle; 0.2.1 adds checklists. Read the version changelog and benchmark study and evidence limits.
Install and verify
Node.js 20.10+, Git, tar and POSIX process groups. Trusted-host workspace hygiene, not a filesystem/account sandbox. Harness setup remains caller-owned.
Stable is the default. The latest published prerelease is an explicit choice, not a stable upgrade. Source version 0.2.1 is tracked separately from publication.
Registry channels checked . A prerelease channel may be older than stable; compare the exact versions above.
Stable 0.2.1
Reviewed reusable cases
Prepare historical Ledger tasks or supplied prompts from an explicitly reviewed pre-solution base.
Ledger draft reads completed requirements without completion responses; it never guesses the historical base.
Fresh source snapshots exclude future Git history, controller metadata and declared answer paths.
Boundary: Review is required; no adversarial host isolation or automatic answer detection is claimed.
Capability evidence
Published 0.2.1 npm archive README inspected; SHA-512 integrity and README SHA-256 retained. No live-provider test is claimed.
Verified exact releases: 0.2.1.
Package-owned source files: src/v2/workspace.ts src/v2/cli.ts README.md . Their fingerprints are checked by the frontend generator.
Stable 0.2.1
Delegated task and workflow attempts
Compare models, harnesses and configurations in independent fresh starting repositories.
YYLO Pi, command adapters and existing Workflow Runner are supported.
Selected-step comparisons execute an independent prefix through that step for each variant, then stop.
Workflows own sessions and dependencies. Errors are retained; no automatic compatibility repair, retry or resume.
Boundary: Prefix results measure earlier-step effects as well. Delegation grants no production authority.
Capability evidence
Published 0.2.1 npm archive README inspected; SHA-512 integrity and README SHA-256 retained. No live-provider test is claimed.
Verified exact releases: 0.2.1.
Package-owned source files: src/v2/adapters.ts src/v2/harness.ts README.md . Their fingerprints are checked by the frontend generator.
Stable 0.2.1
Independent later evaluations
New checks, judges and human assessments consume retained outputs without rerunning candidates.
No prebound evaluator catalog is required.
Evaluators receive copies of retained files; original attempts and previous evaluations stay unchanged.
Execution, checks, judge disagreements and evaluator errors are separate; no automatic combined verdict.
Boundary: Malformed, oversized or unavailable evidence is an evaluation error, not a capability verdict.
Capability evidence
Published 0.2.1 npm archive README inspected; SHA-512 integrity and README SHA-256 retained. No live-provider test is claimed.
Verified exact releases: 0.2.1.
Package-owned source files: src/v2/evaluators.ts README.md . Their fingerprints are checked by the frontend generator.
Stable 0.2.1
Traceable comparison rows
Display partial results, individual assessments, cost coverage, integrity errors and disqualifications honestly.
JSON and Markdown tables identify the treatment, attempt, execution status and each evaluator.
Known answer exposure is recorded as disqualification without deleting prior evidence.
Legacy evidence is preserved but not automatically migrated or interpreted by the new runtime.
Boundary: One-shot results are not reliability estimates. Reported cost is not billing; unknowns remain unknown.
Capability evidence
Published 0.2.1 npm archive README inspected; SHA-512 integrity and README SHA-256 retained. No live-provider test is claimed.
Verified exact releases: 0.2.1.
Package-owned source files: src/v2/cli.ts README.md . Their fingerprints are checked by the frontend generator.
Stable 0.2.1
Frozen criteria and deterministic loss
Evaluate retained outputs against explicit project and task checklists.
Case creation freezes normalized project/task criteria and their digest. Public criteria are visible to candidates; changing an input file cannot revise the case.
Every criterion needs pass, fail or unknown plus evidence. The runner computes equal-weight loss = failed / total. Any unknown produces null loss with insufficient_evidence; malformed assessments produce evaluation_error.
Later evaluation criteria replace the whole inherited checklist. Resupply both files to retain both. Reports mark criteria_changed; earlier attempts and evaluations remain untouched.
Boundary: Loss is partial quality, not production acceptance. Compare matching checklist/evaluator identities and keep execution, errors and disqualification separate.
Capability evidence
Published archive README and package metadata inspected; artifact integrity and README digest retained below. Documentation review is not a live-provider test.
Verified exact releases: 0.2.1.
Package-owned source files: README.md . Their fingerprints are checked by the frontend generator.
Methodology and historical evidence
The current lifecycle is case draft/create → run → evaluate → report, with append-only disqualify. A historical case needs review of original requirements, the pre-solution base and answer exclusions. New checks, judges and human assessments may evaluate retained outputs without candidate reruns or original-catalog prebinding. Execution, check failures, judge disagreement, evaluator errors and disqualification stay separate; there is no automatic winner or combined verdict.
Each workflow variant starts from the same initial input and runs its own prefix through the selected step, then stops. Results include upstream effects; this is not a measurement of that step alone or downstream continuation. The workflow/harness owns sessions, dependencies and errors; cross-harness compatibility is not guaranteed.
Trusted-host workspace/context hygiene is not enforced filesystem/account/network isolation. Retained input is not proof of literal provider-message delivery. Unknown cost remains unknown, and one-shot results are not reliability estimates.
Thin-runner command walkthrough · Fifteen-task implementation study. Old v1/v2 commands and evidence are historical, not migrated or rewritten by the redesign.
Troubleshooting
Start with the executable’s --version and --help. If a command is missing, compare its availability label with your installed version. Do not silently install a prerelease to make an example work.
Ledger and Benchmark can be used standalone. YYLO delegates enforce compatibility separately; newest packages are not automatically a compatible combination. For sandbox, credential or workspace failures, repair the named prerequisite before dispatch. Preserve logs and retained evidence; do not delete state to hide a failure.
Sources and release evidence
Capability review: 2026-10-07. Reviewed version: 0.2.1. Changes by version. Canonical package repository. Published availability below is verified for exact versions, never inferred from the current source version.
- @yylo/benchmark 0.1.0 artifact
Integrity and provenance
sha512-PlaJQis5v9/wyCyXsQsIQyglAnj3Ceb8sGkCLOzxtJ/oGJHsq353+MRNGGtz54DbAJXT+5I3rDBktE0vn6dr4w==
README SHA-256: fc167dfaeddd599a2d23de22ecb545f8b3094375fdeb840513cc710aa4f8eb00
- @yylo/benchmark 0.1.1-rc.2 artifact
Integrity and provenance
sha512-XulCxkBbf9s+P9tdhFBGTYn0+H+IRMpliYLou9odNDuZgTiGGEeQ01RlUdl67CKiv9gWGviewR9vYYKyNSHbJw==
README SHA-256: a34cc4c8ca4125577dde74b557deb3d121d7f4395c4cc287a003f54f13709005
- @yylo/benchmark 0.2.1 artifact
Integrity and provenance
sha512-9zF6P8EZ6wwSemLsYSugC5ttnL4A2awO0ETDdoHECAOPgRIx1gmsp3/Ce0en6LzlMY7VN0ibOgM8kvpTZOMrfg==
README SHA-256: 1e72b97ef16a2a7011d4089c1b0bd156d8f0d2e1b3bc407a57f97d5b08766429
Discover Skills for reusable agent procedures. Skills summarize intent here; the tagged repository owns the complete instructions.