Longitudinal agent evaluation · v0.1.0-rc.5

Compare agent runs with evidence that lasts.

YYLO Benchmark validates cases, runs isolated snapshots, reconciles evidence, and produces bounded reports without hiding source artifacts.

Install / first run

From package to a bounded task.

Install the release candidate, then begin with an outcome small enough to verify.

Install
npm install -g @yylo/benchmark@0.1.0-rc.1
First run
yy benchmark --help

Capabilities / mechanisms

Control is visible, not implied.

01

Validated cases

Reject malformed benchmark definitions before execution.

02

Isolated snapshots

Keep candidate runs independent and reproducible.

03

Immutable evidence

Reconcile artifacts and receipts before producing reports.

04

First-class delegate

Invoke the same benchmark contract through yy or yylo.

Next step

Use this product inside a YYLO workflow.

Explore YYLO