2026-08-29 · 12 min read
A copyable library of five agent-task dependency shapes — chain, fan-out, fan-in, diamond, and a cross-project hand-off held open by a forward reference — each drawn as a diagram, declared with the exact yylo-ledger commands, and proven by a recorded 2026-08-29 run: what ready admits at each stage, what the priority scores rank, and how wide each shape may safely run.
Authority · YYLO Ledger
yylo-ledger · dependencies
2026-08-29 · 9 min read
Model task dependencies an agent runtime can execute: one edge direction resolved only by terminal status, fan-in as an all-of join, admission as a computed query, cycles refused at the write, and scheduling rules drawn from readiness — including the sharp edge where the shipped topological sort stops answering mid-wave.
Solution · YYLO Ledger
yylo-ledger · dependencies
2026-08-29 · 10 min read
Drive YYLO Ledger from scripts and CI with contracts that hold: honest exit codes, parseable stdout in NDJSON and JSON, projection and cursor discipline, stdin-safe task input, and a receipt for every mutation.
Solution · YYLO Ledger
yylo-ledger · automation · cli
2026-08-29 · 9 min read
A reference architecture for AI development systems whose every change is auditable by construction — audit points that write evidence at each transition, admission boundaries that gate mutation, and rollback seams that retreat without destroying records — with YYLO as the worked implementation.
Authority · Ecosystem
audit-architecture · system-design · yylo
2026-08-29 · 8 min read
Select a coding agent with a decision matrix scored on repository size, risk, and change surface — one representative benchmark case per cell, a written reliability bar, and a default-plus-escalation portfolio with explicit economics — instead of brand loyalty or a borrowed leaderboard standing.
Solution · Ecosystem
benchmark · evaluation · selection
2026-08-29 · 10 min read
A versioned, citable taxonomy of coding-agent failure: nine stable classes with permanent identifiers, each carrying the symptoms that announce it, the detection that confirms it from evidence rather than testimony, and the recovery its evidence state permits — with attribution across the stack layers and a stability policy that keeps citations from drifting.
Authority · Ecosystem
failure-modes · recovery · taxonomy
2026-08-29 · 8 min read
A versioned, citable taxonomy for the coding-agent stack: six terms with stable identifiers, each with one definition, an owns-and-excludes split, and boundary criteria that settle real disputes — plus the stability policy and changelog that keep a citation from drifting.
Authority · Ecosystem
agent-harness · taxonomy
2026-08-29 · 8 min read
The citable ownership definition of the YYLO system — which component owns execution and admission, which owns task truth, which owns evaluation evidence — with the boundaries, failure domains, and interfaces between them.
Awareness · Ecosystem
yylo · yylo-ledger · benchmark · system-ownership
2026-08-29 · 12 min read
The explicit routing contract for agent task boards that span repositories: an opt-in per-project allowlist, one validated user registry, exact process replacement into the destination wrapper, and refusals that never fall back to the source board.
Solution · YYLO Ledger
yylo-ledger · routing · multi-repo
2026-08-29 · 11 min read
Feeding a pool of agent workers is a selection problem sitting on top of a readiness answer: scope the ready set so claims stay out, spend scarce slots where they unlock downstream work, and close the window where two dispatchers hand the same task to two agents.
Solution · YYLO Ledger
yylo-ledger · parallel
2026-08-29 · 11 min read
The artifact taxonomy that makes agent work provable after the fact — receipts that bind claims to bytes, manifests that vouch for collections, session records that outlive the terminal, and checkpoints that make interruption legible — together with the integrity, retention, and regrade semantics that keep each role trustworthy, quoted from YYLO's committed contracts.
Authority · Ecosystem
evidence · system-design · yylo
2026-08-29 · 8 min read
Follow one change through YYLO's governed lifecycle — task record, isolated worktree, focused validation, read-only preflight, guarded finish, and the serialized merge queue — with the documented command surface at every station.
Awareness · Ecosystem
task-lifecycle · validated-change · yylo
2026-08-29 · 9 min read
The storage contract for agent work that lives in your repository: one safe Markdown file per task with a validated schema, a five-status lifecycle with structural refusals, and completion bound to a commit hash instead of a checkbox.
Awareness · YYLO Ledger
yylo-ledger · git-native
2026-08-29 · 11 min read
A versioned support ledger for YYLO's agent-harness dispatch surface: one dated cell per harness, each with the command that re-verifies it against your own installed release, plus the update policy that keeps a citation honest.
Authority · Ecosystem
agent-harness · support-matrix
2026-08-29 · 10 min read
Move finished agent tasks into sealed, verifiable cold packs without losing a byte of evidence: revision-bound plans, content-addressed packs, read-only resolution, tamper evidence, and a maintenance rhythm that keeps hot state fast.
Solution · YYLO Ledger
yylo-ledger · archives · task-history
2026-08-29 · 10 min read
A downloadable, versioned extract of the public SWE-bench Verified leaderboard: 180 dated coding-agent system entries on one fixed 500-instance task set from 2023-10-10 to 2026-02-26, with extraction methodology, explicit limitations, and a content hash for citation.
Authority · YYLO Benchmark
benchmark · evaluation · dataset
2026-08-29 · 9 min read
Run several coding agents on one repository without ambiguous shared task state: an ownership write matrix, per-agent admission and filesystem scopes, and a sealed merge-admission gate that turns every landing race into a readable refusal.
Solution · YYLO Ledger
yylo-ledger · multi-agent · concurrency
2026-08-29 · 8 min read
Keep coding-agent workflows portable across harnesses and providers: a workflow format you own, declared substitution points that carry identity and paths at run time, and provider failures isolated inside the step that dispatched them.
Solution · Ecosystem
workflow-runner · portability · providers
2026-08-29 · 11 min read
A copyable, citable pre-launch gate for Ralph-style autonomous loops: fifteen permanent check items across five gates — bounds, secrets, isolation, stop conditions, and evidence — each checkable by observation before the first pass starts, each failing the launch when it fails, held honest by an admission rule and stable by a versioning policy.
Authority · YYLO CLI
ralph-loop · bounded-loops · evidence
2026-08-29 · 10 min read
The cross-asset contract behind every recurring research asset this site publishes: the register of eleven shipped assets in three families, the five-declaration methodology template each one states, the append-only changelog rules, the floor every citation must meet, and the ownership every refresh answers to.
Authority · Ecosystem
evidence · methodology
2026-08-29 · 11 min read
The operator flow for re-running governed grading over retained benchmark attempts: which artifacts a regrade may read and which inputs it refuses to accept, how the candidate-success fact is re-derived from retained evidence, how append-only ordering makes the newest grading generation current without erasing its predecessors, and the comparability rules that keep regraded cohorts honest.
Solution · YYLO Benchmark
benchmark · evaluation · grading
2026-08-29 · 7 min read
Observed repeated-attempt variance from the public record: 18 declared coding-agent identities listed two or more times against one fixed 500-instance task set moved by a median 5.2 and up to 39.8 percentage points between listings — two charts, a hash-pinned downloadable JSON summary, and the disclosures that bound what these numbers mean.
Authority · YYLO Benchmark
benchmark · evaluation · variance
2026-08-29 · 10 min read
A downloadable, versioned library of six repository-specific coding-agent benchmark cases mined from the public git histories of the released YYLO ecosystem repositories: each case pins a real base commit and the real later grading commit that carries its deterministic checks, with task, environment, and grading notes sufficient for reproduction and both fail and pass states verified.
Authority · YYLO Benchmark
benchmark · evaluation · case-library
2026-08-29 · 7 min read
Seed agent automation from five tested workflow templates shipped inside Workflow Runner — what each one is for, which seam it demonstrates, and the lint, dry-run, and committed tests behind them.
Solution · Ecosystem
workflow-runner · templates
2026-08-29 · 9 min read
A recurring, evidence-dated report on the coding-agent harness landscape: five scaffold-layer findings recomputed from a pinned 180-entry public leaderboard extract, two product-layer observations drawn from an evidence-dated matrix, and the methodology, changelog, and update ownership that keep it maintainable.
Awareness · Ecosystem
agent-harness · landscape-report
2026-08-29 · 10 min read
Auditing what an agent did and why takes four linked records — the task body that states intent, the response that accounts for the work, the validation it cites, and the commit that makes the diff checkable — and this page walks that chain end to end with recorded transcripts, exact-match lookups, and tamper drills.
Solution · YYLO Ledger
yylo-ledger · evidence
2026-08-29 · 10 min read
The operator's discipline for agent work that runs while you are elsewhere: admit each batch from dependency-ready tasks, make claims visible, requeue stale work, stop on conditions the queue itself can evaluate, and close every iteration with recorded evidence.
Solution · YYLO Ledger
yylo-ledger · autonomous
2026-08-29 · 10 min read
A citable definition of task memory for coding agents: the durable record a fresh process reads to recover work and writes to deposit progress — defined against chat history and vector stores, with the freshness and scope rules that keep it trustworthy.
Awareness · YYLO Ledger
yylo-ledger · task-truth
2026-08-29 · 6 min read
The five workflow YAML templates shipped inside YYLO's Workflow Runner as downloadable files: byte counts, ordered step ids, SHA-256 digests, the lint and dry-run record captured on 2026-08-29 against the committed runner template, and the seed command plus repository anchor behind every file.
Authority · Ecosystem
workflow-runner · templates
2026-08-29 · 8 min read
A shared TODO.md under multiple agents loses completed work, hides who claimed what, and cannot compute what may start — three structural failures, one reproduced deterministically, each mapped to the guarantee a Git-native task ledger provides instead.
Problem · YYLO Ledger
yylo-ledger · concurrency
2026-08-29 · 13 min read
A citable reference architecture for parallel coding agents isolated in Git worktree lanes: six permanent components (records owner, protected trunk, lane, candidate root, fenced arbiter, mirror), a one-writer ownership rule, the lane state machine from frozen-base allocation to reachability-gated cleanup, serialized merge admission by expected-value compare-and-swap with readback, and the configuration surface that binds it — every product behavior verified against committed YYLO source.
Authority · YYLO CLI
worktrees · parallel-agents · system-design
2026-08-28 · 11 min read
A precedence model for what coding agents are allowed to know: the four text layers (environment, instructions, skills, prompt), per-tool discovery and collision rules for Claude Code, Cursor, Kiro, OpenCode, and OpenClaw, and the safety limits that keep curated context cheap to trust.
Solution · YYLO CLI
agent-skills · context-precedence · yylo
2026-08-28 · 12 min read
Run a fair coding-agent benchmark on your own repository — select a case worth measuring, author it as a ledger task, freeze the attempt matrix and its spend ceiling, execute isolated attempts, and grade them with hash-pinned graders.
Solution · YYLO Benchmark
benchmark · evaluation
2026-08-28 · 12 min read
How to grade coding-agent output without bending the verdict: judges pinned to bytes, the two honest meanings of blind, boolean verdicts that fail closed into records, disagreement handled as governed generations instead of overwrites, and the hash-sealed receipt chain a third party can audit without trusting the operator.
Solution · YYLO Benchmark
benchmark · evaluation · grading
2026-08-28 · 10 min read
A reusable design method for long-running agent workflows: numeric and semantic stop conditions, machine-detectable staleness, trustworthy checkpoints, and recovery patterns for stuck or derailed runs.
Problem · YYLO CLI
failure-design · bounded-loops · yylo
2026-08-28 · 10 min read
What model leaderboards actually measure — preference arenas, fixed public suites, indexes, and usage rankings — why their numbers diverge from repository-level coding-agent benchmarks, and which decision each kind of evidence supports.
Awareness · YYLO Benchmark
benchmark · evaluation
2026-08-28 · 10 min read
Keep coding-agent state alive across interruptions: where session evidence persists, how continuation is scoped per shell, how sessions branch and clone, and the handoff manifests that let another operator pick up the thread.
Problem · YYLO CLI
sessions · continuity · yylo
2026-08-28 · 9 min read
Three names, three jobs stacked one above another: the assistant proposes while you drive, the agent executes a goal through tools, the orchestration CLI coordinates many agent runs. Each category defined from the vendors' own pages, with the failure model and selection consequences that follow.
Awareness · YYLO CLI
agent-categories · orchestration · yylo
2026-08-28 · 10 min read
Run a fair bake-off between several coding agents on one identical engineering task — one frozen plan, equal attempts and ceilings, observed identity, hash-pinned judging, and per-agent retained evidence — then read the standings without crowning a tie.
Solution · YYLO Benchmark
benchmark · evaluation · comparison
2026-08-28 · 9 min read
Isolate parallel coding agents with Git worktrees: the one-worktree-per-agent topology, ownership rules that prevent conflicts, recovery when boundaries fail, and how YYLO maps every task to its own worktree.
Solution · YYLO CLI
worktrees · parallel-agents · yylo
2026-08-28 · 11 min read
Four minimal upgrades that leave a Ralph loop recognizably a loop while bounding what each pass may spend, where its memory lives, what stops it, and what it leaves behind: acceptance tests inside the pass, a task board the worker cannot rewrite, stop conditions the loop evaluates itself, and per-pass evidence checkpoints.
Solution · YYLO CLI
ralph-loop · task-truth · yylo
2026-08-28 · 11 min read
The evidence vocabulary honest agent evaluations are built on: four cost-completeness states, a closed terminal-class set for failures, and the null-and-typed-label reporting rules that keep unknown apart from zero, absent apart from failed, and not-applicable apart from both.
Problem · YYLO Benchmark
benchmark · evaluation · evidence
2026-08-28 · 11 min read
How to design a coding-agent evaluation plan that cannot drift: what belongs inside a content-addressed plan object, how the spend ceiling divides itself deterministically across models and attempts, how model aliases resolve to exact identities before the hash, and the forward-only change policy that turns every legitimate edit into a new plan instead of a silent mutation.
Solution · YYLO Benchmark
benchmark · evaluation · planning
2026-08-28 · 11 min read
The migration path from a running prompt loop to an auditable YYLO workflow — inventory the loop's parts, apply three renames (the prompt becomes typed steps, the session becomes receipt-bound evidence, the iteration becomes a named attempt), give the plan file a durable destination, and stage every move so it can roll back.
Solution · YYLO CLI
ralph-loop · workflow-runner · yylo
2026-08-28 · 10 min read
Run the same Ralph-style loop with Claude, Codex, Cursor, and Pi: the four-role loop contract that stays harness-agnostic, what each harness needs before its first pass, and the evidence a mixed run records automatically.
Solution · YYLO CLI
ralph-loop · multi-harness · yylo
2026-08-28 · 10 min read
Run Claude Code, Codex, and Gemini as one canonical workflow: per-service dispatch surfaces and shorthand families, harness-agnostic steps with typed handoffs, and provider failure boundaries for quota, credentials, and evidence.
Solution · YYLO CLI
multi-harness · claude · codex · gemini · yylo
2026-08-28 · 9 min read
A layer-by-layer map for assembling a coding-agent stack from open source — open weights and their two licenses, readable agent loops, bundled harnesses, and the three coordination layers open source keeps forgetting, with live-checked options per layer.
Awareness · Ecosystem
open-source · agent-stack · yylo
2026-08-28 · 9 min read
Carry one batch of independent agent tasks from standing start to closed evidence, stage by stage, with readiness checks, bounded fan-out, worker isolation, and aggregation proof at every step.
Solution · YYLO CLI
parallel-agents · fan-out · yylo
2026-08-28 · 11 min read
A topology decision for Ralph-shaped work, answered with four explicit safety criteria instead of throughput arithmetic: isolation, quota, evidence, and recovery each compared across one board-driven sequential loop and a capped pool of bounded parallel lanes, ending in a four-question pass that serializes any work a criterion cannot vouch for.
Comparison · YYLO CLI
ralph-loop · parallel-agents · yylo
2026-08-28 · 10 min read
The definitive technical reference to the Ralph loop: Geoffrey Huntley's July 2025 origin and the Ralph Wiggum naming, what the bash while-loop does on every pass, the variant family from Claude Code and Amp to Anthropic's plugin, and the hard limits practitioners report.
Awareness · YYLO CLI
ralph-loop · autonomous-agents · yylo
2026-08-28 · 9 min read
A side-by-side run anatomy of one evening of coding-agent work executed as a raw Ralph loop and as a bounded evidence-producing workflow: what each system produces, what each loses, what bounds honestly cost, and the observed triggers that say the trade has flipped.
Comparison · YYLO CLI
ralph-loop · bounded-loops · yylo
2026-08-28 · 12 min read
What to do when a coding-agent benchmark dies midway: the three statuses an interrupted attempt can hold, the resume-or-discard decision tree that never re-bills an indeterminate attempt, how workflow steps reconcile through a private dispatch journal before any resume, and the standing audits that prove the salvaged run still adds up.
Problem · YYLO Benchmark
benchmark · evaluation · recovery
2026-08-28 · 10 min read
Why the same coding agent passes a task once and fails the next run — and how repeated attempts turn that flip into a measured rate: attempt-count arithmetic, Wilson intervals at small samples, typed failure separation, repeat consistency as a majority share, and the selection rules variance-aware evaluators use.
Problem · YYLO Benchmark
benchmark · evaluation · statistics
2026-08-28 · 12 min read
Keep OpenCode's configured agents and add what configuration cannot declare — readiness between tasks, quotas around fan-out, and durable evidence per unit of work — with the agent-to-task mapping plus a headless service wrapper wired through the documented extension point today.
Solution · YYLO CLI
opencode · interoperability · yylo
2026-08-28 · 10 min read
Run Pi as YYLO's documented first-party agent service — tree-structured sessions and Pi's deliberate no-subagent stance mapped onto bounded dispatch, one Pi per task worktree, execution envelopes, and merge admission.
Solution · YYLO CLI
pi · interoperability · yylo
2026-08-28 · 9 min read
Three mechanisms break the transfer from a public coding leaderboard to outcomes inside one specific repository — task distribution, harness effects, and sampling variance — each grounded in what the boards publish about themselves, ending in a three-question rule for when a standing may still guide a decision.
Problem · YYLO Benchmark
benchmark · evaluation · statistics
2026-08-28 · 14 min read
A diagnostic catalog of the ways an unbounded Ralph loop fails — search-blind duplication, stale repetition, drift, context exhaustion, cost blowout, fragile state, the broken morning, and claimed completion — each identified by the observable symptoms it leaves on four surfaces, with the mechanism that produces it, the check that separates it from its neighbors, and the direction its remedy lives in.
Problem · YYLO CLI
ralph-loop · failure-modes · yylo
2026-08-27 · 8 min read
A practical portability audit for teams adopting coding agents: find where lock-in actually concentrates across model, harness, and task truth, then verify four properties — open formats, Git-native task truth, exportable evidence, and switchable harnesses.
Problem · YYLO CLI
vendor-lock-in · portability · yylo
2026-08-27 · 9 min read
A layered blueprint for harness engineering: five boundaries — orchestration, context, evidence, lifecycle, and failure — each with a named contract, using YYLO's runner contracts as the worked example.
Solution · YYLO CLI
harness-engineering · architecture · yylo
2026-08-27 · 8 min read
Move in-flight agent work from one harness to another without losing anything: an evidence-dated switch surface, a four-step procedure that preserves task truth and receipts, and the session commands that keep both sides addressable.
Solution · YYLO CLI
harness-switching · sessions · yylo
2026-08-27 · 9 min read
Keep Cursor as your agent surface and add a control plane above it: the IDE-agent-plus-control-plane pattern, the boundary rules, the Cursor service wiring, and worktree isolation for parallel Cursor agents.
Solution · YYLO CLI
cursor · interoperability · yylo
2026-08-27 · 12 min read
Keep Kiro's spec-driven workflow and add the control plane it lacks: the spec-to-task mapping, dependencies across specs, merge admission, and a headless-CLI service you can wire through the documented extension point today.
Solution · YYLO CLI
kiro · interoperability · yylo
2026-08-27 · 8 min read
Run several coding agents as one validated workflow: put each agent behind a harness-agnostic step contract, hand off typed responses between steps, and land the result through a single admission gate.
Solution · YYLO CLI
multi-agent · workflow-runner · yylo
2026-08-27 · 7 min read
A coding-agent harness is the environment an agent runs inside — prompt, tools, context, session. The exact boundary against the agent, the model, the IDE, and the control plane, with YYLO as the worked example.
Awareness · YYLO CLI
agent-harness · yylo
2026-07-18 · 6 min read
One run directory holds the manifest, per-step responses, session handoff, doctor diagnostics, and fail-closed recovery that survives interruption.
Problem · YYLO CLI
workflow-runner · session-handoff
2026-07-18 · 5 min read
Pick Workflow Runner or Parallel Runner in one keyed pass over step coupling, evidence needs, and failure semantics, with concrete commands for each choice.
Solution · YYLO CLI
workflow-runner · parallel-runner
2026-07-18 · 7 min read
A four-surface injection taxonomy for coding-agent prompts, plus the deterministic bounds YYLO puts on shell entry, command substitution, fan-out context, and secrets.
Problem · YYLO CLI
prompt-safety · yylo
2026-07-18 · 5 min read
Admit only dependency-ready tasks, cap concurrency below provider limits, isolate each worker, and review per-run evidence before closing tasks.
Solution · YYLO CLI
parallel-runner · yylo-ledger
2026-07-18 · 4 min read
Run your first coding-agent loop with one task, one iteration cap, reviewable per-cycle evidence, and explicit stop conditions.
Solution · YYLO CLI
yylo · bounded-loops
2026-07-18 · 7 min read
Give agents durable task memory instead of hand-moved board cards with dependency-aware readiness, required responses, commit evidence, and hash-chained history in Git.
Awareness · YYLO Ledger
yylo-ledger · task-truth