Published field notes

Practical guides

Runnable recommendations for bounded loops, safe fan-out, durable handoff, prompt boundaries, and task truth.

2026-08-29 · 12 min read

Reusable agent-task dependency examples and diagrams

A copyable library of five agent-task dependency shapes — chain, fan-out, fan-in, diamond, and a cross-project hand-off held open by a forward reference — each drawn as a diagram, declared with the exact yylo-ledger commands, and proven by a recorded 2026-08-29 run: what ready admits at each stage, what the priority scores rank, and how wide each shape may safely run.

Authority · YYLO Ledger

yylo-ledger · dependencies

2026-08-29 · 9 min read

Executable agent-task dependency graph design

Model task dependencies an agent runtime can execute: one edge direction resolved only by terminal status, fan-in as an all-of join, admission as a computed query, cycles refused at the write, and scheduling rules drawn from readiness — including the sharp edge where the shipped topological sort stops answering mid-wave.

Solution · YYLO Ledger

yylo-ledger · dependencies

2026-08-29 · 10 min read

Shell and JSON pipelines for agent task automation

Drive YYLO Ledger from scripts and CI with contracts that hold: honest exit codes, parseable stdout in NDJSON and JSON, projection and cursor discipline, stdin-safe task input, and a receipt for every mutation.

Solution · YYLO Ledger

yylo-ledger · automation · cli

2026-08-29 · 9 min read

An auditable AI software-development system architecture

A reference architecture for AI development systems whose every change is auditable by construction — audit points that write evidence at each transition, admission boundaries that gate mutation, and rollback seams that retreat without destroying records — with YYLO as the worked implementation.

Authority · Ecosystem

audit-architecture · system-design · yylo

2026-08-29 · 8 min read

Choose a coding agent by repository and change type

Select a coding agent with a decision matrix scored on repository size, risk, and change surface — one representative benchmark case per cell, a written reliability bar, and a default-plus-escalation portfolio with explicit economics — instead of brand loyalty or a borrowed leaderboard standing.

Solution · Ecosystem

benchmark · evaluation · selection

2026-08-29 · 10 min read

The coding-agent failure and recovery taxonomy, version 1

A versioned, citable taxonomy of coding-agent failure: nine stable classes with permanent identifiers, each carrying the symptoms that announce it, the detection that confirms it from evidence rather than testimony, and the recovery its evidence state permits — with attribution across the stack layers and a stability policy that keeps citations from drifting.

Authority · Ecosystem

failure-modes · recovery · taxonomy

2026-08-29 · 8 min read

The coding-agent harness taxonomy, version 1

A versioned, citable taxonomy for the coding-agent stack: six terms with stable identifiers, each with one definition, an owns-and-excludes split, and boundary criteria that settle real disputes — plus the stability policy and changelog that keep a citation from drifting.

Authority · Ecosystem

agent-harness · taxonomy

2026-08-29 · 8 min read

What the control plane, Ledger, and Benchmark each own

The citable ownership definition of the YYLO system — which component owns execution and admission, which owns task truth, which owns evaluation evidence — with the boundaries, failure domains, and interfaces between them.

Awareness · Ecosystem

yylo · yylo-ledger · benchmark · system-ownership

2026-08-29 · 12 min read

Safe explicit cross-project task routing

The explicit routing contract for agent task boards that span repositories: an opt-in per-project allowlist, one validated user registry, exact process replacement into the destination wrapper, and refusals that never fall back to the source board.

Solution · YYLO Ledger

yylo-ledger · routing · multi-repo

2026-08-29 · 11 min read

Select dependency-ready work for parallel agents

Feeding a pool of agent workers is a selection problem sitting on top of a readiness answer: scope the ready set so claims stay out, spend scarce slots where they unlock downstream work, and close the window where two dispatchers hand the same task to two agents.

Solution · YYLO Ledger

yylo-ledger · parallel

2026-08-29 · 11 min read

Evidence architecture for agentic engineering

The artifact taxonomy that makes agent work provable after the fact — receipts that bind claims to bytes, manifests that vouch for collections, session records that outlive the terminal, and checkpoints that make interruption legible — together with the integrity, retention, and regrade semantics that keep each role trustworthy, quoted from YYLO's committed contracts.

Authority · Ecosystem

evidence · system-design · yylo

2026-08-29 · 8 min read

From coding prompt to validated change

Follow one change through YYLO's governed lifecycle — task record, isolated worktree, focused validation, read-only preflight, guarded finish, and the serialized merge queue — with the documented command surface at every station.

Awareness · Ecosystem

task-lifecycle · validated-change · yylo

2026-08-29 · 9 min read

Git-native task management for coding agents

The storage contract for agent work that lives in your repository: one safe Markdown file per task with a validated schema, a five-status lifecycle with structural refusals, and completion bound to a commit hash instead of a checkbox.

Awareness · YYLO Ledger

yylo-ledger · git-native

2026-08-29 · 11 min read

The YYLO harness support and interoperability matrix, version 2

A versioned support ledger for YYLO's agent-harness dispatch surface: one dated cell per harness, each with the command that re-verifies it against your own installed release, plus the update policy that keeps a citation honest.

Authority · Ecosystem

agent-harness · support-matrix

2026-08-29 · 10 min read

Immutable archives for large agent-task histories

Move finished agent tasks into sealed, verifiable cold packs without losing a byte of evidence: revision-bound plans, content-addressed packs, read-only resolution, tamper evidence, and a maintenance rhythm that keeps hot state fast.

Solution · YYLO Ledger

yylo-ledger · archives · task-history

2026-08-29 · 10 min read

A bounded longitudinal coding-agent benchmark dataset

A downloadable, versioned extract of the public SWE-bench Verified leaderboard: 180 dated coding-agent system entries on one fixed 500-instance task set from 2023-10-10 to 2026-02-26, with extraction methodology, explicit limitations, and a content hash for citation.

Authority · YYLO Benchmark

benchmark · evaluation · dataset

2026-08-29 · 9 min read

Multi-agent coordination without ambiguous shared state

Run several coding agents on one repository without ambiguous shared task state: an ownership write matrix, per-agent admission and filesystem scopes, and a sealed merge-admission gate that turns every landing race into a readable refusal.

Solution · YYLO Ledger

yylo-ledger · multi-agent · concurrency

2026-08-29 · 8 min read

Portable coding-agent workflows without lock-in

Keep coding-agent workflows portable across harnesses and providers: a workflow format you own, declared substitution points that carry identity and paths at run time, and provider failures isolated inside the step that dispatched them.

Solution · Ecosystem

workflow-runner · portability · providers

2026-08-29 · 11 min read

The Ralph-loop safety checklist, version 1

A copyable, citable pre-launch gate for Ralph-style autonomous loops: fifteen permanent check items across five gates — bounds, secrets, isolation, stop conditions, and evidence — each checkable by observation before the first pass starts, each failing the launch when it fails, held honest by an admission rule and stable by a versioning policy.

Authority · YYLO CLI

ralph-loop · bounded-loops · evidence

2026-08-29 · 10 min read

Methodology, changelog, and citation guidance for recurring research

The cross-asset contract behind every recurring research asset this site publishes: the register of eleven shipped assets in three families, the five-declaration methodology template each one states, the append-only changelog rules, the floor every citation must meet, and the ownership every refresh answers to.

Authority · Ecosystem

evidence · methodology

2026-08-29 · 11 min read

Regrade benchmark attempts under new criteria, without rerunning candidates

The operator flow for re-running governed grading over retained benchmark attempts: which artifacts a regrade may read and which inputs it refuses to accept, how the candidate-success fact is re-derived from retained evidence, how append-only ordering makes the newest grading generation current without erasing its predecessors, and the comparability rules that keep regraded cohorts honest.

Solution · YYLO Benchmark

benchmark · evaluation · grading

2026-08-29 · 7 min read

Repeated-attempt variance, observed and downloadable

Observed repeated-attempt variance from the public record: 18 declared coding-agent identities listed two or more times against one fixed 500-instance task set moved by a median 5.2 and up to 39.8 percentage points between listings — two charts, a hash-pinned downloadable JSON summary, and the disclosures that bound what these numbers mean.

Authority · YYLO Benchmark

benchmark · evaluation · variance

2026-08-29 · 10 min read

Repository-specific benchmark cases and reproducibility notes

A downloadable, versioned library of six repository-specific coding-agent benchmark cases mined from the public git histories of the released YYLO ecosystem repositories: each case pins a real base commit and the real later grading commit that carries its deterministic checks, with task, environment, and grading notes sufficient for reproduction and both fail and pass states verified.

Authority · YYLO Benchmark

benchmark · evaluation · case-library

2026-08-29 · 7 min read

Reusable, tested workflow templates for coding agents

Seed agent automation from five tested workflow templates shipped inside Workflow Runner — what each one is for, which seam it demonstrates, and the lint, dry-run, and committed tests behind them.

Solution · Ecosystem

workflow-runner · templates

2026-08-29 · 9 min read

The state of coding-agent harnesses, edition 1

A recurring, evidence-dated report on the coding-agent harness landscape: five scaffold-layer findings recomputed from a pinned 180-entry public leaderboard extract, two product-layer observations drawn from an evidence-dated matrix, and the methodology, changelog, and update ownership that keep it maintainable.

Awareness · Ecosystem

agent-harness · landscape-report

2026-08-29 · 10 min read

Task intent, responses, tests, and commits as durable evidence

Auditing what an agent did and why takes four linked records — the task body that states intent, the response that accounts for the work, the validation it cites, and the commit that makes the diff checkable — and this page walks that chain end to end with recorded transcripts, exact-match lookups, and tamper drills.

Solution · YYLO Ledger

yylo-ledger · evidence

2026-08-29 · 10 min read

Task management for autonomous coding-agent workflows

The operator's discipline for agent work that runs while you are elsewhere: admit each batch from dependency-ready tasks, make claims visible, requeue stale work, stop on conditions the queue itself can evaluate, and close every iteration with recorded evidence.

Solution · YYLO Ledger

yylo-ledger · autonomous

2026-08-29 · 10 min read

Task state as durable agent memory

A citable definition of task memory for coding agents: the durable record a fresh process reads to recover work and writes to deposit progress — defined against chat history and vector stores, with the freshness and scope rules that keep it trustworthy.

Awareness · YYLO Ledger

yylo-ledger · task-truth

2026-08-29 · 6 min read

Downloadable validated workflow YAML examples

The five workflow YAML templates shipped inside YYLO's Workflow Runner as downloadable files: byte counts, ordered step ids, SHA-256 digests, the lint and dry-run record captured on 2026-08-29 against the committed runner template, and the seed command plus repository anchor behind every file.

Authority · Ecosystem

workflow-runner · templates

2026-08-29 · 8 min read

Why shared Markdown TODO files fail under concurrent agents

A shared TODO.md under multiple agents loses completed work, hides who claimed what, and cannot compute what may start — three structural failures, one reproduced deterministically, each mapped to the guarantee a Git-native task ledger provides instead.

Problem · YYLO Ledger

yylo-ledger · concurrency

2026-08-29 · 13 min read

A Git-worktree parallel-agent reference architecture, version 1

A citable reference architecture for parallel coding agents isolated in Git worktree lanes: six permanent components (records owner, protected trunk, lane, candidate root, fenced arbiter, mirror), a one-writer ownership rule, the lane state machine from frozen-base allocation to reachability-gated cleanup, serialized merge admission by expected-value compare-and-swap with readback, and the configuration surface that binds it — every product behavior verified against committed YYLO source.

Authority · YYLO CLI

worktrees · parallel-agents · system-design

2026-08-28 · 11 min read

Agent skills, project instructions, and context precedence

A precedence model for what coding agents are allowed to know: the four text layers (environment, instructions, skills, prompt), per-tool discovery and collision rules for Claude Code, Cursor, Kiro, OpenCode, and OpenClaw, and the safety limits that keep curated context cheap to trust.

Solution · YYLO CLI

agent-skills · context-precedence · yylo

2026-08-28 · 12 min read

Benchmark coding agents on a real repository

Run a fair coding-agent benchmark on your own repository — select a case worth measuring, author it as a ledger task, freeze the attempt matrix and its spend ceiling, execute isolated attempts, and grade them with hash-pinned graders.

Solution · YYLO Benchmark

benchmark · evaluation

2026-08-28 · 12 min read

Blinded and governed grading for coding-agent output

How to grade coding-agent output without bending the verdict: judges pinned to bytes, the two honest meanings of blind, boolean verdicts that fail closed into records, disagreement handled as governed generations instead of overwrites, and the hash-sealed receipt chain a third party can audit without trusting the operator.

Solution · YYLO Benchmark

benchmark · evaluation · grading

2026-08-28 · 10 min read

Bounded failure design for long-running agent workflows

A reusable design method for long-running agent workflows: numeric and semantic stop conditions, machine-detectable staleness, trustworthy checkpoints, and recovery patterns for stuck or derailed runs.

Problem · YYLO CLI

failure-design · bounded-loops · yylo

2026-08-28 · 10 min read

Coding-agent benchmarks vs model leaderboards

What model leaderboards actually measure — preference arenas, fixed public suites, indexes, and usage rankings — why their numbers diverge from repository-level coding-agent benchmarks, and which decision each kind of evidence supports.

Awareness · YYLO Benchmark

benchmark · evaluation

2026-08-28 · 10 min read

Durable coding-agent sessions and handoffs

Keep coding-agent state alive across interruptions: where session evidence persists, how continuation is scoped per shell, how sessions branch and clone, and the handoff manifests that let another operator pick up the thread.

Problem · YYLO CLI

sessions · continuity · yylo

2026-08-28 · 9 min read

Coding agents vs coding assistants vs orchestration CLIs

Three names, three jobs stacked one above another: the assistant proposes while you drive, the agent executes a goal through tools, the orchestration CLI coordinates many agent runs. Each category defined from the vendors' own pages, with the failure model and selection consequences that follow.

Awareness · YYLO CLI

agent-categories · orchestration · yylo

2026-08-28 · 10 min read

Compare multiple coding agents on one task

Run a fair bake-off between several coding agents on one identical engineering task — one frozen plan, equal attempts and ceilings, observed identity, hash-pinned judging, and per-agent retained evidence — then read the standings without crowning a tie.

Solution · YYLO Benchmark

benchmark · evaluation · comparison

2026-08-28 · 9 min read

Git worktrees for parallel coding agents

Isolate parallel coding agents with Git worktrees: the one-worktree-per-agent topology, ownership rules that prevent conflicts, recovery when boundaries fail, and how YYLO maps every task to its own worktree.

Solution · YYLO CLI

worktrees · parallel-agents · yylo

2026-08-28 · 11 min read

Harden a Ralph loop with tests, task truth, and checkpoints

Four minimal upgrades that leave a Ralph loop recognizably a loop while bounding what each pass may spend, where its memory lives, what stops it, and what it leaves behind: acceptance tests inside the pass, a task board the worker cannot rewrite, stop conditions the loop evaluates itself, and per-pass evidence checkpoints.

Solution · YYLO CLI

ralph-loop · task-truth · yylo

2026-08-28 · 11 min read

Honest cost, failure, and zero evidence semantics

The evidence vocabulary honest agent evaluations are built on: four cost-completeness states, a closed terminal-class set for failures, and the null-and-typed-label reporting rules that keep unknown apart from zero, absent apart from failed, and not-applicable apart from both.

Problem · YYLO Benchmark

benchmark · evaluation · evidence

2026-08-28 · 11 min read

Build an immutable coding-agent evaluation plan

How to design a coding-agent evaluation plan that cannot drift: what belongs inside a content-addressed plan object, how the spend ceiling divides itself deterministically across models and attempts, how model aliases resolve to exact identities before the hash, and the forward-only change policy that turns every legitimate edit into a new plan instead of a silent mutation.

Solution · YYLO Benchmark

benchmark · evaluation · planning

2026-08-28 · 11 min read

Migrate a Ralph loop to an auditable YYLO workflow

The migration path from a running prompt loop to an auditable YYLO workflow — inventory the loop's parts, apply three renames (the prompt becomes typed steps, the session becomes receipt-bound evidence, the iteration becomes a named attempt), give the plan file a durable destination, and stage every move so it can roll back.

Solution · YYLO CLI

ralph-loop · workflow-runner · yylo

2026-08-28 · 10 min read

One canonical multi-harness Ralph workflow

Run the same Ralph-style loop with Claude, Codex, Cursor, and Pi: the four-role loop contract that stays harness-agnostic, what each harness needs before its first pass, and the evidence a mixed run records automatically.

Solution · YYLO CLI

ralph-loop · multi-harness · yylo

2026-08-28 · 10 min read

One multi-harness workflow for Claude Code, Codex, and Gemini

Run Claude Code, Codex, and Gemini as one canonical workflow: per-service dispatch surfaces and shorthand families, harness-agnostic steps with typed handoffs, and provider failure boundaries for quota, credentials, and evidence.

Solution · YYLO CLI

multi-harness · claude · codex · gemini · yylo

2026-08-28 · 9 min read

The open-source AI coding stack, layer by layer

A layer-by-layer map for assembling a coding-agent stack from open source — open weights and their two licenses, readable agent loops, bundled harnesses, and the three coordination layers open source keeps forgetting, with live-checked options per layer.

Awareness · Ecosystem

open-source · agent-stack · yylo

2026-08-28 · 9 min read

A complete parallel coding-agent workflow

Carry one batch of independent agent tasks from standing start to closed evidence, stage by stage, with readiness checks, bounded fan-out, worker isolation, and aggregation proof at every step.

Solution · YYLO CLI

parallel-agents · fan-out · yylo

2026-08-28 · 11 min read

Parallel vs sequential Ralph loops

A topology decision for Ralph-shaped work, answered with four explicit safety criteria instead of throughput arithmetic: isolation, quota, evidence, and recovery each compared across one board-driven sequential loop and a capped pool of bounded parallel lanes, ending in a four-question pass that serializes any work a criterion cannot vouch for.

Comparison · YYLO CLI

ralph-loop · parallel-agents · yylo

2026-08-28 · 10 min read

The Ralph loop: origin, mechanics, variants, and limits

The definitive technical reference to the Ralph loop: Geoffrey Huntley's July 2025 origin and the Ralph Wiggum naming, what the bash while-loop does on every pass, the variant family from Claude Code and Amp to Anthropic's plugin, and the hard limits practitioners report.

Awareness · YYLO CLI

ralph-loop · autonomous-agents · yylo

2026-08-28 · 9 min read

Ralph loop vs a bounded evidence-producing workflow

A side-by-side run anatomy of one evening of coding-agent work executed as a raw Ralph loop and as a bounded evidence-producing workflow: what each system produces, what each loses, what bounds honestly cost, and the observed triggers that say the trade has flipped.

Comparison · YYLO CLI

ralph-loop · bounded-loops · yylo

2026-08-28 · 12 min read

Recover incomplete benchmark runs without losing evidence

What to do when a coding-agent benchmark dies midway: the three statuses an interrupted attempt can hold, the resume-or-discard decision tree that never re-bills an indeterminate attempt, how workflow steps reconcile through a private dispatch journal before any resume, and the standing audits that prove the salvaged run still adds up.

Problem · YYLO Benchmark

benchmark · evaluation · recovery

2026-08-28 · 10 min read

How repeated attempts reveal coding-agent variance

Why the same coding agent passes a task once and fails the next run — and how repeated attempts turn that flip into a measured rate: attempt-count arithmetic, Wilson intervals at small samples, typed failure separation, repeat consistency as a majority share, and the selection rules variance-aware evaluators use.

Problem · YYLO Benchmark

benchmark · evaluation · statistics

2026-08-28 · 12 min read

Use OpenCode with YYLO

Keep OpenCode's configured agents and add what configuration cannot declare — readiness between tasks, quotas around fan-out, and durable evidence per unit of work — with the agent-to-task mapping plus a headless service wrapper wired through the documented extension point today.

Solution · YYLO CLI

opencode · interoperability · yylo

2026-08-28 · 10 min read

Use the Pi coding agent with YYLO

Run Pi as YYLO's documented first-party agent service — tree-structured sessions and Pi's deliberate no-subagent stance mapped onto bounded dispatch, one Pi per task worktree, execution envelopes, and merge admission.

Solution · YYLO CLI

pi · interoperability · yylo

2026-08-28 · 9 min read

Why public leaderboards may not predict repository outcomes

Three mechanisms break the transfer from a public coding leaderboard to outcomes inside one specific repository — task distribution, harness effects, and sampling variance — each grounded in what the boards publish about themselves, ending in a three-question rule for when a standing may still guide a decision.

Problem · YYLO Benchmark

benchmark · evaluation · statistics

2026-08-28 · 14 min read

Why unbounded Ralph loops fail

A diagnostic catalog of the ways an unbounded Ralph loop fails — search-blind duplication, stale repetition, drift, context exhaustion, cost blowout, fragile state, the broken morning, and claimed completion — each identified by the observable symptoms it leaves on four surfaces, with the mechanism that produces it, the check that separates it from its neighbors, and the direction its remedy lives in.

Problem · YYLO CLI

ralph-loop · failure-modes · yylo

2026-08-27 · 8 min read

Avoid coding-agent vendor lock-in

A practical portability audit for teams adopting coding agents: find where lock-in actually concentrates across model, harness, and task truth, then verify four properties — open formats, Git-native task truth, exportable evidence, and switchable harnesses.

Problem · YYLO CLI

vendor-lock-in · portability · yylo

2026-08-27 · 9 min read

Harness engineering architecture

A layered blueprint for harness engineering: five boundaries — orchestration, context, evidence, lifecycle, and failure — each with a named contract, using YYLO's runner contracts as the worked example.

Solution · YYLO CLI

harness-engineering · architecture · yylo

2026-08-27 · 8 min read

Switch coding-agent harnesses in YYLO

Move in-flight agent work from one harness to another without losing anything: an evidence-dated switch surface, a four-step procedure that preserves task truth and receipts, and the session commands that keep both sides addressable.

Solution · YYLO CLI

harness-switching · sessions · yylo

2026-08-27 · 9 min read

Use Cursor with YYLO

Keep Cursor as your agent surface and add a control plane above it: the IDE-agent-plus-control-plane pattern, the boundary rules, the Cursor service wiring, and worktree isolation for parallel Cursor agents.

Solution · YYLO CLI

cursor · interoperability · yylo

2026-08-27 · 12 min read

Use Kiro with YYLO

Keep Kiro's spec-driven workflow and add the control plane it lacks: the spec-to-task mapping, dependencies across specs, merge admission, and a headless-CLI service you can wire through the documented extension point today.

Solution · YYLO CLI

kiro · interoperability · yylo

2026-08-27 · 8 min read

Use multiple coding agents in one workflow

Run several coding agents as one validated workflow: put each agent behind a harness-agnostic step contract, hand off typed responses between steps, and land the result through a single admission gate.

Solution · YYLO CLI

multi-agent · workflow-runner · yylo

2026-08-27 · 7 min read

What is an AI coding-agent harness?

A coding-agent harness is the environment an agent runs inside — prompt, tools, context, session. The exact boundary against the agent, the model, the IDE, and the control plane, with YYLO as the worked example.

Awareness · YYLO CLI

agent-harness · yylo

2026-07-18 · 6 min read

Build auditable agent workflows with handoff

One run directory holds the manifest, per-step responses, session handoff, doctor diagnostics, and fail-closed recovery that survives interruption.

Problem · YYLO CLI

workflow-runner · session-handoff

2026-07-18 · 5 min read

Choose Workflow Runner or Parallel Runner

Pick Workflow Runner or Parallel Runner in one keyed pass over step coupling, evidence needs, and failure semantics, with concrete commands for each choice.

Solution · YYLO CLI

workflow-runner · parallel-runner

2026-07-18 · 7 min read

Keep prompts, shell substitutions, and task context safe

A four-surface injection taxonomy for coding-agent prompts, plus the deterministic bounds YYLO puts on shell entry, command substitution, fan-out context, and secrets.

Problem · YYLO CLI

prompt-safety · yylo

2026-07-18 · 5 min read

Run kanban tasks safely in parallel

Admit only dependency-ready tasks, cap concurrency below provider limits, isolate each worker, and review per-run evidence before closing tasks.

Solution · YYLO CLI

parallel-runner · yylo-ledger

2026-07-18 · 4 min read

Start a bounded AI-agent development loop

Run your first coding-agent loop with one task, one iteration cap, reviewable per-cycle evidence, and explicit stop conditions.

Solution · YYLO CLI

yylo · bounded-loops

2026-07-18 · 7 min read

Use YYLO Ledger as the source of truth for agent work

Give agents durable task memory instead of hand-moved board cards with dependency-aware readiness, required responses, commit evidence, and hash-chained history in Git.

Awareness · YYLO Ledger

yylo-ledger · task-truth