2026-07-18 · Updated 2026-09-10 · 7 min read
Keep prompts, shell substitutions, and task context safe
A four-surface injection taxonomy for coding-agent prompts, plus the deterministic bounds YYLO puts on shell entry, command substitution, fan-out context, and secrets.
By Juno AI INC · prompt-safety · yylo
An agent never receives only what you typed. Before the backend dispatches a request, the prompt has crossed a shell that can expand it, a template layer that may rewrite it, task and integration text that can steer it, and an environment that carries real credentials. Prompt safety is not one filter you add — it is knowing which surface each piece of text arrived on and giving every crossing a bounded, auditable contract.
YYLO makes those crossings deterministic: prompt-time substitutions are timeout-bounded and run with stdin closed, the task template and the resolved payload are both printed before dispatch, oversized prompts move to managed files instead of argv, and secrets stay in scoped environment variables rather than task prose. The taxonomy below works in any harness; the commands and contracts are YYLO's.
Name the four injection surfaces
Every input reaches the agent through one of four surfaces, and each fails differently:
- Prompt. Instruction-bearing text inside the request itself: a kanban task body, an issue or chat message an integration fetched, or the output of a substitution command concatenated into your prompt. Injected text here speaks with the authority of your own instruction.
- Memory. Text that persists and is trusted later: task responses recorded into the ledger, session transcripts that
yy cccontinues in this shell, instruction files that accumulate agent learnings. A payload that survives one run can steer the next one. - Context. Ambient material the agent treats as standing policy: project instruction files, explicitly installed skills in
.claude/skills/,.agents/skills/, and.pi/skills/, and environment variables. Whoever can write those files steers every future run — which is why skills must be reviewed like code and the shipped hook examples warn when an instruction file such asCLAUDE.mdgrows past 40 KB of accreted text. - Shell. The invoking shell executing or mangling the prompt before the CLI receives it — backticks and
$(...)under double quotes — and prompt-time substitution commands running with your credentials on every iteration.
The defense is not to trust harder. It is to declare, for each input, which surface it crosses and what bounds the crossing. The rest of this guide is that contract applied to the shell, the substitution layer, the dispatch view, fan-out, and secrets.
Keep the invoking shell out of the prompt
The first crossing happens before YYLO sees anything. Under double quotes, your own shell executes backticks and $(...) and substitutes the result into the prompt — the command already ran with your credentials, and the agent receives only its output. Deliver literal text through one of the three safe patterns instead:
Single quotes, a prompt file, and a heredoc with a quoted delimiter all hand YYLO the exact bytes you wrote. The same rule governs rich task Markdown: fences, backticks, and $VARIABLES belong in files, not in inline arguments:
The Kanban wrapper documents --body-file and --response-file (both accept - for stdin) for exactly this class of text.
Size is a safety property too. yylo protects shell-backend runs from OS E2BIG spawn failures by switching prompts larger than JUNO_PROMPT_ARG_MAX_BYTES (default 65,536 bytes) away from argv transport into managed prompt files under /tmp/yylo/, cleaned up where safe — you never need to hand-roll temp files for long prompts. In Pi live mode an oversized prompt keeps a bounded beginning in argv and writes the continuation to a managed file the agent is told to read next.
Bound every prompt-time substitution
YYLO's explicit substitution markers run a command in the working directory and splice its stdout into the prompt — freshly, on every engine iteration:
The marker forms are !'command' and a triple-backtick block. Unlike shell backtick expansion, which ran once before the CLI started, these resolve immediately before each subagent call, so retries see fresh output instead of a stale one-time capture. The execution contract is deliberately tight:
- Timeout-bounded. A substitution is killed after 30 seconds by default;
YYLO_PROMPT_SUBSTITUTION_TIMEOUT_MSchanges the ceiling. A wedged command fails the iteration with a distinct timeout error instead of hanging a headless worker forever. - Stdin closed. Each command runs as
(command) </dev/null, so nothing in a prompt can sit waiting for terminal input that will never arrive. - Bounded capture. Only stdout is spliced in, under a 1 MiB buffer; a failure names the failed command and carries its stderr detail.
Treat the commands themselves as reviewed code. They execute through your login shell with your ordinary environment — credentials included — so !'env' or !'cat ~/.netrc' inside a prompt is exfiltration with extra steps. Keep only commands whose output you would paste into the prompt by hand.
@@key prompt macros share the crossing and the review. Dictionaries live in .juno_task/config.json, they resolve before substitutions by default or after them per promptMacros.order, and an unresolved or circular macro is left as its literal token with a printed warning. A macro warning is a stop condition, not noise — the token you meant to expand is now text the model will read verbatim.
Review the resolved payload, not the template
A normal run shows you both ends of the crossing. Before dispatch, the CLI prints the task header — 📋 Task Template: when substitution or macro syntax is present, plain 📋 Task: otherwise — with a hint naming the resolution that is about to happen. When resolution actually changes the text, or a macro warns, the 🧩 Resolved Task (iteration N): preview prints as well:
Previews are capped at 200 characters and deduplicated per iteration, and --quiet suppresses all of it — so audit runs that were not quiet. Read the resolved line as the thing you are accountable for: it is the request the backend will dispatch, after the shell, macros, and substitutions have all had their turn. And because resolution repeats per iteration, the payload you reviewed at iteration 1 is no guarantee for iteration 3; when a substitution command is retry-sensitive, re-check the preview on the iteration that matters.
Make fan-out context explicit
When Parallel Runner fans work out, each worker's prompt is assembled from a template plus one untrusted task body. The declared crossing is the placeholder: custom prompts must include {{task_id}} or {{item}}, because generic prose does not automatically inject the task:
A missing placeholder fails quietly — the worker runs with no task context, or with context you did not choose — so the Parallel Runner documentation makes the placeholder a prerequisite rather than a suggestion. That is the prompt-safety rule for fan-out: the task body will enter some prompt somewhere, and the placeholder is what makes it enter one you designed.
Keep secrets out of every layer
Secrets have one crossing point: scoped environment variables. Tokens such as SLACK_BOT_TOKEN (xoxb-…) and GITHUB_TOKEN (ghp_…) belong in the project env file — .env.yylo, auto-created on yylo init and loaded before execution so hooks and subagent processes receive it — never in prompts, substitution commands, task bodies, or recorded responses. Two properties make anything else a leak:
- Task text is durable. Bodies and responses persist as task truth and re-enter later prompts and sessions as memory-surface input, so a token pasted there outlives the run that needed it.
- Substitution commands inherit your environment. The same power that makes
!'git status --short'convenient makes any secret-readable command a one-line exfiltration path.
Preview every external write before enabling it. The Slack and GitHub integrations ship dry-run modes for exactly this: slack_respond.sh --dry-run --verbose and github.py respond --dry-run show what would be posted without posting it. Start with least-privilege token scopes, narrow channels and labels, and leave response-and-close behavior off until the preview has been reviewed.
Two neighboring guides own what this one deliberately does not restate: for the run-directory evidence that survives a dispatched prompt — manifests, failure artifacts, recovery — see Build auditable agent workflows with handoff, and for choosing between ordered and fan-out execution see Choose Workflow Runner or Parallel Runner. Prompt boundaries are the input side of the same contract: name the surface, bound the crossing, keep secrets in one place, and review the resolved payload you are about to send.