A practical model-selection study · October 1, 2026

How Much AI Does Your Work Actually Need?

Use YYLO CLI and YYLO Benchmark to find the least expensive model that reliably gets your jobs done.

By Juno AI INC · Evaluated

Choosing an AI model often feels like choosing insurance: pick the strongest one, pay more, and hope you needed it.

But fixing a small test is not the same job as changing how an application loads its settings. Why pay the same price for both?

The useful question is not “Which model is smartest?”

It is “What is the least expensive model that meets my quality standard for this kind of work?”

YYLO CLI lets you run coding agents with a chosen model. YYLO Benchmark lets you compare them on the same prepared tasks and keep their work for later review. Together, they help turn model selection from a guess into a decision you can measure.

We tried four models on five real jobs

We selected five past engineering tasks from our own CLI project and gave each to four models: GPT-5.6 Sol, GPT-5.6 Luna, GPT-5.6 Terra and GPT-5.5.

That gave us 20 retained coding outputs. For each task, the models received the same starting code and requirements, with the same reasoning setting and time limit.

The jobs included changing configuration rules, removing old code, making settings work across different parts of an application, preserving user settings, and correcting a test.

We asked two separate questions about each result:

  • Did it pass our focused checks? Does the selected behavior work without introducing new compiler problems compared with our reviewed solution?
  • Did it fulfill the request? An independent AI reviewer, GPT-6 Astra, examined whether important requirements were missing or incorrectly implemented.

Here are the final results after correcting problems in our evaluation process:

The results: checks and completeness are different

Final task-level results. Each model attempted the same five tasks.
ModelPassed focused checksJudged complete
GPT-5.6 Sol3 of 53 of 5
GPT-5.6 Luna3 of 52 of 5
GPT-5.6 Terra3 of 53 of 5
GPT-5.53 of 53 of 5

These are separate measures, not a combined score. “Complete” is the AI reviewer’s assessment, not a guarantee. Names are the model selections recorded for this experiment; availability and pricing depend on your provider.

Download the 20 results (CSV). The download contains the same task-level scores shown here, not private source code or conversation logs.

The five jobs—and what happened on each

In each cell, the first result is the focused check outcome; the second is the completeness judgment.

Checks / completeness, one retained implementation per task and model.
JobGPT-5.6 SolGPT-5.6 LunaGPT-5.6 TerraGPT-5.5
Add a simpler configuration modePass / CompletePass / IncompletePass / CompletePass / Complete
Remove obsolete command-line codeFail / IncompleteFail / IncompleteFail / IncompleteFail / Incomplete
Connect shared settings to the running applicationFail / IncompleteFail / IncompleteFail / IncompleteFail / Incomplete
Discover skills while preserving user settingsPass / CompletePass / CompletePass / CompletePass / Complete
Correct a workspace-role testPass / CompletePass / CompletePass / CompletePass / Complete
Add a simpler configuration mode
Accept valid settings and reject contradictory ones without breaking existing users.
Remove obsolete command-line code
Delete unused code without leaving broken connections or new compiler problems.
Connect shared settings to the running application
Make shared agent preferences actually reach the places that use them.
Discover skills while preserving user settings
Find the intended skills without overwriting existing settings.
Correct a workspace-role test
Update the intended test expectation while preserving its other checks.

The configuration-mode result explains the split between the columns: Luna passed our focused checks but its implementation could reject valid existing settings when the application loaded them. On the old-code removal job, every model added compiler problems; two also left connections to deleted code. For the shared-settings job, all four implementations failed to connect the settings to the running application correctly. These were assessed implementation problems, not failed evaluation launches.

How we ran the comparison

  • Source of tasks: five randomly selected eligible historical tasks from our own CLI project. The eligible pool contained 131 tasks with usable local code history and test changes—not a random sample of all software work. A case with contradictory reference evidence was excluded before implementation.
  • Coding runs: September 24–28, 2026. Four models, one intended scored attempt per task/model, medium reasoning and a 20-minute limit including setup. Two runs could proceed at once, with model order rotated.
  • Tools used in this study: YYLO CLI 0.2.10, YYLO Benchmark 0.2.0 and Node 22. Coding and AI review used YYLO’s Pi runner. These are our installed study versions, not a claim about the latest public release.
  • Independent review: GPT-6 Astra, medium reasoning, checked the saved work against the original requirements. The task-specific checks and AI judgments were kept separate.
  • Final evaluation: October 1, 2026. All 20 outputs received both assessments, with no remaining evaluation errors or unknown results. We used separately prepared evaluators alongside Benchmark after fixing evaluation problems; this was not a one-command automatic process.
  • Fairness checks: our checks had to reject broken examples and accept correct alternatives—not just recognize our original solution. Skill discovery was tested against Pi 0.84.4; the original historical Pi version was not preserved.

An earlier preparation run for Luna on the old-code removal job was cancelled; it was not a second completed answer from which we chose a winner. Two coding runs were completed later after disk space was restored. We then reassessed the saved implementations without rerunning the coding models.

This is a first-party exploratory study. The authors knew which models produced the work, the checks were revised after initial results, and an AI reviewer can miss things. GPT-5.5’s simpler-configuration result changed from incomplete to complete under revised review without a change to its code. The public tables and download expose our summary results; they do not by themselves provide everything needed to reproduce the experiment independently.

“Passed focused checks” does not mean the whole application passed every test. Some older starting versions already had compiler problems; we checked for additional problems and recorded those existing failures separately.

This is a new five-task study, not an updated score for our earlier ten-task experiment. It is also not a price ranking. These results alone do not establish which model was cheapest or how much money a switch would save.

The useful result was the pattern—not the winner

All four models completed two of the jobs: preserving user settings and correcting a test. None was judged complete on two others: removing old code and connecting configuration across the application.

On the remaining job, all four passed the focused checks, but one missed an important requirement in review.

That gives us three useful lessons:

  • When several models meet the same standard, compare their total cost. You may not need the most expensive option for that category of work.
  • When every model struggles, spending more is not automatically the answer. Try a clearer request, a smaller job, better context or closer human involvement—and measure again.
  • Passing tests is not the same as finishing the job. A low-cost result that misses a requirement may become expensive once someone has to repair it.

Here is the aha moment: you do not need one “best model” for your entire project. You need a good match for each kind of work.

Our small experiment identifies questions worth testing. It does not prove that any model will reliably handle a whole category after one successful example.

Build a small benchmark from work you already do

You do not need a giant public leaderboard. Start with tasks that resemble next week's work.

1. Choose a few representative jobs

Pick examples of repeated work: a small bug fix, a test update, a settings change, or a feature that touches several parts of your application.

Use the code as it looked before the solution existed. Keep the finished answer out of the model's input.

Include difficult examples too. A benchmark made entirely of easy wins will tell you very little about where your cheaper option stops being useful.

2. Define “good enough” before comparing models

Write down what must work, what must stay unchanged, and what review is required.

For a settings change, “the new setting appears” might not be enough. Existing settings may also need to survive untouched.

Your quality bar should reflect the consequences of failure. A small internal utility and a payment feature should not have the same review policy.

3. Compare a small set of options fairly

Use YYLO Benchmark to run prepared cases with explicitly chosen models through YYLO CLI. Keep the task, starting code and evaluation rules the same.

Set a spending and time budget. Retain unsuccessful attempts rather than repeating them until you get an attractive score.

Our study kept the reasoning setting fixed. If you want to find the minimum thinking effort you need, make that a separate comparison: keep the model fixed and vary its supported reasoning setting. Do not change everything at once.

To explore the commands available in your installed version, start with:

bash
yy benchmark --help

Benchmark runs the comparisons you configure. It does not automatically select the cheapest model or change which model your live work uses.

4. Check quality, then compare the real cost

Keep a simple scorecard for each option:

  • Did it meet the requirements?
  • Did it pass the relevant checks?
  • How much did the attempt cost?
  • How much checking, correction and human review did it need?

The useful measure is total cost per accepted job, not price per response.

Include failed attempts, additional reviews and repairs. If spending information is missing or only estimated, label it that way. An incomplete cost total is not evidence that a model is cheaper.

5. Test a lower-cost default before relying on it

Repeat the comparison on more examples and reserve some fresh tasks to check your choice. One success is promising; repeated success is evidence.

Then adopt a policy such as:

Use the lowest-cost option that repeatedly meets our standard for routine work. Use a more capable option or human help when the work is unfamiliar, higher-risk, or fails our checks.

You choose that policy and the quality gates. Do not assume upgrading the model will rescue every difficult job.

What the savings could look like

Here is a hypothetical example—not a measured result from our study.

Suppose you have 100 jobs. Your current approach uses a $5 model attempt for each one: $500 in model spending.

Now suppose your benchmark supports using a $1 option for 80 routine jobs, while the other 20 still go directly to the $5 option:

  • 80 × $1 = $80
  • 20 × $5 = $100
  • Total: $180 instead of $500

That is 64% less model spending if the jobs meet the same quality standard without extra attempts. It excludes evaluation, infrastructure and human time. Additional checking or repair reduces the savings and can erase them.

The point is not the percentage. It is the decision: reserve expensive capability for work where it earns its keep.

A broken benchmark can make you spend more

We also had to fix our measuring tools.

One check rejected solutions because they did not use the same internal name as our reference answer—even though the original request did not require that name. Other evaluations failed before they could judge the work at all.

Treating those problems as model failures could push you toward a more expensive model for the wrong reason.

We corrected the evaluations and assessed the saved implementations again. We did not ask the coding models to redo their work. All 20 implementations ultimately received both assessments.

That is another useful way to control costs: when your checks improve, review the work you already paid for instead of automatically paying to generate it again. New evaluation still has a cost, but it does not require another coding attempt.

Start small, then earn confidence

Five tasks cannot establish a universal winner or a dependable success rate. Our evaluations were revised after the first results, the AI reviewer is not infallible, and one verdict changed during reassessment. One compatibility check also used a specifically chosen current tool version rather than recreating the historical environment exactly.

Those limits matter. Use this study as a practical example of the process—not a buying recommendation for a particular model.

Start with a handful of real jobs. Define success. Compare a few options. Measure the cost of work you can actually accept. Expand the evidence before making a model your default.

The goal is not to buy the least intelligence possible. It is to stop paying for capability your task does not need—without giving up the quality it does.

Put the comparison into practice

Use the real-repository benchmarking guide to prepare a comparison, and the YYLO Benchmark documentation for version-specific commands. The earlier ten-task experiment remains a separate historical study, not the source of these scores.

To cite this study: Juno AI INC, “How Much AI Does Your Work Actually Need?”, October 1, 2026. Use this article’s canonical link and specify that the results cover five tasks and four models.