01 / Delegate
Run independently
The tested deployment met every required rule for this bounded workflow.
Enterprise vertical / Coding agents
Veyl finds recurring agent mistakes, tests changes to the deployment, and evaluates the exact setup on real past work. The result shows what can run without routine review, what still needs an engineer, and what must stop.
What is coding-agent evaluation?
Coding-agent evaluation tests whether a specific agent setup can complete representative engineering work under the rules that matter in a real codebase. A useful evaluation covers the model, agent, instructions, context, tools, permissions, environment, and final handoff. It should reveal what can run alone, what still needs review, and what is not ready.
Repository tests are necessary, but they rarely encode every business rule, ownership boundary, recovery condition, or data-handling requirement. An agent can make a plausible change, keep the suite green, and still produce the wrong operational result.
Veyl uses private checks that the agent does not see. Those checks are established before execution and tested against empty, known-good, and deliberately plausible-but-wrong solutions before they are allowed to grade an agent.
| Public study | Ordinary result | What deeper evaluation showed |
|---|---|---|
| Ramp CLI reconstruction | 18/18 passed | Seven repairs violated held-out business rules |
| Brex Substation reconstruction | 18/18 passed | Three narrow workflow lanes qualified, nothing broader |
| Moov ACH mutation study | 20 broken variants stayed green | The 71-package suite missed material payment behavior |
These are independent public-repository studies, not customer engagements or claims about private production systems. The full methods, limitations, checks, and source artifacts are linked from Veyl’s evidence page.
A model leaderboard cannot tell an engineering leader whether one deployed agent may own one company workflow. The surrounding system changes the result. Veyl freezes the exact deployment so the evidence has a stable identity:
Change any material part of that setup and the old answer may no longer apply. Veyl records the change and identifies which evidence must run again.
01 / Delegate
The tested deployment met every required rule for this bounded workflow.
02 / Supervise
The workflow is useful to automate, but the evidence does not support removing review.
03 / Stop
The deployment repeatedly failed a material rule or the evidence is not strong enough.
Veyl runs the evaluation program for the customer. Teams may begin with one recurring job whose review burden matters and expand across additional workflows as evidence is established. Veyl finds repeated failure patterns, tests improvements to the deployment, and measures whether the change removes human work.
| Your team supplies | Veyl handles |
|---|---|
| Repeated jobs that consume review time | Failure discovery and workflow reconstruction |
| Approved repository history and source access | Improvements to agent instructions, context, tools, or setup |
| A named engineer who confirms the rules | Private checks and repeated runs of the improved deployment |
| Security, provider, and final approval | Review routing, ongoing checks, and value measurement |
A passing workflow does not mean the agent is generally safe. It means the tested deployment qualified for the tested work under the recorded conditions. A failing workflow does not mean the agent is useless. It shows exactly where review or further development is still required.
This narrowness is deliberate. The goal is to expand useful autonomy without turning a local result into a company-wide promise the evidence cannot support.
Inspect before buying
Public results · Evaluation method · Private-code boundaries · About Veyl