Veyl

Enterprise vertical / Coding agents

Prove which agent work still needs an engineer.

Veyl finds recurring agent mistakes, tests changes to the deployment, and evaluates the exact setup on real past work. The result shows what can run without routine review, what still needs an engineer, and what must stop.

What is coding-agent evaluation?

Coding-agent evaluation tests whether a specific agent setup can complete representative engineering work under the rules that matter in a real codebase. A useful evaluation covers the model, agent, instructions, context, tools, permissions, environment, and final handoff. It should reveal what can run alone, what still needs review, and what is not ready.

Why passing CI is not enough

Repository tests are necessary, but they rarely encode every business rule, ownership boundary, recovery condition, or data-handling requirement. An agent can make a plausible change, keep the suite green, and still produce the wrong operational result.

Veyl uses private checks that the agent does not see. Those checks are established before execution and tested against empty, known-good, and deliberately plausible-but-wrong solutions before they are allowed to grade an agent.

Public studyOrdinary resultWhat deeper evaluation showed
Ramp CLI reconstruction18/18 passedSeven repairs violated held-out business rules
Brex Substation reconstruction18/18 passedThree narrow workflow lanes qualified, nothing broader
Moov ACH mutation study20 broken variants stayed greenThe 71-package suite missed material payment behavior

These are independent public-repository studies, not customer engagements or claims about private production systems. The full methods, limitations, checks, and source artifacts are linked from Veyl’s evidence page.

Evaluate the deployment, not the model name

A model leaderboard cannot tell an engineering leader whether one deployed agent may own one company workflow. The surrounding system changes the result. Veyl freezes the exact deployment so the evidence has a stable identity:

Change any material part of that setup and the old answer may no longer apply. Veyl records the change and identifies which evidence must run again.

Three workflow decisions

01 / Delegate

Run independently

The tested deployment met every required rule for this bounded workflow.

02 / Supervise

Keep human review

The workflow is useful to automate, but the evidence does not support removing review.

03 / Stop

Do not automate yet

The deployment repeatedly failed a material rule or the evidence is not strong enough.

What a managed Veyl evaluation looks like

Veyl runs the evaluation program for the customer. Teams may begin with one recurring job whose review burden matters and expand across additional workflows as evidence is established. Veyl finds repeated failure patterns, tests improvements to the deployment, and measures whether the change removes human work.

Your team suppliesVeyl handles
Repeated jobs that consume review timeFailure discovery and workflow reconstruction
Approved repository history and source accessImprovements to agent instructions, context, tools, or setup
A named engineer who confirms the rulesPrivate checks and repeated runs of the improved deployment
Security, provider, and final approvalReview routing, ongoing checks, and value measurement

What the customer receives

What the result does not claim

A passing workflow does not mean the agent is generally safe. It means the tested deployment qualified for the tested work under the recorded conditions. A failing workflow does not mean the agent is useless. It shows exactly where review or further development is still required.

This narrowness is deliberate. The goal is to expand useful autonomy without turning a local result into a company-wide promise the evidence cannot support.

Common questions

Can Veyl evaluate Cursor, Claude Code, Codex, or an internal coding agent?
Yes. The unit of evaluation is the complete deployment, not a preferred model vendor. The exact agent, model, tools, permissions, and environment are frozen and named in the result.
Does Veyl need a prepared benchmark?
No. Veyl reconstructs candidate scenarios from approved engineering history and asks a named engineer to confirm the small set that matters.
Does customer code have to leave its environment?
No. Veyl supports a customer-controlled runner. Any model provider that receives repository context must be explicitly approved.
How is this different from SWE-bench or an agent leaderboard?
Public benchmarks compare systems on shared tasks. Veyl answers a private operating question about one company workflow and one exact deployment.
What happens after the first evaluation?
The decision stays tied to its dependencies. Relevant changes to the model, tools, permissions, environment, or code trigger the required rerun.

Inspect before buying

Public results · Evaluation method · Private-code boundaries · About Veyl

Reduce code review