TL;DR: An agent harness needs clear limits, trustworthy completion evidence and a useful handoff when a run stops. Define those boundaries before widening tool access. The JavaScript example demonstrates step limits, simulated cost reservations and repeated-failure handling through scripted outcomes; it does not run a model or enforce a sandbox.
Consider an illustrative test-repair agent working in a disposable checkout. Its job is to fix one named failing test and return a patch with a verification report. The trace below is a script of possible attempts. For each one, reserve is the assumed maximum cost and cost is the charge to simulate if it runs. Both use illustrative units. patch-a and patch-b are fixture labels, not content hashes.
candidate 1: artifact patch-a, command test-one, failed, reserve 3, cost 2
candidate 2: artifact patch-a, command test-one, failed, reserve 3, cost 2
candidate 3: artifact patch-b, command test-one, passed, reserve 3, cost 2
This policy stops after two consecutive failures with the same command and artifact label. Candidate 3 never runs, even though the script gives it a changed patch and a passing result. The owner receives the last observation and decides whether to authorize another run.
A flaky test or changed environment could justify trying again. This small controller cannot establish that; its repeated-failure rule is a reason to review the run, not proof that more work would be useless.
An AI agent harness is software that manages the model and tool loop, working context, and execution state. This article develops its stop-and-handoff behavior through a small JavaScript simulation. It assumes familiarity with functions, objects, and basic tests. The example needs no model credentials or packages. Jump to the runnable JavaScript if you want to try it first.
Anthropic's note on building effective agents distinguishes predetermined workflows from agents that direct their own process, and it treats stopping conditions as a control rather than an afterthought (Building effective agents).
Define the run contract
A contract is useful only where a checker can enforce it. Write the allowed work, the completion evidence, the limits, and the stop destination before widening tool access.
| Field | Illustrative value | Reader decision |
|---|---|---|
| Work | Fix one named failing test | Define a bounded outcome |
| Output | Patch plus verification report | Specify inspectable evidence |
| Completion | Trusted test result tied to the current patch | Do not accept self-reported success |
| Allowed tools | Named test runner and patch operation | Enforce through trusted adapters or runtime |
| Workspace | Disposable checkout | Test actual filesystem isolation separately |
| Network | Explicit destinations only | Configure and verify runtime policy |
| Steps | 6 attempts | Calibrate against representative runs |
| Cost | 12 illustrative units | Replace with accounted costs and reliable reservations |
| Stop destination | Named human owner and last observation | Decide who can authorize another run |
This article's controller checks steps, simulated reservations, and repeated labels, then handles the supplied completion outcome. Workspace isolation, network policy, timeouts, and restart recovery require a separate executor. NVIDIA's OpenShell tutorial describes a default-deny network policy: traffic is denied unless a rule allows it (First network policy). The example here does not configure or test network denial.
Completion evidence must come from a trusted executor. A model assertion that the test passed is not success. Artifact labels in the script are fixture identifiers. A real executor needs trustworthy content identity, because renaming identical bytes can look like progress, and reusing a label for changed bytes can look like a repeat.
Replay the contract
Save both JavaScript blocks below in one file named harness.mjs, then run node harness.mjs with a current Node.js installation. The example requires no packages, model credentials, or network access.
// Offline teaching example: scripted observations, no model, shell, or sandbox.
// Costs are integer illustrative units, not tokens or currency.
// Input is a candidate script, not a historical log. Unconsumed rows never run.
export function replay(attempts, { maxSteps = 6, budget = 12 } = {}) {
if (!Number.isSafeInteger(maxSteps) || maxSteps < 1 ||
!Number.isSafeInteger(budget) || budget < 0) throw new Error('Invalid limits');
let spent = 0;
let previousFailure = null;
const events = [];
const finish = (status, reason) => ({
status, reason, spent, steps: events.length,
lastObservation: events.at(-1) ?? null,
// An observation is evidence to inspect, not authorization to resume.
nextAction: status === 'succeeded' ? 'Review patch and test evidence' :
'Inspect last observation; authorize a new bounded run if appropriate',
events,
});
for (const attempt of attempts) {
if (events.length >= maxSteps) return finish('needs_review', 'step_limit');
const { artifact, command, outcome, reserve, cost } = attempt;
if (typeof artifact !== 'string' || typeof command !== 'string' ||
!['passed', 'failed'].includes(outcome) ||
!Number.isSafeInteger(reserve) || reserve < 0 ||
!Number.isSafeInteger(cost) || cost < 0) throw new Error('Invalid observation');
// reserve represents a trusted upper bound in this single-run simulation.
if (reserve > budget - spent) return finish('needs_review', 'budget_preflight');
spent += cost;
events.push({ artifact, command, outcome, reserve, cost });
// A bad reservation invalidates the cap assumption even if the test passed.
if (cost > reserve) return finish('needs_review', 'reservation_exceeded');
if (outcome === 'passed') return finish('succeeded', 'verification_passed');
const failure = JSON.stringify([artifact, command]);
if (failure === previousFailure) return finish('needs_review', 'repeated_failure');
previousFailure = failure;
}
return finish('needs_review', 'observations_exhausted');
}
Append this usage example to the same file:
const unchangedFailure = {
artifact: 'patch-a',
command: 'test-one',
outcome: 'failed',
reserve: 3,
cost: 2,
};
const verifiedPass = {
...unchangedFailure,
artifact: 'patch-b',
outcome: 'passed',
};
const result = replay([unchangedFailure, unchangedFailure, verifiedPass]);
console.log(JSON.stringify(result, null, 2));
The key fields in the output form this illustrative handoff summary:
Status: needs_review
Reason: repeated_failure
Attempts consumed: 2
Cost consumed: 4 illustrative units
Last artifact label: patch-a
Next action: inspect the evidence before authorizing a new bounded run
The third candidate is unread, so its passing outcome never runs and does not count toward cost. Replacing the call with replay([unchangedFailure, verifiedPass]) returns succeeded with verification_passed and spent: 4: a simulated pass after a changed label, still inside the reservation.
Proposed runtime; the executor is not included in this example:
Candidate -> Controller admission check -> Trusted executor
|
Controller result check <- Test result and cost
|
v
Handoff -> Owner inspects evidence and authorizes any new run
The file implements the controller decisions for one simulated run. It does not call a model, execute a test, or enforce sandbox permissions. A real executor must tie its test result to the actual artifact; concurrent workers also need atomic shared budget reservations.
Decide whether another attempt is useful
The controller consumes candidates in order. Unread rows never run. Only consumed rows contribute cost.
The repeated-failure rule is consecutive only. It compares the current failure's command and artifact label with the immediately previous failure. Two identical failures stop the run with repeated_failure before a third candidate is considered. Alternating failures (A fails, B fails, A fails again) do not trip that heuristic. They continue until the step cap, the budget preflight, or the end of the script. The heuristic is not general loop detection. Flaky tests and environmental drift need different policies.
A changed artifact label permits another attempt of the same command within the remaining limits. That is a counterexample to "the same command always means stop," not proof of useful progress. The label comparison cannot see bytes. If two different patches share a label, the controller may stop as if the work were unchanged. If identical content is relabeled, the controller may allow another try as if something new had been produced.
Successful simulated verification is narrow. When a consumed observation reports passed, and the incurred cost does not exceed its reservation, the run returns succeeded with reason verification_passed. That result means the scripted executor said the current patch passed inside the accounting rules. It does not mean a model ran, a repository changed, or a real test process executed.
The step cap is checked before starting another candidate. At maxSteps: 1, a failing one-row script ends with observations_exhausted; a two-row script stops with step_limit before consuming its second row. An empty script also returns observations_exhausted, with zero spend. Ending the script without a pass never establishes success.
Costs are integer teaching units. Each reserve represents a trusted upper bound on the next attempt's cost. If it exceeds budget - spent, budget_preflight refuses that candidate without consuming it or charging its cost.
Once an attempt has run, its cost is already incurred. If cost exceeds reserve, the controller returns reservation_exceeded even if the supplied test outcome passed. The excess remains in spent: stopping cannot undo it. Real spending caps therefore depend on reliable upper bounds reserved before dispatch; a check after the charge can only detect a broken assumption.
There is no automatic recovery. A returned observation is evidence to inspect, not authorization to resume. In a real system, merge permission is separate from producing a patch or handoff packet. After a stop, a person inspects the artifact and reauthorizes a new bounded run if that is still the right work.
Leave a recoverable result
Every finish path returns the same shape: status, reason, spent, steps, lastObservation, nextAction, and events. Use the reason to decide what to inspect. Do not treat a non-success status as a prompt to loop the same script.
| Reason | Candidate consumed? | Cost accounted? | Inspect |
|---|---|---|---|
repeated_failure | Yes; the following candidate never runs | Both identical failures | Last observation; same label and command |
verification_passed | Yes | Includes the passing attempt | Patch and trusted test evidence |
budget_preflight | No | Prior attempts only | Remaining budget versus the unread reserve |
reservation_exceeded | Yes | Includes the excess charge | The broken upper bound; charge is already counted |
step_limit | No | Prior attempts only | Whether the cap was hit before a new candidate |
observations_exhausted | All scripted attempts that ran | All consumed attempts | Last observation; no inferred pass |
nextAction on success is to review the patch and test evidence. On every needs_review path it is to inspect the last observation and authorize a new bounded run if that is still appropriate. The controller does not resume, retry, or widen the contract.
Map each instruction to enforcement
Map each instruction to the code or runtime policy that enforces it. A contract field without a checker remains an instruction.
| Instruction | Responsible enforcement | Demonstration status |
|---|---|---|
| Stop after the attempt limit | Controller preflight | Tested in replay |
| Do not exceed the cost budget | Pre-dispatch reservation with trustworthy upper bounds | Simulated for one run |
| Stop repeating an unchanged failure | Controller compares command and artifact | Consecutive-label heuristic tested in replay |
| Work only inside the checkout | OS or runtime filesystem policy | Not implemented |
| Stop hung commands | Executor timeout and process-tree termination | Not implemented |
| Resume after a worker crash | Durable state, artifact verification, and retry authority | Not implemented |
Isolation and outbound network access need runtime policy. Cancelling hung commands needs process lifecycle control. Crash recovery needs durable state, artifact verification, and retry authority. These controls sit outside the replay. Calibrate step and cost numbers against representative runs of the real workflow. Do not copy 6 and 12 as a universal threshold.
Apply this to one workflow
Copy the contract table and replace every illustrative value with the actual outcome, tools, owner, and limits for one blocked workflow. Replay these failure cases:
- The same command fails again against the unchanged artifact label.
- A changed artifact label still produces a failing result.
- The next reservation will not fit in the remaining budget.
- An attempt costs more than its reservation.
- Changing artifacts reach the step cap without passing.
- The script ends without passing evidence.
Name the human who can authorize another run, then verify the controls you intend to trust. A worksheet row marked "not implemented" still needs an enforcing component before you rely on it.
Keep the first baseline on a fixed model, fixed tool permissions, and a fixed success rubric. Routing, including cheaper fallbacks or bounded escalation, is a later experiment. The Switchyard integration with NeMo Relay documents routing calls, token use, and latency overhead that need accounting. Measure that cost only after the harness already stops a failed run and leaves evidence.
For repository rules and merge authority, read How I Use Coding Agents Without Giving Up Architectural Control. The AI Demo to Production checklist covers the broader workflow, ownership, and rollout decisions.
The harness job is small on purpose: bound one run, refuse another attempt when its stop policy fires, and leave an owner with status, reason, consumed budget, and the last observation. Use that packet on the next blocked workflow before adding tools, concurrency, or a larger budget. Teams that need a structured review of one workflow's production-readiness constraints can use the Production Readiness Diagnostic.