Engineering guide

AI Agent Harnesses: Stop Failed Runs and Hand Off the Evidence

Bound an agent run, stop repeated failures, and hand off useful evidence with a small JavaScript example you can run locally.

François Guéguen 10 min read

TL;DR: An agent harness needs clear limits, trustworthy completion evidence and a useful handoff when a run stops. Define those boundaries before widening tool access. The JavaScript example demonstrates step limits, simulated cost reservations and repeated-failure handling through scripted outcomes; it does not run a model or enforce a sandbox.


Consider an illustrative test-repair agent working in a disposable checkout. Its job is to fix one named failing test and return a patch with a verification report. The trace below is a script of possible attempts. For each one, reserve is the assumed maximum cost and cost is the charge to simulate if it runs. Both use illustrative units. patch-a and patch-b are fixture labels, not content hashes.

candidate 1: artifact patch-a, command test-one, failed, reserve 3, cost 2
candidate 2: artifact patch-a, command test-one, failed, reserve 3, cost 2
candidate 3: artifact patch-b, command test-one, passed, reserve 3, cost 2

This policy stops after two consecutive failures with the same command and artifact label. Candidate 3 never runs, even though the script gives it a changed patch and a passing result. The owner receives the last observation and decides whether to authorize another run.

A flaky test or changed environment could justify trying again. This small controller cannot establish that; its repeated-failure rule is a reason to review the run, not proof that more work would be useless.

An AI agent harness is software that manages the model and tool loop, working context, and execution state. This article develops its stop-and-handoff behavior through a small JavaScript simulation. It assumes familiarity with functions, objects, and basic tests. The example needs no model credentials or packages. Jump to the runnable JavaScript if you want to try it first.

Anthropic's note on building effective agents distinguishes predetermined workflows from agents that direct their own process, and it treats stopping conditions as a control rather than an afterthought (Building effective agents).

Define the run contract

A contract is useful only where a checker can enforce it. Write the allowed work, the completion evidence, the limits, and the stop destination before widening tool access.

Illustrative run contract
FieldIllustrative valueReader decision
WorkFix one named failing testDefine a bounded outcome
OutputPatch plus verification reportSpecify inspectable evidence
CompletionTrusted test result tied to the current patchDo not accept self-reported success
Allowed toolsNamed test runner and patch operationEnforce through trusted adapters or runtime
WorkspaceDisposable checkoutTest actual filesystem isolation separately
NetworkExplicit destinations onlyConfigure and verify runtime policy
Steps6 attemptsCalibrate against representative runs
Cost12 illustrative unitsReplace with accounted costs and reliable reservations
Stop destinationNamed human owner and last observationDecide who can authorize another run

This article's controller checks steps, simulated reservations, and repeated labels, then handles the supplied completion outcome. Workspace isolation, network policy, timeouts, and restart recovery require a separate executor. NVIDIA's OpenShell tutorial describes a default-deny network policy: traffic is denied unless a rule allows it (First network policy). The example here does not configure or test network denial.

Completion evidence must come from a trusted executor. A model assertion that the test passed is not success. Artifact labels in the script are fixture identifiers. A real executor needs trustworthy content identity, because renaming identical bytes can look like progress, and reusing a label for changed bytes can look like a repeat.

Replay the contract

Save both JavaScript blocks below in one file named harness.mjs, then run node harness.mjs with a current Node.js installation. The example requires no packages, model credentials, or network access.

// Offline teaching example: scripted observations, no model, shell, or sandbox.
// Costs are integer illustrative units, not tokens or currency.
// Input is a candidate script, not a historical log. Unconsumed rows never run.
export function replay(attempts, { maxSteps = 6, budget = 12 } = {}) {
  if (!Number.isSafeInteger(maxSteps) || maxSteps < 1 ||
      !Number.isSafeInteger(budget) || budget < 0) throw new Error('Invalid limits');
  let spent = 0;
  let previousFailure = null;
  const events = [];
  const finish = (status, reason) => ({
    status, reason, spent, steps: events.length,
    lastObservation: events.at(-1) ?? null,
    // An observation is evidence to inspect, not authorization to resume.
    nextAction: status === 'succeeded' ? 'Review patch and test evidence' :
      'Inspect last observation; authorize a new bounded run if appropriate',
    events,
  });

  for (const attempt of attempts) {
    if (events.length >= maxSteps) return finish('needs_review', 'step_limit');
    const { artifact, command, outcome, reserve, cost } = attempt;
    if (typeof artifact !== 'string' || typeof command !== 'string' ||
        !['passed', 'failed'].includes(outcome) ||
        !Number.isSafeInteger(reserve) || reserve < 0 ||
        !Number.isSafeInteger(cost) || cost < 0) throw new Error('Invalid observation');
    // reserve represents a trusted upper bound in this single-run simulation.
    if (reserve > budget - spent) return finish('needs_review', 'budget_preflight');
    spent += cost;
    events.push({ artifact, command, outcome, reserve, cost });
    // A bad reservation invalidates the cap assumption even if the test passed.
    if (cost > reserve) return finish('needs_review', 'reservation_exceeded');
    if (outcome === 'passed') return finish('succeeded', 'verification_passed');
    const failure = JSON.stringify([artifact, command]);
    if (failure === previousFailure) return finish('needs_review', 'repeated_failure');
    previousFailure = failure;
  }
  return finish('needs_review', 'observations_exhausted');
}

Append this usage example to the same file:

const unchangedFailure = {
  artifact: 'patch-a',
  command: 'test-one',
  outcome: 'failed',
  reserve: 3,
  cost: 2,
};

const verifiedPass = {
  ...unchangedFailure,
  artifact: 'patch-b',
  outcome: 'passed',
};

const result = replay([unchangedFailure, unchangedFailure, verifiedPass]);
console.log(JSON.stringify(result, null, 2));

The key fields in the output form this illustrative handoff summary:

Status: needs_review
Reason: repeated_failure
Attempts consumed: 2
Cost consumed: 4 illustrative units
Last artifact label: patch-a
Next action: inspect the evidence before authorizing a new bounded run

The third candidate is unread, so its passing outcome never runs and does not count toward cost. Replacing the call with replay([unchangedFailure, verifiedPass]) returns succeeded with verification_passed and spent: 4: a simulated pass after a changed label, still inside the reservation.

Proposed runtime; the executor is not included in this example:

Candidate -> Controller admission check -> Trusted executor
                                               |
Controller result check <- Test result and cost
          |
          v
Handoff -> Owner inspects evidence and authorizes any new run

The file implements the controller decisions for one simulated run. It does not call a model, execute a test, or enforce sandbox permissions. A real executor must tie its test result to the actual artifact; concurrent workers also need atomic shared budget reservations.

Decide whether another attempt is useful

The controller consumes candidates in order. Unread rows never run. Only consumed rows contribute cost.

The repeated-failure rule is consecutive only. It compares the current failure's command and artifact label with the immediately previous failure. Two identical failures stop the run with repeated_failure before a third candidate is considered. Alternating failures (A fails, B fails, A fails again) do not trip that heuristic. They continue until the step cap, the budget preflight, or the end of the script. The heuristic is not general loop detection. Flaky tests and environmental drift need different policies.

A changed artifact label permits another attempt of the same command within the remaining limits. That is a counterexample to "the same command always means stop," not proof of useful progress. The label comparison cannot see bytes. If two different patches share a label, the controller may stop as if the work were unchanged. If identical content is relabeled, the controller may allow another try as if something new had been produced.

Successful simulated verification is narrow. When a consumed observation reports passed, and the incurred cost does not exceed its reservation, the run returns succeeded with reason verification_passed. That result means the scripted executor said the current patch passed inside the accounting rules. It does not mean a model ran, a repository changed, or a real test process executed.

The step cap is checked before starting another candidate. At maxSteps: 1, a failing one-row script ends with observations_exhausted; a two-row script stops with step_limit before consuming its second row. An empty script also returns observations_exhausted, with zero spend. Ending the script without a pass never establishes success.

Costs are integer teaching units. Each reserve represents a trusted upper bound on the next attempt's cost. If it exceeds budget - spent, budget_preflight refuses that candidate without consuming it or charging its cost.

Once an attempt has run, its cost is already incurred. If cost exceeds reserve, the controller returns reservation_exceeded even if the supplied test outcome passed. The excess remains in spent: stopping cannot undo it. Real spending caps therefore depend on reliable upper bounds reserved before dispatch; a check after the charge can only detect a broken assumption.

There is no automatic recovery. A returned observation is evidence to inspect, not authorization to resume. In a real system, merge permission is separate from producing a patch or handoff packet. After a stop, a person inspects the artifact and reauthorizes a new bounded run if that is still the right work.

Leave a recoverable result

Every finish path returns the same shape: status, reason, spent, steps, lastObservation, nextAction, and events. Use the reason to decide what to inspect. Do not treat a non-success status as a prompt to loop the same script.

Stop decisions and their effects
ReasonCandidate consumed?Cost accounted?Inspect
repeated_failureYes; the following candidate never runsBoth identical failuresLast observation; same label and command
verification_passedYesIncludes the passing attemptPatch and trusted test evidence
budget_preflightNoPrior attempts onlyRemaining budget versus the unread reserve
reservation_exceededYesIncludes the excess chargeThe broken upper bound; charge is already counted
step_limitNoPrior attempts onlyWhether the cap was hit before a new candidate
observations_exhaustedAll scripted attempts that ranAll consumed attemptsLast observation; no inferred pass

nextAction on success is to review the patch and test evidence. On every needs_review path it is to inspect the last observation and authorize a new bounded run if that is still appropriate. The controller does not resume, retry, or widen the contract.

Map each instruction to enforcement

Map each instruction to the code or runtime policy that enforces it. A contract field without a checker remains an instruction.

Instructions and their enforcement
InstructionResponsible enforcementDemonstration status
Stop after the attempt limitController preflightTested in replay
Do not exceed the cost budgetPre-dispatch reservation with trustworthy upper boundsSimulated for one run
Stop repeating an unchanged failureController compares command and artifactConsecutive-label heuristic tested in replay
Work only inside the checkoutOS or runtime filesystem policyNot implemented
Stop hung commandsExecutor timeout and process-tree terminationNot implemented
Resume after a worker crashDurable state, artifact verification, and retry authorityNot implemented

Isolation and outbound network access need runtime policy. Cancelling hung commands needs process lifecycle control. Crash recovery needs durable state, artifact verification, and retry authority. These controls sit outside the replay. Calibrate step and cost numbers against representative runs of the real workflow. Do not copy 6 and 12 as a universal threshold.

Apply this to one workflow

Copy the contract table and replace every illustrative value with the actual outcome, tools, owner, and limits for one blocked workflow. Replay these failure cases:

  • The same command fails again against the unchanged artifact label.
  • A changed artifact label still produces a failing result.
  • The next reservation will not fit in the remaining budget.
  • An attempt costs more than its reservation.
  • Changing artifacts reach the step cap without passing.
  • The script ends without passing evidence.

Name the human who can authorize another run, then verify the controls you intend to trust. A worksheet row marked "not implemented" still needs an enforcing component before you rely on it.

Keep the first baseline on a fixed model, fixed tool permissions, and a fixed success rubric. Routing, including cheaper fallbacks or bounded escalation, is a later experiment. The Switchyard integration with NeMo Relay documents routing calls, token use, and latency overhead that need accounting. Measure that cost only after the harness already stops a failed run and leaves evidence.

For repository rules and merge authority, read How I Use Coding Agents Without Giving Up Architectural Control. The AI Demo to Production checklist covers the broader workflow, ownership, and rollout decisions.

The harness job is small on purpose: bound one run, refuse another attempt when its stop policy fires, and leave an owner with status, reason, consumed budget, and the last observation. Use that packet on the next blocked workflow before adding tools, concurrency, or a larger budget. Teams that need a structured review of one workflow's production-readiness constraints can use the Production Readiness Diagnostic.

All articles