Test an agent in CI with zero API calls

Record an agent's model calls once against the live API, commit the cassette, and replay it on every CI run after that: no API key in CI, no network calls, no per-run token cost, and a deterministic pass/fail instead of a flaky one.

Goal

By the end of this recipe your CI job runs your agent's test suite with CASSETTE_MODE=replay and no provider credentials configured anywhere in the workflow, and a prompt or tool-schema change fails that job loudly instead of silently passing against a stale fixture.

Prerequisites

  • An agent built on the Vercel AI SDK (generateText / streamText) behind wrapLanguageModel.
  • @nkwib/tapedeck installed as a dev dependency, with ai (>=6.0.0 <8) and vitest (>=2.0.0) as peers.
  • A model call you're willing to make once, live, to produce the initial recording.
npm install -D @nkwib/tapedeck

Steps

1. Wrap the model

Put cassetteMiddleware at the wrapLanguageModel layer, reading the mode from an environment variable so the same code runs live in development and replays in CI:

import { wrapLanguageModel } from 'ai';
import { cassetteMiddleware } from '@nkwib/tapedeck';
import type { CassetteMode } from '@nkwib/tapedeck';

export function cassetteModel(model: Parameters<typeof wrapLanguageModel>[0]['model']) {
  const mode = (process.env.CASSETTE_MODE as CassetteMode | undefined) ?? 'live';
  return wrapLanguageModel({
    model,
    middleware: cassetteMiddleware({ mode, cassetteDir: './cassettes' }),
  });
}

As of tapedeck 0.4.0 cassetteMiddleware is typed structurally (TapedeckMiddleware), so this is assignable to wrapLanguageModel under both ai@6 (spec v3) and ai@7 (spec v4) with no cast on either side.

2. Understand the three modes you'll actually use

Env varModeWhat happens
unset / livelivePassthrough, hits the real API. Use in development.
CASSETTE_MODE=recordrecordCalls the real model, writes ./cassettes/*.json. Run this once, locally, with a real key.
CASSETTE_MODE=replayreplayServes the committed cassette. No network. This is what CI runs.

cassetteMiddleware also has a fourth mode, compare, for scheduled drift checks against the live model rather than CI's hot path; see Drift handling below.

3. Pin the test to a named cassette with withCassette

For a multi-step agent, use withCassette from @nkwib/tapedeck/vitest instead of relying on the ambient CASSETTE_MODE. It forces replay for the duration of the test body and records every model call the agent makes into one named, multi-interaction cassette file:

import { describe, expect, it } from 'vitest';
import { withCassette } from '@nkwib/tapedeck/vitest';
import { runAgent } from '../src/agent.js';
import { cassetteModel } from '../src/model.js';
import { realModel } from '../src/real-model.js';

describe('support agent', () => {
  it('looks up the order and answers', async () => {
    const model = cassetteModel(realModel());

    const result = await withCassette('order-lookup.json', () =>
      runAgent({ model, prompt: 'Where is my order A-1001?' }),
    );

    expect(result.text).toContain('A-1001');
  });
});

withCassette(name, testFn, options?) defaults options.mode to 'replay'; pass { mode: 'record' } in a one-off script to (re)record that same test body without touching the test file itself (see the record script below).

4. Record the cassette once

Write a small script that runs the exact same agent call as the test, under withCassette(..., { mode: 'record' }), against a real (or scripted) model:

// scripts/record.ts
import { withCassette } from '@nkwib/tapedeck/vitest';
import { runAgent } from '../src/agent.js';
import { cassetteModel } from '../src/model.js';
import { realModel } from '../src/real-model.js';

await withCassette(
  'order-lookup.json',
  () => runAgent({ model: cassetteModel(realModel()), prompt: 'Where is my order A-1001?' }),
  { mode: 'record' },
);
ANTHROPIC_API_KEY=sk-... npx tsx scripts/record.ts

This is the one and only live call. Everything after this step is offline.

5. Commit the cassette

git add cassettes/order-lookup.json
git commit -m "record order-lookup cassette"

Cassettes are pretty-printed JSON, so the diff in a re-record PR is the change review: a reviewer reads what the agent's tool calls or answer changed to, the same way they'd read any other diff.

Check for secrets before committing
Redaction runs at record time against the built-in matchers (apiKey, authorization, x-api-key, bearer, token) plus anything you pass via redact: (string | RegExp)[]. If your provider or tools pass secrets under a different field name, add a matcher before the first record.

6. Point CI at replay, with no secrets configured

# .github/workflows/ci.yml
name: ci
on:
  push:
    branches: [main]
  pull_request:

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm
      - run: npm ci
      - run: npm run typecheck
      - run: npm test

Nothing here sets CASSETTE_MODE, ANTHROPIC_API_KEY, OPENAI_API_KEY, or any other provider secret. If cassetteModel defaults to 'live' when the env var is unset, either default it to 'replay' in non-development environments, or set CASSETTE_MODE=replay explicitly in the workflow so a missing env var can't silently fall through to a live call in CI.

Verification

Run the test suite locally with replay before you push, so CI is confirming what you already saw:

CASSETTE_MODE=replay npm test

You should see the cassette hit in the test output and no network activity. Then push and confirm the CI job runs green with no secrets in the job's environment.

Drift handling

replay never falls back to the live API on a miss, and it never serves a stale response: when the request hash (modelProvider, modelId, prompt, toolSchemas, maxOutputTokens, temperature, topP) doesn't match anything in the cassette, withCassette / the middleware throws CassetteMissError and the test fails.

That happens whenever:

  • the prompt or system instructions change,
  • a tool's input schema changes (its description does not: schemas are normalized before hashing),
  • a sampling parameter (temperature, topP, maxOutputTokens) changes,
  • or the agent makes a call the cassette never recorded.
import { CassetteMissError } from '@nkwib/tapedeck';

await expect(
  withCassette('order-lookup.json', () =>
    runAgent({ model, prompt: 'a prompt that was never recorded' }),
  ),
).rejects.toBeInstanceOf(CassetteMissError);

To fix a real miss, re-record and review the diff:

ANTHROPIC_API_KEY=sk-... npx tsx scripts/record.ts
git diff cassettes/order-lookup.json
git add cassettes/order-lookup.json && git commit -m "re-record order-lookup cassette"
A miss is a signal, not a flake
CassetteMissError means the request your code makes today no longer matches the committed fixture: re-record, read the diff, commit. See why a miss throws in CI for the reasoning.

There is a second, quieter kind of drift this recipe's replay gate cannot catch: the model changing its answer to the same recorded request. replay never calls the model, so it can't notice. If you want that check, run tapedeck's compare mode on a schedule (not on every push, since it calls the live API and costs tokens): npx tapedeck compare pnpm test runs your suite with CASSETTE_MODE=compare, compares each live response against its cassette on tool-call trajectory, finish reason, and text, and propagates the child process's exit code, so a scheduled workflow can fail on drift the way the weekly cron in tapedeck's own CI does. See the drift detection section of the guide.

Full runnable example

nkwib/agent-ci-zero-api is a complete, runnable template of this exact recipe: one agent with a lookupOrder tool, generateText with stopWhen: stepCountIs(3), a cassette recorded against a scripted mock model (so the repo works with no API key out of the box), a CassetteMissError drift test, and a record:real script for re-recording against the live Anthropic API. Its ci.yml is npm ci && npm test with no secrets configured, which is the whole point.