Skip to content

Early access is open. Request a spot

The agent arena

Where AI agents
earn production.

Clawsseum is the platform for building agents, wiring them to any model, tool or channel through one adapter contract, and making every release beat the last one before it reaches a customer.

  • Free to start
  • Bring your own keys
  • No token markup

support-resolver / arena

Live

Bout series 2481

v13 challenger vs v12 champion

suite support/refunds, 412 scenarios replayed from production

Win rate

61.3%

176 wins67 ties44 losses287 / 412 played
TranscriptVerdict
  • Asks for a manager after denialv12 wins
  • Refund over $500, needs approvalv13 wins
  • Partial refund for unused seatsv13 wins
  • Chargeback threat, hostile toneTie
  • Refund requested on day 34 of 30v13 wins
  • Duplicate charge after plan upgradev13 wins

Illustrative Arena run. A challenger replays real conversations against the current champion before promotion.

102 adapters for the stack you already run.

  • Anthropic
  • Gemini
  • Mistral
  • Llama
  • Zendesk
  • Stripe
  • HubSpot
  • GitHub
  • Linear
  • Jira
  • Notion
  • Postgres
  • Snowflake
  • WhatsApp
  • Gmail
  • Shopify
  • MCP
  • OpenAPI

Why Clawsseum

Demos are easy.
Production is a bloodsport.

Any team can get an agent to work once. Keeping it working across new models, more systems and real customers is where agent projects stall.

I

Models change monthly.

Every new model is better on average and worse somewhere specific. Teams pin old versions out of fear, or upgrade blind and find out from customers.

II

Integrations are glue.

Each CRM, ticket queue and chat channel gets wired by hand, with its own auth, retries and failure modes. The agent is ten percent reasoning and ninety percent plumbing.

III

Regressions are silent.

Agents get worse quietly. A prompt tweak fixes one flow and breaks three others, and nobody has a test that would have caught it.

The platform

Four parts.
One contract.

Forge builds the agent, Armory wires it to your systems, Arena proves it is better than what is live, and Gate runs it with guardrails. Use them together or adopt one at a time.

IForgeBuild

Agents as code, not as canvases.

Define agents in TypeScript or Python with typed tools, durable memory and plans you can read in a diff. Run them locally against real adapters or recorded fixtures.

  • Typed tools with zod or pydantic
  • Durable, resumable runs
  • Local runtime with hot reload
agents/support-resolver.ts
import { agent, tool, z } from "@clawsseum/forge";
import {
  anthropic, openai, zendesk, stripe, slack,
} from "@clawsseum/armory";

const refund = tool({
  name: "issue_refund",
  input: z.object({
    chargeId: z.string(),
    amount: z.number().positive(),
  }),
  sideEffect: true, // Gate pauses for approval
  run: (args, { adapters }) =>
    adapters.stripe.refunds.create(args),
});

export default agent({
  name: "support-resolver",
  model: anthropic("claude-sonnet-5-5"),
  fallback: [openai("gpt-5")],
  adapters: [zendesk(), stripe(), slack()],
  tools: [refund],
});
IIArmoryConnect

One adapter contract for everything.

Models, channels, apps and databases all speak the same typed contract. Swap Claude for another model, or Slack for Teams, in config. Every adapter gets least-privilege scopes by default.

  • Scoped credentials per adapter
  • Retries, rate limits and fallbacks built in
  • Any MCP server or OpenAPI spec
support-resolver / adapters6 healthy
  • AnthropicModel

    claude-sonnet-5-5

    primary
  • O

    OpenAIModel

    gpt-5

    fallback
  • S

    SlackChannel

    #support-escalations

    threads
  • ZendeskApp

    tickets:read, macros:read

    scoped
  • StripeApp

    charges:read, refunds:write

    approval
  • PostgreSQLData

    replica, billing schema

    read-only
$ claw adapters swap model anthropic:claude-opus-5-5 --arena
IIIArenaProve

Every release fights its predecessor.

Record real conversations as scenarios, replay them against a challenger, and judge both with rubric models and hard assertions. Versions earn an Arena rating instead of a vibe check.

  • Replay production traffic safely
  • Rubric judges plus deterministic checks
  • Promotion gates in CI
Arena rating, support/refunds6 versions
v8v9v10v11v12v131612

Win rate

63.8%

Ties

22.1%

Regressions

0 critical

IVGateRun

Ship with brakes, not just an engine.

A durable runtime that pauses side effects for approval, enforces budgets, redacts PII at the adapter boundary and traces every step down to the token.

  • Approval policies for risky actions
  • Budgets and kill switches per agent
  • OpenTelemetry traces
run 8f12 / support-resolver@v13Paused
  • zendesk.tickets.get
    118ms
  • model.plan
    842ms
  • stripe.charges.list
    204ms
  • model.decide
    611ms
  • stripe.refunds.create
    awaiting approval

Refund of $740.00 exceeds the $500 policy limit.

ApproveDenybudget $0.011 of $0.05

Workflow

From init to production in four commands.

The CLI, the SDK and the dashboard share one model of your agent. Whatever you do locally is exactly what runs in the Arena and behind the Gate.

  1. 01

    Scaffold an agent

    A typed project with an agent, a tools folder, a starter suite and a local runtime.

  2. 02

    Arm it

    Connect systems with OAuth or keys. Scopes are least-privilege until you widen them.

  3. 03

    Make it fight

    Replay recorded traffic against the current champion and judge every scenario.

  4. 04

    Send it through the Gate

    Deploy with approvals, budgets and tracing on. Roll back to the champion in one command.

~/acme/agents

$ claw init support-resolver --template support

created agents/support-resolver.ts

created arena/support/refunds.suite.ts (24 seed scenarios)

ok local runtime ready on http://localhost:7420

$ claw adapters add zendesk stripe slack

zendesk connected tickets:read macros:read

stripe connected charges:read refunds:write (approval)

slack connected #support-escalations

$ claw arena run --suite support/refunds --vs @v12

412 bouts 263 wins 91 ties 58 losses

win rate 63.8% critical regressions 0 cost -22%

ok gate passed: v13 is eligible for promotion

$ claw deploy --env production

ok support-resolver@v13 live in production (us-east)

approval policy refunds>500 -> #support-leads

budget $0.05 per run, kill switch armed

$

IIIThe Arena

Agents don't ship
until they win.

Evals that live in a notebook never block a release. The Arena sits in your CI and on your deploy path, so a worse agent cannot reach customers by accident.

01

Record

Capture real conversations and tool calls from production, scrubbed of PII at the adapter boundary, and promote the interesting ones to scenarios.

02

Replay

Run every scenario against the challenger and the champion side by side. Side-effecting adapters are sandboxed or served from recordings.

03

Rule

Rubric judges score outcomes, deterministic checks enforce policy, and humans review the close calls. The gate decides who ships.

Arena ratings

suite support/refunds, 412 scenarios

#VersionModelRatingStatus
1support-resolver@v13claude-sonnet-5-51612+32challenger
2support-resolver@v12claude-sonnet-5-51580champion
3support-resolver@v13-opusclaude-opus-5-51574+28experiment
4support-resolver@v11claude-sonnet-5-51544retired
5support-resolver@v13-gptgpt-51531-49experiment
pull request 1184checks
  • build

    Successful in 41s

  • unit tests

    312 passed

  • clawsseum / arena

    v13 beat v12: 63.8% win rate, 0 critical regressions

  • clawsseum / policy

    refund limits, tone, PII: all clear

Ready to merge

IIThe Armory

102 adapters.
Zero glue code.

Every adapter speaks the same typed contract: actions with schemas, scoped credentials, retries, rate limits and redaction hooks. Agents never see a raw API key, and every call is recorded for Arena replay.

Bring your own systems.

Point the Armory at an MCP server or an OpenAPI spec and get a typed adapter with an allowlist, scoped auth and recording for free. Internal services get the same guardrails as Stripe or Salesforce.

adapters/internal.ts
import { adapter } from "@clawsseum/armory";

// Any MCP server becomes a scoped, replayable adapter.
export const github = adapter.fromMCP(
  "npx -y @modelcontextprotocol/server-github",
  { scope: ["issues:write", "pulls:read"] },
);

// So does any internal API with an OpenAPI spec.
export const billing = adapter.fromOpenAPI(
  "https://billing.internal/openapi.json",
  {
    auth: adapter.auth.oauth2({ secret: "BILLING" }),
    allow: ["GET /invoices/*", "POST /credits"],
  },
);

IVThe Gate

Built for agents that touch real systems.

Refunds, account changes, outbound email. When an agent can act, the runtime has to be able to say no.

Security overview

Bring your own keys

Use your provider contracts and rate limits. We never mark up tokens.

Scoped credentials

Each adapter gets least-privilege scopes from a vault. Agents never hold raw secrets.

Approvals

Side-effecting tools pause for a human in Slack or the dashboard, by policy.

PII redaction

Sensitive fields are masked at the adapter boundary, before any model sees them.

Budgets and kill switches

Cap cost per run, per agent and per day. Stop an agent everywhere in one click.

Traces

Every step, prompt and tool call, searchable, with OpenTelemetry export.

Data residency

Run the control plane and traces in the US or the EU.

Your cloud

Self-host the runtime in your own VPC when data cannot leave it.

Use cases

What teams send into the arena.

The pattern is the same everywhere: an agent that reads from a few systems, acts on one, and has to be provably better than the last version before it is allowed to.

01
S

Support resolution

Resolve billing and account tickets end to end, with refunds above a threshold sent to a human.

gate resolution rate at least 85%, 0 policy violations

02

Engineering triage

Turn alerts and error spikes into deduplicated, assigned issues with a first diagnosis attached.

gate correct owner in 90% of bouts, no duplicate issues

03
S

Revenue operations

Keep the CRM honest: enrich accounts, log calls, draft follow-ups and flag stalled deals.

gate zero writes outside allowed fields, human send on all email

04
N

Finance operations

Match invoices to payments, chase missing receipts and prepare month-end reconciliations.

gate 100% match accuracy on the golden set before any write

Early access

Enter the arena.

We are onboarding a small number of teams each week. Tell us where to reach you and what your first agent should do.

Weekly cohorts. No spam, one onboarding email.

Need a security review, a custom adapter or your own cloud? Talk to the team