What does a sales agent built on Claude and Gong transcripts cost at production scale?

- Your pilot economics are honest. An agent on Claude reads Gong transcripts and CRM and writes a prep note for very little. You cached the long prompts, which is good engineering.
- But the cost that grows is the recommendations reps never run. Adoption is the denominator of your unit cost. At a one-in-three run rate, each recommendation a rep acts on costs three times the invoice.
- And your guidance multiplies the rest. Every agent and skill carries its own copy, and every strategy change rewrites them all. On an illustrative 400-rep estate, engineer time is most of the real cost.
- That guidance belongs in one place: a behavior graph. It holds your best reps' judgment, elicited through FDE-run interviews and RLHF, including what never shows up in a transcript. It adds your CRO's authorized strategy, versioned as it moves.
- So keep your harness and buy the judgment. Your agents query one governed Behavior instead of re-sending the playbook. Then measure cost per useful action.
Your sales agent runs on Claude. It reads Gong transcripts and CRM, writes the prep note, and drafts the follow-up. Ten reps ran the pilot. The token bill was a rounding error.
You did the sensible things early. Long prompts are cached. Simple asks go to a smaller model. Anthropic reports that caching cuts cost by up to 90% on long prompts, and your bill shows it.
Then finance asks for the production number.
Run the pilot at production volume
Ten reps become four hundred, so you multiply the bill by forty. That number is wrong, and not because of the price per token.
Your agents don't just read the deal. They also carry your guidance on how to sell: the competitor section, the pricing rules, the discovery standard.
Say five agents and twelve skills, each holding its own slice of that guidance.
Caching makes re-sending that text cheap. It does nothing about rewriting it. The cache only hits on an identical prefix. In week three, your CRO changes the pricing motion for the new AI add-on. The guidance is now wrong in seven skills. You edit seven files. Every cached prefix holding the old text misses. Evals re-run across all twelve, because you can't tell which answers moved.
That is one change. By our rough estimate, your guidance turns over every three to six months.
The pilot measured one term. Production has three:
- The tokens to generate a recommendation. The invoice, kept small by caching.
- The work to keep its guidance current. Edits, cache misses, eval re-runs, your hours.
- The share of recommendations a rep runs. The denominator.
Your spend dashboard tracks the first. It has no field for the other two.
Watch one recommendation go unused
A rep is on a renewal in APAC. The buyer says your biggest competitor offered a bundled AI add-on at no extra cost. She opens your agent on the second screen.
The agent loads the whole competitive section, plus three past calls, to answer a question about one rival. Thousands of tokens in, one recommendation out.
The answer is accurate and "insighty": a sharp read of the deal, with no move she can run. She closes the tab and pings a colleague.
The tokens were billed. Nothing was run. Ameeth Dubey at Atomicwork saw this on an earlier build:
"This is what we tried to do with ChatGPT. The only problem is adoption."
Ameeth Dubey, Atomicwork
Which term actually moves the total?
Bisect the cost. Generation is not the failing step. The call was cheap, and the answer was right.
The context layer is not the failing step either. Foundation Capital called context graphs AI's trillion-dollar opportunity in December 2025. Whether you run one today or are still choosing, keep that layer.
The break is one step later, when the rep decides whether to act. If she runs one recommendation in three, each one she acts on costs three times the invoice. Maintenance widens the gap every time strategy moves.
So you fix the output with more context. Each fix adds tokens to every call. None makes her more likely to act.
Root cause: the judgment was never in the data
Your agent reasons over transcripts, CRM, and email. It knows only what someone said on a recorded call. The behaviors were never broken down, because nobody has to write down a reflex.
Each of the three terms breaks down what is missing from that data:
- Generation grows because the agent compensates. With no move to hand over, it hauls in more context.
- Maintenance grows because the guidance is hand-written into files, so every change is a sweep across copies.
- The denominator shrinks because the rep checks what the data can't show. Does this fit her deal? Is it what her CRO wants this quarter? Has a rep she trusts won with it?
Chitresh Yadav runs that check live:
"If I hear something like our competitors, Fortinet or Palo Alto or Zscaler, I'm thinking whether my rebuttal to those were inline to our battle card."
Chitresh Yadav, VP and Global Head of Sales Engineering, Versa Networks
No retrieval budget recovers what was never logged. You are paying tokens to search for judgment that isn't there.
The fix: hold the judgment in one place
A behavior graph sits on top of the context layer and holds three inputs in one model specific to your company:
- Your best reps' judgment, elicited through FDE-run interviews
- Your CRO's forward-looking, authorized strategy
- The entity context of your past deals
It produces behaviors: the move to run for this motion, role, product, and competitor. Your agents query it for the one Behavior a deal needs, instead of carrying the playbook in every prompt.
Reps run those Behaviors, and running them is adoption: the denominator from the load test.
Your agent recommends too. Both systems recommend a course of action. The difference is upstream: one reasons from patterns in past data, the other from your reps' judgment and your CRO's authorized strategy.
Now replay the renewal. Same buyer, same bundled add-on. The agent asks for the move. One Behavior comes back with a receipt from the deal where your best rep ran it and won. She runs it.
Then replay the week-three pricing change. It lands as one versioned update, approved by the CRO. Your seven skills stay as they are.

Where Zime fits. That layer is what Zime ships.
- Zime runs on Claude, so the harness you built stays yours.
- Your CRO approves every strategy update before an agent sees it.
- It goes live in 7 days on your own data.
- The adoption bar it is held to: 80% of eligible reps within 90 days.
The first four rows are your harness today. The last four are where guidance has to change.

The bold row carries it. You can version guidance yourself; you cannot mine judgment that was never in the data.
Build the single source yourself?
You could build that single source yourself.
The question is what you sign up to maintain. The Signal profiled a CRO who runs his revenue org from Claude Code. Eight months in, the judgment is still hand-built: a Sales Bible from 409 calls of one rep. He did not mine it. He sat with the rep and wrote it down.
That is one rep. Your estate needs a version per role, competitor, product and region, each going stale on the reset clock above. You would rebuild the copies you meant to stop paying for. The full in-house math is in Build vs Buy Sales AI and Your skill.md Is Killing Your GTM AI.
Own the harness. Buy the judgment.
Put adoption in the unit cost
Your spend dashboard is green because it graded the invoice. Put all three terms into one number instead:
cost per acted recommendation = (generation + maintenance)/recommendations reps run
Run it on an illustrative estate. Four hundred reps, five recommendations a day, 250 selling days: 500,000 recommendations a year. Guidance changes about once a month. Prices are Claude Sonnet 5.5 list prices from Anthropic's pricing page, October 2026. Swap in your own numbers.

The invoice is about a quarter of the total. Engineer time is most of the rest, and it never shows on the spend dashboard.
Now divide by the share of recommendations reps run.

The 35% and 60% points match the modelled execution rates below. Every point of adoption you lose raises the cost of the recommendations you keep.
Below are three dimensions in a reference scenario built on a context layer, against Zime's modelled scenario and targets.

Modelled end to end, that is 2.4X as many useful actions performed and 58% lower cost per useful action.
Is a well-cached prompt ever enough?
Honestly, sometimes it is. One product, fewer than twenty reps, a strategy that changes once a year: a cached prompt and a small eval suite will serve you. It stops holding once the guidance changes faster than you can edit the copies.
What changes once the judgment has one home
At MyAdvice, checklist adoption rose from 66% to 87%, and fatal objections fell by 89%.
Martin Mackay, CRO at Versa Networks, said what he wanted:
"I wanted Zime to understand my playbooks and coach reps. No other tool does this."
Martin Mackay, CRO, Versa Networks
Versa reported a 10% increase in win rates, an outcome HFS Research cited in its July 2026 profile. The harness stayed. The judgment stopped living in twelve places.
Stop re-sending the judgment.
The agents, skills, and workflows you built on Claude stay yours. Zime adds the layer underneath: judgment elicited from your best reps, your CRO's authorized strategy, and Behaviors your agents query instead of carrying the playbook. Live in 7 days on your own data, enough to price one Behavior against your cost per useful action.
FAQ
Build a sales agent with Claude and Gong transcripts vs buy a sales execution platform: cost, governance, and adoption compared? Keep the build and buy the judgment layer. On cost, the build is cheap per call but grows with every copy of your guidance and every strategy change. On governance, you own every edit and eval re-run. On adoption, reps run what fits judgment your data never held, which a behavior graph holds once.
Isn't token cost a solved problem with prompt caching? Per call, mostly yes, and every production harness should cache. But sales guidance changes all quarter, so the cache keeps emptying. And caching can't raise the share of recommendations a rep runs, which is what sets your real unit cost.
Doesn't my agent already tell reps what to do? Yes. Both systems recommend a course of action. The difference is upstream: one reasons from patterns in past data, the other from your reps' judgment and your CRO's authorized strategy.
Are the cost figures in this post measured results? No. The cost-per-useful-action, execution, and stale-action figures are modelled projections and targets from a reference scenario, not customer results. Customer outcomes here are reported separately.



