The pivot from “we need to get everyone in the org using AI” to “we need to figure out how to get a handle on our token spend” happened quickly.
If usage has scaled across your R&D department and beyond, you're now tasked with budgeting and governing it. But that's tough when AI spend is more opaque than any other line on your software bill.
Every month you get bills from your LLM providers that only show the simple math of total consumption against per-token costs. You can dig into your dashboards for model usage information, but that doesn’t help you figure out what work drove the token usage or whether it was worth the cost.
I see this topic come up regularly in the Braintrust forum for F Suite members, and it’s a common discussion topic at live events. How can you improve AI token management when usage is so opaque?
There’s no one right answer to this problem, so I pulled together ideas and tips from F Suite members and others to help you find a solution that fits your needs.
1. Start Tagging AI Spend with Unique Attributes
The starting point for effective AI token management is proper tagging of the spend. Without it, you have no way of tracing costs back to people, teams, or tasks.
Assign every workflow an owner and a home in your budget before you try optimizing token spend. At a minimum, track:
Who owns the workflow
Which function or department it belongs to
Whether it's product-facing or internal
What P&L line it hits
Which model and provider are running it
Without that structure, you’ll always be guessing at the breakdown of token spend in your LLM invoices.
Think about it the same way you think about corporate credit cards. Handing everyone the same card makes for an easy rollout and a miserable reconciliation process later. Split spend by task or department from the start, and attribution gets much easier. The same logic applies to AI access. Assign API keys per department or per task instead of letting usage pool under one shared account.
2. Create Room for Internal Experimentation with AI
Not every dollar of internal AI spend needs to prove its ROI immediately. Treating it that way will make people afraid to use the tools that you’ve asked them to adopt. You’ll hurt team morale while missing out on the AI-powered productivity.
Treat early internal usage as a learning cost instead of a governance problem. It’s inevitable that people experimenting with AI for the first time will waste tokens. They’ll run tasks that don’t need AI at all, retry prompts that were poorly framed, or let an agent loop longer than it should have.
Focus more on the tail end of that experimentation when processes become standard workflows. You don’t want a poorly designed agentic workflow left unattended overnight to rack up tens of thousands of dollars in token spend before anyone catches it (a real example I heard from another F Suite member). Put limits in place for standard workflows so mistakes become small errors rather than major budget outliers.
Seat-based plans can help here too. A flat monthly rate per user tends to be cheaper than API billing for people who are still discovering what the tool is good for, and it removes the anxiety of watching a meter run while you experiment.
Save the tighter, usage-based controls for workflows that have moved past exploration and into production.
3. Route Tasks to the Cheapest Possible Model
Frontier models aren't automatically the right choice for every task (even if the LLM providers move you to the latest model by default).
BCG's research on managing AI costs puts routing among the highest-impact levers available, estimating that sending each task to the cheapest model that meets the required intelligence can improve bill efficiency by 3-8%, with output length and reasoning settings adding another 5-10% on top of that. Most default settings overprovide, sending simple lookups and routine drafts to the same model handling your hardest reasoning work.
Test the assumption in your own tasks.
Run the same workflow on your default frontier model and a lighter one, then compare the output. Plenty of finance and back-office work doesn't show a meaningful quality gap between the two, even when engineering workflows genuinely need the frontier option.
The savings compound fast once you know where the line actually sits for your AI use cases.
4. Consider Architecture as Well as Model Choice
Model choice has a clear impact on AI costs because each one comes with a different per-token cost. But it’s often just as important to consider workflow architecture to manage token spend.
A simple chat interaction is the cheapest way to use AI. Ask a question, get a response, done. Agentic workflows, where the model plans steps, calls tools, and checks its own work, run closer to 4 times that cost. Multi-agent setups, where several models coordinate on a task, can run 15 times higher than simple chat, driven by how many loops it takes to reach an answer rather than which model is doing the work.
Engineering teams may instinctively reach for the most sophisticated architecture available, and multi-agent systems are genuinely impressive. Some tasks earn that complexity. Plenty of others get built as multi-agent workflows when a single agent call, or even a simple chat prompt with the right context, would have done the job just as well.
Ask what a new agentic or multi-agent build would cost as a simpler workflow before you approve it, and whether the added complexity solves a problem the simpler version couldn't.
It’s an opportunity to partner more closely with the R&D department and collaborate on token management.
5. Measure for Business Outcomes, Not Token Efficiency
Tokens are an input, not a result. Measuring how efficiently you're using them tells you nothing about whether the work they produced was worth doing.
This idea is clearest in the R&D department. Lines of code written is an activity metric. Lines of code shipped and still in the codebase after review is an outcome metric. The same workflow can look wildly productive or barely worth running, depending on which one you're tracking.
Instead of watching token counts, watch what token counts are supposed to produce:
Cycle time on a specific process
Support tickets closed without escalation
Hours of manual work removed from a workflow
Cost per completed outcome, not cost per request
A cheap workflow that produces nothing useful is still a bad deal. An expensive one that reliably delivers the outcome you need might be the highest AI ROI you have.
6. Look at Your Infrastructure Bill as an Element of AI Token Management
The bill from your LLM providers isn’t the only factor in the token management conversation. As engineers get more productive with AI tools, they ship more, test more, and spin up more environments, and your AWS, Azure, and database costs climb right along with your token spend.
One F Suite member put a rough number on it: for every $2 to $3 spent on AI, expect another $1 to $1.50 in infrastructure spend.
A team writing and shipping code faster is also running more builds, more tests, and more deployments, and each of those touches infrastructure you're already paying for. If your forecasting only accounts for the token line, you'll miss a real chunk of what AI adoption is actually costing you.
But there’s no clean lever to pull here yet. Hardware investment can bring the ratio down over time, but it takes 12 to 18 months to source at scale, which isn't a near-term fix.
The more useful step right now is simply forecasting the two together instead of separately, so the infrastructure increase doesn't show up as a surprise a quarter after the token increase already did.
Staying on top of AI Token Management as Cost Structures Continue to Evolve
The tactics in this piece make sense against today's pricing and today's usage patterns. But LLM providers keep changing how they price tokens, models keep getting cheaper and then get replaced by more expensive ones, and the bill will likely stay opaque from the vendor side.
Your job is less about definitively solving AI token management today and more about treating it as something you own and revisit.
Build the system, then keep checking it as pricing and usage shift underneath it.
You don't have to figure this out alone either. The finance leaders furthest ahead on this are the ones comparing notes with peers doing the same work.
If you want in on those conversations, join The F Suite as a member.