Repository navigation
A hard dollar cap across a whole AgentChat conversation, via the model client #8316
domondi1
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
One
team.run(...)can turn into many model calls: every turn of a RoundRobinGroupChat or Swarm, each agent's tool-calling loop, reflection, a termination check that keeps going longer than expected. Token usage is visible per call, but nothing stops the whole conversation at a dollar amount, and a loop between two agents can spend a lot before a human notices.OpenAIChatCompletionClientpassesdefault_headersstraight to the underlyingAsyncOpenAI, so you can put a hard per-conversation budget in front of the model without changing your agents: point the client at an OpenAI-compatible gateway that enforces the budget, and tag every call with the conversation's id and budget. Share one client across the agents and they all draw from the same budget.Run the gateway next to your app (it uses your provider key):
Afterwards:
shows how many calls the conversation made, the total tokens, and what it cost.
Each call reserves its worst-case cost (input +
max_tokens) before it's sent, so when agents make calls concurrently (parallel tool calls, a Swarm) they can't overspend together. Once the conversation's budget can't cover the next call, it's refused with HTTP 402 before it reaches the provider. In AgentChat that comes back as anAPIStatusError(status 402) and ends the run (wrapped in aGroupChatError, so catch it where you callteam.run). The two headers aren't forwarded to the provider.What I tested: autogen-agentchat / autogen-ext 0.7.5, inferrail 0.4.13, against a local stub upstream standing in for OpenAI. A two-agent RoundRobinGroupChat with a generous budget ran to its message limit and all calls were attributed to the one id. With a budget below the conversation's cost, the first calls succeeded and the next was refused before reaching the upstream, ending the run. I haven't run it against the real OpenAI API from AutoGen yet, so reports are welcome.
Caveats: chat completions only (keep embeddings/other clients separate); set
max_tokens; one gateway (the budget is a SQLite file on its host); and the model needs a price in Inferrail (gpt-4o-mini,gpt-4.1-mini,gpt-4.1are built in,inferrail modelslists others, and you can add your own). To budget per agent instead of per conversation, give each agent a client with its own work-id.Setup details: guide. I maintain Inferrail (open source, Apache-2.0), so take the suggestion with that in mind. Curious how people here currently bound the cost of a long multi-agent run.
All reactions