Skip to chapter content

Chapter 7 5 min read

Don’t Blow the Budget — Max Tokens

What a token is; cost & length control

Temperature controls how the model chooses its words. This chapter is about a different dial entirely: how much it’s allowed to write, and what that costs.

What a token actually is

A token isn’t a word, though it’s close enough for rough estimates. In English, one token works out to roughly four characters, or about three-quarters of a word, so “WordPress hooks are powerful” comes out to around five or six tokens, not four. Other languages tokenize differently, and code, URLs, and unusual punctuation tend to use more tokens per character than plain English prose does.

Every AI provider bills by tokens, and both directions count: the prompt you send in and the response you get back. This matters for two separate reasons, and this chapter is really about both of them: how much you pay, and how long the answer is allowed to run before something has to give.

Capping the response with using_max_tokens()

->using_max_tokens( 3000 )

This caps how many tokens the model is allowed to generate for that call. What happens if the model needs more room than that to finish answering is the part worth getting right, because it’s not what most people assume the first time they meet this setting.

What actually happens when the cap gets hit

On a lot of raw provider APIs, hitting a token limit just means you get truncated text back, and it’s on you to notice the cutoff yourself. The AI Client doesn’t do that. If the model hits the ceiling before it’s done, you don’t get a half-finished sentence, you get a WP_Error, telling you plainly that generation was stopped by the token limit. Here’s the actual message from a real test against OpenAI, hit while writing this chapter:

Generation stopped due to token limit with reason "max_output_tokens".

That’s a deliberate design choice, and a sensible one once you think about what a silently truncated response could mean for the rest of this book. A cut-off sentence is merely annoying. A cut-off JSON response, the kind Chapter 8 builds toward, isn’t just incomplete, it’s invalid JSON that json_decode() can’t parse at all. Failing loudly here means you always know when an answer is incomplete, instead of finding out later because something downstream broke on malformed data.

Seeing both outcomes side by side

Create chapter-07-max-tokens.php inside includes:

<?php
/**
* Chapter 7: Don't Blow the Budget, Max Tokens
* Usage: add [ai_course_ch7] to any page or post to see the output.
*/
if ( ! defined( 'ABSPATH' ) ) {
exit; // No direct access.
}
function ai_course_ch7_max_tokens() {
$prompt = 'Explain how WordPress hooks work, in detail, for a developer who has never used them before.';
$short = wp_ai_client_prompt( $prompt )
->using_max_tokens( 30 )
->generate_text();
if ( is_wp_error( $short ) ) {
$short_result = 'Generation failed: ' . esc_html( $short->get_error_message() );
} else {
$short_result = wp_kses_post( $short );
}
$long = wp_ai_client_prompt( $prompt )
->using_max_tokens( 3000 )
->generate_text();
if ( is_wp_error( $long ) ) {
return 'Could not generate the second example: ' . esc_html( $long->get_error_message() );
}
$output = '<p><strong>Max tokens: 30</strong><br>' . $short_result . '</p>';
$output .= '<p><strong>Max tokens: 3000</strong><br>' . wp_kses_post( $long ) . '</p>';
return $output;
}
add_shortcode( 'ai_course_ch7', 'ai_course_ch7_max_tokens' );

Notice the first block deliberately doesn’t return early on error, the way every other shortcode in this book has so far. It captures the failure message into $short_result and keeps going, specifically so you can see both outcomes on the same page: a call that failed because 30 tokens wasn’t enough room, next to one that succeeded because 3000 was.

Add [ai_course_ch7] to a page and load it. The first line shows the failure message. The second shows a full, finished explanation. There isn’t a middle ground here, no partial paragraph to read, just “didn’t fit” or “fit fine.”

Note

Worth knowing: the error message you just saw, “stopped due to token limit”, isn’t worded the same on every provider, but the underlying cause is more common than it looks.

On Gemini, you may see a different-looking error instead, something like “Unexpected Google API response: Missing the ‘candidates[0].content’ key.”

On OpenAI, if you’re on one of the reasoning models (the GPT-5 family, o1, o3), the failure above is likely happening for a related reason. Reasoning models spend part of their token budget on invisible internal reasoning before writing anything you actually see. A cap that would comfortably fit a normal model’s response can get eaten entirely by reasoning tokens, with nothing left for the visible answer.

Switching to a non-reasoning model for this chapter (an early preview of using_model_preference(), properly covered in Chapter 19) tends to make token usage far more predictable, if you’d rather not chase a bigger number.

That’s also why this chapter’s cap on the successful call is set as high as 3000. Earlier drafts of this example tried 400, then 900, and both still failed against a real reasoning model, the invisible reasoning-token overhead on these models isn’t a small, predictable tax, it can be substantial and it varies between runs.

If you’re on a non-reasoning model, you’ll likely find you need far less than 3000 to succeed. Whichever provider you’re on, the fix is the same, and it’s the subject of the rest of this chapter: don’t trust a number until you’ve actually tested it against your own setup.

So how do you pick a number that won’t just fail

Since there’s no graceful degradation to lean on, the number you choose has to comfortably fit the answer you’re actually expecting, not just roughly gesture at it. On a non-reasoning model, a short excerpt or a single tagline can live comfortably under 100 tokens, a full blog post draft needs a couple thousand.

On a reasoning model, add a generous, unpredictable buffer on top of that. Invisible reasoning tokens can eat far more of the budget than the visible answer itself, and how much varies run to run, not just model to model.

If you’re not sure, generate a real example first with a high cap, see roughly how many tokens a good answer actually takes, and set your real limit a reasonable margin above that, not right up against it. This chapter’s own numbers took three attempts to land on something that actually worked reliably, that’s normal, not a sign you’re doing something wrong.

This is also exactly why is_wp_error() earns its place in every single function in this book, not just as a defensive habit against network failures, but because “the model needed more room than I gave it” is a completely normal, expected failure mode that your code needs to handle every time it sets a cap at all.

Try it yourself

Pick something you’d actually build, a meta description generator, a comment reply suggester, whatever fits your own project. Write the prompt, then guess a token limit before testing it. Run it. If it fails, read the error, raise the cap, and try again until you find a number that reliably succeeds with some headroom to spare. That process, guess, fail, adjust, is the actual skill here, not landing on the right number on the first try.

This book is created with Chapterwright