Chapter 11
Fine-Tuning Generation, Top-P, Top-K, and Stop Sequences
Alternative sampling controls, plain-English, when to reach for these vs. temperature
Temperature has been the dial for the whole of this module, and for most of what you’ll build, it’s still the only one you need. Three more settings exist underneath it, worth knowing about even if you reach for them less often: using_top_p(), using_top_k(), and using_stop_sequences().
Top-p and top-k narrow the options before temperature touches them
At each step, the model is choosing its next word from a large set of candidates, each with a likelihood attached. Temperature, from Chapter 6, controls how willing the model is to pick a less-likely candidate from that set. Top-p and top-k work earlier in the process: they shrink the set itself, before temperature gets involved.
using_top_p() keeps only the smallest group of candidates whose combined likelihood adds up to the number you give it. A top-p of 0.9 means “only consider candidates that together make up the top 90% of likely options,” cutting off the long tail of unlikely, weird word choices entirely. A top-p of 1.0 means no cutoff at all, the full set stays in play.
using_top_k() is blunter. Instead of a percentage, it keeps a fixed number of the most likely candidates, regardless of how their likelihoods add up. A top-k of 40 always considers exactly 40 candidates, no more. Support for this one varies more between providers than top-p does, some ignore it if the underlying model doesn’t expose that control, so treat it as available but not guaranteed everywhere.
In practice, most people reach for temperature alone, and it’s enough for nearly every case in this book. Top-p is worth adding when you want creative output that still avoids genuinely strange word choices, pairing a moderate-to-high temperature with a top-p around 0.9 is a common combination. Top-k comes up less often and is more provider-dependent, so I’d treat it as something to know exists rather than something to reach for by default.
Stop sequences: telling the model exactly where to stop
using_stop_sequences() is the most immediately useful of the three. You give it a string, and the moment the model is about to generate that string, it stops, right there, nothing after it.
->using_stop_sequences( '4.' )
Seeing a generation cut off cleanly
Create chapter-11-fine-tuning-generation.php inside includes:
<?php
/**
* Chapter 11: Fine-Tuning Generation, Top-P, Top-K, and Stop Sequences
* Usage: add [ai_course_ch11] and/or [ai_course_ch11_topp] to any page or post.
*/
if ( ! defined( 'ABSPATH' ) ) {
exit; // No direct access.
}
function ai_course_ch11_stop_sequences() {
$prompt = 'List 5 WordPress security best practices as a numbered list, formatted like "1. ...", "2. ...", and so on.';
$full = wp_ai_client_prompt( $prompt )->generate_text();
if ( is_wp_error( $full ) ) {
return 'Could not generate the first example: ' . esc_html( $full->get_error_message() );
}
$stopped = wp_ai_client_prompt( $prompt )
->using_stop_sequences( '4.' )
->generate_text();
if ( is_wp_error( $stopped ) ) {
return 'Could not generate the second example: ' . esc_html( $stopped->get_error_message() );
}
$output = '<p><strong>No stop sequence:</strong></p><p>' . wp_kses_post( nl2br( $full ) ) . '</p>';
$output .= '<p><strong>Stopped at "4.":</strong></p><p>' . wp_kses_post( nl2br( $stopped ) ) . '</p>';
return $output;
}
add_shortcode( 'ai_course_ch11', 'ai_course_ch11_stop_sequences' );
Add [ai_course_ch11] to a page and load it. The first block gives you all 5 security tips. The second stops the instant the model is about to write “4.”, leaving you with a clean 3-item list and nothing trailing after it. Notice it stops before the number, not after, the sequence itself never makes it into the output.
This is genuinely useful whenever you need output to end at a precise structural boundary, not just a rough length, the way using_max_tokens() gives you. If a prompt is generating a single list item, a single paragraph, or one turn of a conversation, a stop sequence is meant to enforce that boundary instead of just hoping the model respects your instructions in the prompt itself. On most models, it does exactly that.
Warning: if you’re on OpenAI, “meant to” and “does” aren’t the same thing on every model. If the “Stopped at” block doesn’t actually stop, if you get a full 5-item list back again instead of a clean 3-item one, that’s not an error, and that’s exactly what makes it easy to miss. using_stop_sequences() can be silently ignored on OpenAI’s reasoning models (the GPT-5 family, o1, o3), the same family that ignores temperature and top_p.
Some reasoning models reject the stop parameter outright with a visible error (o1-preview has been documented doing exactly that). Others just ignore it with no error at all, the request succeeds, the model just doesn’t honor the boundary. The second case is worse to debug, since nothing tells you it failed, you have to actually read and compare both blocks of output to notice.
The workaround is the same as the rest of this chapter: switch to a different connected provider, or wait for Chapter 19’s using_model_preference().
Top-p, side by side
The same file also has a smaller example for top-p, since it’s harder to describe in words than to see:
function ai_course_ch11_top_p() {
$prompt = 'Describe a walk through an autumn forest in one sentence.';
$narrow = wp_ai_client_prompt( $prompt )
->using_temperature( 0.9 )
->using_top_p( 0.1 )
->generate_text();
if ( is_wp_error( $narrow ) ) {
return 'Could not generate the first example: ' . esc_html( $narrow->get_error_message() );
}
$wide = wp_ai_client_prompt( $prompt )
->using_temperature( 0.9 )
->using_top_p( 1.0 )
->generate_text();
if ( is_wp_error( $wide ) ) {
return 'Could not generate the second example: ' . esc_html( $wide->get_error_message() );
}
$output = '<p><strong>Top-p 0.1:</strong><br>' . wp_kses_post( $narrow ) . '</p>';
$output .= '<p><strong>Top-p 1.0:</strong><br>' . wp_kses_post( $wide ) . '</p>';
return $output;
}
add_shortcode( 'ai_course_ch11_topp', 'ai_course_ch11_top_p' );
Temperature is fixed at 0.9 in both calls on purpose, so top-p is the only thing changing. Add [ai_course_ch11_topp] to a page and load it. The 0.1 version tends to stay close to safe, predictable wording even at a high temperature, because the candidate pool it’s choosing from was already narrow. The 1.0 version has the full range available and can wander further.
Warning, if you’re on OpenAI: the same issue that can show up in Chapter 6’s temperature demo applies here too, and for the same reason.
OpenAI’s reasoning models (the GPT-5 family, o1, o3) don’t support top_p any more than they support temperature. If both [ai_course_ch11_topp] runs come back identical, that’s almost certainly why.
The AI Provider for OpenAI connector doesn’t currently expose a model picker, just an API key field, so switching to a different connected provider (Anthropic or Gemini’s standard models both support top_p normally) is the practical workaround for this chapter until Chapter 19 covers using_model_preference(), a code-level way to request a specific model.
Try it yourself
Take a prompt where the model tends to keep going past where you want it to, a list, a set of steps, a single Q&A turn, and add a stop sequence that matches whatever marks the start of the next unwanted section. Confirm the output ends exactly where you told it to, not just roughly where you asked in the prompt wording.