Chapter 14
Text and Images Together, Multimodal Output
Mixed-part responses
Every chapter in this module so far has asked for one kind of output at a time: text, or an image. This chapter is about asking for both in the same response, genuinely mixed, not two separate calls stitched together afterwards.
A different shape of request
Everything before this chapter called a specific generation method, generate_text(), generate_image(), matched to a specific kind of output. Multimodal output doesn’t work that way, since the response itself is a mix; there’s a new, more general method for it:
use WordPress\AiClient\Messages\Enums\ModalityEnum;
$result = wp_ai_client_prompt( 'Create a recipe for a chocolate cake and include photos for the steps.' )
->as_output_modalities( ModalityEnum::text(), ModalityEnum::image() )
->generate_result();
ModalityEnum::text() and ModalityEnum::image() are the two modalities this book actually uses, but they’re not the only two that exist. Audio and video are modalities too, in principle as_output_modalities() could mix those in the same way. This book doesn’t cover them here for the same reason Chapter 17 holds off on speech and video generation entirely, no provider currently supports them, so there’s nothing to test. Module 10 revisits this exact method once that changes, rather than treating it as new all over again.
as_output_modalities() tells the model which kinds of content you’re willing to receive, text, image, or both together. generate_result() is new too, it’s not tied to one modality the way generate_text() or generate_image() are, since the response might contain either, or a mix of both in sequence.
Reading a mixed response
The result isn’t a plain string or a single File object this time, it’s a message made up of parts, and each part could be text or a file. Each part has a getType() method returning an enum, plus two nullable getters, getText(): ?string and getFile(): ?File. Only one of those two is ever populated on a given part, so checking which one isn’t null tells you exactly what you’re holding:
foreach ( $result->toMessage()->getParts() as $part ) {
$text = $part->getText();
$file = $part->getFile();
if ( null !== $text ) {
echo wp_kses_post( $text );
} elseif ( null !== $file && $file->isImage() ) {
echo '<img src="' . esc_url( $file->getDataUri(), array( 'data' ) ) . '">';
}
}
That’s the pattern this chapter actually uses from here on.
Worth a quick honest note here: the documentation showed a different approach at the time of writing, one that didn’t actually work when tested. This version came from checking the underlying code directly instead. File::isImage() wasn’t independently confirmed the same way getText()/getFile() were, but it’s a reasonable bet since it wasn’t the method that caused trouble, worth keeping an eye on if you hit a different error there.
Seeing it work
Create chapter-14-multimodal-output.php inside includes:
<?php
/**
* Chapter 14: Text and Images Together, Multimodal Output
* Usage: add [ai_course_ch14] to any page or post to see the output.
*/
if ( ! defined( 'ABSPATH' ) ) {
exit; // No direct access.
}
use WordPress\AiClient\Messages\Enums\ModalityEnum;
function ai_course_ch14_multimodal_output() {
$result = wp_ai_client_prompt( 'Write a short 3-step recipe for banana bread, and for each step include a photo illustrating it.' )
->as_output_modalities( ModalityEnum::text(), ModalityEnum::image() )
->generate_result();
if ( is_wp_error( $result ) ) {
return 'Could not generate this right now: ' . esc_html( $result->get_error_message() );
}
$output = '';
try {
foreach ( $result->toMessage()->getParts() as $part ) {
$text = $part->getText();
$file = $part->getFile();
if ( null !== $text ) {
$output .= wp_kses_post( $text );
} elseif ( null !== $file && $file->isImage() ) {
$output .= '<img src="' . esc_url( $file->getDataUri(), array( 'data' ) ) . '" style="max-width:100%; height:auto; margin: 12px 0;">';
}
}
} catch ( Throwable $e ) {
return 'Something went wrong reading the response: ' . esc_html( $e->getMessage() );
}
return $output;
}
add_shortcode( 'ai_course_ch14', 'ai_course_ch14_multimodal_output' );
Add [ai_course_ch14] to a page and load it. Instead of a wall of text or a single picture, you get something that reads like an actual illustrated recipe, a step written out, a photo for that step, the next step, another photo.
Notice the try/catch ( Throwable $e ) wrapped around the loop. is_wp_error() only catches errors the AI Client reports itself, not a genuine PHP error like calling a method that doesn’t exist.
That kind of error is a Throwable, not an Exception. A plain catch ( Exception $e ) wouldn’t catch it, Throwable is the one that does.
Note the array( 'data' ) on esc_url(). The official documentation’s own example for this feature doesn’t include it, same gap Chapter 12 ran into with a plain image. Left off here, any image part would break silently the exact same way. Consistency matters more than matching the docs exactly when the docs have the same gap we already found and fixed.
Warning, if you’re on OpenAI or Gemini: this chapter generates images as part of a mixed response, so the same provider-specific issues from Chapters 12 and 13 (OpenAI’s ~30 second timeout, Gemini’s restrictions on generating more than one image per call) can show up here too, on the image parts specifically. If the text comes through but images don’t, or the whole thing fails, check those chapters’ Warnings first before assuming this chapter’s code is different somehow, it isn’t.
A third, different failure is also possible on OpenAI specifically: “No models found that support text_generation for this prompt.” Worth telling apart from the others, this one happens before any request even reaches OpenAI’s servers.
The AI Client checks locally which of your connected models are registered as supporting the combination you asked for, text and image together in one response. If none are, it fails immediately rather than making a call that would fail anyway.
This doesn’t mean OpenAI is incapable of this, GPT-4o and newer models genuinely do support native combined text-and-image output. It more likely means the connector plugin’s current model catalog doesn’t yet have a model listed with that specific combined capability, the kind of gap that tends to close as the plugin matures, not a permanent limitation of OpenAI itself.
Try it yourself
Write your own prompt that naturally wants both text and images together, a short how-to, a product description with a mockup, a story with an illustration. Run it through this chapter’s shortcode and see how the model splits its response into parts. Count how many text parts and how many image parts come back, it won’t always be a perfectly even alternation the way the recipe example suggests.