Multimodal AI

Models that read images, watch video, hear audio, and generate media back.

Sign in to track your progress across sections.

14 articles

Multimodal AI

Listed inMultimodal AIMultimodal AIon

Images are billed as input tokens and share the context budget. Plus the provider’s own list of what vision reads badly.

Beginner4 min
#multimodal

Text to Speech

Listed inText to SpeechMultimodal AIon

Uncompressed formats start playing sooner, streams fail after playback begins, and disclosing that the voice is synthetic is a policy requirement.

Intermediate4 min
#multimodal
#audio

Image Generation

Listed inImage GenerationMultimodal AIon

Rewriting the prompt gives you a different picture, not an edit of the one you had. Since no provider promises the same prompt returns the same image, store the result along with the prompt, the model and the interaction ID.

Intermediate3 min
#vision

Nano Banana API

Listed inNano Banana APIMultimodal AIon

The nickname is the documented product name for Gemini's native image generation. Multi-turn conversation is the recommended way to change an image, reference budgets are capped by category and in total, and every result carries a SynthID watermark.

Beginner3 min
#api