Multimodal AI
Listed inMultimodal AIMultimodal AIon
Images are billed as input tokens and share the context budget. Plus the provider’s own list of what vision reads badly.
Models that read images, watch video, hear audio, and generate media back.
Sign in to track your progress across sections.
14 articles
Listed inMultimodal AIMultimodal AIon
Images are billed as input tokens and share the context budget. Plus the provider’s own list of what vision reads badly.
Listed inText to SpeechMultimodal AIon
Uncompressed formats start playing sooner, streams fail after playback begins, and disclosing that the voice is synthetic is a policy requirement.
Listed inAudio ProcessingMultimodal AIon
Transcribing first meters you by file size and sending the audio meters you by tokens. An hour of audio is about 115,000 of them.
Listed inLangChain for Multimodal AppsMultimodal AIon
Message content becomes a list of typed blocks. The shapes for images, audio and files, and when the abstraction is worth its dependency.
Listed inImage GenerationMultimodal AIon
Rewriting the prompt gives you a different picture, not an edit of the one you had. Since no provider promises the same prompt returns the same image, store the result along with the prompt, the model and the interaction ID.
Listed inNano Banana APIMultimodal AIon
The nickname is the documented product name for Gemini's native image generation. Multi-turn conversation is the recommended way to change an image, reference budgets are capped by category and in total, and every result carries a SynthID watermark.
Listed inMultimodal Use CasesMultimodal AIon
Document extraction, visual QA, accessibility, and media generation.
Listed inOpenAI Vision APIMultimodal AIon
Sending images in a request and controlling detail and cost.
Listed inSpeech to TextMultimodal AIon
Streaming and batch transcription, and where word error rate bites.
Listed inLlamaIndex for Multimodal AppsMultimodal AIon
Indexing and retrieving across text and images together.
Listed inVideo UnderstandingMultimodal AIon
Sampling frames vs native video input, and the token cost of both.
Listed inWhisper APIMultimodal AIon
OpenAI's transcription model, hosted and self-hosted.
Listed inImage UnderstandingMultimodal AIon
Passing images to a model and getting reliable structured answers back.
Listed inDALL·E APIMultimodal AIon
Generating and editing images programmatically.