Visual Intelligence
Lyria 3.5 Accepts Image Prompts; Music Editing Stays Single-Turn
Google opened Lyria 3.5 in the Gemini API on September 3 and expanded it through Gemini on September 4. The preview can take text plus up to 10 images and generate 44.1 kHz stereo music. It remains a single-turn, variable-output system; image conditioning does not establish faithful scene understanding or reliable creative control.
Citation-ready: Google's Lyria 3.5 preview accepts text and up to 10 images as input and generates 44.1 kHz stereo music, while multi-turn editing is not supported in the current API version.

What happened and why it matters
The API uses images as creative conditioning alongside text, but Google does not claim a literal or deterministic visual translation, and the current version cannot refine a generated track through a multi-turn edit conversation.
Official Gemini API music-generation guide and launch update
Primary reference: Google AI for Developers: Generate music with Lyria 3.5. Kaleido Field checked the event date and the article's attributed facts against this source.
| Source date | September 3-4, 2026 |
|---|---|
| Checked by Kaleido Field | September 5, 2026, 09:15 CST |
| Source function | current visual-intelligence analysis separating API preview, image inputs, audio output, model variants, duration, prompt control, safety filtering, SynthID, variability, and independent image-to-music evaluation |
An image is a conditioning signal, not a verified description
The guide says the model composes music inspired by visual content. A sunset can influence mood, color associations, instrumentation, or pace, but the documentation does not define which visual features are extracted or how faithfully they survive in the track.
A useful test pairs the same image with controlled prompts, removes or swaps one image at a time, and has reviewers compare mood, subject, palette, setting, and unwanted stereotype without being told which input produced which output.
Clip and Pro serve different iteration loops
The Clip endpoint always returns 30 seconds, while Pro can generate a song lasting a couple of minutes with duration and structure influenced by the prompt. Google recommends exploring with Clip before committing to the longer model.
The current single-turn boundary means revision requires another generation request rather than a conversational edit of the same musical object. Keep prompt, image hashes, endpoint, response steps, audio file, lyrics, duration, and human selection together.
Provenance is present but not a rights decision
Google says every generated audio file carries an imperceptible SynthID watermark and blocks requests for specific artist voices or copyrighted lyrics. Those controls can support disclosure and policy enforcement.
They do not decide ownership, license compatibility, similarity to protected works, or whether the watermark survives every edit, recompression, stem separation, or platform upload. Those questions need separate legal and technical evidence.
Evidence boundary
Official API facts: `lyria-3.5-clip-preview` produces fixed 30-second clips; `lyria-3.5-pro-preview` produces prompt-controlled songs lasting a couple of minutes; both accept text and images through the Interactions API and output MP3 audio at 44.1 kHz. The guide permits up to 10 images, documents safety filtering, includes SynthID in generated audio, states that results vary between calls, and says multi-turn editing is unsupported. Official app update: Lyria 3.5 is available globally in Gemini web and mobile and through additional Google creative surfaces. Google positioning: improved vocals, arrangements, fidelity, and creative control. Not established: independent audio quality, visual faithfulness, cultural interpretation, copyright outcome, false-positive filter rate, watermark robustness after transformation, deterministic reproduction, or editability comparable with a digital audio workstation.
FAQ
How many images can Lyria 3.5 use?
The Gemini API guide says a request can include up to 10 images alongside text.
What audio does it return?
The documented preview endpoints return MP3 audio at 44.1 kHz stereo, plus generated lyrics or structure in the response steps.
Can a user edit the same track over several prompts?
No. Google documents music generation as single-turn in the current Lyria 3.5 version.