Visual Intelligence

Lyria 3.5 Accepts Image Prompts; Music Editing Stays Single-Turn

By Kaleido Field Staff ยท September 5, 2026

Images can steer a track, but they do not become an editable score

Google opened Lyria 3.5 in the Gemini API on September 3 and expanded it through Gemini on September 4. The preview can take text plus up to 10 images and generate 44.1 kHz stereo music. It remains a single-turn, variable-output system; image conditioning does not establish faithful scene understanding or reliable creative control.

Citation-ready: Google's Lyria 3.5 preview accepts text and up to 10 images as input and generates 44.1 kHz stereo music, while multi-turn editing is not supported in the current API version.

Google Lyria 3.5 launch artwork with a record player and the words Turn it up with Lyria 3.5
Image source: Google. Used for editorial coverage of image-to-audio generation desk.

What happened and why it matters

The API uses images as creative conditioning alongside text, but Google does not claim a literal or deterministic visual translation, and the current version cannot refine a generated track through a multi-turn edit conversation.

Official Gemini API music-generation guide and launch update

Primary reference: Google AI for Developers: Generate music with Lyria 3.5. Kaleido Field checked the event date and the article's attributed facts against this source.

Source check
Source dateSeptember 3-4, 2026
Checked by Kaleido FieldSeptember 5, 2026, 09:15 CST
Source functioncurrent visual-intelligence analysis separating API preview, image inputs, audio output, model variants, duration, prompt control, safety filtering, SynthID, variability, and independent image-to-music evaluation

An image is a conditioning signal, not a verified description

The guide says the model composes music inspired by visual content. A sunset can influence mood, color associations, instrumentation, or pace, but the documentation does not define which visual features are extracted or how faithfully they survive in the track.

A useful test pairs the same image with controlled prompts, removes or swaps one image at a time, and has reviewers compare mood, subject, palette, setting, and unwanted stereotype without being told which input produced which output.

Clip and Pro serve different iteration loops

The Clip endpoint always returns 30 seconds, while Pro can generate a song lasting a couple of minutes with duration and structure influenced by the prompt. Google recommends exploring with Clip before committing to the longer model.

The current single-turn boundary means revision requires another generation request rather than a conversational edit of the same musical object. Keep prompt, image hashes, endpoint, response steps, audio file, lyrics, duration, and human selection together.

Provenance is present but not a rights decision

Google says every generated audio file carries an imperceptible SynthID watermark and blocks requests for specific artist voices or copyrighted lyrics. Those controls can support disclosure and policy enforcement.

They do not decide ownership, license compatibility, similarity to protected works, or whether the watermark survives every edit, recompression, stem separation, or platform upload. Those questions need separate legal and technical evidence.

Evidence boundary

Official API facts: `lyria-3.5-clip-preview` produces fixed 30-second clips; `lyria-3.5-pro-preview` produces prompt-controlled songs lasting a couple of minutes; both accept text and images through the Interactions API and output MP3 audio at 44.1 kHz. The guide permits up to 10 images, documents safety filtering, includes SynthID in generated audio, states that results vary between calls, and says multi-turn editing is unsupported. Official app update: Lyria 3.5 is available globally in Gemini web and mobile and through additional Google creative surfaces. Google positioning: improved vocals, arrangements, fidelity, and creative control. Not established: independent audio quality, visual faithfulness, cultural interpretation, copyright outcome, false-positive filter rate, watermark robustness after transformation, deterministic reproduction, or editability comparable with a digital audio workstation.

Reader briefing

Keep the source trail in view.

One concise email when a model, benchmark, or visual-intelligence claim materially changes.

FAQ

How many images can Lyria 3.5 use?

The Gemini API guide says a request can include up to 10 images alongside text.

What audio does it return?

The documented preview endpoints return MP3 audio at 44.1 kHz stereo, plus generated lyrics or structure in the response steps.

Can a user edit the same track over several prompts?

No. Google documents music generation as single-turn in the current Lyria 3.5 version.