GOO-395 · local stack · 4 Oct 2026

ElevenLabs text-to-speech in Galleria: end-to-end test evidence

Eleven v4 and Eleven Turbo v2.5 generate, play, bill per character and produce captions, from the real UI through to the credit ledger. Two results were not a clean pass; neither is a defect in this work.

After the first run the change set was audited for scope and workarounds, trimmed, and the affected paths were run again. This page shows the final state.

Run against the feature branch's own binaries on the make up stack. Every claim below was checked in the database or a service log, not taken from the screen.

18real generations on the new models, 15 completed and 3 failed by design, plus 1 Seed Audio regression run
13 / 13completed generations charged exactly their quote, one ledger row each; no failed one charged
65 creditstotal charged across the runs, $0.82 of provider cost
0page or React errors in the browser console during the run

What each hop was shown to do

Quote equals charge

For every completed generation the credits charged in credit_ledger equal what the estimate returned for the same character count, on the cost_plus pricing path, with exactly one usage row.

ModelCharactersQuoteChargedProvider costTimeDriven by
Eleven v46011$0.0048003.8 sAPI
Eleven v46211$0.0049604.4 sBrowser
Eleven v41,01366$0.08104016.6 sAPI
Eleven v45,0002727$0.40000062.8 sAPI
Eleven Turbo v2.56011$0.0030002.9 sAPI
Eleven Turbo v2.56211$0.0031003.4 sBrowser
Eleven Turbo v2.51,01344$0.0506506.3 sAPI
Eleven Turbo v2.55,0001717$0.25000012.6 sAPI

Four more short generations were each quoted and charged 1 credit: a 22-character Japanese sentence and a 47-character 192 kbps run on v4, and one 39-character run per model and a final 38-character v4 run after the scope passes. Turbo accepted 5,000 characters, which fal does not document; bytes per character matched the 1,013-character run, so nothing was truncated.

In the browser

CheckResultWhat was observed
Audio models in the pickerpartialSeed Audio 1.0 (default), Eleven v4, Eleven Turbo v2.5 are listed. The option list shows no per-model logos; the Image picker is the same, so this is the shared picker's existing behaviour. The logo appears on the selected model's card.
Eleven v4 settingspassVoice (21 options, Rachel default), Stability 0.5, Similarity Boost 0.75, Language (Auto + 32), Text Normalization, Output Format (MP3 128 / 192 kbps), Timestamps. Every list control is a dropdown.
Counter and live quotepassThe count is visible at every length on both new models: 0/5000, 62/5000 (1 credit), 1000/5000 (6 credits) in the floating composer, and Max 5000 characters · n/5000 in the sidebar layout. Each estimate request carried characters equal to the counter.
Over the limitpass5001/5000, toast "Prompt too long", no generation request sent, generation count unchanged in the database.
Generate, play, captions on v4partialRow appeared without a reload, audio played (current time advanced 0.106 s to 1.222 s), captions panel listed two cues, SRT and VTT downloaded and are valid. The last cue ends 31 ms after the audio file does; see below.
Eleven Turbo v2.5passSame shared controls plus Style and Speed, no Output Format. 1,000 characters quotes 4. Generate, play, captions and both downloads pass, including the duration check.
Rows without timestampspassThe captions toggle is on exactly the five rows whose asset has timestamps in the database. Failed rows show only recreate and delete.
Seed Audio 1.0passIdentical to main. Max 3000 characters in the sidebar, the count appearing only past 2,500 (absent at 2,500, 2501/3000 at 2,501) and counting the raw text, spaces included (2504/3000 with five trailing spaces). 3,001 characters is blocked with its original message, "Seed Audio prompts must be 3000 characters or fewer.", and no request. Quote is 10 credits.
Switching modelspassSettings reset to defaults on every switch. No v4 key reached a Turbo request or the reverse.
Galleria Audio view with Eleven v4 selected and its settings card open
Eleven v4 selected. Seven controls, the counter at 0/5000 and the 1-credit fallback quote on an empty prompt.
Composer with a 1,000 character prompt showing 1000/5000 and Generate 6 credits
1,000 characters on v4: counter 1000/5000, quote 6 credits.
Composer on Eleven Turbo v2.5 with a 1,000 character prompt showing Generate 4 credits
The same text on Turbo v2.5 quotes 4 credits. Style and Speed replace Output Format.
Prompt too long toast with the counter at 5001/5000
5,001 characters: blocked before any request leaves the browser.
Audio row with its captions panel open showing two cues and SRT and VTT links
Captions panel on the new v4 row: two cues with start times, SRT and VTT downloads.
Captions panel on a Japanese audio row
Japanese text breaks at its own punctuation, with no spaces inserted.
Model picker open on the Audio view listing three models
The picker lists the three audio models. Option rows carry no logos for any media type.
Seed Audio 1.0 with a 2,501 character prompt showing 2501/3000 in the composer footer
Seed Audio 1.0 at 2,501 characters: the count appears only near the limit, as before.
Sidebar layout with Seed Audio 1.0 and a short prompt showing Max 3000 characters
Sidebar layout, Seed Audio 1.0, short prompt: "Max 3000 characters" and no count.
Sidebar layout with Eleven v4 showing Max 5000 characters and 62/5000
Sidebar layout, Eleven v4: the count is always shown, because the price follows it.

The two generations made in the browser

Prompt for both: "The tide is turning now. Bring the small boats in before dark." (62 characters). Press play to hear what was stored.

Eleven v4, voice Aria

57df9972 · 4.049 s · mp3 44.1 kHz 128 kbps

Eleven Turbo v2.5, voice Rachel

05312794 · 3.474 s · mp3 44.1 kHz 128 kbps

Request bodies captured from the browser

POST /api/generations  ->  201   (Eleven v4)
{"user_id":"7c1e96fa-…","model_id":"4b89989b-…","prompt":"The tide is turning now. Bring the small boats in before dark.",
 "config":{"voice":"Aria","stability":0.5,"similarity_boost":0.75,"language_code":"","apply_text_normalization":"auto",
           "output_format":"mp3_44100_128","timestamps":true},
 "asset_type":"audio","action":"generate","organization_id":"036bb424-…"}

POST /api/generations  ->  201   (Eleven Turbo v2.5)
{"user_id":"7c1e96fa-…","model_id":"2263156a-…","prompt":"The tide is turning now. Bring the small boats in before dark.",
 "config":{"voice":"Rachel","stability":0.5,"similarity_boost":0.75,"style":0,"speed":1,"language_code":"",
           "apply_text_normalization":"auto","timestamps":true},
 "asset_type":"audio","action":"generate","organization_id":"036bb424-…"}

Live quote requests

POST /api/credits/estimate   {"model_identifier":"eleven-v4-tts", … ,"characters":62,   "config":{…}}
  -> {"credits":"1","cogs_usd":"0.00496","pricing_path":"cost_plus","estimated":false}
POST /api/credits/estimate   {"model_identifier":"eleven-v4-tts", … ,"characters":1000, "config":{…}}
  -> {"credits":"6","cogs_usd":"0.08","pricing_path":"cost_plus","estimated":false}
POST /api/credits/estimate   {"model_identifier":"eleven-turbo-v2.5-tts", … ,"characters":1000, "config":{…}}
  -> {"credits":"4","cogs_usd":"0.05","pricing_path":"cost_plus","estimated":false}
Empty prompt (no characters key)
  -> {"credits":"1","pricing_path":"flat_fallback","estimated":false}

What vertex-ai sent to fal

Auto language put an empty language_code in the request, and vertex-ai dropped it: neither line below carries the key.

[FAL Audio] model=eleven-v4-tts path=elevenlabs/tts/eleven-v4 chars=62
  params=map[apply_text_normalization:auto output_format:mp3_44100_128 similarity_boost:0.75 stability:0.5 timestamps:true voice:Aria]
[FAL Audio] stored audio path=gs://goodtake_ai_dev/audios/fal-tts/fal_tts_91851304_20261004_162728.mp3 content_type=audio/mpeg

[FAL Audio] model=eleven-turbo-v2.5-tts path=fal-ai/elevenlabs/tts/turbo-v2.5 chars=62
  params=map[apply_text_normalization:auto similarity_boost:0.75 speed:1 stability:0.5 style:0 timestamps:true voice:Rachel]
[FAL Audio] stored audio path=gs://goodtake_ai_dev/audios/fal-tts/fal_tts_c8fd8df8_20261004_163025.mp3 content_type=audio/mpeg

Database

GenerationModelStatusCharsTiming entriesUsage rowsChargedBilled costBalance after
57df9972-e603-…eleven-v4-ttscompleted621211.00000.0049609,999,841
05312794-242e-…eleven-turbo-v2.5-ttscompleted621211.00000.0031009,999,840

Concatenating the 12 stored word values of either asset gives the prompt back exactly. The go-credits spend lines for both read path=cost_plus; go-queue logged one start and one completed line per id.

Downloaded caption files

Eleven v4, SRT
1
00:00:00,000 --> 00:00:01,720
The tide is turning now.

2
00:00:01,920 --> 00:00:04,080
Bring the small boats in before dark.
Eleven v4, VTT
WEBVTT

00:00:00.000 --> 00:00:01.720
The tide is turning now.

00:00:01.920 --> 00:00:04.080
Bring the small boats in before dark.
Eleven Turbo v2.5, SRT
1
00:00:00,070 --> 00:00:01,277
The tide is turning now.

2
00:00:01,486 --> 00:00:03,181
Bring the small boats in before dark.
Japanese row (API-made), SRT
1
00:00:00,000 --> 00:00:02,320
今日は、良い天気ですね。

2
00:00:02,320 --> 00:00:03,520
散歩に行きましょう。

Driven through the API

CheckResultEvidence
Limit counts characters, before the gatepass5,001 characters: HTTP 400 text is 5001 characters; the limit for this model is 5000, no generation row, no ledger row. 5,000 Japanese characters (15,000 bytes) passed the limit.
Short generation with timestamps, both modelspass12 entries of {word, start_s, end_s}, identical in the API response and the database.
Unknown voice fails once, no chargepassFailed within 2 s. One attempt in the queue log; the task is archived, not in the retry set. No ledger row. The stored message is fal's raw "Voice not found" error inside the worker's wrapper; the UI cannot send an unknown voice today.
No timestamps requested, none storedpassThe asset's provider_metadata has no timestamps key, in the database and through the API. With timestamps requested, 7 entries were stored that rebuild the prompt.
Quote equals charge at 1,013 and 5,000 characterspassSee the table above.
v4 optionspasslanguage_code: ja works. mp3_44100_192 returns a 192 kbps MP3 (ffprobe).
Not listed publiclypassGET /public/v1/models?type=audio: 200, listing only seed-audio-1.0. MCP list_models {type: "audio"}: the same. type=image (10) and type=video (19) are unchanged. The in-app model list includes both new models.
Seed Audio 1.0 unchangedpassAccepted, routed to /bytedance/generate-audio, fails with "not configured" because the local BytePlus key is unset, exactly as before. No charge.
Migration 123 rollback on seeded datapassDown: version 122, both per-character rows deleted, constraint restored, 99 other cost rows byte-identical, price falls back to 1 credit. Up and re-seed: price back to 6 credits.

Scope passes

The owner's rule: leave existing behaviour alone unless the new models needed it changed, keep the implementation simple, and remove workarounds. Three independent audits of the diff, and a line-by-line check by the reviewer, led to these changes. All are re-verified above.

Minimal-diff audit, 6 October. A fresh independent reviewer read every changed hunk against main. Two things touched existing material without need and were undone: a helper that had been extracted out of Seed Audio's enqueue function (that function is byte-identical to main again; the new function carries its own copy of the call, as the image and video ones do), and another ticket's route (/dashscope/generate-video) that the regenerated API docs had picked up. The panel now counts characters only for models that carry their own limit, so image, video and Seed Audio renders run none of the new counting code. After the changes one live Turbo generation was charged exactly its quote, and the panel counted and quoted as before in both layouts.

Audit of the added code, 6 October. A second independent reviewer classified every added line as required, convention, optional or redundant. Removed as unneeded: a frontend flag that always travelled with the text limit, a duplicate unknown-model check in the speech route (the service already rejects it), and most of two documentation additions. Kept on purpose: the admin price preview, the quoted amount on the ledger, the splitting of Japanese and Chinese captions, the tests that are the only coverage of a contract, and the Language, Text Normalization and Output Format settings, which the ticket lists. The small test that pins which models are Galleria-only was removed in this pass and restored the same day, because the ticket asks for it. After the removals the panel counted and quoted exactly as before, and one live generation with timestamps was charged its quote.

The change set now modifies 17 lines that existed before it; everything else is addition. Each of the 17 is forced: four trailing commas, the audio handler's provider reject and enqueue call and two comment lines, the credit-gate call, the admin unit list with its message and one comment line, the pricing doc's unit list, one import, two counter conditions, and the Projects chat's model filter.

One difference remains for existing models, and no user can see it: the gateway sends the prompt's character count on every pricing request, not only for the new models. The engine reads it only for per-character rows. An image model's estimate was byte-identical with and without the field on all three pricing paths.

Is it Galleria only?

Yes, by the owner's definition: outside Galleria the two models cannot be selected in any picker and are not listed by the Claude connector. The Projects chat picker was the one place that listed them; it now hides these two models and nothing else. The last three rows are not places where a model is picked or listed for a user, and are left as on main.

SurfaceThese two modelsSame as Seed Audio on main
Galleria Audio tabListed, generate, play. The only UI that can run them.yes
Public API, Claude connector (MCP), CLI: model listsNot listed, for any type. Live list_models: audio returns only Seed Audio; image 9, video 17, default 9, none ElevenLabs.no: Seed Audio is listed under type=audio, as before
Public API, MCP, CLI: generateRefused; only image and video models are accepted, and there is no audio routeyes
WorkflowsNot selectable. The add-node menu lists 10 image, 13 video and 2 video-edit models, no audio model (checked in the browser).yes
Ads Engine, Personas, Character Swap, phone UINot offeredyes
Projects chat model pickerNot listed. The catalogue has 32 enabled models and the picker shows 30; the two missing are the ElevenLabs models. Its Audio tab shows only Seed Audio (checked in the browser, new and existing project).no: Seed Audio is listed there, as before
Public generations list and get, MCP list_generations and get_generation, Assets "All"A clip made in Galleria is readable there by the same organisation: its text, the model identifier and the audio URLyes
Model catalogue endpoints of the web appReadable without login, by design for the signed-out Galleria pickeryes
The in-app create endpointNot tied to the Galleria page; any client with in-app access can call ityes

Two results that were not a clean pass

Not tested