Translate Gemma Image OCR translation was tested as a single-step workflow: give a screenshot to the model, select the source language and English as the target, then inspect the returned text. The practical question was not whether these models can translate a clean paragraph. It was whether they retain the small details that make interface text useful: currency punctuation, serial numbers, delivery windows, a security warning, and an appointment location.
This comparison covers Translate Gemma 4B Image, Translate Gemma 12B Image, and Translate Gemma 27B Image. They belong to Google’s TranslateGemma family, which the model cards describe as Gemma 3-based translation models supporting 55 languages. For image input, the cards describe text extraction and translation through a dedicated chat template rather than a separate OCR product.
Test setup and parameters
The set contains six synthetic UI screenshots: a Turkish ecommerce cart, German warning label, French stock notification, Spanish deal banner, Polish OTP message, and Dutch appointment confirmation. Each image was sent to all three Wiro models with the known source language selected, target language set to English, and maxNewTokens set to 200. That token limit is the documented default for each Wiro integration. The task was to translate the visible screenshot content while preserving numbers, symbols, prices, serial numbers, and placeholders.
This is a narrow test. It does not measure long documents, handwriting, uncertain source-language detection, or translation into languages other than English. It does expose a common production risk: a translated sentence can look fluent while one misread noun reverses a real instruction. The model cards also specify image inputs normalized to 896 by 896 and encoded to 256 tokens, with a 2K-token total input context. Dense or tiny text therefore deserves a manual check, even when the output reads smoothly.
Run time and cost reporting
| Model | Completed test runs | Average run time in this set | Cost shown by Wiro docs |
|---|---|---|---|
| Translate Gemma 4B Image | 6/6 | 20.7 seconds | No current per-output price is listed in the model documentation. |
| Translate Gemma 12B Image | 6/6 | 22.6 seconds | No current per-output price is listed in the model documentation. |
| Translate Gemma 27B Image | 6/6 | 29.2 seconds | No current per-output price is listed in the model documentation. |
The times above are observations from these six runs, not a service guarantee. The documents do not provide a current price for these exact configurations, so no cost has been inferred from model size or from an unrelated example task. That matters: runtime and billing can change with queueing, image size, hardware, and account configuration. The useful comparison here is relative: 4B finished first, 12B was close behind, and 27B took roughly eight and a half seconds longer than 4B on average.
What the six screenshots actually show
1. Turkish cart summary: the note is the hard part

All three outputs recovered the cart total, shipping charge, and 2-3 business-day delivery window. The difference appears in the final note. The 12B and 27B outputs interpreted it as “Do not leave it with the doorman,” while 4B changed that instruction to “Do not leave it at the cashier.” This is not a cosmetic wording choice. It changes the delivery instruction, so a checkout or logistics workflow should show the original text alongside the translation when a short note controls an action.
2. German warning label: structured values survive intact

This was the cleanest structured-text case. Each model preserved DE-77-2048 and translated the 24-month warranty correctly. The 4B wording, “Attention: Read instructions,” is slightly more literal. The larger models returned “Caution: Read the instructions.” Both are serviceable English translations. The important result is that none of the three altered the serial number, which is the field that would create a support problem if it changed.
3. French stock notice: little room for divergence

The stock notice had short, conventional UI language: out of stock, a new delivery on Friday, and an email-notification prompt. The 12B and 27B outputs match exactly. The 4B output omitted “you” in the question, producing “Would like to be notified by email?” It is understandable but less natural. This is a good example of when the smaller model is enough: the meaning landed, and a product team can apply its own approved UI copy after translation.
4. Spanish deal banner: a successful translation is not the whole reliability story

4B translated the banner as “Daily offer,” while 27B used “Deal of the day.” Both retained the $50 free-shipping threshold and the 3-5 business-day estimate. The 12B result for this test failed during inference, so it cannot be treated as a translation result. The broader lesson is practical: a batch pipeline needs retry handling and a visible failure state. A model that delivers good text on five images but returns no output on a sixth needs operational safeguards before it handles customer-facing content.
5. Polish OTP message: do not automate security instructions blindly

12B and 27B returned the intended warning: the password expires in 15 minutes and the code must not be shared. 4B preserved the placeholder but changed the wording to “Do not use this code.” That changes the instruction. This output should not be called wrong because it has a missing comma or different style; it is wrong for a security flow because the action changes. Human review, source-text display, or an approved string catalog is the right choice for OTP, identity, payment, and legal UI.
6. Dutch appointment confirmation: concise output can still be complete

All three models preserved the core appointment facts: Wednesday, 10:00 AM, Office 3B, and an ID reminder. The 4B response is a full sentence version. The larger models remove punctuation and the word “Please,” but keep the information. For compact operational screens, that is often acceptable. The decision should depend on whether the translation feeds a reader-facing screen or a downstream parser that expects exact formatting.
Which Translate Gemma Image model should you pick?
Choose 4B Image when speed matters most and the screenshot contains routine, low-risk product text. It completed every run and was the fastest in this set. Its weak points were short instructions: the cart note and OTP warning show why it should not be the only decision-maker for sensitive messages.
Choose 12B Image when you want more polished wording than 4B without stepping up to the slowest option. Its successful outputs handled the cart note and OTP wording well. One inference failure in the Spanish test means it needs retries and monitoring in a production batch.
Choose 27B Image when natural English phrasing and careful handling of the tested short instructions are worth the extra latency. It gave the strongest overall wording in this small set and preserved the critical values. That is a preference, not proof of universal superiority. The six images are too few to claim an accuracy ranking.
Related reading
- TranslateGemma 4B: 8 Translation Prompts for Real Apps
- Easy OCR: 5 Layout Tests
- dots.ocr-1.5: OCR in 6 Screenshot Tests
Model sources
- Google TranslateGemma 4B Image model card on Hugging Face
- Google TranslateGemma 12B Image model card on Hugging Face
- Google TranslateGemma 27B Image model card on Hugging Face
For routine screenshot localization, these models remove a separate copy-and-paste step. For text that changes a customer’s money, access, safety, or appointment, treat the result as a fast first pass and retain a route to the original screenshot.