Skip to content
Comparisons

Seed-V2 Mini vs Qwen3.5-27B: 5 Small Tests

Seed-V2 Mini vs Qwen3.5-27B is a comparison of two chat models on five deliberately small tasks: structured extraction, a JavaScript utility, translation, constrained summarization, and simple time arithmetic. The point was not to crown a general winner from five prompts. It was to see what happens when an application asks for a narrow output shape and has little room to clean up the answer afterward.

Both models are available on Wiro: Seed-V2 Mini and Qwen3.5-27B. The original five results remain below, with the prompts and returned text preserved. A fresh Wiro run was also submitted for each model using a JSON-only routing prompt. The task responses completed without downloadable text artifacts, so this update does not claim a timing or cost figure that Wiro did not return.

What this Seed-V2 Mini vs Qwen3.5-27B test checks

Small format-sensitive tasks are common in production: pull fields from an email, return a code helper, translate a UI string, summarize an event, or calculate a schedule. In each case, correctness is only half the job. The answer also needs to fit the contract. A valid answer wrapped in commentary can still break a JSON parser or a downstream automation.

The five prompts use the same wording for both models. They ask for JSON only, code only, Turkish only, bullets only, and one time only. That makes the comparison about observable behavior, not about different prompt treatment. It is still a tiny sample. It cannot measure long-context reliability, tool calling, factual recall, safety behavior, or performance on a real codebase.

Seed-V2 Mini vs Qwen3.5-27B comparison illustration
Wiro-generated comparison visual created for this updated test article.

Parameters used and what they mean

The published examples did not record every request parameter, so they should be read as output samples rather than a repeatable benchmark log. The current Wiro model documentation shows that Seed-V2 Mini accepts temperature, topP, frequency and presence penalties, a completion-token limit, reasoning effort, and a thinking mode. Its defaults include temperature 1.0, topP 0.7, medium reasoning effort, and enabled thinking.

Qwen3.5-27B exposes temperature, top_p, top_k, repetition and length penalties, token limits, seed, quantization, and do_sample. Its documented defaults include temperature 0.7, top_p 0.95, seed 123456, quantization enabled, and sampling enabled. For the fresh JSON-routing check, Seed-V2 Mini was sent temperature 0, topP 0.7, thinking disabled, and a 400-token completion ceiling. Qwen3.5-27B was sent temperature 0, top_p 0.95, do_sample false, seed 123456, quantization true, and a 400-token ceiling.

Those settings aim for repeatable, compact output. They do not make the models identical. They simply reduce one avoidable source of variation. Wiro’s completed task records for these two text requests did not expose elapsed seconds, downloadable output files, or total cost. Rather than estimate them, this post leaves run time and cost unreported. Cost also depends on token use and the selected Wiro configuration, so a number from another request would not be a useful substitute.

Results from the five small tests

1. JSON extraction

Prompt: Extract JSON with keys name, email, plan, seats. Return JSON only. Input: Name: Jamie Lee; Email: [email protected]; Plan: Pro Annual; Seats: 12.

Seed-V2 Mini returned:

{
  "name": "Jamie Lee",
  "email": "[email protected]",
  "plan": "Pro Annual",
  "seats": 12
}

Qwen3.5-27B returned:

{
 "name": "Jamie Lee",
 "email": "[email protected]",
 "plan": "Pro Annual",
 "seats": 12
 }

Both answers contain the requested fields and parse as JSON despite the extra whitespace before Qwen’s closing brace. This is a tie on content. The practical lesson is not that either model needs no validation. Any application that accepts model-generated JSON should still parse it and handle a failure path.

2. JavaScript slug utility

Prompt: Write a JavaScript function slugify(str) that lowercases, replaces spaces with hyphens, removes non alphanumeric characters (keep hyphens), collapses multiple hyphens, and trims leading and trailing hyphens. Return code only.

Seed-V2 Mini placed the character-removal replacement before the whitespace replacement. That order removes spaces before they can become hyphens, so it does not satisfy the stated requirement for ordinary multi-word input. Qwen3.5-27B used the safer order: whitespace becomes hyphens first, then non-alphanumeric characters are removed, repeated hyphens collapse, and edges trim. On this exact prompt, Qwen’s implementation is the useful one.

function slugify(str) {
 return str
 .toLowerCase()
 .replace(/\s+/g, '-')
 .replace(/[^a-z0-9-]/g, '')
 .replace(/-+/g, '-')
 .replace(/^-+|-+$/g, '');
}

This is why code-only output is not enough. A short edge case such as "Hello, World!" would catch the difference immediately.

3. English to Turkish

Prompt: Translate to Turkish. Return Turkish only. Text: This model returns thinking text by default. Disable it for clean JSON outputs.

Seed-V2 Mini returned: Bu model varsayılan olarak düşünme metni döndürür. Temiz JSON çıktıları için bunu devre dışı bırakın.

Qwen3.5-27B returned: Bu model varsayılan olarak düşünme metni döndürür. Temiz JSON çıktıları için devre dışı bırakın.

Both are concise Turkish-only answers and preserve the operational point. Seed’s version makes the omitted object explicit with bunu; Qwen’s version is still natural because the object is recoverable from context. This one is a tie for a UI or support translation.

Two terminal windows for an LLM output-format comparison
Second Wiro-generated visual for the output-format comparison.

4. Five-bullet summary

Prompt: Summarize in 5 bullet points. Return bullets only. The source explains Wiro’s asynchronous Run and Task Detail workflow.

Both models returned five points and kept the essential sequence: start a run, poll a task ID, wait for completion, then read output metadata. Seed-V2 Mini compressed its result into one visually dense line in the saved output. Qwen3.5-27B used five separate Markdown bullets, which is easier to scan and better matches the requested format. Neither answer added unsupported details.

5. Time math

Prompt: A meeting starts at 09:15 and lasts 35 minutes. There is a 10 minute break. Then a second meeting lasts 50 minutes. What time does it end? Return only the time in HH:MM.

Both models returned 10:50. The calculation is straightforward: 09:15 plus 35 minutes reaches 09:50; the break reaches 10:00; another 50 minutes reaches 10:50. More importantly, both followed the single-value constraint.

How to choose between the two

Pick Seed-V2 Mini when the request is short, the response needs to stay compact, and the application benefits from its explicit thinking control. It handled extraction, translation, summary content, and time math cleanly in these samples. Test generated code before shipping it: the slug utility shows why that step matters even for a small function.

Pick Qwen3.5-27B when you want the documented Qwen family controls around sampling, seed, quantization, and deterministic code or translation requests. Its slug utility had the correct transformation order, and its five-bullet summary had clearer structure. Qwen’s public materials also describe a switch between thinking and non-thinking modes, a useful distinction when a product needs either deeper work or tightly shaped output. See the Qwen3 GitHub repository and the Qwen3 model card on Hugging Face for the family documentation.

The useful conclusion is narrower than a leaderboard claim. Seed-V2 Mini produced solid compact answers in four of five checks. Qwen3.5-27B produced the stronger code result and more readable constrained summary. For either model, specify the output shape, select deterministic settings for machine-readable work, and validate the returned data before it reaches a customer or a production system.