Benchmark April 2026, updated September 2026 AI-written

Local LLM Benchmark: Gemma 4 vs Qwen 3.5

Head-to-head on a Mac Studio M3 Ultra. 26 prompts, 6 categories, and one surprising finding: the model that decoded faster finished 8.2x slower.

Updated 27 Sep 2026: the results file of the 5 Apr run was written but never committed, and a later run overwrote it, so every measurement below was re-derived from the LM Studio server log and the benchmark harness log of that run. The setup and the total-time medians reproduce. Most quality scores do not survive, because LM Studio does not log the text of streamed answers, so the April quality table (overall 3.77 vs 2.79) and head-to-head tally (Gemma 9, ties 11, Qwen 4) are gone. One speed figure was wrong: the April version reported 8.2 vs 3.4 tok/s, a harness estimate that never counted Qwen's reasoning tokens. The time to first token (1.84 s vs 0.66 s) cannot be confirmed from the logs and is gone too.

Data: per-prompt CSV, per-request CSV, and how it was recovered, claim by claim.

TL;DR

Gemma 4 31B (bf16) was 8.2x faster end-to-end: a median 8.0 s per prompt against 65.6 s for Qwen 3.5 27B (Q8_0).

Not because it generates faster. Qwen decoded 10.60 tokens per second to Gemma's 7.85, but it generated a median 764 tokens per prompt to Gemma's 50.5, most likely because of its thinking mode. Smaller model, much slower in practice. Counterintuitive.

The Setup

I wanted to pick a daily-driver model for my Hermes agent. Two candidates on disk: Google's new Gemma 4 31B at full bf16 precision, and Alibaba's Qwen 3.5 27B at Q8_0 quantization.

Both served locally via LM Studio on port 1234, both loaded with a 32K context window. Hardware: Mac Studio with an M3 Ultra (512 GB unified memory), the same rig I use to evaluate and run a rotating stable of local models.

The Benchmark

26 prompts across six categories: reasoning, coding, math, instruction following, creative writing, and tool use. 24 of them are scored programmatically (exact match, unit tests, constraint checks, tool trace validation); the other two were meant for a Claude judge. No judge result from this run exists, and the April scores left them out.

Deterministic prompts ran at temperature=0, seed=42. Creative ones at temperature=0.8 with three seeds. Two warmup calls per model per prompt discarded, then three scored repeats per model per prompt: 156 scored generations.

Quality: What Survived

The harness wrote every answer to a results file that a later run overwrote before it was ever committed. The server log keeps each request and its timings, not the answer text, so most scores cannot be recomputed. Seven prompts still can: the harness logged failed unit tests and some constraint failures as it ran, and a tool call's arguments reappear in the next request.

Prompt Scored by Gemma4 bf16 Qwen3.5 Q8
Range compactor function unit tests 5 5
Bug fix in a list-rotation function unit tests 5 5
LRU cache class unit tests 5 5
nginx log line parser unit tests 5 5
One sentence of exactly 17 words, no commas constraint check 5 2
Refund a late order per policy (tool call) tool trace 5* 0
Rain check, then book a calendar event (tool call) tool trace 2.5 to 5 2.5*

Scores on a 0 to 5 scale; each cell's recovered value or range is the same in all three repeats. * Read from tool-call arguments that LM Studio shortened in its log, assuming the hidden part is well-formed JSON that repeats neither compared field; without that assumption the score is between 2.5 and 5, which still beats Qwen's 0 on the refund call. Gemma's calendar-event arguments are cut before the compared field, so its score there is only known to be between 2.5 and 5. That prompt's scorer also compares the start time with the exact string "15:00", so a full ISO timestamp can never earn the argument points: its score says little about whether the event was booked correctly.

Of the comparisons the logs can still decide, Gemma lost none: four coding ties and wins on the strict-format sentence and the refund tool call. The calendar event is a tie or a Gemma win if the assumption above holds, and undecided if it does not. The other 17 scored prompts (all of reasoning and math, three of four instruction prompts, the scored creative prompts, two tool-use prompts) cannot be re-scored. One more reason not to miss the old math row: at the time, the answer key of one math prompt was wrong (5/12 where the right answer is 1/2).

Performance: The Surprise

On paper, Qwen 3.5 should be faster. It's smaller (28.6 GB vs 61.4 GB of weights) and most of its layers use linear attention. On Apple Silicon, where decoding is memory-bandwidth-bound, I expected the smaller model to produce tokens faster, and per token it did.

In practice, Gemma 4 was 8.2x faster end-to-end (medians per prompt):

Metric Gemma4 Qwen3.5
Total time per prompt (s) 8.0 65.6
Tokens generated per prompt 50.5 764
Decode speed (tok/s, server) 7.85 10.60
Prompt processing before the first token (s, server) 1.24 0.27

Total time is the harness's own clock. Decode speed and prompt processing come from llama.cpp's server-side timings; prompt processing is not the same measurement as the client-side time to first token. Summed over all 78 scored generations, Qwen needed 5.5x as much wall time as Gemma.

The measured part: Qwen generated 15x more tokens per prompt, and even at a faster per-token rate that is far slower. The likely culprit is thinking mode. Qwen 3.5 emits reasoning tokens in a separate reasoning_content stream before producing the final answer, and those tokens share the max_tokens budget with the answer. The logs count tokens but do not keep the text, so they cannot split reasoning from answer, and I did not run Qwen with thinking turned off.

My first run gave Qwen the same per-prompt budget as Gemma, 256 to 768 tokens. In 62 of 78 answers it used the entire budget, which ends a response with finish_reason: length (inferred from the token counts; the finish reason itself is not in the logs). For the 17-word sentence the harness logged an empty answer in all three repeats, though Qwen still passed three of the four unit-tested coding prompts inside the budget. I multiplied its budget by 8, to 2,048 to 6,144 tokens, for the run reported here, and in 6 of 78 answers it still used all of it.

Even with the fix, the median Qwen prompt took 65.6 seconds against Gemma's 8.0 (means: 87.1 and 15.8). For an agent loop where you're making dozens of calls, this is the whole ballgame.

What Else Stood Out

Strict formats went to Gemma

The one instruction prompt the logs can score asks for a single 17-word sentence with no commas. Gemma scored 5/5 in every repeat. Qwen's answer ran to 235 words and contained a comma: 2/5. The bullet-point prompt I quoted here in April cannot be scored from the logs; all three of Qwen's attempts at it used the full budget.

Coding was a wash (4 ties)

Both models passed the unit tests of all four unit-tested coding prompts, in every repeat: range compactor, a bug fix in a list-rotation function, LRU cache, nginx log parser. The fifth coding prompt, a SQL query, was left for the judge, and no judge score for it survives. Need harder coding prompts to differentiate.

Only Gemma made the refund call

Asked to handle a late order under a refund policy, Gemma looked up the order and called the refund tool in every repeat, with the right order id and amount visible in the logged arguments. Qwen called no tool at all.

Qwen's speed wins: prompt processing and decode

Qwen got through the prompt faster (0.27 s vs 1.24 s before the first token could be generated) and decoded faster (10.60 vs 7.85 tok/s). It lost on volume, not speed.

Caveats

This is a deployment comparison, not a pure architecture comparison. bf16 vs Q8_0 folds together model quality, quantization precision, memory bandwidth, and runtime backend.

Sample size is small: 26 prompts, 4 to 5 per category, and the quality evidence that survived is 7 prompts. Broad claims would be premature. "Gemma 4 is better than Qwen 3.5" is not what this tells us.

What it does tell me: for my specific hardware, my specific workload, and my specific budget of patience, Gemma 4 31B is the better daily driver for Hermes. The speed half of that verdict is backed by the recovered logs; the quality half rests on the seven prompts above.

Code

Full suite, prompts, runner, and scoring code is open source, along with the recovered data for this run and the scripts that recovered it:

github.com/omar16100/llm-benchmark