R0 · frozen V3
Benchmark before training.
24 cases across Polish, instruction following, reasoning, coding, code analysis and context. This is the frozen pre-training reference point for future SZkrO revisions.
87.5%
Correctness
66.7%
Instruction adherence
62.5%
Strict pass
15.5
Median tok/s
Category correctness
R0 comparison
Why Granite won.
| Model | Correctness | Instruction | Strict | Median case | Median tok/s |
|---|---|---|---|---|---|
| Granite 4.1 3B | 87.5% | 66.7% | 62.5% | 4.1 s | 15.5 |
| Gemma 3 4B | 83.3% | 62.5% | 54.2% | 5.9 s | 10.2 |
| Gemma 4 E2B | 79.2% | 79.2% | 62.5% | 40.1 s | 9.5 |
| Phi-4 Mini 3.8B | 66.7% | 62.5% | 45.8% | 5.4 s | 12.5 |
| Llama 3.2 3B | 66.7% | 62.5% | 50.0% | 4.6 s | 11.9 |