BitSharp / research · proof of depth · arXiv 2603.04162 · v2 · 2026-03-05

Six quantization methods, 22 Polish benchmarks, 285 dollars.

22 GB3.26 GB

Bielik-11B-v2.3-Instruct · FP16 → QuIP# E8P12 · 6.7× compression · Sec. 6.4

Claims on a thread · hover or scroll

Every number points to its row.

Bielik-Q2-Sharp compares QuIP#, SpinQuant+GPTQ, ButterflyQuant, QTIP, VPTQ and AQLM — calibrated on Polish (CulturaX-PL), with the caveats written down in the open. This is not what the studio does for a living — it is proof that we can take a subject down to research level.

After quantization the model holds 71.92% across 22 benchmarks — baseline: 72.07%.source: Table 4

The entire experiment: 285 dollars of GPU time on vast.ai — the cost table in the paper accounts for every one of them.source: abstract + Table 3

Caveats, stated plainly: rotation-based methodsfail at generation, even though they hold log-likelihood. The 76.50% figure does exist in the paper — with a methodological caveat attached (Sec. 5.5), which is why it is not in any headline.source: abstract

Full precisionFP16 · 22 GB

Quantization answers the question of how many of the eleven billion parameters in a Polish model you can throw away before it stops understanding Polish — and it answers on 22 benchmarks, not on a hunch.

After quantization~2.4 bpw · 3.26 GB

Quantization measures how many parameters you can throw away before the model stops understanding Polish. 22 benchmarks, not a hunch.

Read the paper on arXiv · external  Models from this paper → models · HF cards

What the same knowledge weighs. Six answers.

figures: Table 1 / Table 13 / Sec. 6.4 · GB after quantization
FP16 (base)22.0 GB
VPTQ5.00 GB
AQLM3.62 GB
QTIP3.27 GB
QuIP# E8P123.26 GB
IQ2_XXS (baseline)~2.6 GB

Bar widths = GB / 22 GB (FP16). QuIP# E8P12:71.92% across 22 tasks at ~2.4 bits per weight (Table 13) — within reach of cards with 4 GB of VRAM (Conclusion). SpinQuant+GPTQ and ButterflyQuant have no bar: the paper does not report their size in GB (Table 1: „--” / decompressed model) — a missing number is a result too.

Second study · PolAgentBench · draft 2026-07-24 · pre-submission

The breaking point is universal. The ways of dying are not.

PolAgentBench: 67 deterministic agentic tasks — Polish prompt, English tool schema — across two Bieliks of different scale and architecture (7B Minitron and 11B), 20 quantization configurations in total, all GGUF.

Both models fall off a cliff between 3 and 2 bits(p < 0.0001 in each) — despite differing in scale, architecture and everything else we measured.source: PolAgentBench, Table 1 (draft)

Same threshold, opposite death: the 7B hangs in a tool loop, the 11B disintegrates after roughly 2.8 steps.source: PolAgentBench, Finding 2 (draft)

Caveats, stated plainly: a rounding artifact moved the apparent threshold by a full bit —the pre-correction result would not have replicated.source: PolAgentBench, Sec. 3.5 (draft)

0.7160.045

11B pass rate · Q3_K_M → Q2_K · 67 tasks · p < 0.0001 · Table 1