BitSharp / research · proof of depth · arXiv 2603.04162 · v2 · 2026-03-05
Six quantization methods, 22 Polish benchmarks, 285 dollars.
Bielik-11B-v2.3-Instruct · FP16 → QuIP# E8P12 · 6.7× compression · Sec. 6.4
Claims on a thread · hover or scroll
Every number points to its row.
Bielik-Q2-Sharp compares QuIP#, SpinQuant+GPTQ, ButterflyQuant, QTIP, VPTQ and AQLM — calibrated on Polish (CulturaX-PL), with the caveats written down in the open. This is not what the studio does for a living — it is proof that we can take a subject down to research level.
After quantization the model holds 71.92% across 22 benchmarks — baseline: 72.07%.source: Table 4
The entire experiment: 285 dollars of GPU time on vast.ai — the cost table in the paper accounts for every one of them.source: abstract + Table 3
Caveats, stated plainly: rotation-based methodsfail at generation, even though they hold log-likelihood. The 76.50% figure does exist in the paper — with a methodological caveat attached (Sec. 5.5), which is why it is not in any headline.source: abstract
Quantization answers the question of how many of the eleven billion parameters in a Polish model you can throw away before it stops understanding Polish — and it answers on 22 benchmarks, not on a hunch.
Quantization measures how many parameters you can throw away before the model stops understanding Polish. 22 benchmarks, not a hunch.
Read the paper on arXiv · external Models from this paper → models · HF cards
What the same knowledge weighs. Six answers.
figures: Table 1 / Table 13 / Sec. 6.4 · GB after quantizationBar widths = GB / 22 GB (FP16). QuIP# E8P12:71.92% across 22 tasks at ~2.4 bits per weight (Table 13) — within reach of cards with 4 GB of VRAM (Conclusion). SpinQuant+GPTQ and ButterflyQuant have no bar: the paper does not report their size in GB (Table 1: „--” / decompressed model) — a missing number is a result too.
Second study · PolAgentBench · draft 2026-07-24 · pre-submission
The breaking point is universal. The ways of dying are not.
PolAgentBench: 67 deterministic agentic tasks — Polish prompt, English tool schema — across two Bieliks of different scale and architecture (7B Minitron and 11B), 20 quantization configurations in total, all GGUF.
Both models fall off a cliff between 3 and 2 bits(p < 0.0001 in each) — despite differing in scale, architecture and everything else we measured.source: PolAgentBench, Table 1 (draft)
Same threshold, opposite death: the 7B hangs in a tool loop, the 11B disintegrates after roughly 2.8 steps.source: PolAgentBench, Finding 2 (draft)
Caveats, stated plainly: a rounding artifact moved the apparent threshold by a full bit —the pre-correction result would not have replicated.source: PolAgentBench, Sec. 3.5 (draft)
11B pass rate · Q3_K_M → Q2_K · 67 tasks · p < 0.0001 · Table 1