2026-03-05 arXiv 2603.04162 · v2 published
2026-02 9 models + 3 datasets on HF
2026-07-24 PolAgentBench: draft complete · 67 tasks
2026-07-23 3 note theses pending approval
openhaft.pl · production · live
Two papers of our own, an agentic benchmark we wrote, nine public Bielik models and deployments you can click. Leading rather than promising — every claim on this page points to its proof.
time to import a 4,000+ product catalog · openhaft.pl · → deployments
Work, not declarations. Pick a registry.
thumbnails: real screenshots from this project · duotone
/researchA paper with a thread to every numberBielik-Q2-Sharp: six 2-bit methods, 22 benchmarks, caveats stated plainly — claims point at the rows of Table 4.
arXiv 2603.04162↗ l. 01–08
/deployments285 dollars, down to the lineThe thread runs from the claim to the „Total $285” row of the cost table; beside it, a live shop importing 4,000+ products.
Table 3 · openhaft.pl↗ kt-total
/notesA registry of thinking, with footnotesWriting about the field, each piece with a thesis and clickable sources; three theses are waiting in the approval queue.
0 posts · 3 theses§ 01–03/modelsNine models, nothing withheldEvery 2-bit Bielik variant with its weights, Hessians and evaluation logs — you can recompute the same numbers.
huggingface.co/Jakubrd43.26 GB · 2.4 bpw/workshopCode that looks the same when nobody is watchingml-serve: model registry, A/B, observability — MIT, readable line by line.
github.com/jakubprejznerregistry.py/contactOne sentence and an addressNo form, no pricing calculator. A project, a question about the paper, joint research — an email written to a person.
kontakt@bitsharp.pl→First-rate AI engineering — on evidence, not slogans.
BitSharp · AI engineering studio · RzeszówAreas of work
Question answering over company knowledge that cites its source down to the paragraph — hybrid retrieval, reranking, relevance evaluation on gold sets.
proof → notes/retrievalTool orchestration, failure handling, the point where a human steps in — measured with our own benchmark of 67 agentic tasks in Polish.
proof → PolAgentBenchInvoices, contracts, correspondence: schema enforced at decoding, deterministic validation, a human approves only the exceptions.
proof → deploymentsServing with full observability, model routing, quantization — from 22 GB down to 3.26 GB without losing the Polish.
proof → workshop · researchWorking standard
- Deterministic: fixed seed, temperature 0, results reproducible bit for bit.
- Commit-stamped: every result carries the hash of the code that produced it.
- Tests before GPU: 204 unit tests before a single run starts.
- Caveats stated plainly: methodological artifacts go in the body text, not a footnote.
- Open artifacts: models, data and evaluation logs are public.
- Deterministic code where it wins — an LLM only where the rules cannot be written down.
Founder
Jakub Prejzner — AI Engineer. Builds RAG systems, process automation and document data extraction; follows the field and ships from it as it moves. The quantization papers are proof he can take a subject down to research level.