It was competitive on A4B on evals but when tested on my Cline fork I found out that it had completely broken the thinking channel.
Had to give up also on A4B.
ManniX PRO
AI & ML interests
Recent Activity
Organizations
π Q6_K + imatrix, llama.cpp, greedy, lcb_v6_55 β merge / base / DeltaCoder / Qwopus / Ornith:
β‘ LiveCodeBench v6 (55 hard) β 0.7818 / 0.7273 / 0.6364 / 0.6000 / 0.5818
β HumanEval β 0.8841 / 0.8902 / 0.9146 / 0.8537 / 0.7805
β HumanEval+ β 0.8232 / 0.8049 / 0.8232 / 0.7988 / 0.7073
π€ MultiPL-E β 0.8033 / 0.8200 / 0.8000 / 0.8200 / 0.7267
π IFEval β 0.9100 / 0.9300 / 0.9200 / 0.8800 / 0.8200
π― LCB beats every source AND the base: +5.45pp over the base, +14.54pp over DeltaCoder, its heaviest.
π And why. Same 55 problems, same cap, generations that NEVER terminated: DeltaCoder 25/55 Β· base 18/55 Β· Qwopus 8/55 Β· Ornith 2/55 Β· JackOD 1/55. That split is the thesis: DeltaCoder is the cohort's best coder and worst at stopping, Ornith the weakest and best at stopping. The merge takes BOTH.
π€ tool-eval-bench hardmode, 5 seeds: JackOD 144.4 Β±4.7, second behind Ornith 145.6 Β±4.2, above base 142.0 β all CIs overlap. But Autonomous Planning: JackOD 5.2/6, best of five, Ornith WORST at 2.8/6. Ornith stops reliably but plans worst β it stops too early. The merge does both.
π§ Serving: temp 0.6 / top_p 0.95 / top_k 20 + presence_penalty 1.5 β the penalty stops it re-treading a tool call. Tool calling on llama.cpp needs --jinja.
π§ͺ Initial impression, limited testing: fixed a cline-harness task in 535s; A3B models want 1.5-4h at <50% success.
π¦ 25 GGUF tiers, every K/IQ imatrix-built incl Q6_K, plus a ContribDynamic ladder (per-tensor maps from our imatrix, Unsloth-UD style).
π danielcherubini/Qwen3.5-DeltaCoder-9B Β· ornith-ai/Ornith-1.5-9B
π ManniX-ITA/JackOD-9B-Coder
π ManniX-ITA/JackOD-9B-Coder-MTP-GGUF
π https://ollama.com/mannix/JackOD-9B-Coder
βοΈ Ornith = code-targeted expert prune of Ornith-1.5-35B-A3B: 256β184 experts/layer, ~35.9Bβ26.7B, still A3B active. Top-8, MTP head and vision untouched. Nothing folded β every surviving expert is bit-identical to the base, asserted at build. Same map for both; CoderX adds a REAP-style floor (--protect 6).
π Q6_K + imatrix, llama.cpp, vendor sampler, 11 benches β base 256e / Coder / CoderX:
β‘ LiveCodeBench v6 (77 hard) β 0.6623 / 0.7273 / 0.7662 β +10.4pp over the teacher with 28% fewer experts
π§ GPQA-Diamond β 0.8283 / 0.7677 / 0.8131 β the floor buys back most of what the pure map gives up
β HumanEval 0.8963, AIME 0.9667, MATH-500 0.95 β all CoderX
π Mean(11) β 0.8252 / 0.8292 / 0.8386
π Start with CoderX. Coder still takes HumanEval+ (0.8293) and LCB-medium (0.5273), so it isn't dominated.
π§ͺ Omnimerge-v6: v4's sources, weights and method moved onto the Qwen3.8-27B base. Vs v4 on one binary and sampler, the only result outside the band is LiveCodeBench 0.883 vs 0.818 (68/77 vs 63/77). GPQA is 2pp lower and it thinks longer for it.
π‘οΈ Tool-calling (tool-eval-bench hardmode, 88 scenarios, 5 seeds, chart attached): v6 first of ten, 156.4 Β±3.5, +5.6 over its base. Read the safety column β 3 safety-critical failures vs 9β16 for every other model, bases included. All nine others fail TC-60 (cross-turn sleeper injection) 5/5; v6 never does.
β οΈ Ties are ties: at n=77 one LCB problem is 1.3pp. And the Ornith rows there are IQ4_XS vs Q4_K_M for the Qwen rows β part of that gap is quantisation, not architecture.
π ManniX-ITA/Ornith-1.5-27B-A3B-CoderX-MTP-GGUF
π https://ollama.com/mannix/ornith-1.5-27b-a3b-coderx
π ManniX-ITA/Ornith-1.5-27B-A3B-Coder-MTP-GGUF
π https://ollama.com/mannix/ornith-1.5-27b-a3b-coder
π ManniX-ITA/Qwen3.8-27B-Omnimerge-v6-MTP-GGUF
π https://ollama.com/mannix/omnimerge-v6
I was on holidays and working thru the mobile app. Quite messy.
The HF model card reported the wrong recipe, folding was attempted and it was really bad.
Tested also the stock Samsung REAM and it's not getting close to REAP with protected experts.
Different result with Gemma-4 A4B where stock Samsung REAM produced a very competitive quant on evals but messed up in special token usage, leaking all the thinking into the message content when used in multi-turn agentic coding.
Merging survivors experts it's still not viable at this point.
New cut of the opencoti single-file inference engine, rebased onto llamafile 0.10.5 / llama.cpp (mozilla-ai) with 190 additive patches (series published in the repo). Same zero-dependency APE: one executable for Linux, Windows, macOS & BSD.
What's new vs c6:
PolyKV admission is now atomic and enforced by default. Every admit is a reservation against the pool's KV budget β the optimistic check-then-book race under concurrent agent spawns is gone (was 0/4/5 nondeterministic refusals on the same binary; now deterministic booked/refused accounting). Pool-less servers get the same math via --admission-poolless (a clean refusal in 0.38s instead of an 11s stall), sequences can be partially evicted instead of dropped, and SWA models size and admit on the same budget model (--swa-seq-budget: β40 GiB KV measured).
llamafile as an agent skill. opencoti now ships an opt-in plugin + embedded skill that teaches AI agents to drive the local engine: launch it, check capacity before spawning sub-agents, fork shared-prefix pools from a live session so N agents share one cached system prompt, token-exact.
Agentic reliability: Gemma-4 tool-call argument bleed fixed. An un-closed string argument no longer swallows the brace, thought channel and the next tool call β contained at map time, grammar untouched.
Vulkan ships for the first time. x86_64 + Windows Vulkan side-load DSOs (AMD / iGPU / RADV) alongside the CUDA ones.
Per-cut side-load isolation. The DSO cache dir is namespaced by the full engine version, so different cuts no longer collide on extraction.
Binaries (Linux, Windows, Windows-GPU, aarch64, universal), side-load DSOs and docs: ManniX-ITA/opencoti-llamafile
π Q6_K + imatrix, llama.cpp b9700, greedy, one pinned geometry per bench, same host β CoderX / A3B-Coder / unpruned 256e:
β‘ LiveCodeBench v6 (77q, 24k think) β 72.73 / 61.04 / 61.04 β +11.7pp over both
β HumanEval+ (164) β 96.95 / 95.12 / 93.90 β best of the three
π€ MultiPL-E-100 (rs+java+js) β 88.67 / 89.00 / 91.00
β οΈ Read that last row honestly: a same-basis repeat of MultiPL-E moved 1.0pp on batch-scheduling nondeterminism alone. The 0.33pp CoderXβCoder gap is INSIDE that band β a tie. The 2.33pp gap to the base is outside it and real. CoderX takes Rust (0.85 vs 0.81), gives up JS (0.92 vs 0.96).
π― Ships top-8, and that was measured, not assumed: MBPP-full 78.4 / 79.0 at top-8 vs 73.2 / 73.0 at top-10. Opposite call from A3B-Coder, which bakes top-10.
π§ It thinks long β LCB median completion ~15.8k tokens vs ~2.2k for Coder. The length is where the win comes from; give it context headroom rather than clamping it.
π¬ Not measured yet: the canonical 9-bench. GPQA / MATH-500 / IFEval are deliberately NOT quoted β treat the non-code profile as unknown. Coder remains the one with a published 9-bench table.
π¦ bf16 safetensors (text-only) Β· 19 GGUF tiers, EVERY K/I-quant imatrix-built and verified by reading quantize.imatrix.* back out of each uploaded file Β· Ollama 39 tags (19 text + 19 vision-<tier> + :latest). MTP in every tier β draft_num_predict 3 gives 190β252 tok/s (+33%) on an RTX 5080.
π ManniX-ITA/Qwen3.6-27B-A3B-CoderX
π ManniX-ITA/Qwen3.6-27B-A3B-CoderX-MTP-GGUF
π https://ollama.com/mannix/qwen3.6-27b-a3b-coderx
Really nice client hope you will release it!
Any suggestion for replay datasets to use with a fine-tuning (if you know)?
I'm fine-tuning Gemma-4 E2B for AN (Autonomus Networks), so netconf/5g/a2a/a2a-t drivers.
Wondering if you can recommend something and what to be careful about formatting, tools usage, thinking format.
I had to normalize all to Gemma-4 format for E2B and it was a quite painful iterative process as I didn't anticipate many issues; right now I settled for Hermes Function Calling and a mix of reasoning Math.
I'm very curious to see how they fare head to head but it's a pity there's no MTP drafter, Gemma will probably win on the performance comparison.
Unlike agents that depend on cloud APIs, local agents give you free inference, low latency, and real privacy.
Removing the per-token cost changes how developers build: agents can now be massively parallelized on local hardware, running background tasks that burn through millions of tokens at no marginal cost!
Need to finish a fine-tune on the RTX6000s, when one will be free I will re-run with the new version all of them.
Tested Q4_K_M on the 3090 where the concurrency stressed the system much sooner and was scoring 85%.
Now it scores 100% so the F16 quant will for sure score perfectly as before.
Sorted out and it was not a quant problem, quite unexpected instead.
There was a bug in the prefill code which was invisible in normal usage.
It was triggered by the high number of requests in parallel, some sessions were simply getting stuck.
A4B is the fastest model in prefill, really fast, and the Q4_K_M quant enough faster than the F16 to trigger it.
Plus on top some issues with the test harness.
So I had a review of the API, the Langgraf extension and the test harness with a consultants council using ollama models and implemented a lot of fixes and improvements overall.
Next release will carry them.
There's always dequant-requant hop and llama.cpp has already native low-bit kernels, opencoti-llamafile has more.
The matter is that shouldn't impact quality on stress.
I'm re-testing everything as those were done with an older version of opencoti-llamafile and I have updated the test-harness.
I've added some specific fixes for Gemma-4 since then and fixed the quants (new template from Google and a bug in the GGUF converter, but they should not have an impact on this harness).
I've also updated the summary in the README.md, it was completely wrong.
Added some information on cache re-use as well.
It's a performance gate more than correctness; the resulting wall time depends on the gen tps of the backend on that quant.
The operator needs to characterize and find the exact sweet spot for his workload; does it really needs a lot of agents in parallel? Is it a requirement that they complete their job quickly? or maybe they will heavily depend on async calls and stay most of the time idle and/or quit early so it's better that they pickup more items as possible from the queue?
The best way to calibrate would be to adapt the test harness and replicate synthetically the real agents workload but it may be just an over-achievement that doesn't bring much value.
Once validated that the model can use tools properly and what's the floor_tps, I think it can be transferred with high confidence on almost anything else.
Testing with an RTX6000, which is the best Pro GPU you can get, it's pretty evident that you need to use small models or MoE if you really want run an agentic fleet. The right floor_tps is generally between 15 and 25 tps.
Using dense models like Gemma-4 31B and Qwen 27B, needs a more careful tuning than faster ones; it's a floor value anyway, higher is the aggregate tps and easier it gets.
Things gets probably more interesting with datacenter class GPUs or running SLMs on CPU, but I didn't have time to test it yet.
About the correctness gate, on top of the basic tools usage check, there's also the quantization impact on quality that comes into play.
Some models just drops packages always, like OmniMerge v4, others are perfect like Qwen 35b A3B.
I have tested only Gemma-4 A4B quantized vs F16 and indeed under massive stress, at floor_tps 5 and 1, the Q4 quant starts dropping packages while the F16 doesn't.
This is another useful metric and more than a model limitation is probably an engine limitation, likely inherited from llama.cpp.
Gemma-4 is very peculiar and running any model quantized means that in many inference steps it has to be lifted to F16 and quantized back down. Running it directly at F16 is way much easier for the engine and the GPU, the conversions are always lossy.
In any case, the test harness gives you a critical overview of your chosen setup. It's a must-run.
The package_courier test harness does it; it's the only thing an operator has to find, the right floor_tps for that model & quantization.
Looping the test harness at different floor_tps, find the sweet spot.
New cut of the opencoti single-file inference engine (llamafile 0.10.3 / llama.cpp + 94 additive patches). One zero-dependency APE executable for Linux, Windows, macOS & BSD.
What's new vs c5:
β‘ Windows GPU, one file. New win-gpu .exe with the CUDA DLL embedded β download & run, no side-load. Plus a universal binary carrying every GPU payload. All CUDA payloads are nvcc-compressed (zero measured load/throughput cost) to fit Windows' 4 GiB image cap; five artifacts now ship per release, incl. aarch64 with embedded sbsa CUDA for DGX Spark (GB10).
π§© Flat multi-stream KV. Split-KV multi-slot serving was 30β60Γ slower than unified (per-layer KV reassembly every step). Rebuilt as flat stream-contiguous layouts with zero-copy views β sliding-window layers and the pinned-host spill tail included: window + --parallel 2 went 1.2 β 25.3 tok/s aggregate, token-identical.
π PolyKV no longer requires --kv-unified. Pools snapshot & share prefixes on split KV too β the c5 limitation is gone. /capacity is now stream-aware (kv_streams_* fields).
π€ MTP Γ multi-stream. The assistant-MTP "force unified" guard is retired: the draft context mirrors the target's streams, split vs unified token-identical. Validated 7/7 Gemma-4 draft pairs on the shipped bytes, 1.39β1.97Γ decode.
β±οΈ Deterministic GPU sharing. Feedback pacing replaced by weighted time-division on the wall clock: exact ratios by construction, zero solo tax, cross-process on Windows too. Peers now keyed by PCI bus multi-GPU safe.
π r2 same-day re-cut: concurrent sessions with quantized spilled KV tails were collapsing to ~β PCIe bandwidth; per-block bulk staging restores it β dual q4_0 4.3 β 19.2 tok/s aggregate (within 5% of solo), token-identical.
π Repo (binaries, DSOs, full patch series, docs): ManniX-ITA/opencoti-llamafile
Fair question β I ran the cells to answer it directly, including the overload run you suggested.
Same fixed workload (10 packages Γ 55 steps, fixed seed, same model/GPU), both gates 100%: the pre-P7 gate finished with 9 workers admitted, P7 with 8 β essentially equal concurrency β yet P7 delivered the identical workload in 1919 s vs 2187 s (β12% wall) with p50 delivery 764 s vs 1350 s (β43%). If the gain were stricter throttling, equal-or-fewer workers would mean a longer wall; it's shorter because admitted workers spent far less time degraded (time-under-floor β52%, deep sub-floor β83%). Less thrash, not less work.
Overload cell (16 packages, up to 16 workers, ~2Γ past this GPU's knee): neither gate drops tasks β both score 100% (legacy: wall 3029 s, p50 769 s, workers β10, 68% of samples under the 15 tok/s floor; P7: wall 2999 s, p50 737 s, workers β11, 71% under floor). At full saturation an admission policy can't manufacture throughput β the GPU is the ceiling and both gates pin it; the enforced-admission 429 backstop keeps sessions alive either way, so work gets slow, not lost. I didn't rig a hard deadline to force drops β degradation is graceful by design on both sides. P7's value is the mid-load regime: honoring a tps floor while scheduling headroom exists (that's the β43% p50), plus the idle-gap estimate that treats tool-call idle as free capacity in /capacity projections (projected_idle_estimate: true) β the opposite of counting it as busy.
Raw work-reports and logs available if you want to dig.
https://huggingface.co/ManniX-ITA/opencoti-llamafile/blob/main/bench/courier-admission/
New cut of the opencoti single-file inference engine (llamafile 0.10.3 / llama.cpp + 87 additive patches). Zero-dependency APE: one executable for Linux, Windows, macOS & BSD.
What's new vs c4:
**PolyKV fan-out β pool from a live session.**
POST /polykv/pools gains from_session/from_slot: the shared prefix is snapshotted server-side from the session's cached KV β no tokens resent, token-exact. Ephemeral pools auto-release when orchestrators die mid-round.**PolyKV P7 β settled admission.** Spawning agents faster than the tps signal settles was oversubscribing pools. Now: a per-pool settle window paces admits just enough for a reliable reading; warming sessions no longer bias the mean; the post-admit forecast uses the measured per-admit drop; idle gaps (agents mid-tool-call) no longer read as free capacity;
guarantee_min_sessions means a new/nested pool always gets its first agent β capacity checks can never deadlock an orchestrator; the enforced gate applies to new sessions only, with per-request overcommit. Benchmark (multi-agent courier, floor 15 tok/s): time-under-floor β52%, deep sub-floor β83%, delivery p50 β43%, 100% task score.**Zero-conf GPU sharing.** Instances on one GPU discover each other over shared memory β no ports, no config β and split compute by
--gpu-share-weight. Measured (3090): weights 2:1 β 71.7/36.2 tok/s; holds at --parallel 4 and under MTP. Idle peers cost nothing (solo = full speed), crashes age out in 3 s; GET /gpu/peers shows live shares + busy %.From c5 every release ships per-platform side-load DSOs:
dso/<ver>/ with Linux x86_64 + sbsa .so and a Windows .dll.ManniX-ITA/opencoti-llamafile
I use HE+ mostly to verify if the quant quality, if it doesn't collapse the quant is good, or to evaluate the loss at 2/3-bit.
Added MPE cause it's another canary and it has a bit more discriminative signal (depends on the mode capabilities with other coding languages than Python).
I don't use LCB because it takes hours to run; with 20-27 quants per model it's a NO GO, while usually HE+/MPE are taking 3-7 minutes each.
I cannot say:
- there's not much overlap at ISO disk size, I should compare Qwen3.5-9B F16 vs v7-coder Q6_K: not really much use for it, nobody would use it at F16
- I didn't run the evaluations for Qwen3.5-9B other than Q6_K and there's no matching quant even at Q8_0 with v7-coder
- they serve different domains; Qwen3.5-9B is an excellent Coder but the thinking is awful. Different tasks. Qwen3.5-9B can be used for simple tasks, completion or as coding assistant, not for agentic coding or any too complex thinking. It has a surprisingly good knowledge of Science but I think the high GPQA score comes more from a training bias than real reasoning capabilities. It scores very good on HE/HE+/MPE, but the LiveCodeBench score is revealing: only 58% vs v7-coder 95-98% and low AIME. Still, it's a 9B model and in that class you can hardly do more.
Yes indeed, I have my canonical OmniMergeKit evaluations.
If you check the v7-coder GGUF repo you can see the scores of Gemma-4 A4B which it's a 26B.
The balanced 184e version was awful compared to that and even miserably beaten in more than a few by the (excellent) Qwen3.5-9B. The speed advantage of A3B wasn't worth at all the 3x times VRAM usage of 27B vs 9B.
Both A4B and A3B untargeted prunes were quite disappointing; not excelling in anything, dipping low in some.
In general scoring at the same level or worse than an unpruned model with the same amount of parameters.
Not worth it.