Post
538
🚀 JackOD-9B-Coder — a 9B merge built to FINISH agentic coding tasks. Four-way omnimerge_v2 over Qwen3.5-9B: Jack = Qwopus3.5-9B-Coder (0.30), O = Ornith-1.5-9B (0.15), D = DeltaCoder (0.55). MTP head kept.
📊 Q6_K + imatrix, llama.cpp, greedy, lcb_v6_55 — merge / base / DeltaCoder / Qwopus / Ornith:
⚡ LiveCodeBench v6 (55 hard) — 0.7818 / 0.7273 / 0.6364 / 0.6000 / 0.5818
✅ HumanEval — 0.8841 / 0.8902 / 0.9146 / 0.8537 / 0.7805
✅ HumanEval+ — 0.8232 / 0.8049 / 0.8232 / 0.7988 / 0.7073
🤝 MultiPL-E — 0.8033 / 0.8200 / 0.8000 / 0.8200 / 0.7267
📋 IFEval — 0.9100 / 0.9300 / 0.9200 / 0.8800 / 0.8200
🎯 LCB beats every source AND the base: +5.45pp over the base, +14.54pp over DeltaCoder, its heaviest.
🛑 And why. Same 55 problems, same cap, generations that NEVER terminated: DeltaCoder 25/55 · base 18/55 · Qwopus 8/55 · Ornith 2/55 · JackOD 1/55. That split is the thesis: DeltaCoder is the cohort's best coder and worst at stopping, Ornith the weakest and best at stopping. The merge takes BOTH.
🤝 tool-eval-bench hardmode, 5 seeds: JackOD 144.4 ±4.7, second behind Ornith 145.6 ±4.2, above base 142.0 — all CIs overlap. But Autonomous Planning: JackOD 5.2/6, best of five, Ornith WORST at 2.8/6. Ornith stops reliably but plans worst — it stops too early. The merge does both.
🔧 Serving: temp 0.6 / top_p 0.95 / top_k 20 + presence_penalty 1.5 — the penalty stops it re-treading a tool call. Tool calling on llama.cpp needs --jinja.
🧪 Initial impression, limited testing: fixed a cline-harness task in 535s; A3B models want 1.5-4h at <50% success.
📦 25 GGUF tiers, every K/IQ imatrix-built incl Q6_K, plus a ContribDynamic ladder (per-tensor maps from our imatrix, Unsloth-UD style).
🙏 danielcherubini/Qwen3.5-DeltaCoder-9B · ornith-ai/Ornith-1.5-9B
🔗 ManniX-ITA/JackOD-9B-Coder
🔗 ManniX-ITA/JackOD-9B-Coder-MTP-GGUF
🔗 https://ollama.com/mannix/JackOD-9B-Coder
📊 Q6_K + imatrix, llama.cpp, greedy, lcb_v6_55 — merge / base / DeltaCoder / Qwopus / Ornith:
⚡ LiveCodeBench v6 (55 hard) — 0.7818 / 0.7273 / 0.6364 / 0.6000 / 0.5818
✅ HumanEval — 0.8841 / 0.8902 / 0.9146 / 0.8537 / 0.7805
✅ HumanEval+ — 0.8232 / 0.8049 / 0.8232 / 0.7988 / 0.7073
🤝 MultiPL-E — 0.8033 / 0.8200 / 0.8000 / 0.8200 / 0.7267
📋 IFEval — 0.9100 / 0.9300 / 0.9200 / 0.8800 / 0.8200
🎯 LCB beats every source AND the base: +5.45pp over the base, +14.54pp over DeltaCoder, its heaviest.
🛑 And why. Same 55 problems, same cap, generations that NEVER terminated: DeltaCoder 25/55 · base 18/55 · Qwopus 8/55 · Ornith 2/55 · JackOD 1/55. That split is the thesis: DeltaCoder is the cohort's best coder and worst at stopping, Ornith the weakest and best at stopping. The merge takes BOTH.
🤝 tool-eval-bench hardmode, 5 seeds: JackOD 144.4 ±4.7, second behind Ornith 145.6 ±4.2, above base 142.0 — all CIs overlap. But Autonomous Planning: JackOD 5.2/6, best of five, Ornith WORST at 2.8/6. Ornith stops reliably but plans worst — it stops too early. The merge does both.
🔧 Serving: temp 0.6 / top_p 0.95 / top_k 20 + presence_penalty 1.5 — the penalty stops it re-treading a tool call. Tool calling on llama.cpp needs --jinja.
🧪 Initial impression, limited testing: fixed a cline-harness task in 535s; A3B models want 1.5-4h at <50% success.
📦 25 GGUF tiers, every K/IQ imatrix-built incl Q6_K, plus a ContribDynamic ladder (per-tensor maps from our imatrix, Unsloth-UD style).
🙏 danielcherubini/Qwen3.5-DeltaCoder-9B · ornith-ai/Ornith-1.5-9B
🔗 ManniX-ITA/JackOD-9B-Coder
🔗 ManniX-ITA/JackOD-9B-Coder-MTP-GGUF
🔗 https://ollama.com/mannix/JackOD-9B-Coder