Instructions to use ManniX-ITA/Qwen3.6-27B-A3B-Coder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ManniX-ITA/Qwen3.6-27B-A3B-Coder with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ManniX-ITA/Qwen3.6-27B-A3B-Coder") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ManniX-ITA/Qwen3.6-27B-A3B-Coder") model = AutoModelForMultimodalLM.from_pretrained("ManniX-ITA/Qwen3.6-27B-A3B-Coder", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ManniX-ITA/Qwen3.6-27B-A3B-Coder with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ManniX-ITA/Qwen3.6-27B-A3B-Coder" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ManniX-ITA/Qwen3.6-27B-A3B-Coder", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ManniX-ITA/Qwen3.6-27B-A3B-Coder
- SGLang
How to use ManniX-ITA/Qwen3.6-27B-A3B-Coder with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ManniX-ITA/Qwen3.6-27B-A3B-Coder" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ManniX-ITA/Qwen3.6-27B-A3B-Coder", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ManniX-ITA/Qwen3.6-27B-A3B-Coder" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ManniX-ITA/Qwen3.6-27B-A3B-Coder", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ManniX-ITA/Qwen3.6-27B-A3B-Coder with Docker Model Runner:
docker model run hf.co/ManniX-ITA/Qwen3.6-27B-A3B-Coder
I really like what I'm seeing here
This model is closing in on a sweet spot. I think the prune and the bump of active experts to 10 together was a good approach.
Would it interest you to do a parallel prune to 184 experts using the same calibrating dataset with REAM instead of REAP, and see how they compare? https://github.com/SamsungSAILMontreal/ream
Can you publish the expert mapping to create compatible merges from existing finetunes?
I really think you're onto something, and multiple specialized prunes of Qwen3.6-35B-A3B and gemma-4-26B-A4B-it are attractive for mixture-of-agents setups!
Definitely yes, I will test it and compare it.
Thanks, very interesting!
Can you publish the expert mapping to create compatible merges from existing finetunes?
Missed it, it's published now.
Look at the GitHub repo: https://github.com/mann1x/omnimergekit
I always publish the scripts and maps there (unless I miss it).
Scripts (the expert-mapping toolchain)
- recipes/qwen3_6_35b_a3b_prune/make_drop_map.py — the mapping script. Reads a competence map, ranks experts per layer (--score tc|wnorm|wmax, --agg sum|wmax|wsum, --cat-weight, --floor-count), drops the lowest N/layer, writes the drop map
{"0":[ids],…,"39":[ids],"mtp":[ids]}. - recipes/qwen3_6_35b_a3b_prune/expert_drop_qwen35b.py — consumes a drop map and performs the actual 256→184 expert prune (keeps vision + shared expert + linear-attn, slices the native MTP head).
- recipes/qwen3_6_35b_a3b_prune/build_calib_corpus_qwen.py — builds the base router calib corpus (Qwen-templated, per-bench balanced).
- recipes/qwen3_6_35b_a3b_prune/gen_lcb_calib_corpus.sh + harvest_lcb_calib_corpus.py — generate and PASS-filter the coder (LCB) calib corpus.
- gemma4/neuron_analysis/expert_neuron_analysis_v5_targeted.py — the competence producer (arch-generic; same script for Gemma-4 and Qwen3.5-MoE, auto-detects layer.mlp.experts). Emits the competence_qwen35b*.json maps.
Maps & calib corpora (recipes/qwen3_6_35b_a3b_prune/results/)
Balanced baseline (already on main):
- drop_map_184e.json — the --agg sum REAP domain-blind cut (the base the Coder is built on).
- competence_qwen35b.json — competence map input.
- router_calib_corpus_qwen.jsonl — base calib corpus.
Coder-targeted (PR #8, published today):
- coder (LCB) — drop_map_184e_coder.json + competence_qwen35b_coder.json; targeting --agg wmax --cat-weight corpus_targeted_lcb=2.0 + floor-clamp.
- coder_lcbmpe — drop_map_184e_coder_lcbmpe.json + competence_qwen35b_coder_lcbmpe.json; adds a MultiPL-E channel.
- coder_lcbmpeife — drop_map_184e_coder_lcbmpeife.json + competence_qwen35b_coder_lcbmpeife.json; adds an IFEval channel (anti-rumination).
- router_calib_corpus_coder_lcbmpe_qwen.jsonl — LCB+MultiPL-E coder calib corpus.
Pipeline: producer → competence_qwen35b*.json → make_drop_map.py → drop_map_184e*.json → expert_drop_qwen35b.py → 184e model. Per-variant details are in that dir's STATE.md (P2 = balanced, P3 = coder).
I ran it. Thanks for the push — it turned into the most useful ablation this model has had.
Setup. Same base, same 184/256 expert budget, same calibration data for every arm, Q6_K + imatrix, greedy, one host, one binary. Arms:
- ours — the published cut (competence-map saliency, selection only)
- REAM — stock Samsung REAM (REAP saliency + group merge,
group_size=16) - REAP-select — REAP saliency, selection only, no merge
- base — the unpruned 256e model, as the anchor
Results
MultiPL-E 100 (300 items) and HumanEval+ (164):
| arm | MPE | HE+ |
|---|---|---|
| ours (published) | 0.730 | 0.970 |
| REAM | 0.720 | 0.957 |
| base 256e | 0.717 | 0.951 |
| REAP-select | 0.717 | 0.970 |
Jitter band on this basis is ±1.0 pp on MPE (two bit-identical builds read 0.730 / 0.740), so everything above is one band wide. On code pass-rate, REAM and our cut are tied.
The separation shows up on hard LiveCodeBench — v6, 77 hard problems, greedy, 12k thinking / 32k total:
| arm | LCB | loops | truncated |
|---|---|---|---|
| REAP-select | 0.649 | 5/77 | 6/77 |
| ours | 0.610 | 12/77 | 6/77 |
| base 256e | 0.610 | 18/77 | 7/77 |
| REAM | 0.468 | 60/77 | 55/77 |
The honest reading
Two things, and the second one matters more than the first.
1. REAM's low LCB number is not a quality number. Its median completion is 32,166 tokens against a 32,768 ceiling — the median answer is sitting on the wall. That score is truncation-taxed and I won't compare it to anything until the cap is moved. A 24k-thinking / 48k-total re-run over all arms on the same 77 problems is running now; I'll post it when it lands.
2. REAM's merge is close to a no-op here, and that's why it's safe. Driven by REAP saliency, the greedy grouping lands on groups where the centroid holds ~64% of the merged weight — so "merge" mostly means "keep the centroid", and REAM ≈ REAP-select. When the merge actually averages, it hurts, monotonically with dose. Holding saliency fixed at ours and sweeping group_size:
| group_size | MPE | HE+ |
|---|---|---|
| 2 | 0.653 | 0.939 |
| 4 | 0.597 | 0.896 |
| 16 | 0.330 | 0.732 |
Same method, same budget, same data — 0.33 MPE at the dose REAM ships with. So on this model the merging ingredient isn't buying capacity; the ordering that keeps the merge from doing anything is what's carrying it.
What is actually different between the two cuts
I recovered both keep sets from the weights (byte-matching survivors against the base, so this isn't read off a log). Ours and REAP-select disagree on 38.5 experts per layer — 20.9% of the keeps, 53% of the 72-expert drop decision — and land on top of each other on every code bench. On LCB, 29 of 77 problems flip, in both directions, for a net +3.
Two criteria with almost nothing in common reach equally good models by half-different routes. That's the real finding, and it's not what I expected going in.
It isn't that expert choice is free, though. A control arm scored with plain row-norm instead of task-attributed competence moves only 16.7 experts per layer off our map — less than half the REAP disagreement — and costs 5 pp of MPE and nearly triples the LCB loop rate (34/77 vs 12/77). Direction matters; distance doesn't.
Still running
- the 24k/48k cap-validation pass over all 9 columns (the REAM row above is the one it exists to settle)
- two hybrid arms: our ranking with a REAP-derived stability floor, at two doses
I'll follow up here with the full assessment once both finish.
That is very helpful to know, thank you! That's good, disciplined investigation.
Follow-up as promised — the 24k/48k cap-validation pass is done, all 9 columns on one basis.
First, the control. Raising the ceiling is inert for arms that were not clipped: our cut, REAP-select, the base and the gs=2 arm each reproduced their 32k solve sets exactly, problem for problem. And the published GGUF vs the same weights rebuilt locally scored 47/77 with zero per-problem disagreements. So anything that moved, moved because of its own truncation.
Hard LiveCodeBench v6, 77 hard problems, greedy, 24k thinking / 48k total
| arm | solved | truncated | median tokens |
|---|---|---|---|
| REAP-select (no merge) | 50/77 | 5 | 1,967 |
| row-norm control | 48/77 | 7 | 6,181 |
| ours (published) | 47/77 | 6 | 2,157 |
| base 256e | 47/77 | 7 | 2,739 |
| stock REAM | 46/77 | 44 | 49,152 |
| ours + merge gs=2 | 41/77 | 17 | 4,554 |
On REAM specifically — I was right to hold the number, and the answer favours REAM
My earlier caveat was correct: REAM's 32k score was truncation-taxed. Unclipped it goes 0.468 → 0.597 (36 → 46 problems). That is a real +13 pp, against a replicate band of zero.
So the honest head-to-head is 46 vs 47 — one problem. On capability, stock REAM and our cut are indistinguishable on this bench. That is more favourable to REAM than my first reply implied, and I want that on the record rather than leaving the 0.468 impression standing.
Two caveats, both real:
0.597 is still a floor, not a score. Its median completion sits exactly on the new ceiling and 44/77 are still length-terminated. REAM was clipped at 32k and is still clipped at 48k. I stopped there deliberately — an arm whose median answer is 49k tokens has already answered the practical question, and a third ceiling would burn a day of GPU to refine a number that cost has already decided.
The cost gap is the whole story: median 49,152 tokens vs 2,157. ~23× the tokens for one fewer solve. That is the same thing the 32k loop counts showed (60/77 vs 12/77), now measured on the axis that actually matters when you serve the thing.
Interesting sub-result: REAP saliency used for selection only, with REAM's merging switched off, is the best non-hybrid arm at 50/77 — ahead of both our cut and the unpruned base. The merging is not what's helping.
One more thing, and it cuts against my own model
Our 184e cut and the unpruned 256e both solve 47/77 — but not the same 47. Only 33 are shared; each solves 14 the other misses. So "the prune costs nothing on hard LCB" is true only as an aggregate. The pruned model is a genuinely different solver that happens to break even, and I'd been quoting that parity more strongly than the data supports.
Thanks again for the nudge toward REAM — the ablation it forced has been the most informative thing I've run on this model.
This is excellent information. I'm looking into a third alternative or complementary strategy: using expert mapping from REAP or REAM to quantize priority experts less, low-priority experts more, with specialized focus for the quantization. Excessive REAM merging is a strong hint that making low-priority experts too noisy is a concern - so REAP to weighted quantization without REAM starts looking like the best way to avoid reasoning loops.
I definitely think that REAP, then weighted quantization has potential as a staged pipeline to make a highly efficient specialist from existing generalist MoEs.