One dense giant. Sixty-four specialist minds. Morpho-180B takes the battle-tested open dense foundation and grows it into a routed collective: every token is served by the two domain experts that know it best, plus a shared generalist that never sleeps. Train each expert to convergence. Leave none behind.
β‘ Why Morpho
- π§ 180B total / ~78B active per request β giant knowledge, agile inference.
- π― 64 domain experts, zero overlap β every training example maps to exactly one expert. Auditable by construction.
- π Fully auditable data pipeline β raw β tagged β tokenized corpora, code included, in
Morpho-72B-MoE-Data. - π Resume-first engineering β every expert checkpoint (weights + loss + step) is versioned in
Morpho-72B-MoE-Checkpoints. Any box can pick up exactly where the last one died. - π 128K context via RoPE YaRN for long documents, codebases, and books.
π Architecture
| Component | Spec |
|---|---|
| Base | 72B dense backbone, 4-bit NF4 QLoRA (frozen) |
| Layers | 80 (40 dense even + 40 MoE odd) |
| Hidden | 8192 Β· 64 Q heads / 8 KV heads (GQA) Β· SwiGLU Β· RMSNorm |
| Router | Learned, top-2 of 8 per MoE layer + 1 shared expert |
| Adapters | LoRA rank-128 / alpha-256 on q,k,v,o,gate,up,down (~1B trainable params) |
| Context | 128K via RoPE YaRN |
| Total / active | ~180B / ~78B per request |
πΊ The 64 experts
Eight sectors β Finance, Law, Medicine, Tech, Science, Business, Arts, Humanities β each with 8 specialists. Full roster with per-expert data counts is logged at every training run; example counts range from ~10K to ~35K per expert, 200K+ examples total, each mapped to exactly one expert.
β»οΈ Data refinery β every byte auditable
π Build pipeline
flowchart LR
A["1 Β· Base<br/>72B dense"] --> B["2 Β· Data<br/>815K raw / 200K+ SFT"]
B --> C["3 Β· MoE structure<br/>64 routed experts"]
C --> D["4 Β· SFT<br/>per-expert to convergence"]
D --> E["5 Β· DPO<br/>preference alignment"]
E --> F["6 Β· Merge<br/>unified MoE"]
F --> G["7 Β· Verify<br/>benchmarks + audit"]
G --> H["8 Β· GGUF<br/>quantized release"]
H --> I["9 Β· Distill<br/>compact students"]
style D fill:#ff2d78,stroke:#00f0ff,color:#fff
Stage 4 Β· SFT is COMPLETE β all 61 data-bearing experts converged (phase-decayed LR 2e-4 β 5e-5 β 1e-5, patience 2000 steps, up to 10 epochs). Remaining: rebuild data for 31/52/54, then stage 5 Β· DPO.
π Convergence gallery β every expert trains to its floor
No early exits. Each of the 61 secured experts trained until its loss stopped dropping for 2000 straight steps β 15 of them reached the loss floor below 1e-3. The chart above is drawn from the verified best_loss.txt of every secured adapter.
SFT checkpoint status (live, 2026-10-11 β Master Plan v2)
61/64 secured and converged. SFT is complete for every expert with training data. The last three (rebuild queue below) need data first β nothing is skipped silently.
Provenance: the 2026-09-21/22 storage incident wiped one batch of
best/weights; every affected expert has since been genuinely retrained to convergence and re-verified byte-exact on HF. Old pre-incident loss numbers are reference only, never claims.
| Expert | Domain | Best loss | Weights |
|---|---|---|---|
| 00 | Trading & Markets | 0.0547 | β secured (fresh retrain, converged) |
| 01 | Investing & Portfolio | 0.2746 | β secured (fresh retrain) |
| 02 | Banking & Lending | 0.1307 | β secured (fresh retrain, converged) |
| 03 | Insurance & Risk | 0.4501 | β secured (fresh retrain, converged) |
| 04 | Accounting & Auditing | 0.1500 | β secured (fresh retrain, converged) |
| 05 | Taxation & Compliance | 0.0754 | β secured (fresh retrain, converged) |
| 06 | Real Estate Finance | 0.4678 | β secured (fresh retrain, converged) |
| 07 | Financial Engineering | 0.2545 | β secured (fresh retrain, converged) |
| 08 | Constitutional Law | 0.0692 | β secured (fresh retrain, converged) |
| 09 | Criminal Law | 0.2838 | β secured (fresh retrain, converged) |
| 10 | Civil Law | 0.0955 | β secured (fresh retrain, converged) |
| 11 | Corporate Law | 0.4710 | β secured (fresh retrain) |
| 12 | Intellectual Property | 0.3093 | β secured (fresh retrain, converged) |
| 13 | International Law | 0.3621 | β secured (fresh retrain, converged) |
| 14 | Tax Law | 0.0086 | β secured (fresh retrain) |
| 15 | Regulatory Compliance | 0.5570 | β secured (fresh retrain, converged) |
| 16 | Clinical Medicine | 0.1744 | β secured (fresh retrain, converged) |
| 17 | Pharmacology | 0.0966 | β secured (fresh retrain, converged) |
| 18 | Diagnostics & Imaging | 0.0015 | β secured (fresh retrain, converged) |
| 19 | Surgery & Procedures | 0.0035 | β secured (fresh retrain, converged) |
| 20 | Neurology | 0.2621 | β secured (fresh retrain, converged) |
| 21 | Cardiology | 0.3481 | β secured (fresh retrain, converged) |
| 22 | Oncology | 0.5547 | β secured (fresh retrain, converged) |
| 23 | Public Health | 0.1529 | β secured (fresh retrain, converged) |
| 24 | Machine Learning & AI | 0.0002 | β secured (fresh retrain, converged) |
| 25 | Systems & Infrastructure | 0.0016 | β secured (fresh retrain, converged) |
| 26 | Cybersecurity | 0.0781 | β secured (fresh retrain, converged) |
| 27 | Databases & SQL | 0.1779 | β secured (fresh retrain, converged) |
| 28 | Networking & Cloud | 0.2036 | β secured (fresh retrain, converged) |
| 29 | Web Development | 0.0003 | β secured (fresh retrain, converged) |
| 30 | DevOps & MLOps | 0.000059 | β secured (fresh retrain, converged) |
| 31 | Embedded & IoT | β | π§± rebuild queue (261 examples, see below) |
| 32 | Physics | 0.00246 | β secured (fresh retrain, converged) |
| 33 | Chemistry | 0.000000596 | β secured (fresh retrain, loss floor) |
| 34 | Biology & Life Sciences | 0.1974 | β secured (fresh retrain, converged) |
| 35 | Mathematics | 0.3517 | β secured (fresh retrain, converged) |
| 36 | Mechanical Engineering | 0.1394 | β secured (fresh retrain, converged) |
| 37 | Electrical Engineering | 0.0123 | β secured (fresh retrain, converged) |
| 38 | Civil Engineering | 0.1018 | β secured (fresh retrain, converged) |
| 39 | Environmental Science | 0.6756 | β secured (fresh retrain, converged) |
| 40 | Strategic Management | 0.0318 | β secured (fresh retrain, converged) |
| 41 | Marketing & Sales | 0.000285 | β secured (fresh retrain, converged) |
| 42 | Human Resources | 0.00213 | β secured (fresh retrain, converged) |
| 43 | Operations & Logistics | 0.0000397 | β secured (fresh retrain, converged) |
| 44 | Supply Chain | 0.0000859 | β secured (fresh retrain, converged) |
| 45 | Project Management | 0.00004 | β secured (fresh retrain, converged) |
| 46 | Entrepreneurship | 0.0931 | β secured (fresh retrain, converged) |
| 47 | Consulting | 0.2452 | β secured (fresh retrain, converged) |
| 48 | Creative Writing | 0.2855 | β secured (box-24 run, converged) |
| 49 | Graphic Design | 0.0003436 | β secured (box-24 run, loss floor) |
| 50 | Music & Audio | 0.00138 | β secured (box-25 run, converged) |
| 51 | Film & Video | 0.0006447 | β secured (box-25 run, loss floor) |
| 52 | Translation & Localization | β | π§± rebuild queue (246 examples, see below) |
| 53 | Linguistics | 0.0138 | β secured (box-26 run, converged) |
| 54 | Storytelling & Narrative | β | π§± rebuild queue (240 examples, see below) |
| 55 | Content Creation | 0.0002578 | β secured (box-27 run, loss floor) |
| 56 | Pedagogy | 0.4577 | β secured (box-27 run, converged) |
| 57 | Psychology | 0.5472 | β secured (box-27 run, converged) |
| 58 | Sociology | 0.00000834 | β secured (box-27 run, loss floor) |
| 59 | Economics | 0.00013 | β secured (box-30 run, loss floor, bf16) |
| 60 | Philosophy | 0.4239 | β secured (box-39 run, converged, bf16) |
| 61 | History | 0.3632 | β secured (box-39 run, converged, bf16) |
| 62 | Political Science | 0.0000000397 | β secured (box-40 run, loss floor, bf16) |
| 63 | Critical Analysis & Reasoning | 0.0000229 | β secured (box-40 run, loss floor, bf16) |
Precision note: adapters 00β58 are stored fp32 (6.7GB); from 59 on, best-saves are bf16 (3.4GB, max cast diff 6.1e-05) so pushes survive box egress. Merge casts everything to a common dtype β no effect on the final model.
Storage note (Oct 2026): the quota wall that stalled experts 59β61 is resolved β 8.28TB of orphaned LFS objects purged (byte-exact keep-set of the 57 live adapters verified before/after), history squashed, namespace trimmed to the 4 project repos. All 59 adapters live in this repo's checkpoint companion; nothing was lost.
π§± Rebuild queue β who needs more data and why
Three experts came out of the v5 tagging with too little data to train a stable LoRA adapter. The pipeline enforces a 500-example minimum: below that, a rank-128 adapter memorizes instead of generalizing, and the "converged" loss would be a lie. So these three are parked β not skipped β until their data is rebuilt:
| Expert | Domain | Today | Why it's short | Rebuild plan |
|---|---|---|---|---|
| 31 | Embedded & IoT | 261 ex. | Niche hardware Q&A is thin across the 17 sources | Targeted collection (Arduino / IoT StackExchange-style Q&A) β tag β tokenize β upload, then train |
| 52 | Translation & Localization | 246 ex. | Parallel-sentence pairs under-tagged in v5 pass | Targeted collection (gated parallel corpora, e.g. Tatoeba-style) β tag β tokenize β upload, then train |
| 54 | Storytelling & Narrative | 240 ex. | Long-form stories split across sectors in v5 pass | Targeted collection (prompt + story pairs, e.g. WritingPrompts-style) β tag β tokenize β upload, then train |
Rules for the rebuild: real data only (no synthesis), same bluemorpholimited/Morpho-72B-MoE-Data repo, same audit trail (raw β tagged β tokenized + code), same one-example-one-expert mapping, same convergence bar (patience 2000). When each crosses 500 verified examples it joins the training queue like any other expert.
π Collection running now (box 41): Arduino StackExchange + microcontroller multiturn β 31; OPUS-100 bitext (es/fr/de/it-en) β 52; WritingPrompts β 54. Tag β tokenize β upload β train follows automatically.
π₯ Live training dashboard
Checkpoints (weights + config + loss): Morpho-72B-MoE-Checkpoints Β· Data + code: Morpho-72B-MoE-Data
π Evaluation
End-to-end benchmarks (MMLU, GSM8K, HumanEval, MT-Bench + per-expert domain suites) run at stage 7 Β· Verify, after the merge. No scores are published before then β any numbers you see elsewhere are not from this project. This card will be updated with the full scoreboard on merge.
π Usage (after merge)
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("bluemorpholimited/Morpho-72B-MoE", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
"bluemorpholimited/Morpho-72B-MoE",
device_map="auto", trust_remote_code=True,
)
out = model.generate(**tok("Explain covered calls like a trader:", return_tensors="pt"), max_new_tokens=256)
print(tok.decode(out[0]))
β οΈ Today this repo hosts the model card + visual assets. Weights land here at stage 6 (merge). Per-expert SFT adapters are already available in the Checkpoints repo for research use.
𧬠Reproduce it
All data, tags, tokenized batches, and training code are public in the Data repo β 815K raw examples across 17 source datasets, deterministically mapped to 64 experts. The SFT loop (train_sft_experts.py) is resume-safe: kill the box mid-run and the next one continues from the exact best checkpoint.
βοΈ Limitations & safety
- SFT-stage adapters are domain specialists, not yet aligned (DPO pending) β expect raw, unfiltered completions in research use.
- 4-bit base + LoRA trades some precision for trainability on free-tier GPUs.
- Training data is web-scale; domain tags are heuristic β audit logs are provided for verification, not as a guarantee.
π License & credit
- Weights/code: Apache 2.0.
- Built by Blue Morpho β trained on free public compute (Hugging Face + molab GPU boxes), proving industrial-grade MoE training doesn't need a datacenter.
- If you use Morpho artifacts, cite:
bluemorpholimited/Morpho-72B-MoE (2026).
π· NO EXPERT LEFT BEHIND π·