Noema Overfit

Paged MoE model bundles for Noema's Overfit expert-paging runtime. Each subfolder is one .noema-paged package: a resident.gguf (all non-expert weights, always loaded) plus experts-*.bin page files that are streamed and evicted on demand, described by manifest.json.

These are runtime-specific packages, not standalone GGUF files. They require a Noema Overfit runtime compatible with native contract v3.

Models

`gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.noema-paged/`

Base model google/gemma-4-26B-A4B-it-qat-q4_0-unquantized
Source GGUF unsloth/gemma-4-26B-A4B-it-qat-GGUF โ€” UD-Q4_K_XL (14.25 GB)
Architecture gemma4 (128 experts, 8 active, 30 MoE layers, fused gate/up)
Resident weights resident.gguf โ€” 1.40 GB
Expert pages experts-000.bin โ€” 12.96 GB
Alignment 16384 bytes

`Qwen3.6-35B-A3B-UD-Q4_K_M.noema-paged/`

Base model Qwen/Qwen3.6-35B-A3B
Architecture qwen35moe (256 experts, 8 active, 40 MoE layers)
Source GGUF Qwen3.6-35B-A3B-UD-Q4_K_M.gguf (22.13 GB)
Resident weights resident.gguf โ€” 2.57 GB
Expert pages experts-000.bin, experts-001.bin โ€” 19.57 GB total
Alignment 16384 bytes

`Qwen3.5-122B-A10B-Q4_K_M.noema-paged/`

Architecture qwen35moe (256 experts, 8 active, 48 MoE layers)
Source GGUF Qwen3.5-122B-A10B-Q4_K_M (2 shards, 74.2 GB)
Resident weights resident.gguf โ€” 4.0 GB
Expert pages experts-000.bin โ€ฆ experts-004.bin โ€” 65 GB
Alignment 16384 bytes

Packages are generated with Noema's paged-model conversion tooling. File sizes, source fingerprints, and SHA-256 integrity hashes are recorded in each package's manifest.json.

Downloads last month
62
GGUF
Model size
2B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for NoemaAI-labs/Noema-Overfit

Quantized
(147)
this model