Qwen3.8 27B Core AI 128K

Core AI export of Qwen3.8-27B for Apple Silicon on macOS 27 or newer. The shared-weight multifunction bundle uses an S=16 prefill function and an S=1 GPU-pipelined decode function while retaining a 131,072-token dynamic KV-cache bound.

The transformer body uses INT4 linear weights. The vocabulary head uses symmetric INT8 block-32 weights, which improves decode throughput over an INT4 head while keeping the compiled payload smaller than the previous export.

This is an independent conversion of Qwen/Qwen3.8-27B at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0.

Layout and verification

The language bundle is at gpu-pipelined/qwen3_8_27b_decode_int4linh8_pf16. SHA256SUMS covers every published file. The compiled payload is 17,602,327,452 bytes with SHA-256 593a72b86c1f58d6648959ec08487d4caefb4e161726fea0e57259730c4dfc09.

Usage

Add Apple's coreai-models Swift package and point resourcesAt at the bundle directory containing metadata.json and tokenizer/:

import CoreAILanguageModels
import Foundation
import FoundationModels

let bundleURL = URL(
    filePath: "/path/to/gpu-pipelined/qwen3_8_27b_decode_int4linh8_pf16",
    directoryHint: .isDirectory
)
let model = try await CoreAILanguageModel(resourcesAt: bundleURL)
let session = LanguageModelSession(model: model)

session.prewarm()

let answer = try await session.respond(
    to: "Explain why the sky is blue."
).content
print(answer)

session.prewarm() primes both inference functions before the first response. Do not call either function directly or set COREAI_CHUNK_THRESHOLD: the runtime reads the bundle metadata and automatically calls prefill for complete 16-token prompt chunks before using main for the remainder and autoregressive decode.

Performance

Measured on the same Apple M4 Pro with 48 GB unified memory:

Harness Metric Previous export This export
llm-benchmark, 128-token prompt Prefill 10.897 tok/s 55.293 tok/s
llm-benchmark, 128-token generation Decode 10.572 tok/s 12.380 tok/s

The packaged-app canary also returned the expected response. Its streaming and token-accounting boundaries differ from llm-benchmark, so its throughput is used only as a functional check and is not reported here. Performance depends on OS version, prompt length, memory pressure, and thermal state.

Runtime contract

The host runtime must select the prefill function for complete 16-token prompt chunks and the main function for the remaining prompt tokens and generation. Both functions share model state and weights. This release does not include the previous decode-only artifact.

License

The converted model remains subject to the Apache License 2.0 supplied with the source model. See LICENSE. This conversion is independent and is not produced or endorsed by Qwen.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ETeissonniere/Qwen3.8-27B-CoreAI-128K

Base model

Qwen/Qwen3.8-27B
Finetuned
(513)
this model