Qwen3.8 27B Core AI 128K
Core AI export of Qwen3.8-27B for Apple Silicon on macOS 27 or newer. The shared-weight multifunction bundle uses an S=16 prefill function and an S=1 GPU-pipelined decode function while retaining a 131,072-token dynamic KV-cache bound.
The transformer body uses INT4 linear weights. The vocabulary head uses symmetric INT8 block-32 weights, which improves decode throughput over an INT4 head while keeping the compiled payload smaller than the previous export.
This is an independent conversion of Qwen/Qwen3.8-27B at revision
1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0.
Layout and verification
The language bundle is at
gpu-pipelined/qwen3_8_27b_decode_int4linh8_pf16. SHA256SUMS covers every
published file. The compiled payload is 17,602,327,452 bytes with SHA-256
593a72b86c1f58d6648959ec08487d4caefb4e161726fea0e57259730c4dfc09.
Usage
Add Apple's coreai-models Swift package and point resourcesAt at the bundle
directory containing metadata.json and tokenizer/:
import CoreAILanguageModels
import Foundation
import FoundationModels
let bundleURL = URL(
filePath: "/path/to/gpu-pipelined/qwen3_8_27b_decode_int4linh8_pf16",
directoryHint: .isDirectory
)
let model = try await CoreAILanguageModel(resourcesAt: bundleURL)
let session = LanguageModelSession(model: model)
session.prewarm()
let answer = try await session.respond(
to: "Explain why the sky is blue."
).content
print(answer)
session.prewarm() primes both inference functions before the first response.
Do not call either function directly or set COREAI_CHUNK_THRESHOLD: the runtime
reads the bundle metadata and automatically calls prefill for complete
16-token prompt chunks before using main for the remainder and autoregressive
decode.
Performance
Measured on the same Apple M4 Pro with 48 GB unified memory:
| Harness | Metric | Previous export | This export |
|---|---|---|---|
llm-benchmark, 128-token prompt |
Prefill | 10.897 tok/s | 55.293 tok/s |
llm-benchmark, 128-token generation |
Decode | 10.572 tok/s | 12.380 tok/s |
The packaged-app canary also returned the expected response. Its streaming and
token-accounting boundaries differ from llm-benchmark, so its throughput is
used only as a functional check and is not reported here. Performance depends on
OS version, prompt length, memory pressure, and thermal state.
Runtime contract
The host runtime must select the prefill function for complete 16-token prompt
chunks and the main function for the remaining prompt tokens and generation.
Both functions share model state and weights. This release does not include the
previous decode-only artifact.
License
The converted model remains subject to the Apache License 2.0 supplied with the
source model. See LICENSE. This conversion is independent and is not produced
or endorsed by Qwen.
Model tree for ETeissonniere/Qwen3.8-27B-CoreAI-128K
Base model
Qwen/Qwen3.8-27B