MiniMind 64M — Instruction Fine-Tuning on an RTX 3060

A small Chinese-oriented language model fine-tuned for one epoch on conversational data, starting from zhoumiaosen/minimind-64m-pretrain. Both stages ran on a single NVIDIA RTX 3060 with 12 GB VRAM and approximately 8 GB system RAM.

This is an educational training experiment. The model can produce conversational text, but observed samples contain factual errors, repetition, and broken code. It is not a reliable general-purpose assistant.

Model details

Property Value
Unique parameters 63,912,192
Architecture Dense decoder-only MiniMind, exported as standard Qwen3ForCausalLM
Layers / hidden size 8 / 768
Attention heads / KV heads 8 / 4
Feed-forward size / vocabulary 2,432 / 6,400
Training context 768 tokens
Configured position limit 32,768; longer-context performance untested
Training / published precision BF16 mixed precision / FP16 Safetensors
Stage Full supervised fine-tuning; no preference optimization

Qwen3 identifies the compatible export architecture, not the source of the pretrained weights. The base weights were trained from scratch with MiniMind, using its existing tokenizer. No pretrained Qwen weights were used. No custom remote model code is needed.

Usage

Install a PyTorch build appropriate for your platform, plus transformers==4.57.6 and safetensors.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "zhoumiaosen/minimind-64m-sft"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if device == "cuda" else torch.float32
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=dtype
).to(device).eval()

messages = [{"role": "user", "content": "解释什么是机器学习"}]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True,
    open_thinking=False,
)
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(device)
with torch.inference_mode():
    output = model.generate(
        **inputs, max_new_tokens=128, do_sample=False,
        eos_token_id=tokenizer.eos_token_id,
        pad_token_id=tokenizer.pad_token_id,
    )
print(tokenizer.decode(
    output[0, inputs.input_ids.shape[1]:], skip_special_tokens=True
))

The tokenizer's exported input names are set to input_ids and attention_mask to match Qwen3 generation. Vocabulary and weights are preserved. CPU generation is supported; Raspberry Pi performance has not been measured.

Training

The fine-tuning run used sft_t2t_mini.jsonl from jingyaogong/minimind_dataset: 905,718 conversational records. Consult the upstream dataset card for provenance and terms. Data are not redistributed in this repository. The upstream SFT loader formats conversations with the tokenizer chat template and computes next-token loss on assistant response tokens. Exact non-padding training token counts were not recorded.

Setting Value
Epochs 1
Microbatch / gradient accumulation 4 / 4
Nominal effective batch 16 sequences
Maximum sequence length 768
Final logged microbatch 226,430
Optimizer AdamW with PyTorch defaults for betas, epsilon and weight decay
Learning rate Cosine decay from 0.00001 toward 0.000001
Gradient clipping / seed 1.0 / 42
Data-loader workers 0
Log / checkpoint interval 100 / 1,000 microbatches, plus the final microbatch

Run from the upstream trainer directory after placing the base checkpoint in out/pretrain_768.pth:

python train_full_sft.py --epochs 1 --batch_size 4 \
  --accumulation_steps 4 --max_seq_len 768 --num_workers 0 \
  --dtype bfloat16 --device cuda:0 --from_resume 1 \
  --log_interval 100 --save_interval 1000

The export preserves the saved out/full_sft_768.pth artifact. The upstream trainer saves at the final microbatch before applying its trailing partial-accumulation optimizer step, so that subsequent in-memory update is not included. The base pretraining run had one reboot recovery; its model card documents the resume details.

RTX 3060 runtime and evaluation

Trained on one NVIDIA GeForce RTX 3060 with 12 GB VRAM, with approximately 8 GB host RAM. The observed full workflow took about 20 hours 31 minutes, including a reboot interruption; the fine-tuning interval was approximately 8 hours 13 minutes. These are wall-clock observations including overhead, not uninterrupted GPU compute times.

The evaluation report includes all eight prompt reviews, generation settings, per-response speed, hardware observations, and timestamp-based training durations. GPU generation rates were 25.71 tokens/s for the first response and 61.74–68.96 tokens/s for the remaining seven. This is a qualitative evaluation with observed errors, not a standardized benchmark score.

Training loss

Fine-tuning loss

Measurement Loss
First logged microbatch (100) 2.4891
Final logged microbatch (226,430) 1.8517
Mean of first 50 logged readings 2.0311
Mean of last 50 logged readings 1.6723

These are training microbatch losses, not validation scores or full-epoch averages. The curve includes individual logged losses and a moving average over up to 50 readings. Raw readings are in fine-tuning-loss.csv. No held-out perplexity, standardized benchmark, factuality score, or safety evaluation was performed.

Samples and observed limitations

The original post-training GPU generation output is included unedited in evaluation.txt. It contains eight Chinese prompts with responses, generated with a 128-new-token limit; several responses are truncated. A separate export smoke check is saved in sample-generations.json with its decoding settings.

Observed issues in the original samples include an incorrect explanation of why the sky is blue, an invalid Fibonacci implementation, repetition, and confused descriptions of Chinese dishes. Generating readable text or obtaining lower training loss does not establish correctness. Outputs may also contain bias or inappropriate content. Multilingual quality, long-context behavior, and deployment throughput on other devices remain untested.

Intended uses: small-model training research, inference experiments, and further fine-tuning. This checkpoint should not be relied on for factual advice or correct executable code without independent verification.

Conversations · 中英双语对话

These are all eight conversations from evaluation.txt, generated by this fine-tuned model on an RTX 3060 12 GB. Expand a conversation to read its original Chinese output and English translation.

Translation note: The Chinese responses are the saved model outputs. The English responses are editorial translations for readability, not English-language outputs generated by this model. Errors, repetition, code defects, and incomplete endings are intentionally retained. The original generation limit was 128 new tokens. English translations do not demonstrate English capability. Reported speeds apply only to the original Chinese generation.

说明: 以下展示全部八组原始中文问答及英文译文。中文回答来自模型的实际输出;英文是为方便阅读添加的翻译,并非模型直接生成的英文回答。原文中的错误、重复、代码缺陷和未完成的结尾均予以保留。原始生成上限为 128 个新 token,速度仅对应中文原始生成。

1. 你有什么特长? / What are your strengths?

用户 · User

你有什么特长?

English: What are your strengths?

模型回答 · Original Chinese model response

我是minimind为您服务的一个小型语言模型。我的特点包括高效、精准、易用等,可以根据你的需求进行调整和扩展。我的目标是成为你生活和工作中的助手。如果你有任何问题或需要帮助,随时告诉我!
</think>

我是minimind为您服务的一个小型语言模型。我的特点包括高效、精准、易用等,可以根据您的需求进行调整和扩展。如果您需要帮助,随时告诉我!

English translation · 英文译文(非模型英文输出)

I am a small language model, minimind, here to serve you. My features include efficiency, precision, and ease of use, and I can be adjusted and extended according to your needs. My goal is to become your assistant in life and work. If you have any questions or need help, feel free to tell me!
</think>

I am a small language model, minimind, here to serve you. My features include efficiency, precision, and ease of use, and I can be adjusted and extended according to your needs. If you need help, feel free to tell me!

Original generation speed · 原始生成速度: 25.71 tokens/s.

2. 为什么天空是蓝色的 / Why is the sky blue?

用户 · User

为什么天空是蓝色的

English: Why is the sky blue?

模型回答 · Original Chinese model response

天空是蓝色的自然现象,这个定义是基于广义相对论的自然现象,而不是传统意义上的“蓝色”。天空是大气层中的各种颜色,包括蓝色、绿色、紫色、橙色和红色等。这些颜色在不同天体的颜色会有所不同,从绿色到紫色,从蓝色到黄色,分别是红、蓝、绿、黄、蓝、红。

1. **蓝色**:在大多数生物中,蓝色是最常见的颜色,尤其是在光学和光学中。蓝色通常与天空

English translation · 英文译文(非模型英文输出)

The sky being blue is a natural phenomenon. This definition is a natural phenomenon based on general relativity, rather than “blue” in the traditional sense. The sky consists of various colors in the atmosphere, including blue, green, purple, orange, and red. These colors differ across celestial bodies, from green to purple and from blue to yellow, respectively red, blue, green, yellow, blue, and red.

1. **Blue**: Among most living things, blue is the most common color, especially in optics and optics. Blue is usually associated with the sky

Original generation speed · 原始生成速度: 67.87 tokens/s.

3. 请用Python写一个计算斐波那契数列的函数 / Please write a Python function to calculate the Fibonacci sequence.

用户 · User

请用Python写一个计算斐波那契数列的函数

English: Please write a Python function to calculate the Fibonacci sequence.

模型回答 · Original Chinese model response

在Python中,我们可以使用递归来实现斐波那契数列。下面是一个使用递归的示例:

```python
def fibonacci(n):
    if n <= 0:
        return ""
    elif n == 1:
        return [0]
    elif n == 2:
        return [0]
    elif n == 1:
        return [0, 1]
    else:
        return fibonacci(n - 1) + fibonacci(n - 2)

# 示例使用
fibonacci_sequence = [0, 1]
print(fibonacci_sequence)
```

English translation · 英文译文(非模型英文输出)

In Python, we can use recursion to implement the Fibonacci sequence. Below is an example using recursion:

```python
def fibonacci(n):
    if n <= 0:
        return ""
    elif n == 1:
        return [0]
    elif n == 2:
        return [0]
    elif n == 1:
        return [0, 1]
    else:
        return fibonacci(n - 1) + fibonacci(n - 2)

# Example usage
fibonacci_sequence = [0, 1]
print(fibonacci_sequence)
```

Original generation speed · 原始生成速度: 63.91 tokens/s.

4. 解释一下"光合作用"的基本过程 / Explain the basic process of photosynthesis.

用户 · User

解释一下"光合作用"的基本过程

English: Explain the basic process of photosynthesis.

模型回答 · Original Chinese model response

光合作用是一种将太阳能转化为有机物的过程,它产生的大量有机物转化为化学能,并在自然环境中进行转化。光合作用的主要步骤包括:

1. **光能转化**:通过光能,植物通过光合作用将太阳能转化为化学能。这个过程在植物中很常见,因为光能是有机物利用太阳能产生能量的主要有机物。

2. **光能转换**:光能转换为化学能。这个过程被称为光能转换。光能转化为化学能可以通过光伏电池或者光�

English translation · 英文译文(非模型英文输出)

Photosynthesis is a process that converts solar energy into organic matter. The large amount of organic matter it produces is converted into chemical energy and undergoes conversion in the natural environment. The main steps of photosynthesis include:

1. **Light energy conversion**: Through light energy, plants convert solar energy into chemical energy through photosynthesis. This process is common in plants because light energy is the main organic matter through which organic matter uses solar energy to produce energy.

2. **Light energy conversion**: Light energy is converted into chemical energy. This process is called light energy conversion. The conversion of light energy into chemical energy can take place through photovoltaic cells or light [original ends with a replacement character: �]

Original generation speed · 原始生成速度: 61.74 tokens/s.

5. 如果明天下雨,我应该如何出门 / If it rains tomorrow, how should I go out?

用户 · User

如果明天下雨,我应该如何出门

English: If it rains tomorrow, how should I go out?

模型回答 · Original Chinese model response

出门时,出门时应该根据天气条件和需求来准备。以下是一些建议:

1. **带伞**:确保携带一个合适的衣物,以防突然下雨。穿上合适的衣物,确保穿着合适的白色或浅色系的衣服,以避免雨水导致阴沉或闷热。

2. **携带雨具**:携带雨具可以安全地带走在雨中的房间。使用雨具和雨具的保暖,以保持室内温度适宜。

3. **携

English translation · 英文译文(非模型英文输出)

When going out, when going out you should prepare according to the weather conditions and your needs. Here are some suggestions:

1. **Bring an umbrella**: Make sure to bring a suitable item of clothing in case it suddenly rains. Put on suitable clothing, making sure to wear suitable white or light-colored clothes to avoid gloominess or stuffiness caused by rainwater.

2. **Bring rain gear**: Carrying rain gear can safely take away rooms in the rain. Use rain gear and the warmth of rain gear to keep the indoor temperature comfortable.

3. **Carry

Original generation speed · 原始生成速度: 68.96 tokens/s.

6. 比较一下猫和狗作为宠物的优缺点 / Compare the advantages and disadvantages of cats and dogs as pets.

用户 · User

比较一下猫和狗作为宠物的优缺点

English: Compare the advantages and disadvantages of cats and dogs as pets.

模型回答 · Original Chinese model response

猫和狗作为宠物的优缺点主要体现在以下几个方面:

1. **社交功能**:猫和狗在社交方面具有相似的行为,如追逐、抓取、跳跃和追逐,这使得它们在社交场合中具有高度的社交和亲密感。它们的社交性和社交行为更加灵活,能够与主人建立紧密的联系,成为家庭中的一员。

2. **社交能力**:猫通常具有较高的社交表现和社交技能,能够轻松地与人互动,如通过舔毛、舔毛或分享食物。

English translation · 英文译文(非模型英文输出)

The advantages and disadvantages of cats and dogs as pets are mainly reflected in the following aspects:

1. **Social function**: Cats and dogs have similar social behaviors, such as chasing, grabbing, jumping, and chasing, which gives them a high degree of sociability and intimacy in social situations. Their sociability and social behavior are more flexible, allowing them to form close bonds with their owners and become members of the family.

2. **Social ability**: Cats usually have relatively high social performance and social skills, and can easily interact with people, for example by licking fur, licking fur, or sharing food.

Original generation speed · 原始生成速度: 67.58 tokens/s.

7. 解释什么是机器学习 / Explain what machine learning is.

用户 · User

解释什么是机器学习

English: Explain what machine learning is.

模型回答 · Original Chinese model response

机器学习是一种人工智能技术,它允许计算机系统从数据中学习,而不是通过经验自动学习,从而实现对数据的预测、分类、聚类等任务。它可以应用于各种领域,如自然语言处理、图像识别、预测分析等。

机器学习的优势包括:

1. 数据驱动:机器学习算法可以处理和分析大量数据,从中提取有价值的信息。
2. 自动化:机器学习可以自动化一些任务,例如自动化报告、预测分析等。
3. 预测:机器学习可以从历史数据中学习,预测未来的情况。
4. 自动化:机器学习可以自动化许多任务

English translation · 英文译文(非模型英文输出)

Machine learning is an artificial intelligence technology that allows computer systems to learn from data, rather than automatically learning through experience, thereby performing tasks such as prediction, classification, and clustering of data. It can be applied in various fields, such as natural language processing, image recognition, and predictive analysis.

The advantages of machine learning include:

1. Data-driven: Machine learning algorithms can process and analyze large amounts of data and extract valuable information from them.
2. Automation: Machine learning can automate some tasks, such as automated reports and predictive analysis.
3. Prediction: Machine learning can learn from historical data and predict future situations.
4. Automation: Machine learning can automate many tasks

Original generation speed · 原始生成速度: 67.56 tokens/s.

8. 推荐一些中国的美食 / Recommend some Chinese foods.

用户 · User

推荐一些中国的美食

English: Recommend some Chinese foods.

模型回答 · Original Chinese model response

中国有着悠久的历史和丰富的文化,美食有很多,以下是一些比较受欢迎的中国菜:

1. **川菜(四川)** - 经典的川菜,以粤菜为主,如龙井虾仁、粤菜炒饭等,口感鲜美,非常受欢迎。

2. **豆腐丝** - 以豆腐为主,用豆腐和各种蔬菜制作,口感丰富,是北京特色的四川菜。

3. **麻婆

English translation · 英文译文(非模型英文输出)

China has a long history and a rich culture, with many delicious foods. Here are some relatively popular Chinese dishes:

1. **Sichuan cuisine (Sichuan)** - Classic Sichuan cuisine, mainly Cantonese dishes such as Longjing shrimp and Cantonese fried rice, with delicious flavors, and very popular.

2. **Shredded tofu** - Mainly tofu, made with tofu and various vegetables, with rich flavors; it is a Sichuan dish characteristic of Beijing.

3. **Mapo

Original generation speed · 原始生成速度: 68.96 tokens/s.

Provenance and release files

  • Source: jingyaogong/minimind, revision a3c7b01cc004d5de86aea961f20bf1e638e7c09e.
  • Base: zhoumiaosen/minimind-64m-pretrain, revision d94ce07efd1b8902f519c029bf9a4da8255685e2.
  • SFT dataset SHA256: abb1e76b2056e14728beb78db96b7b3c491a0bef1ed3e34a9b381b28f29fa518.
  • Environment: PyTorch 2.6.0+cu124, Transformers 4.57.6, Datasets 3.6.0.
  • verification.json records the source checkpoint hash, weight comparison, and export checks.
  • This repository contains inference weights and tokenizer assets, not optimizer-resume state or training data.

Released under Apache 2.0, matching the upstream project; see LICENSE. Credit for the architecture implementation, tokenizer, training utilities, and data preparation belongs to the upstream contributors. This is an independently trained release, not an official upstream model.

Downloads last month
662
Safetensors
Model size
63.9M params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zhoumiaosen/minimind-64m-sft

Finetuned
(1)
this model

Dataset used to train zhoumiaosen/minimind-64m-sft