Instructions to use fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2:Q4_K_M # Run inference directly in the terminal: llama cli -hf fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2:Q4_K_M # Run inference directly in the terminal: llama cli -hf fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2:Q4_K_M
Use Docker
docker model run hf.co/fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2 with Ollama:
ollama run hf.co/fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2 with Docker Model Runner:
docker model run hf.co/fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2:Q4_K_M
- Lemonade
How to use fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull fox3000foxy/Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2:Q4_K_M
Run and chat with the model
lemonade run user.Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2-Q4_K_M
List all available models
lemonade list
- Atomic Chat
⚠️ Luna-Protocol-1.5B-Discord-Dialogues-200k-base-v2 — EXPÉRIMENTAL / INSTABLE
Ce modèle n'est pas recommandé pour un usage en production. Il s'agit d'une itération corrective sur 200k-base, toujours basée sur la variante base (non-Instruct) de Qwen2.5-1.5B. Voir Luna-Protocol-1.5B-Discord-Dialogues (basé sur Qwen2.5-1.5B-Instruct) pour une version plus stable.
Ce qui change par rapport à la v1 (200k-base)
Suite au rapport qualitatif sur la v1, deux ajustements ont été apportés :
- Rôle
assistant→luna: le mot "assistant" porte une forte connotation IA-serviable dans les poids pré-entraînés, ce qui va à l'encontre de l'objectif de Luna Protocol (un modèle confondu avec un humain sur Discord, pas un assistant IA). Le dataset a été retraité pour renommer chaque tour<|im_start|>assistanten<|im_start|>luna, et le prompt de génération se termine désormais par<|im_start|>luna\nplutôt que<|im_start|>assistant\n. - Nouveau repo, pas de reprise de checkpoint : pour ne pas hériter d'un checkpoint entraîné avec l'ancien rôle
assistant, l'entraînement est reparti de zéro sur ce nouveau repo, avec un dataset préprocessé distinct (Discord-Dialogues-Preprocessed-Luna-Protocol-role-luna).
Le reste de la configuration est inchangé par rapport à la v1 : QLoRA r=16/lora_alpha=16, ~200 000 exemples filtrés 8–512 tokens, template ChatML configuré manuellement (le modèle base n'a ni chat_template ni EOS de conversation natifs).
❌ Problème connu : le renommage de rôle n'a pas résolu la non-terminaison
Malgré la correction du rôle, le problème de génération infinie/instable observé en v1 persiste. Le renommage assistant → luna corrige un biais de ton (moins de réponses "IA-serviable" génériques), mais n'agit pas sur la cause racine : un modèle base n'a jamais été entraîné à reconnaître la fin d'un tour de conversation, et 200k exemples en LoRA (1.5 epoch) ne suffisent pas à faire émerger ce comportement de façon fiable à partir de zéro.
Conclusion de cette série d'expériences
L'hypothèse initiale (repartir d'un modèle base pour un style Discord plus marqué) n'a pas été validée dans des conditions utilisables : le gain de ton, s'il existe, ne compense pas l'instabilité de génération. Le développement reprend donc sur une base Instruct, qui a déjà un comportement d'arrêt fiable avant même le fine-tuning — voir le notebook et le repo 200k-instruct.
Statut
Conservé à titre de trace expérimentale / comparaison. Non recommandé pour un usage réel.
Crédits
- Base model : Qwen2.5-1.5B (Qwen team, Alibaba Cloud)
- Training framework : Unsloth
- Dataset : mookiezi/Discord-Dialogues
- Downloads last month
- 648
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit