Instructions to use anthracite-org/magnum-v4-12b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anthracite-org/magnum-v4-12b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="anthracite-org/magnum-v4-12b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("anthracite-org/magnum-v4-12b") model = AutoModelForCausalLM.from_pretrained("anthracite-org/magnum-v4-12b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use anthracite-org/magnum-v4-12b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "anthracite-org/magnum-v4-12b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "anthracite-org/magnum-v4-12b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/anthracite-org/magnum-v4-12b
- SGLang
How to use anthracite-org/magnum-v4-12b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "anthracite-org/magnum-v4-12b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "anthracite-org/magnum-v4-12b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "anthracite-org/magnum-v4-12b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "anthracite-org/magnum-v4-12b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use anthracite-org/magnum-v4-12b with Docker Model Runner:
docker model run hf.co/anthracite-org/magnum-v4-12b
repetitive
I must say this has been a very interesting model to work with and get to know. its smart. but it has an issue of starting to get repetitive a ways into a roleplay as if it cant handle anything more than 5k context tokens.
It always starts off strong. then gradually and slowly loses its mind. if it kept up with its strong start up to 8k context id absolutely love this mode. cranking up the repeat penalty tokens helps a tiny bit but not much.
Using mistral instruct template and i keep the context tokens to 5k. but even at 5k it gets a little repetitive but not as bad as 8k.
I want to love this model, so i hope the repetition is addressed eventually
I must say this has been a very interesting model to work with and get to know. its smart. but it has an issue of starting to get repetitive a ways into a roleplay as if it cant handle anything more than 5k context tokens.
It always starts off strong. then gradually and slowly loses its mind. if it kept up with its strong start up to 8k context id absolutely love this mode. cranking up the repeat penalty tokens helps a tiny bit but not much.
Using mistral instruct template and i keep the context tokens to 5k. but even at 5k it gets a little repetitive but not as bad as 8k.
I want to love this model, so i hope the repetition is addressed eventually
This is more or less an issue with every model. It really helps to write more detailed answers yourself that the model can make something out of because giving short repetitive answers yourself will lead to the model doing the same. You can always tweak settings in silly tavern or whatever you're using but the most important thing is to give detailed answered yourself in roleplays and usually the model will be more creative as well when it gets more input.
I must say this has been a very interesting model to work with and get to know. its smart. but it has an issue of starting to get repetitive a ways into a roleplay as if it cant handle anything more than 5k context tokens.
It always starts off strong. then gradually and slowly loses its mind. if it kept up with its strong start up to 8k context id absolutely love this mode. cranking up the repeat penalty tokens helps a tiny bit but not much.
Using mistral instruct template and i keep the context tokens to 5k. but even at 5k it gets a little repetitive but not as bad as 8k.
I want to love this model, so i hope the repetition is addressed eventually
This is more or less an issue with every model. It really helps to write more detailed answers yourself that the model can make something out of because giving short repetitive answers yourself will lead to the model doing the same. You can always tweak settings in silly tavern or whatever you're using but the most important thing is to give detailed answered yourself in roleplays and usually the model will be more creative as well when it gets more input.
thats not quite what ive found. ive seen models handle things just fine, moving the story or situation along just fine without things becoming too repetitive. for instance, Undi95__Lumimaid-Magnum-v4-12B-GGUF__Lumimaid-Magnum-v4-12B.q8_0.gguf, this model is handling the repetitive nature considerably better. cant say why, dont really understand it. this one however handles 8k context quite nicely with very little repetition.
Mind if i ask what quantization you used and what did you use to inference it? i can try and troubleshoot if this is an actual problem, as my own testing with EXL2 went very well without any issues.