Instructions to use rzzhan/ThinMQM-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use rzzhan/ThinMQM-8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="rzzhan/ThinMQM-8B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("rzzhan/ThinMQM-8B") model = AutoModelForCausalLM.from_pretrained("rzzhan/ThinMQM-8B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use rzzhan/ThinMQM-8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "rzzhan/ThinMQM-8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rzzhan/ThinMQM-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/rzzhan/ThinMQM-8B
- SGLang
How to use rzzhan/ThinMQM-8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "rzzhan/ThinMQM-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rzzhan/ThinMQM-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "rzzhan/ThinMQM-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "rzzhan/ThinMQM-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use rzzhan/ThinMQM-8B with Docker Model Runner:
docker model run hf.co/rzzhan/ThinMQM-8B
Update model card: correct license, add pipeline_tag and library_name
#1
by nielsr HF Staff - opened
This PR aims to significantly improve the model card by making the following updates:
- Corrected License: The
licensein the metadata has been updated frommittoapache-2.0to accurately reflect the license stated in both the current model card content and the associated GitHub repository. - Added Pipeline Tag: The
pipeline_tag: text-generationhas been added. This categorizes the model as a language model performing text generation tasks (generating reasoning/evaluations for MT) and improves its discoverability on the Hugging Face Hub. - Added Library Name: The
library_name: transformershas been included. Evidence fromconfig.json(transformers_version) and the GitHub README's acknowledgments (mentioningtransformersfor model/data loading) confirms compatibility. This will enable an automated "how to use" code snippet on the model page, enhancing user experience. - Enriched Content: The model card content has been expanded to include the paper abstract and key sections from the GitHub README, such as "Introduction", "Quick Start", "Configuration" (with detailed model templates and decoding recommendations), and "Meta-Evaluation". This provides a more comprehensive overview of the model, its usage, and its performance.
Existing relevant information, such as the paper link, GitHub repository link, and original badges, has been preserved.
Thank you for your service. I will check other repositories later this week.
rzzhan changed pull request status to merged