Instructions to use AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL") model = AutoModelForCausalLM.from_pretrained("AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL
- SGLang
How to use AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL with Docker Model Runner:
docker model run hf.co/AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL
Overview
The AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL model is derived from the DeepSeek-R1 Qwen Distill 7B base model, optimized through confidence-based Reinforcement Learning using GRPO for enhanced intrinsic confidence calibration.
It relates to the paper Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence.
Purpose and Capabilities
This calibration process significantly improves the reliability of the model’s internal confidence signals. The model is optimized for use with the Guided by Gut (GG) framework, a self-guided test-time scaling (TTS) strategy that leverages these intrinsic confidence signals to perform complex reasoning tasks efficiently—without costly external verifier models.
Guided by Gut (GG) Framework
Traditional TTS methods often require substantial computational resources due to their reliance on external verifier models like Process Reward Models (PRMs) or extensive sampling strategies (e.g., Best-of-N). The GG framework provides a powerful yet computationally efficient alternative:
- Intrinsic signals: Utilizes token-level confidence and step novelty from the model itself.
- Confidence Calibration via RL: Refines intrinsic confidence through targeted RL fine-tuning.
- Efficiency: Offers significant reductions in GPU memory and inference speed, enabling smaller models (1.5B parameters) to compete with or outperform significantly larger models (32B–70B parameters).
Key Advantages of GG:
- Up to 10× less GPU memory compared to traditional PRM-based methods.
- 8× faster inference speed than PRM-based approaches.
- 50% lower KV cache memory usage compared to Best-of-N strategies.
Calibration and Training
The RL fine-tuning is computationally efficient and minimal:
- Dataset: Fine-tuned on the LIMO dataset for 3 epochs.
- Hardware: Completed using 2 NVIDIA A100 80GB GPUs.
- Time: Approximately three day.
More Information
- GitHub Repository: 👨💻 Amirhosein-gh98/Guided-by-Gut
- Research Paper: 📄 "Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence" on arXiv.
- Downloads last month
- 15
