Instructions to use Skywork/Skywork-13B-Base-3.1TB with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Skywork/Skywork-13B-Base-3.1TB with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Skywork/Skywork-13B-Base-3.1TB", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Skywork/Skywork-13B-Base-3.1TB", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Skywork/Skywork-13B-Base-3.1TB with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Skywork/Skywork-13B-Base-3.1TB" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Skywork/Skywork-13B-Base-3.1TB", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Skywork/Skywork-13B-Base-3.1TB
- SGLang
How to use Skywork/Skywork-13B-Base-3.1TB with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Skywork/Skywork-13B-Base-3.1TB" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Skywork/Skywork-13B-Base-3.1TB", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Skywork/Skywork-13B-Base-3.1TB" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Skywork/Skywork-13B-Base-3.1TB", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Skywork/Skywork-13B-Base-3.1TB with Docker Model Runner:
docker model run hf.co/Skywork/Skywork-13B-Base-3.1TB
error "ModuleNotFoundError: No module named 'transformers_modules.Skywork.Skywork-13B-Base-3'"
Did you try other versions of Skywork-13B?
Did you try other versions of Skywork-13B?
13b math worked (that's the only other one I tested)
It appears that Hugging Face splits the model name based on dots. You can download the 3.1TB model parameters, rename the model to 3TB, and then use the model dictionary as input.
How do I exactly download the model parameters? sorry I'm new to this field
How do I exactly download the model parameters? sorry I'm new to this field
There is a tab "Files and versions", where you can click to download model parameter files, e.g. pytorch_model-00001-of-00053.bin
