Instructions to use upstage/SOLAR-0-70b-16bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use upstage/SOLAR-0-70b-16bit with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="upstage/SOLAR-0-70b-16bit")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("upstage/SOLAR-0-70b-16bit") model = AutoModelForCausalLM.from_pretrained("upstage/SOLAR-0-70b-16bit", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use upstage/SOLAR-0-70b-16bit with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "upstage/SOLAR-0-70b-16bit" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "upstage/SOLAR-0-70b-16bit", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/upstage/SOLAR-0-70b-16bit
- SGLang
How to use upstage/SOLAR-0-70b-16bit with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "upstage/SOLAR-0-70b-16bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "upstage/SOLAR-0-70b-16bit", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "upstage/SOLAR-0-70b-16bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "upstage/SOLAR-0-70b-16bit", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use upstage/SOLAR-0-70b-16bit with Docker Model Runner:
docker model run hf.co/upstage/SOLAR-0-70b-16bit
Lacking documentation of datasets used, architecture, fine-tuning procedures, source code
Intriguing to see Solar as "a great example of the progress enabled by open source". However, for a model claiming this, Solar is remarkably silent about its own sources. The dataset details are given as:
Orca-style dataset
Alpaca-style dataset
It would be helpful to document exactly which datasets and which versions have been used, and to specify the instruction tuning process and overall architecture in more detail.
Currently, this model trails the very bottom of the openness leaderboard: it is more closed and less documented than even Llama2 itself. Hoping to see this improve!
Well, they were enabled by open source, but they are clearly NOT open source. No dataset, no code, and no commercial use. "Weights available for evaluation/personal use" is not open. One might assume that since they said "alpaca-style dataset" that they are using an actual Alpaca variant - and as Alpaca is CC-by-NC, they may feel they must then restrict to CC-by-NC; but cc-by-nc also requires attribution and "alpaca-style" is NOT an attribution and it would mean, imo, that they were breaching the Alpaca terms. Strange days.