Instructions to use nold/WSB-GPT-7B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use nold/WSB-GPT-7B-GGUF with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="nold/WSB-GPT-7B-GGUF", filename="WSB-GPT-7B_Q2_K.gguf", )
output = llm( "Once upon a time,", max_tokens=512, echo=True ) print(output)
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use nold/WSB-GPT-7B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf nold/WSB-GPT-7B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf nold/WSB-GPT-7B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf nold/WSB-GPT-7B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf nold/WSB-GPT-7B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf nold/WSB-GPT-7B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf nold/WSB-GPT-7B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf nold/WSB-GPT-7B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf nold/WSB-GPT-7B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/nold/WSB-GPT-7B-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use nold/WSB-GPT-7B-GGUF with Ollama:
ollama run hf.co/nold/WSB-GPT-7B-GGUF:Q4_K_M
- Unsloth Studio
How to use nold/WSB-GPT-7B-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for nold/WSB-GPT-7B-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for nold/WSB-GPT-7B-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for nold/WSB-GPT-7B-GGUF to start chatting
- Atomic Chat new
- Docker Model Runner
How to use nold/WSB-GPT-7B-GGUF with Docker Model Runner:
docker model run hf.co/nold/WSB-GPT-7B-GGUF:Q4_K_M
- Lemonade
How to use nold/WSB-GPT-7B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull nold/WSB-GPT-7B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.WSB-GPT-7B-GGUF-Q4_K_M
List all available models
lemonade list
Model Card for WSB-GPT-7B
This is a Llama 2 7B Chat model fine-tuned with QLoRA on 2017-2018ish /r/wallstreetbets subreddit comments and responses, with the hopes of learning more about QLoRA and creating models with a little more character.
Model Description
Developed by: Sentdex
Shared by: Sentdex
GPU Compute provided by: Lambda Labs
Model type: Instruct/Chat
Language(s) (NLP): Multilingual from Llama 2, but not sure what the fine-tune did to it, or if the fine-tuned behavior translates well to other languages. Let me know!
License: Apache 2.0
Finetuned from Llama 2 7B Chat
Demo [optional]: [More Information Needed]
Uses
This model's primary purpose is to be a fun chatbot and to learn more about QLoRA. It is not intended to be used for any other purpose and some people may find it abrasive/offensive.
Bias, Risks, and Limitations
This model is prone to using at least 3 words that were popularly used in the WSB subreddit in that era that are much more frowned-upon. As time goes on, I may wind up pruning or find-replacing these words in the training data, or leaving it.
Just be advised this model can be offensive and is not intended for all audiences!
How to Get Started with the Model
Prompt Format:
### Comment:
[parent comment text]
### REPLY:
[bot's reply]
### END.
Use the code below to get started with the model.
from transformers import pipeline
# Initialize the pipeline for text generation using the Sentdex/WSB-GPT-7B model
pipe = pipeline("text-generation", model="Sentdex/WSB-GPT-7B")
# Define your prompt
prompt = """### Comment:
How does the stock market actually work?
### REPLY:
"""
# Generate text based on the prompt
generated_text = pipe(prompt, max_length=128, num_return_sequences=1)
# Extract and print the generated text
print(generated_text[0]['generated_text'].split("### END.")[0])
Example continued generation from above:
### Comment:
How does the stock market actually work?
### REPLY:
You sell when you are up and buy when you are down.
Despite </s> being the typical Llama stop token, I was never able to get this token to be generated in training/testing so the model would just never stop generating. I wound up testing with ### END. and that worked, but obviously isn't ideal. Will fix this in the future maybe(tm).
Hardware
This QLoRA was trained on a Lambda Labs 1x H100 80GB GPU instance.
Citation
- Llama 2 (Meta AI) for the base model.
- Farouk E / Far El: https://twitter.com/far__el for helping with all my silly questions about QLoRA
- Lambda Labs for the compute. The model itself only took a few hours to train, but it took me days to learn how to tie everything together.
- Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke Zettlemoyer for QLoRA + implementation on github: https://github.com/artidoro/qlora/
- @eugene-yh and @jinyongyoo on Github + @ChrisHayduk for the QLoRA merge: https://gist.github.com/ChrisHayduk/1a53463331f52dca205e55982baf9930
Model Card Contact
harrison@pythonprogramming.net
Vanilla Quantization by nold, Model by WSB-GPT-7B
- Downloads last month
- 99
2-bit
4-bit
5-bit
8-bit