Instructions to use aayanmishra-ml/Hermes-A1-20B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use aayanmishra-ml/Hermes-A1-20B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="aayanmishra-ml/Hermes-A1-20B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("aayanmishra-ml/Hermes-A1-20B") model = AutoModelForCausalLM.from_pretrained("aayanmishra-ml/Hermes-A1-20B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use aayanmishra-ml/Hermes-A1-20B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "aayanmishra-ml/Hermes-A1-20B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aayanmishra-ml/Hermes-A1-20B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/aayanmishra-ml/Hermes-A1-20B
- SGLang
How to use aayanmishra-ml/Hermes-A1-20B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "aayanmishra-ml/Hermes-A1-20B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aayanmishra-ml/Hermes-A1-20B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "aayanmishra-ml/Hermes-A1-20B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aayanmishra-ml/Hermes-A1-20B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Studio
How to use aayanmishra-ml/Hermes-A1-20B with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for aayanmishra-ml/Hermes-A1-20B to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for aayanmishra-ml/Hermes-A1-20B to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for aayanmishra-ml/Hermes-A1-20B to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="aayanmishra-ml/Hermes-A1-20B", max_seq_length=2048, ) - Docker Model Runner
How to use aayanmishra-ml/Hermes-A1-20B with Docker Model Runner:
docker model run hf.co/aayanmishra-ml/Hermes-A1-20B
Hermes-A1-20B
Hermes-A1-20B is a 20-billion parameter multilingual large language model (LLM) built on top of GPT-OSS-20B. Hermes-A1-20B extends the capabilities of the original model with enhanced multilingual understanding, generation, and reasoning, making it suitable for research and production applications across diverse languages.
The model is designed to perform a wide range of tasks, including natural language understanding, code completion, translation, summarisation, and complex reasoning, all with multilingual support.
Model Highlights
| Feature | Description |
|---|---|
| Base Model | GPT-OSS-20B |
| Parameters | 20B |
| Architecture | Transformer-based causal language model |
| Training Objective | Autoregressive causal language modeling |
| Multilingual Support | Enhanced embeddings for multiple languages (see metadata for full list) |
| Applications | Chatbots, text completion, translation, code generation, reasoning tasks |
Technical Overview
Hermes-A1-20B builds on GPT-OSS-20B while introducing several key enhancements:
Multilingual Tokenization and Embeddings
- Improved tokenization and embedding layers to handle multiple languages.
- Optimized for high-frequency languages as well as low-resource languages (coverage listed in metadata).
Architecture
- 20B parameters, 64 attention layers (example, adjust per your actual config), causal self-attention.
- Supports long-context sequences with memory-efficient attention.
Training Details
- Initialized from GPT-OSS-20B weights.
- Fine-tuned on a curated multilingual corpus.
- Mixed-precision training with distributed GPU clusters for efficiency.
Inference Optimization
- Supports batch and streaming generation.
- Can be deployed on GPU and CPU for research or production applications.
Supported Languages
Hermes-A1-20B supports multiple languages for both comprehension and generation. For the full list of languages, please check the model metadata on Hugging Face.
Example language families:
- English, Spanish, French, German, Portuguese
- Chinese (Simplified & Traditional), Japanese, Korean
- Hindi, Arabic, Russian, Turkish
- Other regional languages with partial coverage
Performance may vary depending on language resources and training data coverage.
Use Cases
Conversational AI and Multilingual Chatbots
- Engage in context-aware conversations across supported languages.
Text Generation and Completion
- Story writing, creative content generation, and automated summarization.
Code Generation & Comprehension
- Supports programming languages and natural language code prompts.
Multilingual Translation & Summarization
- Translate text between supported languages.
- Summarize documents in multiple languages.
Reasoning and Knowledge Tasks
- Handles multi-step reasoning queries, QA systems, and educational tasks.
Example Usage
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-generation", model="Spestly/Hermes-A1-20B")
messages = [
{"role": "user", "content": "Who are you?"},
]
pipe(messages)
Limitations
- Performance varies by language and domain; low-resource languages may be less accurate.
- May generate plausible but incorrect or biased outputs. Human oversight recommended.
- Not recommended for safety-critical applications without evaluation.
Citation
@misc{hermes-a1-20b,
title={Hermes-A1-20B: A Multilingual Large Language Model},
author={Aayan mishra},
year={2025},
url={https://huggingface.co/Spestly/Hermes-A1-20B/}
}
- Downloads last month
- 9