gemma4-12b-bioinfo GGUF

This repository contains GGUF files for gemma4-12b-bioinfo, a fine-tuned Gemma 4 12B model for bioinformatics and computational biology.

Use this repository for local inference with llama.cpp, LM Studio, Ollama-compatible workflows, or llama-cpp-python.

The original Hugging Face transformers model is available at: yashm/gemma4-12b-bioinfo.

Files

File Description Recommended use
gemma4-12b-bioinfo-Q4_K_M.gguf 4-bit quantized GGUF Recommended for most local GPU/CPU inference
gemma4-12b-bioinfo-BF16.gguf BF16 GGUF Higher fidelity, much larger memory requirement

Important Prompt Format

Use the Gemma 4 turn format below. Do not add <bos> manually for the GGUF prompt.

<|turn>user
Your question here
<|turn>model

Download

huggingface-cli download yashm/gemma4-12b-bioinfo-GGUF gemma4-12b-bioinfo-Q4_K_M.gguf --local-dir .

For the BF16 file:

huggingface-cli download yashm/gemma4-12b-bioinfo-GGUF gemma4-12b-bioinfo-BF16.gguf --local-dir .

Quick Start: llama.cpp CLI

cat > prompt.txt <<'EOF'
<|turn>user
Explain the role of CRISPR-Cas9 in genome editing in two concise sentences.
<|turn>model
EOF

./llama.cpp/build/bin/llama-cli \
  -m ./gemma4-12b-bioinfo-Q4_K_M.gguf \
  -f prompt.txt \
  -n 512 \
  -c 2048 \
  --temp 0.2 \
  --top-p 0.9 \
  -ngl 99

Use the BF16 GGUF by changing only the model path:

./llama.cpp/build/bin/llama-cli \
  -m ./gemma4-12b-bioinfo-BF16.gguf \
  -f prompt.txt \
  -n 512 \
  -c 2048 \
  --temp 0.2 \
  --top-p 0.9 \
  -ngl 99

Quick Start: llama-cpp-python

from llama_cpp import Llama

llm = Llama(
    model_path="./gemma4-12b-bioinfo-Q4_K_M.gguf",
    n_ctx=2048,
    n_gpu_layers=-1,   # use GPU offload when available; set 0 for CPU-only
    verbose=False,
)

question = "Explain the significance of CRISPR-Cas9 in functional genomics."
prompt = f"<|turn>user\n{question}<|turn>model\n"

output = llm(
    prompt,
    max_tokens=512,
    temperature=0.2,
    top_p=0.9,
    repeat_penalty=1.1,
    stop=["<|turn>user", "<eos>"],
    echo=False,
)

print(output["choices"][0]["text"].strip())

To use the BF16 GGUF in Python, change only:

model_path="./gemma4-12b-bioinfo-BF16.gguf"

Suggested Settings

Setting Value
Context length 2048
Temperature 0.2
Top-p 0.9
Repeat penalty 1.1
Stop strings `["<
GPU layers -ngl 99 in llama.cpp or n_gpu_layers=-1 in llama-cpp-python

Intended Use and Limitations

This model is intended for research, education, and computational biology assistance. It is not a medical device and should not be used for clinical diagnosis, treatment decisions, or professional medical advice. Always verify outputs against trusted databases, literature, and qualified experts.

Citation

@misc{gemma4-12b-bioinfo_2026_gguf,
  author       = {yashm},
  title        = {gemma4-12b-bioinfo GGUF: Fine-Tuned Gemma 4 12B for Bioinformatics},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/yashm/gemma4-12b-bioinfo-GGUF}}
}
Downloads last month
136
GGUF
Model size
12B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yashm/gemma4-12b-bioinfo-GGUF

Quantized
(2)
this model