Model Card: Darkroom-R1-1.1BModel OverviewModel Name: Darkroom-R1-1.1B Base Model: TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T Architecture: 1.1B-parameter dense transformer (Llama-based architecture), 32,000 fixed vocabulary Tuning Method: 4-bit QLoRA Cold-Start SFT + GRPOTrainer (Reinforcement Learning via Verifiable Rewards) Target Behavior: Structured ...... chain-of-thought generation without expanding tokenizer vocabularies Intended Use & DomainPrimary Scope: Experimental proof-of-concept testing whether sub-2B parameter models can internalize dual-stage reasoning protocols (distillation + verifier RL) on single-GPU hardware. Out-of-Scope: Production arithmetic, critical logical deduction, factual reference, or unverified automated decision systems.Technical SpecificationsSpecificationConfigurationBase Quantization4-bit NF4, double quantization, FP16 compute PEFT AdapterLoRA ($r=16, \alpha=32$, dropout=0.0) Target Modulesq_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj Trainable Weights12,615,680 / 1,112,664,064 (1.1338%) Training Budget~10.5 hours on 1x Tesla T4 (14.56 GB VRAM) Context Length1,024 tokens Known LimitationsArithmetic Hallucination: Struggles with direct calculation beyond single-digit integers without explicit multi-step decomposition.Ungrounded Entity Induction: Tends to map abstract relational problems into hallucinated physical properties (such as assigning heights or colors to non-physical premises).Greedy Attention Traps: Susceptible to degenerate repetitive token cycling under deterministic greedy decoding (do_sample=False).

Downloads last month
9
Safetensors
Model size
1B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for toomuchcomputedan/darkroom-r1-1.1b

Finetuned
(109)
this model

Dataset used to train toomuchcomputedan/darkroom-r1-1.1b