--- license: apache-2.0 datasets: - agentica-org/DeepScaleR-Preview-Dataset language: - en base_model: - deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B --- # McGill-NLP/longcot-24k-1.5b

### TL;DR - **Markovian Thinking** for RL in reasoning LLMs: replace the trivial MDP where state = prompt + all past thinking tokens (quadratic compute) with a bounded, fixed-size state, yielding **linear compute** in thinking tokens and constant memory by design. - Delethink RL trains a model to “think” in **fixed-size chunks** with bounded state.. - This 1.5B model uses an effective **thinking budget of about 24K tokens** while only requiring an **8K active context** at any time via chunked rollouts and short carryovers. - Initialized from **`deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B`**, trained with the Delethink RL paradigm. See the paper for full details. ### Links - Repo: https://github.com/McGill-NLP/the-markovian-thinker - Paper: https://arxiv.org/abs/2510.06557v1 - Collection: [The Markovian Thinker](https://huggingface.co/collections/McGill-NLP/the-markovian-thinker-68debd2919c4ae47f50706cd) ## Model Summary - Base model: `deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B` - Objective: Reinforcement Learning using standard LongCoT, trained for 1000 steps. - Thinking 24K budget; uses the entire context. - Intended use: Math/logic reasoning with step-by-step derivations; final answer typically formatted inside LaTeX `\boxed{}`. - Library compatibility: Works well with SGLang for chunked inference; also usable with Transformers for standard generation. ## Intended Uses and Limitations - Intended uses: - Long-form reasoning on math and related tasks. - Bounded-context rollouts with repeated chunking and short carryovers. - Not intended for: - Safety-sensitive applications without human oversight. - Use cases requiring faithful, verifiable citations to external sources. - Limitations: - May hallucinate, make arithmetic/algebraic mistakes, or produce inconsistent plans. - The chunked rollout procedure is needed to realize Delethink’s efficiency advantages. ## Prompting - Use the model’s chat template and request a step-by-step solution with a final boxed answer: - “Please reason step by step, and put your final answer within \boxed{}.” ## Quickstart (SGLang, chunked Delethink rollout) ```python import asyncio import sglang as sgl def main(): llm = sgl.Engine( model_path="McGill-NLP/longcot-24k-1.5b", dtype="bfloat16", attention_backend="flashinfer", mem_fraction_static=0.8, log_level="WARNING", ) prompt = ( r"There exist real numbers $x$ and $y$, both greater than 1, such that " r"$\log_x\left(y^x\right)=\log_y\left(x^{4y}\right)=10$. Find $xy$." "\n\nPlease reason step by step, and put your final answer within \\boxed{}." ) tok = llm.tokenizer_manager.tokenizer query_ids = tok.apply_chat_template( [{"role": "user", "content": prompt}], tokenize=True, add_generation_prompt=True, ) params = {"temperature": 0.6, "max_new_tokens": 24576} ids = llm.generate(input_ids=query_ids, sampling_params=params, return_logprob=True) print(tok.decode(ids, skip_special_tokens=False)) if __name__ == "__main__": main() ``` ### Suggested generation settings - temperature: 0.6 - top_p: 1.0 - top_k: -1 ## Safety and Use - This model can produce incorrect or misleading reasoning steps and answers. Always verify results. - Do not deploy in high-stakes domains without human oversight. ## Citation ```bibtex @misc{Aghajohari2025:TheMarkovianThinker, title={The Markovian Thinker}, author={Milad Aghajohari and Kamran Chitsaz and Amirhossein Kazemnejad and Sarath Chandar and Alessandro Sordoni and Aaron Courville and Siva Reddy}, year={2025}, eprint={...}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={...}, } ```