kundeshwar20 commited on
Commit
d184429
ยท
verified ยท
1 Parent(s): 3525bbd

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +25 -15
README.md CHANGED
@@ -3,18 +3,28 @@ language:
3
  - hi
4
  - en
5
  pipeline_tag: text-generation
 
 
 
 
 
 
 
 
6
  ---
7
 
8
 
9
  <div align="center">
10
- <img src="https://huggingface.co/bharatgenai/Param-1-2.9B-Instruct/resolve/main/BharatGen%20Logo%20(1).png" width="60%" alt="BharatGen" />
 
 
11
  </div>
12
  <hr>
13
  <div align="center">
14
- <!-- <a href="https://arxiv.org/abs/2507.13390" target="_blank" style="margin: 4px;">
15
- <img alt="Paper" src="https://img.shields.io/badge/%20Paper-arxiv-0033ad?style=flat&logo=arxiv&logoColor=white" />
16
- </a> -->
17
- <a href="https://huggingface.co/bharatgenai/Param-1-5B/blob/main/LICENSE" target="_blank" style="margin: 4px;">
18
  <img alt="License" src="https://img.shields.io/badge/License-yellow.svg" />
19
  </a>
20
  </div>
@@ -27,7 +37,7 @@ The model is pretrained from scratch and designed to serve as a strong foundatio
27
 
28
  ---
29
 
30
- ## โœจ Key Highlights
31
 
32
  * **5B parameter** dense Transformer model
33
  * **Bilingual**: English and Hindi
@@ -37,7 +47,7 @@ The model is pretrained from scratch and designed to serve as a strong foundatio
37
 
38
  ---
39
 
40
- ## ๐Ÿš€ Model Inference
41
 
42
  ```python
43
  from transformers import AutoTokenizer, AutoModelForCausalLM
@@ -73,7 +83,7 @@ print(tokenizer.decode(output[0], skip_special_tokens=True))
73
 
74
  ---
75
 
76
- ## ๐Ÿ“Š Benchmarks
77
 
78
  | Task | **Param-1-5B (PT)** |
79
  | ----------------- | ------------------- |
@@ -95,7 +105,7 @@ print(tokenizer.decode(output[0], skip_special_tokens=True))
95
 
96
  ---
97
 
98
- ## ๐Ÿง  Model Architecture
99
 
100
  * Architecture: Transformer (Decoder-only)
101
  * Number of parameters: **~5B**
@@ -112,7 +122,7 @@ print(tokenizer.decode(output[0], skip_special_tokens=True))
112
 
113
  ---
114
 
115
- ## ๐Ÿ“š Training Data
116
 
117
  The model is pretrained on a large-scale bilingual corpus with a strong focus on **English and Hindi**, along with **dedicated Math and Code data**.
118
 
@@ -120,8 +130,8 @@ The model is pretrained on a large-scale bilingual corpus with a strong focus on
120
 
121
  - **English Natural Language:** `3.6T`
122
  - **Hindi Natural Language:** `2.77T`
123
- - **Math & Code:** `238.4B`
124
- - Math: **40%**
125
  - Code: **60%**
126
 
127
  Compared to **Param-1-2.9B**, **Param-1-5B** includes:
@@ -132,7 +142,7 @@ Compared to **Param-1-2.9B**, **Param-1-5B** includes:
132
 
133
  ---
134
 
135
- ## ๐Ÿ—๏ธ Training Details
136
 
137
  * Training framework: `NVIDIA NeMo`
138
  * Training infrastructure: `Yotta's Shakti Cloud`
@@ -141,7 +151,7 @@ Compared to **Param-1-2.9B**, **Param-1-5B** includes:
141
 
142
  ---
143
 
144
- ## โš ๏ธ Limitations
145
 
146
  * This is a **pretrained base model** and may require fine-tuning for instruction-following or chat use cases.
147
  * The model may reflect biases present in large-scale web and code data.
@@ -149,7 +159,7 @@ Compared to **Param-1-2.9B**, **Param-1-5B** includes:
149
 
150
  ---
151
 
152
- ## ๐Ÿ“œ License
153
 
154
  This model is released under the **BharatGen non-commercial license**.
155
 
 
3
  - hi
4
  - en
5
  pipeline_tag: text-generation
6
+ tags:
7
+ - bharatgen
8
+ - bilingual
9
+ - hindi
10
+ - english
11
+ - causal-lm
12
+ - transformers
13
+ license: other
14
  ---
15
 
16
 
17
  <div align="center">
18
+ <a href="https://bharatgen.com" target="_blank">
19
+ <img src="https://huggingface.co/bharatgenai/Param-1-2.9B-Instruct/resolve/main/BharatGen%20Logo%20(1).png" width="60%" alt="BharatGen" />
20
+ </a>
21
  </div>
22
  <hr>
23
  <div align="center">
24
+ <a href="https://bharatgen.com" target="_blank" style="margin: 4px;">
25
+ <img alt="Homepage" src="https://img.shields.io/badge/Homepage-Visit-blue.svg" />
26
+ </a>
27
+ <a href="./LICENSE" target="_blank" style="margin: 4px;">
28
  <img alt="License" src="https://img.shields.io/badge/License-yellow.svg" />
29
  </a>
30
  </div>
 
37
 
38
  ---
39
 
40
+ ## Key Highlights
41
 
42
  * **5B parameter** dense Transformer model
43
  * **Bilingual**: English and Hindi
 
47
 
48
  ---
49
 
50
+ ## Model Inference
51
 
52
  ```python
53
  from transformers import AutoTokenizer, AutoModelForCausalLM
 
83
 
84
  ---
85
 
86
+ ## Benchmarks
87
 
88
  | Task | **Param-1-5B (PT)** |
89
  | ----------------- | ------------------- |
 
105
 
106
  ---
107
 
108
+ ## Model Architecture
109
 
110
  * Architecture: Transformer (Decoder-only)
111
  * Number of parameters: **~5B**
 
122
 
123
  ---
124
 
125
+ ## Training Data
126
 
127
  The model is pretrained on a large-scale bilingual corpus with a strong focus on **English and Hindi**, along with **dedicated Math and Code data**.
128
 
 
130
 
131
  - **English Natural Language:** `3.6T`
132
  - **Hindi Natural Language:** `2.77T`
133
+ - **Math & Code:** `238.4B`
134
+ - Math: **40%**
135
  - Code: **60%**
136
 
137
  Compared to **Param-1-2.9B**, **Param-1-5B** includes:
 
142
 
143
  ---
144
 
145
+ ## Training Details
146
 
147
  * Training framework: `NVIDIA NeMo`
148
  * Training infrastructure: `Yotta's Shakti Cloud`
 
151
 
152
  ---
153
 
154
+ ## Limitations
155
 
156
  * This is a **pretrained base model** and may require fine-tuning for instruction-following or chat use cases.
157
  * The model may reflect biases present in large-scale web and code data.
 
159
 
160
  ---
161
 
162
+ ## License
163
 
164
  This model is released under the **BharatGen non-commercial license**.
165