Instructions to use BAAI/bge-large-en-v1.5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use BAAI/bge-large-en-v1.5 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("BAAI/bge-large-en-v1.5") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Transformers
How to use BAAI/bge-large-en-v1.5 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="BAAI/bge-large-en-v1.5")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("BAAI/bge-large-en-v1.5") model = AutoModel.from_pretrained("BAAI/bge-large-en-v1.5", device_map="auto") - Inference
- Notebooks
- Google Colab
- Kaggle
Fine tune larger dataset lower eval result.
Hello,
I'm fine tuning the model using proprietary patent dataset with 512 max length and 2 set of data, one 218 rows, one 5800 rows. Each query+instruction has 3 passages to form a positive pair. I didn't include negative pairs in the dataset.
Epoch is 3, and for the larger dataset, use 1e-05 learning rate.
I use Sentence transformer to do fine tuning.
The evaluation is done by calculating correlation between anchor/target cosine similarities and ground truth scores in https://www.kaggle.com/competitions/us-patent-phrase-to-phrase-matching
Surprisingly, the model fine tuned by 5800 rows has a lower score compared to the one with 200 rows.
Any idea what could go wrong?