Feature Extraction
sentence-transformers
PyTorch
Core ML
ONNX
Safetensors
English
bert
sentence-similarity
mteb
custom_code
Eval Results (legacy)
text-embeddings-inference
Instructions to use jinaai/jina-embeddings-v2-base-en with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use jinaai/jina-embeddings-v2-base-en with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("jinaai/jina-embeddings-v2-base-en", trust_remote_code=True) sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
Could you please disclose the full list of training data for embedding supervised finetuning?
#32
by kwang2049 - opened
Although there is a general mention in the paper about it as
"Dataset with annotated negatives: We have prepared retrieval datasets, such as MSMarco [Bajaj et al., 2016] and Natural Questions
(NQ) [Kwiatkowski et al., 2019], in addition to multiple non-retrieval datasets like the Natural Language Inference (NLI) dataset [Bowman et al.,2015]. "
Could you please disclose the full list of the dataset names? This is very important for research work that wants to use Jina or follows it. Thanks in advance.
hi @kwang2049 yes, we used
- snli data from simcse: https://github.com/princeton-nlp/SimCSE#training, 1 hard negative + random negatives.
- msmarco, nq, quora-qa, hotpotqa and fever with mined hard negatives.
- cc news title description pairs with random negatives. https://huggingface.co/datasets/cc_news
each row consist of 17 items, including 1 anchor, 1 positive and 15 negatives.
Thanks❤️!
kwang2049 changed discussion status to closed