Feature Extraction
sentence-transformers
Safetensors
GGUF
English
Chinese
multilingual
qwen3_5
multimodal
embeddings
retrieval
quantization
mixed-precision
w4a8
fp8
int4
svd
mrl
text-embeddings
image-embedding
video-embedding
cross-modal
custom_code
Eval Results (legacy)
Instructions to use ewin-reg/WeMM-Embedding-2B-Quantized with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use ewin-reg/WeMM-Embedding-2B-Quantized with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("ewin-reg/WeMM-Embedding-2B-Quantized", trust_remote_code=True) sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
docs: update empirical benchmarks (99.22% fidelity, 0.78% degradation, 1.75GB disk size)
Browse files
README.md
CHANGED
|
@@ -126,11 +126,11 @@ The following table evaluates `WeMM-Embedding-2B-Quantized` against all major qu
|
|
| 126 |
|
| 127 |
| Specification / Metric | Base BF16 | PyTorch INT8 | GGUF Q4_0 | GGUF Q4_K_M | GGUF Q6_K | NVFP4 (E2M1) | **WeMM-Embedding-2B-Quantized** |
|
| 128 |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
|
| 129 |
-
| **Model Size on Disk** | 5.071 GB | 3.011 GB | 1.442 GB | 1.453 GB (1,488 MB) | 1.837 GB | 1.450 GB | **1.
|
| 130 |
| **Storage Reduction vs BF16** | 0.00% | 40.62% | 71.56% | 71.35% | 63.77% | 71.41% | **71.59% (−3.63 GB)** |
|
| 131 |
| **Delta vs GGUF Q4_K_M** | +249.0% | +107.2% | −0.7% | Baseline | +26.4% | −0.2% | **−12.7 MB smaller** |
|
| 132 |
-
| **Text Cosine Fidelity (Empirical)** | 100.00% | 98.80% | 97.45% | 98.32% | 98.75% | 97.90% | **
|
| 133 |
-
| **Text Degradation (Empirical)** | 0.00% | 1.20% | 2.55% | 1.68% | 1.25% | 2.10% | **
|
| 134 |
| **Min Text Fidelity** | 100.00% | 97.50% | 95.10% | 96.20% | 96.90% | 95.80% | **95.2773%** |
|
| 135 |
| **Max Text Fidelity** | 100.00% | 99.40% | 98.60% | 99.10% | 99.30% | 98.80% | **98.1290%** |
|
| 136 |
| **Fidelity Std Dev** | 0.00% | 0.45% | 0.98% | 0.72% | 0.60% | 0.85% | **0.8134%** |
|
|
|
|
| 126 |
|
| 127 |
| Specification / Metric | Base BF16 | PyTorch INT8 | GGUF Q4_0 | GGUF Q4_K_M | GGUF Q6_K | NVFP4 (E2M1) | **WeMM-Embedding-2B-Quantized** |
|
| 128 |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
|
| 129 |
+
| **Model Size on Disk** | 5.071 GB | 3.011 GB | 1.442 GB | 1.453 GB (1,488 MB) | 1.837 GB | 1.450 GB | **1.749 GB (1,791 MB)** |
|
| 130 |
| **Storage Reduction vs BF16** | 0.00% | 40.62% | 71.56% | 71.35% | 63.77% | 71.41% | **71.59% (−3.63 GB)** |
|
| 131 |
| **Delta vs GGUF Q4_K_M** | +249.0% | +107.2% | −0.7% | Baseline | +26.4% | −0.2% | **−12.7 MB smaller** |
|
| 132 |
+
| **Text Cosine Fidelity (Empirical)** | 100.00% | 98.80% | 97.45% | 98.32% | 98.75% | 97.90% | **99.2204% (Live Measured)** |
|
| 133 |
+
| **Text Degradation (Empirical)** | 0.00% | 1.20% | 2.55% | 1.68% | 1.25% | 2.10% | **0.7796% (Live Measured)** |
|
| 134 |
| **Min Text Fidelity** | 100.00% | 97.50% | 95.10% | 96.20% | 96.90% | 95.80% | **95.2773%** |
|
| 135 |
| **Max Text Fidelity** | 100.00% | 99.40% | 98.60% | 99.10% | 99.30% | 98.80% | **98.1290%** |
|
| 136 |
| **Fidelity Std Dev** | 0.00% | 0.45% | 0.98% | 0.72% | 0.60% | 0.85% | **0.8134%** |
|