Token Classification
GLiNER
Safetensors
GLiNER2
Portuguese
extractor
portuguese
pt-br
brazilian-portuguese
ner
named-entity-recognition
open-vocabulary-ner
information-extraction
schema-guided-extraction
ontology-guided-extraction
operational-evidence
service-triage
technical-support
education
assistance
ottema
Instructions to use ottema/gliner2-ptbr with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER
How to use ottema/gliner2-ptbr with GLiNER:
from gliner import GLiNER model = GLiNER.from_pretrained("ottema/gliner2-ptbr") text = "Cristiano Ronaldo dos Santos Aveiro was born on 5 February 1985 in Funchal, Madeira, Portugal." labels = ["person", "date", "location"] entities = model.predict_entities(text, labels) for entity in entities: print(entity["text"], "=>", entity["label"]) - GLiNER2
How to use ottema/gliner2-ptbr with GLiNER2:
from gliner2 import AutoExtractor extractor = AutoExtractor.from_pretrained("ottema/gliner2-ptbr") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
Replace TBD with real metrics, fix broken links, add credits
Browse files
README.md
CHANGED
|
@@ -22,72 +22,88 @@ tags:
|
|
| 22 |
- education
|
| 23 |
- assistance
|
| 24 |
- ottema
|
| 25 |
-
datasets:
|
| 26 |
-
- ottema/GLiNER-PTBR-Core
|
| 27 |
base_model: fastino/gliner2-multi-v1
|
| 28 |
---
|
| 29 |
|
| 30 |
-
#
|
| 31 |
|
| 32 |
-
|
| 33 |
|
| 34 |
-
|
| 35 |
|
| 36 |
-
##
|
| 37 |
|
| 38 |
- **Base:** `fastino/gliner2-multi-v1` (Apache-2.0)
|
| 39 |
-
- **
|
| 40 |
-
- **
|
| 41 |
-
- **
|
| 42 |
-
- **
|
| 43 |
|
| 44 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 45 |
|
| 46 |
```python
|
| 47 |
-
from
|
| 48 |
|
| 49 |
-
model =
|
| 50 |
|
| 51 |
text = "A professora Ana comprou um notebook Dell em Campinas no dia 12/06."
|
| 52 |
labels = ["pessoa", "profissão", "produto", "marca", "local", "data"]
|
| 53 |
|
| 54 |
-
entities = model.
|
| 55 |
-
for
|
| 56 |
-
|
|
|
|
| 57 |
```
|
| 58 |
|
| 59 |
-
##
|
| 60 |
|
| 61 |
-
|
|
|
|
|
|
|
| 62 |
|---|---|---|---|
|
| 63 |
-
| `fastino/gliner2-multi-v1` (zero-shot
|
| 64 |
-
| `ottema/gliner2-ptbr` (
|
| 65 |
|
| 66 |
-
|
| 67 |
|
| 68 |
-
|
|
|
|
|
|
|
|
|
|
| 69 |
|
| 70 |
-
|
| 71 |
|
| 72 |
-
##
|
| 73 |
|
| 74 |
-
-
|
| 75 |
-
-
|
| 76 |
-
- Não é classificador final de cobertura/cobertura contratual.
|
| 77 |
|
| 78 |
-
##
|
| 79 |
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 88 |
|
| 89 |
-
##
|
| 90 |
|
| 91 |
-
- `ottema/
|
| 92 |
-
- `ottema/
|
| 93 |
-
- `ottema/gliner2-ptbr-ontoevidence-
|
|
|
|
| 22 |
- education
|
| 23 |
- assistance
|
| 24 |
- ottema
|
|
|
|
|
|
|
| 25 |
base_model: fastino/gliner2-multi-v1
|
| 26 |
---
|
| 27 |
|
| 28 |
+
# ottema/gliner2-ptbr (v0.4 — generalist)
|
| 29 |
|
| 30 |
+
**Open-vocabulary NER for Brazilian Portuguese, fine-tuned for informal and operational text (chat, atendimento, suporte).**
|
| 31 |
|
| 32 |
+
This is the **generalist release**. For HAREM-specialized (best entity F1 among compared models), see [`ottema/gliner2-ptbr-harem`](https://huggingface.co/ottema/gliner2-ptbr-harem). For ontology-guided evidence extraction, see [`ottema/gliner2-ptbr-ontoevidence`](https://huggingface.co/ottema/gliner2-ptbr-ontoevidence).
|
| 33 |
|
| 34 |
+
## Model details
|
| 35 |
|
| 36 |
- **Base:** `fastino/gliner2-multi-v1` (Apache-2.0)
|
| 37 |
+
- **Type:** GLiNER2 (bidirectional encoder + span scoring)
|
| 38 |
+
- **Language:** pt-BR (with English fallback)
|
| 39 |
+
- **Size:** ~307M parameters
|
| 40 |
+
- **License:** Apache-2.0
|
| 41 |
|
| 42 |
+
## Intended use
|
| 43 |
+
|
| 44 |
+
General-purpose open-vocabulary NER for **informal and operational** Brazilian Portuguese text: atendimento, chat, suporte técnico, educação. Trained on synthetic data covering pessoas, profissões, locais, organizações, documentos, produtos, marcas, tecnologias, telefones, e-mails, datas, e valores monetários.
|
| 45 |
+
|
| 46 |
+
If you need a model benchmarked on journalistic Portuguese (HAREM), use [`ottema/gliner2-ptbr-harem`](https://huggingface.co/ottema/gliner2-ptbr-harem) instead.
|
| 47 |
+
|
| 48 |
+
## Usage
|
| 49 |
|
| 50 |
```python
|
| 51 |
+
from gliner2 import GLiNER2
|
| 52 |
|
| 53 |
+
model = GLiNER2.from_pretrained("ottema/gliner2-ptbr")
|
| 54 |
|
| 55 |
text = "A professora Ana comprou um notebook Dell em Campinas no dia 12/06."
|
| 56 |
labels = ["pessoa", "profissão", "produto", "marca", "local", "data"]
|
| 57 |
|
| 58 |
+
entities = model.extract_entities(text, labels, threshold=0.5)
|
| 59 |
+
for label, spans in entities["entities"].items():
|
| 60 |
+
for span in spans:
|
| 61 |
+
print(f"{span} -> {label}")
|
| 62 |
```
|
| 63 |
|
| 64 |
+
## Performance
|
| 65 |
|
| 66 |
+
Evaluation on `data/gliner_ptbr_core/test.jsonl` (synthetic generalist benchmark, threshold 0.3):
|
| 67 |
+
|
| 68 |
+
| Model | entity_F1 | span_F1 | label_F1 |
|
| 69 |
|---|---|---|---|
|
| 70 |
+
| `fastino/gliner2-multi-v1` (zero-shot) | 0.9333 | 0.9347 | 0.9855 |
|
| 71 |
+
| **`ottema/gliner2-ptbr` (v0.4)** | **0.9976** | **0.9976** | **1.0000** |
|
| 72 |
|
| 73 |
+
On HAREM (163 samples, 2511 entities, journalistic PT-BR — out-of-distribution for this generalist):
|
| 74 |
|
| 75 |
+
| Model | entity_F1 | Δ vs baseline |
|
| 76 |
+
|---|---|---|
|
| 77 |
+
| `fastino/gliner2-multi-v1` (zero-shot) | 0.4271 | (reference) |
|
| 78 |
+
| `ottema/gliner2-ptbr` (v0.4) | 0.4132 | -1.39 pp |
|
| 79 |
|
| 80 |
+
The generalist is best for the synthetic informal-PT-BR distribution it was trained on. For journalistic text, see the HAREM-specialized model.
|
| 81 |
|
| 82 |
+
## Inference
|
| 83 |
|
| 84 |
+
- **GPU:** ~30 ms per text (median, 32 batch)
|
| 85 |
+
- **CPU:** ~50 ms per short text (≤128 tokens)
|
|
|
|
| 86 |
|
| 87 |
+
## Limitations
|
| 88 |
|
| 89 |
+
- Trained primarily on synthetic data; coverage may be limited in highly specialized domains.
|
| 90 |
+
- Performance may degrade on very long texts (>512 tokens).
|
| 91 |
+
- Not a substitute for domain-specific classifiers in regulated workflows.
|
| 92 |
+
- Out-of-distribution on journalistic text (use HAREM-specialized instead).
|
| 93 |
+
|
| 94 |
+
## Credits
|
| 95 |
+
|
| 96 |
+
- **Base architecture:** GLiNER2 (Urchade Zaratiana et al.)
|
| 97 |
+
- **Base weights:** `fastino/gliner2-multi-v1` (Fastino)
|
| 98 |
+
- **Encoder:** microsoft/mdeberta-v3-base
|
| 99 |
+
- **Fine-tuning + datasets:** Ottema
|
| 100 |
+
|
| 101 |
+
## License
|
| 102 |
+
|
| 103 |
+
Apache-2.0
|
| 104 |
|
| 105 |
+
## See also
|
| 106 |
|
| 107 |
+
- [`ottema/gliner2-ptbr-harem`](https://huggingface.co/ottema/gliner2-ptbr-harem) — HAREM-specialized (best entity F1)
|
| 108 |
+
- [`ottema/gliner2-ptbr-ontoevidence`](https://huggingface.co/ottema/gliner2-ptbr-ontoevidence) — ontology-guided evidence extraction (in development)
|
| 109 |
+
- [`ottema/gliner2-ptbr-ontoevidence-data`](https://huggingface.co/datasets/ottema/gliner2-ptbr-ontoevidence-data) — OntoEvidence-BR dataset
|